Method for obtaining sample audio data, speech recognition method and related apparatus
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2026-08-11
AI Technical Summary
需要花费大量精力去采集样本音频数据,对样本音频数据的获取效率较低
[0011]上述方案,通过获取目标音频数据的标注文本和至少两个参考文本,标注文本是基于目标音频数据的字幕确定的,各参考文本是分别利用不同的语音识别模型对目标音频数据进行识别得到的,基于标注文本和至少两个参考文本之间的比对结果,确定目标音频数据的类型,由于类型表征标注文本的准确性或者目标音频数据的语音识别难度,根据类型可以得到目标音频数据对于语言识别模型的语言识别难度或者目标音频数据的字幕的准确度,以此对目标音频数据执行与类型匹配的处理,并将经处理后的目标音频数据作为用于对目标语音识别模型进行训练的样本音频数据,能够提高获取对目标语言识别模型训练的样本音频数据的效率,筛选出对目标语音识别模型训练更有利的样本音频数据,提高对于训练目标语音识别模型的弱监督数据获取的准确率和召回率,从而提高训练后的目标语音识别模型对语言识别的准确度。
Smart Images

Figure CN117894300B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition technology, and in particular to a method for acquiring sample audio data, a speech recognition method, and related apparatus. Background Technology
[0002] Language is the most important carrier of human thought, and speech recognition technology is the technology that enables machines to receive, recognize, and understand speech signals and convert them into corresponding digital signals. With the development of artificial intelligence technology, speech recognition technology has been widely used in many fields, such as smart mobile terminals, smart homes, and in-vehicle devices.
[0003] Currently, speech recognition involves multiple fields and diverse scenarios. Training a speech recognition model typically requires collecting a large amount of audio data, encompassing a sufficient number of scenarios. This necessitates significant effort in collecting sample audio data, resulting in relatively low efficiency. Summary of the Invention
[0004] The main technical problem addressed by this application is to provide a method for acquiring sample audio data, a speech recognition method, and related devices, which can improve the efficiency of acquiring sample audio data for training target language recognition models.
[0005] To address the aforementioned issues, the first aspect of this application provides a method for acquiring sample audio data. This method includes: acquiring labeled text and at least two reference texts for target audio data, wherein the labeled text is determined based on subtitles of the target audio data, and each reference text is obtained by recognizing the target audio data using a different speech recognition model; determining the type of the target audio data based on the comparison results between the labeled text and the at least two reference texts, wherein the type characterizes the accuracy of the labeled text or the speech recognition difficulty of the target audio data; performing type matching processing on the target audio data, and using the processed target audio data as sample audio data for training a target speech recognition model.
[0006] To address the aforementioned issues, a second aspect of this application provides a speech recognition method, comprising: acquiring audio data to be recognized; performing speech recognition on the audio data to be recognized using a target speech recognition model to obtain a speech recognition result; wherein the target speech recognition model is trained using sample audio data obtained by the aforementioned method for acquiring sample audio data.
[0007] To address the aforementioned problems, a third aspect of this application provides a sample audio data acquisition device, comprising: an acquisition module, a comparison module, and a processing module. The acquisition module acquires labeled text and at least two reference texts of the target audio data. The labeled text is determined based on the subtitles of the target audio data, and each reference text is obtained by recognizing the target audio data using a different speech recognition model. The comparison module determines the type of the target audio data based on the comparison results between the labeled text and the at least two reference texts. The type characterizes the accuracy of the labeled text or the speech recognition difficulty of the target audio data. The processing module performs type matching processing on the target audio data and uses the processed target audio data as sample audio data for training the target speech recognition model.
[0008] To address the aforementioned problems, a fourth aspect of this application provides a speech recognition device, comprising an acquisition module and a recognition module. The acquisition module acquires audio data to be recognized; the recognition module performs speech recognition on the audio data using a target speech recognition model to obtain a speech recognition result; wherein the target speech recognition model is trained using sample audio data obtained through the aforementioned sample audio data acquisition method.
[0009] To address the aforementioned problems, a fifth aspect of this application provides a computer device comprising a memory and a processor coupled to each other, the memory storing program data, and the processor executing the program data to implement any step of the aforementioned method for acquiring sample audio data and / or the speech recognition method.
[0010] To address the aforementioned problems, a sixth aspect of this application provides a computer-readable storage medium storing program data executable by a processor, the program data being used to implement any step of the above-described method for acquiring sample audio data and / or the speech recognition method.
[0011] The above scheme obtains labeled text and at least two reference texts from the target audio data. The labeled text is determined based on the subtitles of the target audio data, and each reference text is obtained by recognizing the target audio data using a different speech recognition model. Based on the comparison results between the labeled text and at least two reference texts, the type of the target audio data is determined. Since the type represents the accuracy of the labeled text or the speech recognition difficulty of the target audio data, the speech recognition difficulty of the target audio data for the speech recognition model or the accuracy of the subtitles of the target audio data can be obtained according to the type. The target audio data is then processed for type matching, and the processed target audio data is used as sample audio data for training the target speech recognition model. This improves the efficiency of obtaining sample audio data for training the target speech recognition model, filters out sample audio data that is more conducive to training the target speech recognition model, and improves the accuracy and recall of weakly supervised data acquisition for training the target speech recognition model, thereby improving the accuracy of the trained target speech recognition model for speech recognition.
[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in this application, the accompanying drawings required in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them:
[0014] Figure 1 This is a flowchart illustrating the first embodiment of the method for obtaining sample audio data in this application;
[0015] Figure 2 This is a flowchart illustrating an embodiment of step S12 of this application;
[0016] Figure 3 This is a flowchart illustrating an embodiment of step S13 of this application;
[0017] Figure 4 This is a flowchart illustrating the second embodiment of the method for obtaining sample audio data in this application;
[0018] Figure 5 This is a flowchart illustrating the third embodiment of the method for obtaining sample audio data in this application;
[0019] Figure 6 This is a flowchart illustrating an embodiment of forced alignment of reference text and speech signals;
[0020] Figure 7 This is a flowchart illustrating an embodiment of step S36 of this application;
[0021] Figure 8 This is a flowchart illustrating an embodiment of the speech recognition method of this application;
[0022] Figure 9 This is a schematic diagram of an embodiment of the device for acquiring sample audio data in this application;
[0023] Figure 10 This is a schematic diagram of the structure of an embodiment of the speech recognition device of this application;
[0024] Figure 11 This is a schematic diagram of the structure of an embodiment of the computer device of this application;
[0025] Figure 12 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0027] The terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0028] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0029] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0030] This application provides the following embodiments, and each embodiment will be described in detail below.
[0031] It is understood that the method for acquiring sample audio data and the speech recognition method in this application can be executed by a computer device, which can be any device with processing capabilities, such as a mobile phone, computer, server, etc., and this application does not impose any restrictions on this.
[0032] Please see Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the method for acquiring sample audio data according to this application. The method may include the following steps:
[0033] S11: Obtain the labeled text and at least two reference texts of the target audio data. The labeled text is determined based on the subtitles of the target audio data, and each reference text is obtained by recognizing the target audio data using a different speech recognition model.
[0034] This involves acquiring target audio data. The target audio data can be audio data with subtitles, or audio data with subtitles corresponding to the target audio data. In the case of unsubtitled target audio data, subtitles can be obtained by having an annotator listen to the target audio data and annotate it. Subtitles can represent the audio content of the target audio data; for example, subtitles for a speaker's audio data can represent the speaker's speech, such as a passage of text. For example, the target audio data can be audio data from video data obtained from film and television works, audio data from broadcast audio, audio data corresponding to video data, etc. This application does not limit the method of acquiring the target audio data and its corresponding subtitles.
[0035] In some implementations, the subtitles for the target audio data can be external subtitles or embedded subtitles. External subtitles are contained in the subtitle file of the target audio data's peripheral device; this method provides a separate subtitle file containing timestamp information and subtitle information. Embedded subtitles are the subtitle text in the subtitle display area of the image frame of the corresponding video data. This method does not require an additional subtitle file; the subtitles are automatically displayed within the image frame of the corresponding video data, such as at the bottom, left, or right of the screen.
[0036] In some implementations, the annotation text can be obtained based on the subtitles of the target audio data. For example, the subtitles can be used directly as the annotation text, or the annotation text can be obtained by extracting the subtitles from the subtitle display area of the image frame of the video data corresponding to the target audio data. This application is not limited to these methods.
[0037] In some implementations, different speech recognition models can be used to recognize the target audio data, resulting in at least two reference texts. The number of different speech recognition models is at least two, and these multiple different models can all have speech recognition functionality in at least the same language, such as language recognition functionality for the pronunciation language of the target audio data.
[0038] For example, different speech recognition models can include monolingual speech recognition models (Automatic Speech Recognition, ASR) and multilingual speech recognition models. Monolingual speech recognition models, also known as monolingual ASR models, can be used to recognize speech in a single language, such as Mandarin Chinese. Multilingual speech recognition models, also known as multilingual ASR models, can be used to recognize speech in multiple languages, such as Mandarin Chinese or multiple regional dialects. Large-scale multilingual ASR models are typically trained on large-scale multilingual data and exhibit good robustness to accents, background noise, and technical jargon.
[0039] In some implementations, the different speech recognition models mentioned above may include speech recognition models of the same type as the target language recognition model, such as models that are the same language or models that are all in the same language. Other models that are not part of the target language recognition model may be speech recognition models that have at least the same type of speech recognition function as the target speech recognition model, and these models may be speech recognition models that have been pre-trained using audio data samples. Of course, it is understood that all the different speech recognition models mentioned above can be speech recognition models that have been pre-trained using various audio data samples. For example, the target language recognition model may be a monolingual ASR model, and the multilingual ASR model may be a pre-trained speech recognition model, or both the monolingual ASR model and the multilingual ASR model may be pre-trained speech recognition models, etc. This application does not impose any restrictions on this.
[0040] S12: Based on the comparison results between the labeled text and at least two reference texts, determine the type of the target audio data. The type represents the accuracy of the labeled text or the speech recognition difficulty of the target audio data.
[0041] The labeled text of the target audio data obtained above is compared with at least two reference texts obtained by different speech recognition models to obtain the comparison results between the labeled text and at least two reference texts. Based on the comparison results, the accuracy of the labeled text or the speech recognition difficulty of different speech recognition models on the target audio data can be obtained to determine the type of target audio data.
[0042] For example, the type of target audio data can be determined by comparing the labeled text with at least two pairs of reference texts. The comparison results can be similarity scores. If the similarity between the labeled text and at least two reference texts is less than a first preset threshold, it indicates that the accuracy of the labeled text is low. If the similarity between the labeled text and at least one of the at least two reference texts is not less than a second preset threshold, it indicates that the accuracy of the labeled text is relatively correct, and the speech recognition model of at least one of the reference texts has a high difficulty in recognizing the target audio data. If the similarity of each pair of comparison results is greater than a third preset threshold, it indicates that the accuracy of the labeled text is relatively correct, and the speech recognition model of at least one of the reference texts has a relatively moderate difficulty in recognizing the target audio data. This application is not limited to this method of determining the type.
[0043] In some embodiments, please refer to Figure 2 This embodiment can further extend step S12 of the above embodiment. Based on the comparison results between the labeled text and at least two reference texts, the type of the target audio data is determined. This embodiment may include the following steps:
[0044] S121: Using the comparison results, identify the abnormal text that meets the preset rule conditions from the labeled text and at least two reference texts.
[0045] The comparison results may include similarity scores, which can be denoted as SCORE. Specifically, the annotated text and at least two reference texts can be grouped into pairs, and the similarity score of each pair can be obtained to generate comparison results. Then, using the comparison results, abnormal texts that meet preset rule conditions are identified from the annotated text and at least two reference texts. The preset rule conditions correspond to the conditions that indicate anomalies in the annotated text and at least two reference texts.
[0046] In some embodiments, the aforementioned at least two reference texts may include a first reference text recognized by a first speech recognition model and a second reference text recognized by a second speech recognition model. The first speech recognition model and the target speech recognition model may be of the same type. The following description uses the example of the first speech recognition model and the target speech recognition model being monolingual ASR models, and the second speech recognition model being a multilingual ASR model; however, this application does not impose any limitations on this.
[0047] In some implementations, the recognition accuracy of the second speech recognition model can be higher than that of the first speech recognition model, so that the second reference text and the labeled text can be used as an accuracy reference for the first reference text recognized by the first speech recognition model.
[0048] In some implementations, to facilitate the description of the similarity between groups, the labeled text can be referred to as the "first label," the first reference text as the "second label," and the second reference text as the "third label." Thus, the similarity between the labeled text and the first reference text is denoted as SCORE. 12 The similarity between the annotated text and the second reference text is denoted as SCORE. 13 The similarity between the first reference text and the second reference text is denoted as SCORE. 23 Then, by combining the similarity of each group, anomalous text can be identified from the labeled text and at least two reference texts.
[0049] In some implementations, it can be determined whether the comparison result meets a first preset rule condition. The first preset rule condition may include: the similarity between the labeled text and each reference text is less than a first preset threshold; the similarity between each reference text is greater than a second preset threshold; and the first preset threshold is less than or equal to the second preset threshold. If the comparison result meets the first preset rule condition, the labeled text can be identified as abnormal text in response to this condition.
[0050] To facilitate understanding, the following example illustrates the point:
[0051] First note: Cai Wenji, you haven't developed yet. Focus on scouting and wait for me to use my ultimate to support you.
[0052] Second note: My character, Wenji, is still underdeveloped. Please provide good vision in the bushes later.
[0053] Third note: I haven't even fully developed yet, so Cai Wenji should provide good vision in the bushes later.
[0054] The similarity of each group is obtained through the above annotations, for example, the similarity score (SCORE). 12 SCORE similarity 13 All are less than the first preset threshold, and the similarity score between each reference text is SCORE. 23 If the value exceeds the second preset threshold, it indicates that the reference texts for speech recognition of the target audio data by different speech recognition models are similar, while the annotated text of the subtitles differs significantly from the reference texts of the speech recognition results of each speech recognition model. In this case, the annotated text is determined to be incorrect, i.e., abnormal. The annotated text can be identified as abnormal text in response to the comparison result meeting the first preset rule condition.
[0055] In some implementations, if the comparison result does not meet the first preset rule condition, the labeled text can be determined to be correct text. Then, it is determined whether the comparison result meets the second preset rule condition. The second preset rule condition may include: the similarity between the labeled text and the first reference text is less than a third preset threshold; the similarity between the labeled text and the second reference text is greater than a fourth preset threshold; and the third preset threshold is less than or equal to the fourth preset threshold. If the comparison result does not meet the first preset rule condition but meets the second preset rule condition, in response to the comparison result not meeting the first preset rule condition but meeting the second preset rule condition, the first reference text among at least two reference texts can be determined as abnormal text.
[0056] For example, if the similarity SCORE12 is less than the third preset threshold, while the similarity SCORE13 is greater than the fourth preset threshold, it indicates that the reference texts for speech recognition of the target audio data by different speech recognition models are quite different, and the above-mentioned labeled text is not abnormal text. In this case, the first reference text obtained by the first speech recognition model is incorrect, that is, abnormal. It can be seen that the first speech recognition model identifies abnormalities and can determine the first reference text among at least two reference texts as abnormal text in response to the comparison result not meeting the first preset rule condition and meeting the second preset rule condition.
[0057] Similarly, if the similarity between the labeled text and the first reference text is greater than a fourth preset threshold, and the similarity between the labeled text and the second reference text is less than a third preset threshold, the second reference text is identified as abnormal text. This application does not limit the method for identifying abnormal text.
[0058] S122: Determine the type of the target audio data based on the abnormal text.
[0059] Optionally, the abnormal text can be annotation text, first reference text, or first reference text, etc., and the type of target audio data can be determined based on the text of the abnormal text.
[0060] In some implementations, in response to the annotation text being abnormal text, indicating that the subtitles of the target audio data are incorrect or abnormal, the type of the target audio data is determined to be dirty data type, which indicates that the annotation text is inaccurate.
[0061] In some implementations, in response to the first reference text being abnormal text, the type of the target audio data is determined to be a first recognition difficulty type, which means that the first speech recognition model or the target speech recognition model has a high difficulty in recognizing the target audio data and is prone to recognition errors.
[0062] In some implementations, in response to the fact that neither the labeled text nor the first reference text is abnormal text, the type of the target audio data is determined to be the second recognition difficulty type. The speech recognition difficulty of the first recognition difficulty type is higher than that of the second recognition difficulty type, which means that the first speech recognition model or the target speech recognition model has a lower recognition difficulty for the target audio data and the recognition is more accurate.
[0063] The above scheme can identify abnormal texts that meet preset rules by comparing the results with the labeled text and at least two reference texts. Then, based on the abnormal texts, the type of target audio data can be determined. This can more accurately determine the accuracy of the labeled text or the speech recognition difficulty of the target audio data.
[0064] Continue reading Figure 1 The steps following step S12 are as follows:
[0065] S13: Perform type matching processing on the target audio data, and use the processed target audio data as sample audio data for training the target speech recognition model.
[0066] Different types have their own corresponding processing methods. Based on the type of the target audio data determined above, the target audio data can be processed according to the type, and the processed target audio data can be used as sample audio data for training the target speech recognition model.
[0067] In some implementations, the subtitles of the target audio data can be considered as relatively accurate language recognition text of the target audio data. When the target audio data is used as sample audio data for training the target speech recognition model, during the training process, the labeled text can be used as relatively accurate labeled text of the audio content of the sample audio data. That is, the labeled text and the target audio data are used as a data pair as sample audio data.
[0068] In some embodiments, please refer to Figure 3 This embodiment can further extend step S13 of the above embodiment. Performing type matching processing on the target audio data and using the processed target audio data as sample audio data for training the target speech recognition model, this embodiment may include at least one step:
[0069] S131: In response to the target audio data being of the dirty data type, discard the target audio data. The dirty data type indicates that the annotation text is inaccurate.
[0070] When the target audio data is determined to be of the dirty data type, which indicates that the annotation text is inaccurate, the target audio data can be removed in response to the dirty data type.
[0071] S132: In response to the target audio data being of the first recognition difficulty type, adjust the training weights of the target audio data and use the adjusted target audio data as sample audio data.
[0072] When the target audio data is classified as "Type 1 Difficulty," the first reference text identified by the first speech recognition model is classified as "abnormal text," indicating an anomaly in the model's recognition of the target audio data. Since both the first and target speech recognition models are of the same type, the target audio data presents a significant challenge for the model's recognition, making it beneficial sample audio data for training. Therefore, in response to the target audio data being classified as "Type 1 Difficulty," the training weights of the target audio data can be adjusted, such as by increasing its weight, and the adjusted target audio data can be used as sample audio data.
[0073] S133: In response to the target audio data being of the second recognition difficulty type, the target audio data is directly used as the sample audio data. The speech recognition difficulty of the first recognition difficulty type is higher than that of the second recognition difficulty type.
[0074] In response to the target audio data being of the second recognition difficulty type, it can be said that the first speech recognition model is more accurate in recognizing the target audio data, and the target audio data is directly used as the sample audio data. The speech recognition difficulty of the first recognition difficulty type is higher than that of the second recognition difficulty type.
[0075] The above scheme obtains labeled text and at least two reference texts from the target audio data. The labeled text is determined based on the subtitles of the target audio data, and each reference text is obtained by recognizing the target audio data using a different speech recognition model. Based on the comparison results between the labeled text and at least two reference texts, the type of the target audio data is determined. Since the type represents the accuracy of the labeled text or the speech recognition difficulty of the target audio data, the speech recognition difficulty of the target audio data for the speech recognition model or the accuracy of the subtitles of the target audio data can be obtained according to the type. The target audio data is then processed for type matching, and the processed target audio data is used as sample audio data for training the target speech recognition model. This improves the efficiency of obtaining sample audio data for training the target speech recognition model, filters out sample audio data that is more conducive to training the target speech recognition model, and improves the accuracy and recall of weakly supervised data acquisition for training the target speech recognition model, thereby improving the accuracy of the trained target speech recognition model for speech recognition.
[0076] In some embodiments, in step S11 described above, the subtitles of the target audio data can be peripheral subtitles or embedded subtitles. Peripheral subtitles are contained in the subtitle file of the peripheral device of the target audio data. Embedded subtitles are the subtitle text in the subtitle display area of the image frame of the video data corresponding to the target audio data. See the following embodiments for details.
[0077] For cases where the subtitles for the target audio data can be external subtitles, this application provides the following embodiments for illustration. Please refer to [link / reference]. Figure 4 , Figure 4 This is a flowchart illustrating a second embodiment of the method for acquiring sample audio data according to this application. The method may include the following steps:
[0078] S21: Use the peripheral subtitles of the target audio data as annotation text.
[0079] In some implementations, when the subtitles for the target audio data can be external subtitles, step S11 of the above embodiment may include steps S21 to S23 of this embodiment.
[0080] In this step, the external subtitles are included in the subtitle file of the target audio data peripheral. The external subtitles can be a separate subtitle file set separately for the target audio data. The subtitle file includes subtitle information (i.e., the external subtitles) and the first timestamp corresponding to the external subtitles. The first timestamp refers to the time mark that is synchronized with the external subtitles in the target audio data. The time node corresponding to the external subtitles in the target audio data can be determined by the first timestamp.
[0081] The subtitle file is processed to extract subtitle information and first timestamp information, resulting in peripheral subtitles and their corresponding first timestamps. The peripheral subtitles of the target audio data can then be used as annotation text for the target audio data.
[0082] S22: Segment the target audio data according to the first timestamp to obtain multiple sub-audio data.
[0083] Based on the first timestamp corresponding to the above-mentioned annotated text, the target audio data is segmented to divide the long audio data into multiple short audio data, resulting in multiple sub-audio data.
[0084] S23: Use different speech recognition models to recognize multiple sub-audio data to obtain at least two reference texts.
[0085] Multiple different speech recognition models can be used to recognize multiple sub-audio data separately, obtaining at least two reference texts. For example, each speech recognition model can be used to recognize multiple sub-audio data separately to obtain the reference text for each sub-audio data, and then the reference texts of each sub-audio data can be combined to obtain the reference text for the target audio data corresponding to each speech recognition model.
[0086] For example, the language recognition model may include a monolingual ASR model and a multilingual ASR model. By using a monolingual ASR model to recognize multiple sub-audio data, a first reference text can be obtained. By using a multilingual ASR model to recognize multiple sub-audio data, a second reference text can be obtained.
[0087] The above scheme segments the target audio data according to the first timestamp to obtain multiple sub-audio data, that is, after the long audio data is segmented into multiple short audio data, different speech recognition models are used to recognize the multiple sub-audio data to obtain at least two reference texts, which can improve the accuracy of each speech recognition model in recognizing the target audio data.
[0088] In some implementations, after obtaining the labeled text and at least two reference texts of the target audio data in the above steps, the labeled text and at least two reference texts can be aligned, and / or the time drift of the labeled text can be determined. After step S23, the following steps may also be included:
[0089] S24: Align the annotated text and each reference text to obtain the aligned annotated text and each reference text.
[0090] In this step, the minimum edit distance between each annotated text and each reference text can be obtained respectively, and alignment processing is performed according to the minimum edit distance to obtain the aligned annotated text and each reference text.
[0091] Among them, the edit distance refers to the number of editing operations required to change one text to another text, including three operations: inserting text, deleting text, and replacing text. The minimum edit distance is the minimum number of editing operations required to change one text to another text.
[0092] S25: In response to the aligned annotated text satisfying the time drift condition, use the evaluation speech recognition model to recognize the target audio data to obtain the evaluation text and the second timestamp corresponding to the evaluation text.
[0093] After obtaining the aligned annotated text and each reference text, it can be determined whether the aligned annotated text satisfies the time drift condition. Among them, the time drift condition includes: there is an abnormal character insertion at the preset position of the aligned annotated text relative to the aligned reference text, and the preset position is the head position or the tail position.
[0094] To facilitate the understanding of the time drift condition, the following is an example for illustration. The specific texts after alignment are as follows:
[0095] Annotated text: Cai Wenji hasn't developed yet. Keep a good lookout. Wait for me to use my ultimate skill and then support you.
[0096] First reference text: I haven't developed yet. Cai Wenji, keep a good lookout for me in the grass later.
[0097] Second reference text: I haven't developed yet. Cai Wenji, keep a good lookout in the grass later.
[0098] For the above texts, obtain the alignment result ALIGN obtained by comparing the aligned annotated text and the first reference text 12 , and, obtain the alignment result ALIGN obtained by comparing the aligned annotated text and the second reference text 13 . It is known that in the alignment result ALIGN 12 and the alignment result ALIGN 13 , insertion errors occur at the beginning or end of the sentence in the annotated text. Taking the above text as an example, the character 'I' is inserted at the beginning of the sentence, that is, there is an abnormal insertion of the character 'I' at the head position, and it is determined that the annotated text has time drift, and it is determined that the aligned annotated text satisfies the time drift condition.
[0099] If the aligned annotation text satisfies the time drift condition, the target audio data is recognized using the evaluation speech recognition model to obtain the evaluation text and its corresponding second timestamp. Alternatively, if the aligned annotation text does not satisfy the time drift condition, step S12 and subsequent steps are performed directly using the aligned annotation text and each reference text.
[0100] In some implementations, the speech recognition model is evaluated and selected by testing different speech recognition models using a test dataset of a preset type of speech. The different speech recognition models refer to the aforementioned different speech recognition models, including a first speech recognition model and a second speech recognition model, such as a monolingual ASR model and a multilingual ASR model. The test dataset of the preset type of speech can be a test dataset for the language corresponding to the first speech recognition model. For example, on the test dataset corresponding to the monolingual ASR model, each speech recognition model is tested to obtain test results. Then, based on the test results, the speech recognition model with higher accuracy and better recognition performance is selected as the evaluation speech recognition model.
[0101] The above scheme uses an evaluation speech recognition model to identify the target audio data, obtains the evaluation text and the corresponding second timestamp, so that a more accurate timestamp can be obtained to segment the target audio data.
[0102] S26: Divide the target audio data according to the second timestamp to obtain multiple new sub-audio data.
[0103] Based on the second timestamp corresponding to the evaluation text obtained by the evaluation speech recognition model, the target audio data is segmented to obtain multiple new sub-audio data. Then, the steps of recognizing each sub-audio data using different speech recognition models to obtain at least two reference texts, as well as subsequent steps, are re-executed on these new sub-audio data.
[0104] The step of recognizing multiple sub-audio data to obtain at least two reference texts can be referred to the specific implementation process of steps S23 to S24 above, and will not be repeated here.
[0105] In some implementations, after obtaining the labeled text and each reference text in this embodiment, the above steps S12 and subsequent steps can be performed.
[0106] The above solution aligns the labeled text and each reference text to obtain aligned labeled text and each reference text. It then determines whether the labeled text has time drift based on whether there are character insertion anomalies relative to the aligned reference text at a preset position. If the aligned labeled text meets the time drift condition, an evaluation speech recognition model is used to recognize the target audio data, obtaining the evaluation text and its corresponding second timestamp. The target audio data is then segmented based on the second timestamp to obtain multiple new sub-audio data. The process of re-executing the multiple sub-audio data using different speech recognition models to obtain at least two reference texts, along with subsequent steps, yields more accurate reference texts and solves the problem of time drift in peripheral subtitles.
[0107] In addition, it can improve the accuracy of weakly supervised data for peripheral subtitles, and at the same time, it can filter out sample audio data that is more beneficial to the training of the target speech recognition model, thereby improving the training quality of sample audio data.
[0108] For cases where the subtitles for the target audio data can be embedded, this application provides the following embodiments for illustration. Please refer to [link / reference]. Figure 5 , Figure 5 This is a flowchart illustrating a third embodiment of the method for acquiring sample audio data according to this application. The method may include the following steps:
[0109] S31: Use the first speech recognition model to recognize the target audio data to obtain the first reference text and the third timestamp corresponding to the first reference text.
[0110] In some implementations, when the subtitles for the target audio data can be external subtitles, step S11 of the above embodiment may include steps S31 to S34 of this embodiment.
[0111] In this embodiment, at least two reference texts include a first reference text recognized by a first speech recognition model and a second reference text recognized by a second speech recognition model. The first speech recognition model and the target speech recognition model are of the same type.
[0112] In this step, a first speech recognition model can be used to recognize the target audio data to obtain a first reference text and a corresponding third timestamp. For example, the first speech recognition model can be a monolingual ASR model, which is used to recognize the target audio data to obtain the first reference text and the corresponding third timestamp.
[0113] S32: Extract image frames of video data corresponding to the target audio data according to the third timestamp, and perform text recognition on the subtitle display area of each image frame to obtain the labeled text.
[0114] In this embodiment, the embedded subtitles are the subtitle text in the subtitle display area of the image frame of the video data corresponding to the target audio data. The display area may be the bottom of the screen, the left side of the screen, the right side of the screen, etc., but this application is not limited to these.
[0115] In this context, the target audio data is extracted from the corresponding video data, or there is a corresponding relationship between the video data and the target audio data. The video data consists of multiple consecutive image frames, each representing a single moment in the video data. Multiple image frames are arranged in chronological order to form a dynamic video frame.
[0116] Based on the determined third timestamp, the image frames in the video data can be segmented, that is, the image frames corresponding to the target audio data can be extracted to obtain multiple image frames corresponding to the third timestamp. Then, text recognition is performed on the subtitle display area of each image frame to obtain the labeled text.
[0117] In some implementations, when performing text recognition on the subtitle display area of each image frame, OCR (Optical Character Recognition) technology can be used. The subtitle display area may include areas such as the bottom, top, left, and right sides of the image frame. For each possible location of the embedded subtitle within the image frame, the corresponding recognized text for different areas in each image frame can be obtained. These recognized texts are then sorted according to their confidence level, and the text with the highest confidence level is selected as the subtitle annotation text. Here, confidence level, also known as reliability, is used to measure the accuracy and reliability of the recognized text.
[0118] S33: Segment the target audio data according to the third timestamp to obtain multiple sub-audio data.
[0119] The target audio data can be segmented based on the third timestamp determined above, so as to divide the long audio data into multiple short audio data and obtain multiple sub-audio data.
[0120] S34: Use the second speech recognition model to recognize multiple sub-audio data to obtain the second reference text.
[0121] By using a second speech recognition model to recognize multiple sub-audio data, a second reference text can be obtained. For example, the second speech recognition model can be a multilingual ASR model. By using a multilingual ASR model to recognize multiple sub-audio data, a second reference text can be obtained.
[0122] In the above process, a second speech recognition model can be used to recognize the target audio data to obtain the second reference text and the corresponding fourth timestamp. Then, image frames of the video data corresponding to the target audio data are extracted according to the fourth timestamp corresponding to the second reference text, and text recognition is performed on the subtitle display area of each image frame to obtain the labeled text. Additionally, the target audio data is segmented according to the fourth timestamp corresponding to the second reference text to obtain multiple sub-audio data. Then, the first speech recognition model is used to recognize the multiple sub-audio data to obtain the first reference text. This application is not limited to this.
[0123] In some implementations, after obtaining the annotation text and at least two reference texts of the target audio data in the above steps, the annotation text and at least two reference texts may be aligned and / or corrected.
[0124] Following step S34 above, the following steps may also be included:
[0125] S35: Align the annotation text, at least two reference texts, and the speech signal of the target audio data to obtain the aligned annotation text and at least two reference texts.
[0126] Forced alignment algorithms can be used to align the labeled text, at least two reference texts, and the speech signal of the target audio data. These algorithms can include FA (Force Alignment) and MFA (Montreal Forced Aligner) algorithms, and this application is not limited to these. The following explanation uses the FA algorithm as an example.
[0127] The Formal Alignment (FA) algorithm is a common forced alignment method in speech recognition. It more accurately determines word boundaries and pronunciation information in the speech signal by forcibly aligning the text with the speech signal. Specifically, the forced alignment refers to, given a speech recognition model, inputting a speech sequence and a correct labeled text (which can be a sequence of words or phonemes, etc.), providing the correspondence between speech frames and text, such as frames 20 to 25 corresponding to the phoneme 'a'. In other words, it obtains the start and end times of each factor (or word, etc.) in the speech signal.
[0128] For example, please refer to Figure 6The process involves acquiring the speech signal of the target audio data. For example, a speech recognition system can be used to perform Voice Activity Detection (VAD) on the speech signal of the target audio data to obtain the start and end points of the speech signal. VAD separates the speech signal from the non-speech signal within a signal containing speech and determines the start and end points of the speech signal. VAD helps extract the prosodic information of the audio's pronunciation duration. Based on the FA algorithm, the actual start position of each word in the target audio data can be obtained for subsequent alignment processing.
[0129] Then, the reference text and speech signal obtained from the target audio data are aligned. For example, the words of the recognized reference text "I am Chinese" are aligned with the speech signal, thus obtaining at least two aligned reference texts. Alternatively, the labeled text and at least two reference texts can be aligned to obtain aligned labeled text and at least two reference texts.
[0130] S36: Identify the text to be corrected that contains character anomalies in the aligned annotation text and at least two reference texts, perform correction processing on the text to be corrected, and obtain the corrected text.
[0131] The aligned annotation text can be compared with at least two reference texts. Based on the alignment comparison results, the text to be corrected that contains character anomalies in the aligned annotation text and at least two reference texts can be identified. The text to be corrected is then processed to obtain the corrected text.
[0132] In some implementations, character anomalies include at least one of the following: extra character anomalies, misspelling anomalies, and missing character anomalies. An extra character anomaly indicates the presence of redundant characters in the text to be corrected; a misspelling anomaly indicates the presence of incorrect characters in the text to be corrected; and a missing character anomaly indicates the absence of characters in the text to be corrected. Other texts can be referenced to correct each text to be corrected, resulting in a corrected text. These other texts refer to the aligned annotation text and at least two reference texts, excluding the text to be corrected.
[0133] In some embodiments, please refer to Figure 7 The above embodiment can be further extended to step S36. To perform correction processing on the text to be corrected to obtain corrected text, this embodiment may include at least one of the following steps:
[0134] S361: In response to the presence of a multi-word anomaly, truncate the text from the position of the multi-word in the text to be corrected to obtain the corrected text.
[0135] In the case of multi-character anomalies in the text to be corrected, in response to the existence of multi-character anomalies, truncation can be performed at the multi-character position of the text to be corrected according to the timestamp determined by the FA algorithm (such as the start time of the word or phrase), to obtain the corrected text.
[0136] For ease of understanding, the following will be described by way of some examples. For example, the texts are as follows:
[0137] Annotated text: I'm not developed yet. Cai Wenji, pick me up later in the grass. Keep a good view.
[0138] First reference text: Coo coo coo. I'm not developed yet. Cai Wenji, give me a good view later in the grass.
[0139] Second reference text: I'm not developed yet. Cai Wenji, give me a good view later in the grass.
[0140] The above texts are aligned according to the sentence start position to ensure the confirmation of the starting position, and the timestamps are intercepted according to the results of FA. In the first reference text, "I'm" is confirmed as the start position of the clause, ensuring that the reference text of the starting ASR model is aligned with the start position recognized by OCR, facilitating subsequent processing. There is a multi-character anomaly of “Coo coo coo” in the first reference text. Taking the first reference text as the text to be corrected, the multi-character “Coo coo coo” is intercepted, and “I'm” is confirmed as the start position of the clause to obtain the corrected text.
[0141] For example, the texts are as follows:
[0142] Annotated text: I'm not developed yet. Cai Wenji will pick me up later in the grass. Keep a good view.
[0143] First reference text: I'm not developed yet. Cai Wenji, give me a good view later in the grass. Mmm.
[0144] Second reference text: I'm not developed yet. Cai Wenji, give me a good view later in the grass.
[0145] For each of the above texts, there are extra characters "Mmm" at the end of the first reference text. Taking the first reference text as the text to be corrected, the extra phrase is truncated, that is, the "Mmm" at the end is discarded, and the truncation timestamp is obtained according to the previous FA timestamp to obtain the corrected text.
[0146] S362: In response to the existence of misspelled character anomalies, use other texts to modify the misspelled character positions of the text to be corrected to obtain the corrected text.
[0147] When there are misspelled character anomalies in the text to be corrected, the misspelled character positions of the text to be corrected can be modified by other texts to obtain the corrected text.
[0148] In some embodiments, the typo anomaly is a homophone anomaly, and the first reference text and the second reference text can be corrected by referring to the annotated text. For example, the texts are as follows:
[0149] Annotated text: I'm not developed yet. Cai Wenji, later in the grass, pick me up and open the field of vision well for me.
[0150] First reference text: I'm not developed yet. Cai Wenji, later in the grass, open the field of vision well for me.
[0151] Second reference text: I'm not developed yet. Cai Wenji, later in the grass, open the field of vision well for me.
[0152] In the above text, there is a homophone anomaly "Cai" in the first reference text. The "Cai Wenji" in the first reference text can be corrected to "Cai Wenji" by referring to other texts to obtain the corrected text.
[0153] In some embodiments, the typo anomaly is a homograph anomaly, and the annotated text can be corrected by referring to the first reference text and the second reference text. For example, the texts are as follows:
[0154] Annotated text: I'm not developed yet. Cai Wenji, later in the grass, pick me up and open the field of vision well for me.
[0155] First reference text: I'm not developed yet. Cai Wenji, later in the grass, open the field of vision well for me.
[0156] Second reference text: I'm not developed yet. Cai Wenji, later in the grass, open the field of vision well for me.
[0157] In the above text, there is a homograph anomaly "pick me up" in the annotated text. The "pick me up" in the annotated text can be corrected to "open the field of vision well for me" by referring to other texts to obtain the corrected text.
[0158] S363: In response to the existence of a missing character anomaly, use other texts to complete the missing character position of the text to be corrected to obtain the corrected text.
[0159] When there is a missing character anomaly in the text to be corrected, other texts can be used to complete the missing character position of the text to be corrected to obtain the corrected text.
[0160] In some embodiments, when the text to be corrected is the first reference text or the second reference text, the first reference text or the second reference text can be completed by referring to the annotated text to obtain the corrected text. For example, the texts are as follows:
[0161] Annotated text: I'm not developed yet. Cai Wenji, later in the grass, pick me up and open the field of vision well for me.
[0162] First reference text: I'm not developed yet, Cai Wenji. Give me a good view in the grass later.
[0163] Second reference text: I'm not developed yet, Cai Wenji. Give me a good view in the grass later.
[0164] In the above text, referring to the reference marked text, complete "developed" in the first reference text and the second reference text to "developed yet" to obtain the corrected text.
[0165] In some embodiments, when the text to be corrected is the marked text, the marked text can be completed by referring to the first reference text or the second reference text to obtain the corrected text. <{
[0166] For example, the texts are as follows:
[0167] Marked text: I'm not developed yet, Cai Wenji. Pick me up later in the grass and give me a good view.
[0168] First reference text: I'm not developed yet, Cai Wenji. Give me a good view in the grass later.
[0169] Second reference text: I'm not developed yet, Cai Wenji. Give me a good view in the grass later.
[0170] In the above text, referring to the first reference text and the second reference text, complete "later" in the marked text to "later on" to obtain the corrected text.
[0171] After the above step S36, the marked text and each reference text can be aligned (such as aligning according to the minimum edit distance) to obtain the aligned marked text and each reference text. Then, according to the obtained marked text and each reference text, the above step S12 and subsequent steps can be executed.
[0172] In the above solution, by determining the text to be corrected with character anomalies in the aligned marked text and at least two reference texts, and correcting this text to be corrected by referring to other texts to obtain the corrected text, it is possible to align and correct multiple texts, and can correct character anomalies (such as glyph errors) in the marked text recognized by OCR and character anomalies (such as pronunciation errors) in the reference text recognized by the ASR model, etc. It can improve the accuracy of the reference text and the marked text extracted from the target audio data as a whole. In addition, it makes the subsequent obtained comparison results more accurate, thereby improving the accuracy of determining the type of the target audio data.
[0173] In addition, you can improve the recall rate and accuracy rate of embedded subtitles, screen out sample audio data that is more beneficial to the training of the target speech recognition model, and improve the training quality of the sample audio data.
[0174] In some embodiments, the target audio data obtained using the sample audio data acquisition method of the above embodiments can be used as sample audio data for training a target speech recognition model. After training the target speech recognition model, it can be used for speech recognition applications. In this regard, this application also provides a speech recognition method.
[0175] Please see Figure 8 , Figure 8 This is a flowchart illustrating an embodiment of the speech recognition method of this application. The method may include the following steps:
[0176] S41: Obtain the audio data to be recognized.
[0177] The audio data to be recognized can be any audio data that requires speech recognition, such as audio data contained in voice data or video data, and this application does not impose any restrictions on it.
[0178] S42: Use the target speech recognition model to perform speech recognition on the audio data to be recognized, and obtain the speech recognition result; wherein, the target speech recognition model is trained using the sample audio data obtained by the above-mentioned sample audio data acquisition method.
[0179] The target speech recognition model is used to perform speech recognition on the audio data to be recognized, and the speech recognition result is obtained. Since the target speech recognition model is trained using sample audio data obtained by the above-mentioned sample audio data acquisition method, and the sample audio data obtained by the above-mentioned sample audio data acquisition method is weakly supervised data that is more beneficial to the training of the target speech recognition model, it can improve the training effect of the target speech recognition model after training, thereby improving the accuracy of the target speech recognition model in recognizing the audio data to be recognized.
[0180] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
[0181] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0182] The specific implementation of this embodiment can be referred to the implementation process of the above embodiments, and will not be repeated here.
[0183] In addition to the above embodiments, this application also provides a sample audio data acquisition device for implementing the above-described sample audio data acquisition method.
[0184] Please see Figure 9 , Figure 9 This is a schematic diagram of an embodiment of the sample audio data acquisition device 50 of this application. The sample audio data acquisition device 50 may include: an acquisition module 51, a comparison module 52, and a processing module 53. The acquisition module 51, the comparison module 52, and the processing module 53 are interconnected.
[0185] The acquisition module 51 is used to acquire the labeled text of the target audio data and at least two reference texts. The labeled text is determined based on the subtitles of the target audio data, and each reference text is obtained by recognizing the target audio data using a different speech recognition model.
[0186] The comparison module 52 is used to determine the type of the target audio data based on the comparison results between the labeled text and at least two reference texts. The type represents the accuracy of the labeled text or the speech recognition difficulty of the target audio data.
[0187] The processing module 53 is used to perform type matching processing on the target audio data and use the processed target audio data as sample audio data for training the target speech recognition model.
[0188] The specific implementation of this embodiment can be referred to the implementation process of the above embodiments, and will not be repeated here.
[0189] In addition to the above embodiments, this application also provides a speech recognition device for implementing the above speech recognition method.
[0190] Please see Figure 10 , Figure 10 This is a schematic diagram of the structure of an embodiment of the speech recognition device of this application. The speech recognition device 60 may include an acquisition module 61 and a recognition module 62 connected together.
[0191] The acquisition module 61 is used to acquire the audio data to be recognized.
[0192] The recognition module 62 is used to perform speech recognition on the audio data to be recognized using the target speech recognition model to obtain the speech recognition result; wherein, the target speech recognition model is trained using the sample audio data obtained by the above-mentioned sample audio data acquisition method.
[0193] The specific implementation of this embodiment can be referred to the implementation process of the above embodiments, and will not be repeated here.
[0194] Regarding the above embodiments, this application provides a computer device; please refer to [link / reference]. Figure 11 , Figure 11 This is a schematic diagram of the structure of a computer device according to an embodiment of the present application. The computer device 70 includes a memory 71 and a processor 72, wherein the memory 71 and the processor 72 are coupled to each other. The memory 71 stores program data, and the processor 72 is used to execute the program data to implement the steps of any embodiment of the above-described method for acquiring sample audio data and / or the speech recognition method.
[0195] In this embodiment, processor 72 can also be referred to as a CPU (Central Processing Unit). Processor 72 may be an integrated circuit chip with signal processing capabilities. Processor 72 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor, or processor 72 can be any conventional processor.
[0196] The methods described in the above embodiments can be implemented as computer programs; therefore, this application proposes a computer-readable storage medium. Please refer to [link to relevant documentation]. Figure 12 , Figure 12 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. The computer-readable storage medium 80 stores program data 81 that can be executed by a processor. The program data 81 can be executed by the processor to implement the steps of any embodiment of the above-described method for acquiring sample audio data and / or the speech recognition method.
[0197] In this embodiment, the computer-readable storage medium 80 can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or a medium that can store program data 81. Alternatively, it can be a server that stores the program data 81. The server can send the stored program data 81 to other devices for execution, or it can run the stored program data 81 itself.
[0198] In some embodiments, the functions or modules of the apparatus provided in the above embodiments of this application can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0199] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0200] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0201] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0202] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0203] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application.
[0204] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, and thus stored in a computer-readable storage medium for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Therefore, this application is not limited to any particular hardware and software combination.
[0205] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for acquiring sample audio data, characterized in that, include: Obtain labeled text and at least two reference texts for target audio data. The labeled text is determined based on the subtitles of the target audio data. Each reference text is obtained by recognizing the target audio data using a different speech recognition model. The at least two reference texts include a first reference text recognized by a first speech recognition model and a second reference text recognized by a second speech recognition model. The first speech recognition model and the target speech recognition model are of the same type. Based on the comparison results between the labeled text and the at least two reference texts, the type of the target audio data is determined, wherein the type characterizes the accuracy of the labeled text or the speech recognition difficulty of the target audio data, including: Using the comparison results, abnormal text that meets the preset rule conditions is determined from the labeled text and at least two reference texts; Based on the anomalous text, the type of the target audio data is determined as follows: In response to the labeled text being anomalous, the type of the target audio data is determined to be a dirty data type, where the dirty data type indicates that the labeled text is inaccurate; or, in response to the first reference text being anomalous, the type of the target audio data is determined to be a first recognition difficulty type; or, in response to neither the labeled text nor the first reference text being anomalous, the type of the target audio data is determined to be a second recognition difficulty type, where the speech recognition difficulty of the first recognition difficulty type is higher than that of the second recognition difficulty type. The target audio data is processed to match the type, and the processed target audio data is used as sample audio data for training the target speech recognition model.
2. The method according to claim 1, characterized in that, The step of performing type-matching processing on the target audio data and using the processed target audio data as sample audio data for training the target speech recognition model includes at least one step: In response to the target audio data being of the dirty data type, the target audio data is discarded, where the dirty data type indicates that the annotation text is inaccurate; In response to the target audio data being of the first recognition difficulty type, the training weights of the target audio data are adjusted, and the adjusted target audio data is used as the sample audio data; In response to the target audio data being of the second recognition difficulty type, the target audio data is directly used as the sample audio data, where the speech recognition difficulty of the first recognition difficulty type is higher than that of the second recognition difficulty type.
3. The method according to claim 1, characterized in that, The step of using the comparison results to determine the abnormal text that meets the preset rule conditions from the labeled text and at least two reference texts includes any one or more of the following steps: In response to the comparison result satisfying the first preset rule condition, the labeled text is determined to be the abnormal text; In response to the comparison result not meeting the first preset rule condition but meeting the second preset rule condition, the first reference text among the at least two reference texts is determined as the abnormal text.
4. The method according to claim 3, characterized in that, The comparison results include similarity; The first preset rule conditions include: the similarity between the labeled text and each reference text is less than a first preset threshold, the similarity between each reference text is greater than a second preset threshold, and the first preset threshold is less than or equal to the second preset threshold; and / or, The at least two reference texts include a first reference text recognized by a first speech recognition model and a second reference text recognized by a second speech recognition model, wherein the first speech recognition model and the target speech recognition model are of the same type; the second preset rule conditions include: the similarity between the labeled text and the first reference text is less than a third preset threshold, and the similarity between the labeled text and the second reference text is greater than a fourth preset threshold, wherein the third preset threshold is less than or equal to the fourth preset threshold.
5. The method according to claim 1, characterized in that, The subtitles of the target audio data are peripheral subtitles, which are contained in the subtitle file of the peripheral of the target audio data. The subtitle file also contains a first timestamp corresponding to the peripheral subtitles. The annotation text and at least two reference texts for obtaining the target audio data include: Use the peripheral subtitles of the target audio data as the annotation text; The target audio data is segmented according to the first timestamp to obtain multiple sub-audio data; Different speech recognition models are used to recognize the multiple sub-audio data to obtain the at least two reference texts.
6. The method according to claim 5, characterized in that, After obtaining the labeled text and at least two reference texts of the target audio data, the method further includes: The labeled text and each reference text are aligned to obtain the aligned labeled text and each reference text. In response to the aligned annotation text satisfying the time drift condition, the target audio data is identified using an evaluation speech recognition model to obtain the evaluation text and the second timestamp corresponding to the evaluation text; wherein, the evaluation speech recognition model is selected by evaluating different speech recognition models using a test dataset of preset speech types; The target audio data is segmented according to the second timestamp to obtain multiple new sub-audio data, and the step of recognizing the multiple sub-audio data using different speech recognition models to obtain the at least two reference texts is re-executed on the multiple new sub-audio data.
7. The method according to claim 6, characterized in that, The time drift condition includes: the preset position of the aligned annotation text has a character insertion anomaly relative to the aligned reference text, and the preset position is the beginning position or the end position.
8. The method according to claim 1, characterized in that, The subtitles of the target audio data are embedded subtitles, which are subtitle text in the subtitle display area of the image frame of the video data corresponding to the target audio data; the at least two reference texts include a first reference text recognized by a first speech recognition model and a second reference text recognized by a second speech recognition model, wherein the first speech recognition model and the target speech recognition model are of the same type. The annotation text and at least two reference texts for obtaining the target audio data include: The first speech recognition model is used to recognize the target audio data to obtain the first reference text and the third timestamp corresponding to the first reference text; Image frames of video data corresponding to the target audio data are extracted according to the third timestamp, and text recognition is performed on the subtitle display area of each image frame to obtain the labeled text; The target audio data is segmented according to the third timestamp to obtain multiple sub-audio data; The second speech recognition model is used to recognize the multiple sub-audio data to obtain the second reference text.
9. The method according to claim 8, characterized in that, After obtaining the labeled text and at least two reference texts of the target audio data, the method further includes: Align the labeled text, at least two reference texts, and the speech signal of the target audio data to obtain the aligned labeled text and at least two reference texts. The text to be corrected is identified by identifying characters with abnormalities in the aligned annotation text and at least two reference texts. The text to be corrected is then corrected to obtain the corrected text.
10. The method according to claim 9, characterized in that, The character anomalies include at least one of the following: multiple character anomalies, misspelled character anomalies, and missing character anomalies; The process of correcting the text to be corrected to obtain the corrected text includes at least one of the following steps: In response to the presence of the multi-character anomaly, the text to be corrected is truncated from the position of the multi-character in the text to be corrected to obtain the corrected text; In response to the existence of the typo anomaly, the typo positions of the text to be corrected are modified using other texts to obtain the corrected text; In response to the existence of the missing character anomaly, the missing character positions of the text to be corrected are filled in using other texts to obtain the corrected text; The other text refers to the text other than the text to be corrected among the aligned annotation text and at least two reference texts.
11. A speech recognition method, characterized in that, include: Obtain the audio data to be recognized; The target speech recognition model is used to perform speech recognition on the audio data to be recognized to obtain a speech recognition result; wherein the target speech recognition model is trained using sample audio data obtained by the method described in any one of claims 1 to 10.
12. A device for acquiring sample audio data, characterized in that, include: The acquisition module is used to acquire the labeled text and at least two reference texts of the target audio data. The labeled text is determined based on the subtitles of the target audio data. Each reference text is obtained by recognizing the target audio data using a different speech recognition model. The at least two reference texts include a first reference text recognized by a first speech recognition model and a second reference text recognized by a second speech recognition model. The first speech recognition model and the target speech recognition model are of the same type. The comparison module is used to determine the type of the target audio data based on the comparison result between the labeled text and the at least two reference texts. The type characterizes the accuracy of the labeled text or the speech recognition difficulty of the target audio data. The module includes: using the comparison result to identify abnormal text that meets preset rule conditions from the labeled text and the at least two reference texts; and determining the type of the target audio data based on the abnormal text: in response to the labeled text being abnormal, determining the type of the target audio data as a dirty data type, where the dirty data type indicates that the labeled text is inaccurate; or, in response to the first reference text being abnormal, determining the type of the target audio data as a first recognition difficulty type; or, in response to neither the labeled text nor the first reference text being abnormal, determining the type of the target audio data as a second recognition difficulty type, where the speech recognition difficulty of the first recognition difficulty type is higher than that of the second recognition difficulty type. The processing module is used to perform type-matching processing on the target audio data and use the processed target audio data as sample audio data for training the target speech recognition model.
13. A voice recognition device, characterized in that, include: The acquisition module is used to acquire the audio data to be recognized; The recognition module is used to perform speech recognition on the audio data to be recognized using a target speech recognition model to obtain a speech recognition result; wherein the target speech recognition model is trained using sample audio data obtained by the method described in any one of claims 1 to 10.
14. A computer device, characterized in that, It includes a memory and a processor coupled to each other, the memory storing program data, and the processor executing the program data to implement the steps of the method of any one of claims 1 to 10, and / or, to implement the steps of the method of claim 11.
15. A computer-readable storage medium, characterized in that, The system stores program data that can be executed by a processor, the program data being used to implement the steps of the method according to any one of claims 1 to 10, and / or to implement the steps of the method according to claim 11.
Citation Information
Patent Citations
Data set cleaning method of speech recognition model
CN114187901A
Training data generation method and system for voice recognition and electronic equipment
CN116524906A
Weak supervision data generation method, speech recognition model training method and related equipment
CN117334185A
Generating a task-adapted acoustic model from one or more supervised and / or unsupervised corpora
US20030182120A1