Method, device, electronic device storage medium, and product for generating labeled data

By forcibly aligning and combining the audio and subtitle segments of multimedia files to generate labeled data, the problems of low labeling efficiency and high cost in existing technologies are solved, and high-quality labeled data generation and high-precision speech recognition model training are achieved.

CN114925237BActive Publication Date: 2025-09-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210649393.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-09
Publication Date
2025-09-23
Estimated Expiration
2042-06-09

AI Technical Summary

Technical Problem

In the existing technology, data labeling of machine learning models relies on manual labeling, which leads to low labeling efficiency and high cost, and the accuracy of the labeled data is difficult to guarantee.

Method used

By obtaining the audio and subtitle segments of the multimedia file, using the start and end times of the audio and subtitle segments to force alignment, determining the start and end times of the characters, and intercepting the matching character audio in the audio segment to combine them, the annotation data is generated.

Benefits of technology

It achieves fast and accurate generation of labeled data, improves the quality of labeled data, provides assistance for training high-precision speech recognition models, and saves a lot of manpower costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114925237B_ABST
    Figure CN114925237B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, device, electronic device storage medium and product for generating annotated data, and relates to the field of data processing technology, in particular to the field of machine learning and speech technology. The specific implementation scheme is: obtaining at least one audio segment and at least one subtitle segment corresponding to the target multimedia file; obtaining a combined subtitle segment corresponding to each audio segment according to the start and end time of each audio segment and each subtitle segment; forcibly aligning each audio segment with the combined subtitle segment of each audio segment, and determining the start and end time of each character in each combined subtitle segment; according to the start and end time of each character, intercepting the character audio that matches each character in each audio segment, and combining each character with the matching character audio to obtain the annotated data. The scheme of the present disclosure can quickly and accurately generate annotated data, improve the accuracy of the annotated data, and also save a lot of manpower costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, in particular to the field of machine learning and speech technology, and specifically to a method and device for generating labeled data, an electronic device storage medium, and a product. Background Art

[0002] Data annotation is a crucial process that helps machines learn to identify data features. It processes raw, raw data, including speech, images, text, and video, and converts it into machine-readable information. Currently, the accuracy of machine learning models, such as speech recognition models, relies heavily on the accuracy of data annotation. Summary of the Invention

[0003] The present disclosure provides a method and device for generating annotation data, an electronic device storage medium, and a product.

[0004] According to one aspect of the present disclosure, a method for generating labeled data is provided, comprising:

[0005] Acquire at least one audio segment and at least one subtitle segment corresponding to the target multimedia file;

[0006] Acquire, according to the start and end times of the audio segments and the subtitle segments, the combined subtitle segments corresponding to the audio segments;

[0007] Forcibly aligning each of the audio segments with the combined subtitle segments of each of the audio segments, and determining the start and end time of each character in each of the combined subtitle segments;

[0008] According to the start and end time of each character, the character audio that matches each character is intercepted in each audio segment, and each character is combined with the matching character audio to obtain the annotation data.

[0009] According to another aspect of the present disclosure, there is provided a device for generating labeled data, comprising:

[0010] A segment acquisition module, configured to acquire at least one audio segment and at least one subtitle segment corresponding to a target multimedia file;

[0011] A combined subtitle segment acquisition module is used to acquire the combined subtitle segments corresponding to each of the audio segments according to the start and end times of each of the audio segments and each of the subtitle segments;

[0012] a start and end time determination module, configured to forcibly align each of the audio segments with the combined subtitle segments of the audio segments, and determine the start and end time of each character in each of the combined subtitle segments;

[0013] The annotation data determination module is used to intercept the character audio that matches each character in each audio segment according to the start and end time of each character, and respectively combine each character with the matching character audio to obtain annotation data.

[0014] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0015] at least one processor; and

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in any embodiment of the present disclosure.

[0018] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in any embodiment of the present disclosure.

[0019] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the method according to any one of the embodiments of the present disclosure is implemented.

[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0022] Figure 1 is a schematic diagram of a method for generating labeled data according to an embodiment of the present disclosure;

[0023] Figure 2 is a schematic diagram of another method for generating labeled data according to an embodiment of the present disclosure;

[0024] Figure 3 is a schematic diagram of another method for generating labeled data provided according to an embodiment of the present disclosure;

[0025] Figure 4 is a schematic diagram of another method for generating labeled data according to an embodiment of the present disclosure;

[0026] Figure 5is a schematic diagram of a video file processing flow according to an embodiment of the present disclosure;

[0027] Figure 6 is a schematic diagram of a device for generating annotation data according to an embodiment of the present disclosure;

[0028] Figure 7 It is a block diagram of an electronic device used to implement the method for generating annotation data according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0030] Figure 1 This is a schematic diagram of a method for generating annotated data according to an embodiment of the present disclosure. This embodiment is applicable to the case of automatically generating annotated data for training a speech recognition model based on multimedia files. The method can be executed by a device for generating annotated data, which can be implemented in software and / or hardware and integrated into an electronic device. In this embodiment, the electronic device can be a cloud server, a local server, a server or a computer combined with a blockchain, etc. Specifically, refer to Figure 1 , the method specifically includes the following:

[0031] S110: Acquire at least one audio segment and at least one subtitle segment corresponding to a target multimedia file.

[0032] The target multimedia file may be a video file containing subtitles, an audio file, or other multimedia files with voice information, etc., which is not limited in this embodiment.

[0033] In an optional implementation of this embodiment, before obtaining at least one audio segment and at least one subtitle segment corresponding to the target multimedia file, the obtained target multimedia file can also be decomposed into voice stream data and image stream data, and the voice stream data and image stream data can be processed separately to obtain multiple audio segments and multiple subtitle segments.

[0034] In an example of this embodiment, if the target multimedia file is a video file containing subtitles, the video file can be first decomposed into voice stream data and image stream data, and then a plurality of audio segments can be identified from the voice stream data using a voice activity detection model. Each audio segment identified by the voice activity detection model can include start and end time information of the audio segment. For example, the obtained first audio segment can include the voice content of the first audio segment, and can also include time point information such as the start and end playback time of the first audio segment in the video file. Furthermore, the subtitle segment contained in each frame of the image can be obtained from the image stream data using a character detection model.

[0035] In another example of this embodiment, if the target multimedia file is an audio file, a plurality of audio segments can be identified from the voice stream data through a voice activity detection model, and semantic understanding of each audio segment can be performed through a natural language processing model to obtain a plurality of subtitle segments corresponding to each audio segment.

[0036] S120: Acquire combined subtitle segments corresponding to the respective audio segments according to the start and end times of the respective audio segments and the respective subtitle segments.

[0037] In an optional implementation of this embodiment, after obtaining at least one audio segment and at least one subtitle segment corresponding to the target multimedia file, the combined subtitle segments corresponding to each audio segment can be further obtained based on the start and end time (start time and end time) of each audio segment and each subtitle segment.

[0038] In an example of this embodiment, if the start and end time of the first audio segment is 2-10s, multiple subtitle segments corresponding to the first audio segment can be determined from the subtitle segments based on the start and end time of the first audio segment; exemplarily, the subtitle segments corresponding to the first audio segment can be the first subtitle segment (start and end time is 1-4s), the second subtitle segment (start and end time is 4-7s) and the third subtitle segment (start and end time is 7-12s); further, the first subtitle segment, the second subtitle segment and the third subtitle segment can be spliced ​​according to the order of the start and end time to obtain a combined subtitle segment corresponding to the first audio segment; it can be understood that in this example, the start and end time of the combined subtitle segment corresponding to the first audio segment is 1-12s.

[0039] S130: Forcibly align each audio segment with the combined subtitle segment of each audio segment, and determine the start and end time of each character in each combined subtitle segment.

[0040] In an optional implementation of this embodiment, after respectively obtaining the combined subtitle segments corresponding to the audio segments, the audio segments and their corresponding combined subtitle segments can be further forcibly aligned to determine the start and end time of each character in the combined subtitle segment.

[0041] In the above example, the start and end times of the first audio segment are 2-10 seconds, and the start and end times of the combined subtitle segment corresponding to the first audio segment are 1-12 seconds. By determining the 2nd and 10th seconds in the combined subtitle segment corresponding to the first audio segment, the first audio segment and its combined subtitle segment can be forcedly aligned. Furthermore, the start and end times of each character in the combined subtitle segment can be determined based on the forced alignment results. For example, the start and end times of the first character in the combined subtitle segment can be 1-1.25 seconds, and the start and end times of the second character can be 1.25-1.5 seconds.

[0042] In another optional implementation of this embodiment, after the first audio segment and its combined subtitle segment are forcibly aligned, the start time and duration of each character in the combined subtitle segment may also be determined based on the forced alignment result, which is not limited in this embodiment.

[0043] S140 , according to the start and end time of each character, intercept the character audio that matches each character in each audio segment, and combine each character with the matching character audio to obtain annotation data.

[0044] In an optional implementation of this embodiment, after determining the start and end times of each character in each combined subtitle segment, the character audio that matches each character can be extracted from each audio segment based on the start and end times of each character. Furthermore, each character and the matching character audio are combined to generate annotated data. In this embodiment, the resulting annotated data can be used to train a speech recognition model.

[0045] In the above example, the start and end time of the first character in the combined subtitle segment can be 1-1.25s, then a 1-1.25s audio segment can be intercepted in the first audio segment corresponding to the combined subtitle segment to obtain the first character audio corresponding to the first character, and further, the first character and the first character audio are combined to obtain the first annotation data; the start and end time of the second character in the combined subtitle segment can be 1.25-1.5s, then a 1.25-1.5s audio segment can be intercepted in the first audio segment corresponding to the combined subtitle segment to obtain the second character audio corresponding to the second character, and further, the second character and the second character audio are combined to obtain the second annotation data.

[0046] The solution of this embodiment obtains at least one audio segment and at least one subtitle segment corresponding to the target multimedia file; obtains combined subtitle segments corresponding to each audio segment according to the start and end time of each audio segment and each subtitle segment; forcibly aligns each audio segment with the combined subtitle segment of each audio segment to determine the start and end time of each character in each combined subtitle segment; according to the start and end time of each character, intercepts the character audio matching each character in each audio segment, and combines each character with the matching character audio to obtain annotation data. This can quickly and accurately generate annotation data, improve the accuracy of the annotation data, provide assistance for training a high-precision speech recognition model, and also save a lot of manpower costs.

[0047] Figure 2 This is a schematic diagram of another method for generating annotated data according to an embodiment of the present disclosure. This embodiment is a further refinement of the above technical solution. The technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 2 As shown, the method for generating the labeled data includes the following:

[0048] S210: Acquire at least one audio segment and at least one subtitle segment corresponding to a target multimedia file.

[0049] S220 , comparing the text contents of adjacent subtitle segments in sequence to see if they are the same; if the text content in the first subtitle segment is the same as the text content in the second subtitle segment, merging the first subtitle segment and the second subtitle segment into the same subtitle segment.

[0050] In an optional implementation of this embodiment, after obtaining at least one audio segment and at least one subtitle segment corresponding to the target multimedia file, the text contents of each adjacent subtitle segment can be further compared to determine whether the text contents of each adjacent subtitle segment are the same; if the text content of the first subtitle segment is the same as the text content of the second subtitle segment, the first subtitle segment and the second subtitle segment can be merged to obtain a merged third subtitle segment.

[0051] The advantage of this setting is that the same subtitle segments in each subtitle segment can be filtered out, the start and end time of each subtitle segment can be accurately determined, and the combined subtitle segment corresponding to each audio segment can be accurately obtained, providing a basis for subsequent accurate labeling data.

[0052] S230. According to the start and end times of each audio segment, at least one reference subtitle segment that matches the start and end times of each audio segment is obtained from each subtitle segment; and the reference subtitle segments belonging to the same audio segment are combined in the order of the set start and end times to obtain combined subtitle segments corresponding to each audio segment.

[0053] In an optional implementation of this embodiment, after obtaining at least one audio segment and at least one subtitle segment in the target multimedia file, at least one reference subtitle segment that matches the start and end time of each audio segment can be obtained from each subtitle segment based on the start and end time of each audio segment obtained; further, the reference subtitle segments belonging to the same audio segment can be combined in the order of the set start and end time to obtain combined subtitle segments corresponding to each audio segment.

[0054] The start and end time sequence may be set in chronological order, for example, the reference subtitle segment with the smallest start and end time is placed first, and the reference subtitle segment with the largest start and end time is placed last. This is not limited in this embodiment.

[0055] It can be understood that in this embodiment, the minimum start time of at least one reference subtitle segment matching the audio segment should be less than or equal to the start time of the audio segment, and the maximum end time should be greater than or equal to the end time of the audio segment; for example, if the start and end time of the target audio frequency band is 1-20s, then the minimum start time of each reference subtitle segment corresponding to the start and end time of the audio segment should be less than or equal to 1s, and the maximum end time should be greater than or equal to 20s.

[0056] In an example of this embodiment, if the start and end time of the first audio segment is 2-10s, one or more reference subtitle segments that match the start and end time of the first audio segment can be determined from each subtitle segment based on the start and end time of the first audio segment. For example, the reference subtitle segments that match the start and end time of the first audio segment are divided into "first reference subtitle segment (start and end time is 1-4s), second reference subtitle segment (start and end time is 4-7s) and third reference subtitle segment (start and end time is 7-12s)"; further, the reference subtitle segments can be combined in order of start and end time to obtain a combined subtitle segment; in this example, the combined subtitle segment obtained can be "first reference subtitle segment-second reference subtitle segment-third reference subtitle segment".

[0057] S240: Forcibly align each audio segment with the combined subtitle segment of each audio segment, and determine the start and end time of each character in each combined subtitle segment.

[0058] S250 , according to the start and end time of each character, intercept the character audio that matches each character in each audio segment, and combine each character with the matching character audio to obtain annotation data.

[0059] The solution of this embodiment can obtain at least one reference subtitle segment that matches the start and end time of each audio segment from each subtitle segment based on the acquired start and end time of each audio segment; combine the reference subtitle segments belonging to the same audio segment in the set start and end time order to obtain combined subtitle segments corresponding to each audio segment, and accurately obtain the combined subtitle segments corresponding to each audio segment, providing a basis for the subsequent forced alignment of the audio segment and the combined subtitle segment, and providing help to improve the accuracy of the labeled data.

[0060] Figure 3 This is a schematic diagram of another method for generating annotated data according to an embodiment of the present disclosure. This embodiment is a further refinement of the above technical solution. The technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 3 As shown, the method for generating the labeled data includes the following:

[0061] S310: Acquire at least one audio segment and at least one subtitle segment corresponding to a target multimedia file.

[0062] S320: Acquire combined subtitle segments corresponding to the respective audio segments according to the start and end times of the respective audio segments and the respective subtitle segments.

[0063] S330: Align each audio segment with each combined subtitle segment according to the start and end time of each audio segment and the start and end time of each combined subtitle segment to obtain the start and end time of each character in each combined subtitle segment.

[0064] In an optional implementation of this embodiment, after obtaining the combined subtitle segments corresponding to each audio segment, each audio segment can be further forced to align with each combined subtitle segment based on the start and end time of each audio segment and the start and end time of each combined subtitle segment, so as to obtain the start and end time of each character in each combined subtitle segment.

[0065] In an example of this embodiment, if the start and end time of a first audio segment is 2-20 seconds, and the start and end time of the combined subtitle segment corresponding to the first audio segment is 1-22 seconds, the 2nd and 20th seconds can be determined in the combined subtitle segment corresponding to the first audio segment, thereby achieving forced alignment between the first audio segment and its combined subtitle segment. Furthermore, the start and end time of each character in the combined subtitle segment can be determined based on the forced alignment result. For example, the start and end time of the first character in the combined subtitle segment can be 1-1.25 seconds, and the start and end time of the second character can be 1.25-1.5 seconds.

[0066] In another optional implementation of this embodiment, each audio segment and the combined subtitle segment of each audio segment are forcibly aligned to determine the start and end time of each character in each combined subtitle segment, which may include: inputting each audio segment and the combined subtitle corresponding to each audio segment into a preset forced alignment model to obtain the start and end time of each character in each combined subtitle segment.

[0067] The preset forced alignment model may be a Viterbi model or other alignment models, which are not limited in this embodiment.

[0068] In an example of this embodiment, after obtaining combined subtitle segments corresponding to each audio segment and inputting them into the Viterbi model, the start time and duration of each character in each combined subtitle segment are output; it can be understood that if the start time and duration of the target character are known, the start time of the target character can be added to the duration of the target character to quickly determine the end time of the target character.

[0069] The advantage of this setting is that the start and end time of each character can be quickly determined through the forced alignment model, which improves the execution efficiency of the algorithm and provides a basis for accurately obtaining the labeled data in the subsequent process.

[0070] S340 , according to the start and end time of each character, intercept the character audio that matches each character in each audio segment, and combine each character with the matching character audio to obtain annotation data.

[0071] The solution of this embodiment, after obtaining the combined subtitle segments corresponding to each audio segment, can align each audio segment with each combined subtitle segment according to the start and end time of each audio segment and the start and end time of each combined subtitle, so as to accurately obtain the start and end time of each character in each combined subtitle segment, providing a basis for subsequently determining the character audio that matches each character.

[0072] Figure 4 This is a schematic diagram of another method for generating annotated data according to an embodiment of the present disclosure. This embodiment is a further refinement of the above technical solution. The technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 4 As shown, the method for generating the labeled data includes the following:

[0073] S410: Acquire at least one audio segment and at least one subtitle segment corresponding to a target multimedia file.

[0074] S420: Acquire combined subtitle segments corresponding to the respective audio segments according to the start and end times of the respective audio segments and the respective subtitle segments.

[0075] S430: Forcibly align each audio segment with the combined subtitle segment of each audio segment, and determine the start and end time of each character in each combined subtitle segment.

[0076] S440 , according to the start and end time of each character, intercept the character audio that matches each character in each audio segment, and combine each character with the matching character audio to obtain annotation data.

[0077] In an optional implementation of this embodiment, after determining the start and end time of each character in each combined subtitle segment, the start and end time of each character can be further marked in each audio segment, and each audio segment can be segmented according to the marking result to obtain each character audio.

[0078] In an example of this embodiment, if the start and end time of the first character in the first combined subtitle segment is 1-1.25s, the two times 1s and 1.25s can be marked in the character audio corresponding to the first combined subtitle segment, and the audio segment in the interval of 1-1.25s can be obtained by segmenting to obtain the first character audio corresponding to the first character; if the start and end time of the second character is 1.25-1.6s, the two times 1.25s and 1.6s can be marked in the character audio corresponding to the first combined subtitle segment, and the audio segment in the interval of 1.25-1.6s can be obtained to obtain the second character audio corresponding to the second character.

[0079] It can be understood that, in this embodiment, the first character and the first character audio can constitute the first annotation data; the second character and the second character audio can constitute the second annotation data.

[0080] S450. Determine the duration of each character audio according to the start and end time of each character audio; when the target duration is less than the set time threshold, filter out the target character audio corresponding to the target duration and the target character corresponding to the target character audio in the annotation data set.

[0081] The set time threshold may be 0.2s, 0.25s, or 0.26s, etc., which is not limited in this embodiment.

[0082] In an optional implementation of this embodiment, after obtaining the character audio matching each character, the duration of each character audio can be further determined based on the start and end time of each character audio; when the target duration is less than the set time threshold, the target character audio corresponding to the target duration and the target character corresponding to the target character audio can be filtered out from the annotation data set.

[0083] It can be understood that, in this embodiment, if the start and end time of the first character audio is 1-1.26s, then the duration of the first character audio is 1.26-1=0.26s.

[0084] In an example of this embodiment, if the start and end time of the target character audio is 1-1.2s, and the duration of the target character audio is 0.2s, which is less than the set time threshold (0.25s), it is considered that the pronunciation of the target character audio is unclear, and the target character audio and the target character corresponding to the target character audio are filtered out from the annotation data set.

[0085] The advantage of this setting is that it can optimize the generated annotation data, filter out unclear annotation data, and effectively improve the accuracy of the speech recognition model.

[0086] The solution of this embodiment, after determining the start and end time of each character in each combined subtitle segment, can further intercept the character audio matching each character in each audio segment according to the start and end time of each character, and combine each character with the matching character audio to obtain annotation data. This can accurately obtain annotation data containing the character and the character audio matching therewith, and can accurately obtain training data for the speech recognition model without manual annotation, thus saving a lot of manpower costs.

[0087] In order to enable those skilled in the art to better understand the method for generating the annotation data involved in the present disclosure, a specific example is used below to illustrate: Figure 5 This is a schematic diagram of a video file processing flow according to an embodiment of the present disclosure, with reference to Figure 5 , which mainly include the following:

[0088] S510: Obtain a long video.

[0089] S511. Extract audio stream data from the long video.

[0090] S512: Extract image stream data from the long video.

[0091] S520: Extract each audio segment from the audio stream data using a voice activity detection algorithm.

[0092] S530: Identify each subtitle segment in the image stream data using a character recognition algorithm.

[0093] S540: Determine multiple subtitle segments that match each voice segment based on the time point information of each audio segment and each subtitle segment, and concatenate the subtitle segments to obtain a combined subtitle segment.

[0094] S550: Forcefully align the voice segment with the combined subtitle segment to obtain the start time and duration of each character in the combined subtitle segment.

[0095] S560: Determine the audio that matches each character according to the start time and duration of each character, and combine each character with the audio corresponding to the character to obtain annotation data.

[0096] The solution of this embodiment trains the speech recognition model based on the obtained labeled data, and the model accuracy can reach more than 95%. The accuracy is high and no manual labeling is required, which improves the labeling efficiency and saves a lot of labor costs.

[0097] Figure 6 is a schematic diagram of a device for generating annotated data according to an embodiment of the present disclosure, which can execute a method for generating annotated data involved in any embodiment of the present disclosure; Figure 6 The device 600 for generating annotated data includes: a segment acquisition module 610, a combined subtitle segment acquisition module 620, a start and end time determination module 630, and annotated data determination module 640.

[0098] The segment acquisition module 610 is configured to acquire at least one audio segment and at least one subtitle segment corresponding to the target multimedia file;

[0099] The combined subtitle segment acquisition module 620 is configured to acquire the combined subtitle segments corresponding to the audio segments according to the start and end times of the audio segments and the subtitle segments;

[0100] a start and end time determination module 630, configured to forcibly align each of the audio segments with the combined subtitle segments of the audio segments, and determine the start and end time of each character in each of the combined subtitle segments;

[0101] The annotation data determination module 640 is used to intercept the character audio that matches each character in each audio segment according to the start and end time of each character, and respectively combine each character with the matching character audio to obtain annotation data.

[0102] The solution of this embodiment is to obtain at least one audio segment and at least one subtitle segment corresponding to the target multimedia file through a segment acquisition module; obtain combined subtitle segments corresponding to each audio segment according to the start and end times of each audio segment and each subtitle segment through a combined subtitle segment acquisition module; forcibly align each audio segment with the combined subtitle segment of each audio segment through a start and end time determination module to determine the start and end times of each character in each combined subtitle segment; and intercept character audio matching each character in each audio segment according to the start and end times of each character through a labeling data determination module, and respectively combine each character with the matching character audio to obtain labeling data. This can quickly and accurately generate labeling data, improve the accuracy of the labeling data, provide assistance for training a high-precision speech recognition model, and also save a lot of manpower costs.

[0103] In an optional implementation of this embodiment, the apparatus 600 for generating the annotation data further includes: a subtitle segment merging module;

[0104] The subtitle segment merging module is used to sequentially compare the text contents of adjacent subtitle segments to see if they are the same;

[0105] If the text content in the first subtitle segment is the same as the text content in the second subtitle segment, the first subtitle segment and the second subtitle segment are merged into the same subtitle segment.

[0106] In an optional implementation of this embodiment, the combined subtitle segment acquisition module 62 is specifically configured to

[0107] Obtaining at least one reference subtitle segment that matches the start and end time of each audio segment from each subtitle segment according to the start and end time of each audio segment;

[0108] The reference subtitle segments belonging to the same audio segment are combined according to the set start and end time sequence to obtain combined subtitle segments corresponding to the audio segments.

[0109] In an optional implementation of this embodiment, the start and end time determination module 630 is specifically used to align each audio segment with each combined subtitle segment according to the start and end time of each audio segment and the start and end time of each combined subtitle, so as to obtain the start and end time of each character in each combined subtitle segment.

[0110] In an optional implementation of this embodiment, the start and end time determination module 630 is specifically used to input each of the audio segments and the combined subtitle segments corresponding to each of the audio segments into a preset forced alignment model to obtain the start and end time of each character in each of the combined subtitle segments.

[0111] In an optional implementation of this embodiment, the annotation data determination module 640 is specifically configured to mark the start and end time of each character in each audio segment;

[0112] Each of the audio segments is segmented according to the labeling results to obtain each of the character audios.

[0113] In an optional implementation of this embodiment, the labeled data determination module 640 further includes: a filtering unit;

[0114] The filtering unit is configured to determine the duration of each character audio according to the start and end time of each character audio;

[0115] When the target duration is less than a set time threshold, the target character audio corresponding to the target duration and the target character corresponding to the target character audio are filtered out from the annotation data set.

[0116] The above-mentioned device for generating annotated data can execute the method for generating annotated data provided by any embodiment of the present disclosure, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method for generating annotated data provided by any embodiment of the present disclosure.

[0117] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0118] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0119] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0120] like Figure 7As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 707 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0121] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0122] The computing unit 701 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the method for generating labeled data. For example, in some embodiments, the method for generating labeled data can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the method for generating labeled data described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the method for generating labeled data by any other appropriate means (e.g., by means of firmware).

[0123] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0124] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0125] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0126] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0127] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0128] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0129] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0130] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for generating labeled data, comprising: Acquiring at least one audio segment and at least one subtitle segment corresponding to the target multimedia file; wherein, before acquiring the at least one audio segment and at least one subtitle segment corresponding to the target multimedia file, the method further includes: decomposing the acquired target multimedia file into voice stream data and image stream data; Acquire, according to the start and end times of each audio segment and each subtitle segment, combined subtitle segments corresponding to each audio segment; Forcibly aligning each of the audio segments with the combined subtitle segments of each of the audio segments, and determining the start and end time of each character in each of the combined subtitle segments; According to the start and end time of each character, the character audio that matches each character is intercepted in each audio segment, and each character is combined with the matching character audio to obtain annotation data; the annotation data is used to train the speech recognition model.

2. The method according to claim 1, wherein After obtaining at least one subtitle segment, the method further includes: Compare the text contents of adjacent subtitle segments in sequence to see if they are the same; If the text content in the first subtitle segment is the same as the text content in the second subtitle segment, the first subtitle segment and the second subtitle segment are merged into the same subtitle segment.

3. The method according to claim 1, wherein The step of obtaining, according to the start and end times of the audio segments and the subtitle segments, the combined subtitle segments corresponding to the audio segments respectively includes: Obtaining at least one reference subtitle segment that matches the start and end time of each audio segment from each subtitle segment according to the start and end time of each audio segment; The reference subtitle segments belonging to the same audio segment are combined according to the set start and end time sequence to obtain combined subtitle segments corresponding to the audio segments.

4. The method according to claim 1, wherein The step of forcibly aligning each of the audio segments with the combined subtitle segments of each of the audio segments and determining the start and end time of each character in each of the combined subtitle segments includes: According to the start and end time of each audio segment and the start and end time of each combined subtitle, each audio segment is aligned with each combined subtitle segment to obtain the start and end time of each character in each combined subtitle segment.

5. The method according to claim 1, wherein The step of forcibly aligning each of the audio segments with the combined subtitle segments of each of the audio segments and determining the start and end time of each character in each of the combined subtitle segments includes: Each of the audio segments and the combined subtitle segments corresponding to the audio segments are input into a preset forced alignment model to obtain the start and end time of each character in each of the combined subtitle segments.

6. The method according to claim 1, wherein The step of intercepting the character audio that matches each character in each audio segment according to the start and end time of each character includes: Marking the start and end time of each character in each audio segment; Each of the audio segments is segmented according to the labeling results to obtain each of the character audios.

7. The method according to claim 6, wherein: After obtaining the audio of each character, the method further comprises: Determining the duration of each character audio according to the start and end time of each character audio; When the duration is less than a set time threshold, the target character audio corresponding to the duration and the target character corresponding to the target character audio are filtered out from the annotation data set.

8. A device for generating annotated data, comprising: a segment acquisition module, configured to acquire at least one audio segment and at least one subtitle segment corresponding to a target multimedia file; wherein the apparatus for generating annotated data further comprises: a multimedia file decomposition module, configured to decompose the acquired target multimedia file into voice stream data and image stream data; A combined subtitle segment acquisition module is used to acquire the combined subtitle segments corresponding to each of the audio segments according to the start and end times of each of the audio segments and each of the subtitle segments; a start and end time determination module, configured to forcibly align each of the audio segments with the combined subtitle segments of the audio segments, and determine the start and end time of each character in each of the combined subtitle segments; The annotation data determination module is used to intercept the character audio matching each character in each audio segment according to the start and end time of each character, and respectively combine each character with the matching character audio to obtain annotation data; the annotation data is used to train the speech recognition model.

9. The device according to claim 8, wherein The device further comprises: a subtitle segment merging module; The subtitle segment merging module is used to sequentially compare the text contents of adjacent subtitle segments to see if they are the same; If the text content in the first subtitle segment is the same as the text content in the second subtitle segment, the first subtitle segment and the second subtitle segment are merged into the same subtitle segment.

10. The device according to claim 8, wherein The combined subtitle segment acquisition module is specifically used to Obtaining at least one reference subtitle segment that matches the start and end time of each audio segment from each subtitle segment according to the start and end time of each audio segment; The reference subtitle segments belonging to the same audio segment are combined according to the set start and end time sequence to obtain combined subtitle segments corresponding to the audio segments.

11. The device according to claim 8, wherein The start and end time determination module is specifically used to According to the start and end time of each audio segment and the start and end time of each combined subtitle, each audio segment is aligned with each combined subtitle segment to obtain the start and end time of each character in each combined subtitle segment.

12. The device according to claim 8, wherein The start and end time determination module is specifically used to Each of the audio segments and the combined subtitle segments corresponding to the audio segments are input into a preset forced alignment model to obtain the start and end time of each character in each of the combined subtitle segments.

13. The device according to claim 8, wherein The labeling data determination module is specifically used to Marking the start and end time of each character in each audio segment; Each of the audio segments is segmented according to the labeling results to obtain each of the character audios.

14. The device according to claim 13, wherein The labeled data determination module further includes: a filtering unit; The filtering unit is configured to determine the duration of each character audio according to the start and end time of each character audio; When the duration is less than a set time threshold, the target character audio corresponding to the duration and the target character corresponding to the target character audio are filtered out from the annotation data set.

15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Corpus processing method and device, electronic equipment and computer readable storage medium

    CN112818680A