Simultaneous interpretation training data generation method, related equipment and program product
By generating simultaneous translation training data based on speech recognition and source language alignment rules based on bilingual subtitle audio and video data, the problem of high labeling of simultaneous translation training data is solved, and efficient and low-cost training data acquisition and model optimization are achieved.
Patent Information
- Application Number
- CN202510984570.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-07-17
AI Technical Summary
The annotation of simultaneous translation training data is costly and inefficient, especially when translation in small languages requires professional translators to take time and labor-intensive.
Based on bilingual subtitle audio and video data, simultaneous translation training data is generated through speech recognition and source language alignment rules, and speech recognition is used to use a multilingual speech recognition model to eliminate dirty data with misaligned sounds and words, and weakly supervised simultaneous translation training data is generated.
No manual annotation is required, which saves labor costs, improves the efficiency of training data acquisition, improves the quality of training data and the recognition accuracy of speech recognition models.
Smart Images

Figure CN120493956A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of simultaneous interpretation, and more specifically, to a method for generating simultaneous interpretation training data, related equipment, and program products. Background Art
[0002] Simultaneous interpretation involves multiple scenarios and diverse applications, facing the challenges of different languages, accents, and expressions. Simultaneous interpretation models require extensive training with parallel corpora to improve translation quality. Currently, simultaneous interpretation training data is typically manually annotated, which is labor-intensive and inefficient. Summary of the Invention
[0003] In view of the above problems, this application proposes a method for generating simultaneous interpretation training data, related equipment, and computer program product to reduce the annotation cost of simultaneous interpretation training data and improve data acquisition efficiency. The specific solution is as follows:
[0004] In a first aspect, the present application provides a method for generating simultaneous interpretation training data, comprising:
[0005] Based on bilingual subtitle information in the bilingual subtitle audio and video data, obtaining a first source language text and a first target language text corresponding to each audio segment, wherein the audio segment is a segment included in the audio of the bilingual subtitle audio and video data;
[0006] Performing speech recognition on the audio segment to obtain a second source language text corresponding to the audio segment;
[0007] Aligning the second source language text with the first source language text according to the configured source language alignment rule to obtain a source language aligned text and selecting a target audio segment corresponding to the source language aligned text;
[0008] Simultaneous translation training data is generated based on the target audio segment and the first target language text corresponding to the target audio segment.
[0009] In a second aspect of the present application, a device for generating simultaneous translation training data is provided, comprising:
[0010] a bilingual subtitle information processing unit, configured to obtain a first source language text and a first target language text corresponding to each audio segment based on bilingual subtitle information in the bilingual subtitle audio and video data, wherein the audio segment is a segment of audio contained in the bilingual subtitle audio and video data;
[0011] a speech recognition unit, configured to perform speech recognition on the audio segment to obtain a second source language text corresponding to the audio segment;
[0012] a text alignment unit, configured to align the second source language text with the first source language text according to a configured source language alignment rule, to obtain a source language aligned text, and to select a target audio segment corresponding to the source language aligned text;
[0013] The simultaneous interpretation training data generating unit is configured to generate simultaneous interpretation training data based on the target audio segment and the first target language text corresponding to the target audio segment.
[0014] In a third aspect of the present application, an electronic device is provided, comprising: a memory and a processor;
[0015] The memory is used to store programs;
[0016] The processor is used to execute the program to implement each step of the method for generating simultaneous translation training data described in any one of the first aspects of the present application.
[0017] In a fourth aspect of the present application, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements the various steps of the method for generating simultaneous interpretation training data described in any one of the first aspects of the present application.
[0018] In a fifth aspect of the present application, a computer program product is provided, comprising a computer program. When the computer program is executed by a processor, the computer program implements the various steps of the method for generating simultaneous interpretation training data described in any one of the first aspects of the present application.
[0019] By leveraging the aforementioned technical solution, this application fully utilizes existing large-scale bilingual subtitled audio and video data (such as audio and video data including bilingual subtitles collected from international conferences, press conferences, and other venues) to generate weakly supervised simultaneous interpretation training data. Rather than directly extracting the audio and subtitle text from the bilingual subtitled audio and video data to form the training data, this application considers the potential for interference from the bilingual subtitled audio and video data, such as issues with audio and subtitles not being fully aligned. Based on the bilingual subtitled audio and video data, the first source language text and first target language text corresponding to each audio segment are obtained based on the bilingual subtitled information. Then, the second source language text corresponding to the audio segment is obtained through speech recognition. The first source language text is then aligned with the second source language text according to configured source language alignment rules to obtain the source language aligned text, and the target audio segment corresponding to the source language aligned text is selected. This alignment process yields the source language aligned text, which is then used to select the target audio segment. This eliminates the need for data from the bilingual subtitled audio and video data, where misaligned sounds and characters are present. Consequently, weakly supervised simultaneous interpretation training data is generated based on the target audio segment and its corresponding first target language text, improving the quality of the training data. This application automatically generates simultaneous interpretation training data based on existing bilingual subtitled audio and video data, without the need for manual labeling, saving labor costs and improving the efficiency of training data acquisition. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0021] Figure 1 A schematic diagram of a system architecture for implementing the method for generating simultaneous translation training data provided in an embodiment of the present application;
[0022] Figure 2 A flowchart of a method for generating simultaneous translation training data provided in an embodiment of the present application;
[0023] Figure 3 A schematic diagram of a source language alignment rule processing flow provided in this embodiment;
[0024] Figure 4 A schematic diagram of a target language alignment rule processing flow provided in an embodiment of the present application;
[0025] Figure 5 A flowchart of a method for generating simultaneous interpretation training data for bilingual externally subtitled audio and video provided in an embodiment of the present application;
[0026] Figure 6This is a flowchart of a method for generating simultaneous interpretation training data for a video with embedded bilingual subtitles, provided in an embodiment of the present application;
[0027] Figure 7 A schematic diagram of the structure of a simultaneous translation training data generation device provided in an embodiment of the present application;
[0028] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0030] It is understandable that before using the technical solutions disclosed in the embodiments of this application, the type, scope of use, usage scenarios, etc. of the personal information involved in this application should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0031] The data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0032] Currently, labeling simultaneous interpretation training data is both expensive and time-consuming, especially for lesser-known languages, where finding professional interpreters is time-consuming and labor-intensive. However, there is a large amount of available audio and video data from international conferences, press conferences, and other venues, including both source language audio and bilingual subtitles. This application further processes this bilingual subtitled audio and video data to generate high-quality weakly supervised simultaneous interpretation training data, eliminating the need for manual labeling.
[0033] Bilingual subtitled audio and video data generally falls into two categories: audio and video data with external bilingual subtitle files, and video data with embedded bilingual subtitles. External bilingual subtitle files contain timestamps and subtitle text information corresponding to the source and target languages. Video data with embedded bilingual subtitles lacks a separate subtitle file; the subtitles are embedded within the video frames.
[0034] This application provides a method for generating simultaneous translation training data, which can be applied to Figure 1 The system architecture shown in FIG. 1 may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1(This section includes a server as an example).
[0035] The terminal 100 or the server 200 can be used alone to execute the method for generating simultaneous interpretation training data provided in the embodiment of the present application. In addition, the terminal 100 and the server 200 can also be used together to execute the method for generating simultaneous interpretation training data provided in the embodiment of the present application.
[0036] Next describe Figure 1 The product form of the mid-terminal 100;
[0037] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a learning machine, a translation machine, a teaching screen, a wearable device, a conference terminal, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any restrictions on this.
[0038] The embodiment of the present application provides a method for generating simultaneous translation training data, taking the method as an example of applying the method to a computer device, which can be specifically Figure 1 The terminal 100 or the system consisting of the terminal 100 and the server 200. Figure 2 The method for generating simultaneous translation training data specifically includes the following steps:
[0039] Step S100: Based on the bilingual subtitle information in the bilingual subtitle audio and video data, obtain the first source language text and the first target language text corresponding to each audio segment.
[0040] The audio segment is a segment of the audio contained in the bilingual subtitled audio and video data. An audio stream can be extracted from the bilingual subtitled audio and video data and divided into one or more audio segments. Each audio segment corresponds to a portion of the bilingual subtitle information. In this step, the first source language text and the first target language text corresponding to the audio segment are obtained based on the bilingual subtitle information. That is, the first source language text and the first target language text are obtained with reference to the bilingual subtitle information.
[0041] Bilingual subtitled audio and video data generally falls into two categories: audio and video data with external bilingual subtitle files, and video data with embedded bilingual subtitles. External bilingual subtitle files contain timestamps and subtitle text information for both the source and target languages. The source and target language subtitles comprise the bilingual subtitles. Video data with embedded bilingual subtitles does not have a separate subtitle file; the subtitles are embedded within the video frames.
[0042] For audio and video data with bilingual external subtitle files, the bilingual external subtitle files can be used as bilingual subtitle information, and the source language text (defined as the first source language text) and target language text (defined as the first target language text) corresponding to each audio segment can be obtained based on the bilingual subtitle information.
[0043] For video data with embedded bilingual subtitles, the embedded bilingual subtitles can be identified from the video frame images of the video data as bilingual subtitle information, and the first source language text and the first target language text corresponding to each audio segment are obtained based on the bilingual subtitle information.
[0044] Step S110: Perform speech recognition on the audio segment to obtain a second source language text corresponding to the audio segment.
[0045] In the previous step, the first source language text and the first target language text corresponding to the audio segment were obtained by referring to the bilingual subtitle information in the bilingual subtitled audio and video data. In this step, speech recognition technology, such as a speech recognition model, can be used to perform speech recognition on the audio segment to obtain a recognized text, which serves as the second source language text corresponding to the audio segment.
[0046] The speech recognition model used to perform speech recognition on the audio clip in this embodiment can be any type of speech recognition model. To improve the accuracy of the speech recognition results, this embodiment can use a large multilingual speech recognition model to perform speech recognition on the audio clip. The large multilingual speech recognition model refers to a large model that can understand and transcribe speech in multiple languages, directly mapping speech features to text sequences in an end-to-end manner.
[0047] By adopting a large, multilingual speech recognition model, it supports multiple languages (from mainstream to less-resourced ones), greatly simplifying deployment and maintenance. Furthermore, speech representation knowledge is shared across languages, allowing low-resource languages to leverage knowledge from high-resource languages for improved performance. The large model often matches or surpasses single-language recognition models for mainstream languages with sufficient data, and significantly outperforms traditional speech recognition models for low-resource languages.
[0048] It should be noted that the above steps S100 and S110 can be executed in parallel or in any order. Figure 2This is just an example of an optional execution flow.
[0049] Step S120: Align the second source language text with the first source language text according to the configured source language alignment rule to obtain the source language aligned text and select the target audio segment corresponding to the source language aligned text.
[0050] Since bilingual subtitled audio and video data is weakly supervised, there may be problems with some sounds and words not being aligned. Directly using each audio clip and the corresponding first target language text as simultaneous translation training data is likely to introduce dirty data and reduce the quality of the training data.
[0051] This application filters each audio clip. Specifically, this application can pre-configure the source language alignment rules, and then use the first source language text to align the second source language text according to the alignment rules to obtain the source language aligned text. The audio clip corresponding to the source language aligned text is further filtered out as the target audio clip. Dirty data where the first source language text and the audio clip are not aligned can be effectively eliminated.
[0052] Step S130: Generate simultaneous translation training data based on the target audio segment and the first target language text corresponding to the target audio segment.
[0053] Specifically, the target audio segment is the audio segment corresponding to the source language aligned text. In the previous step, the first target language text corresponding to the target audio segment can be obtained by referring to the bilingual subtitle information. In this step, simultaneous interpretation training data can be generated based on the target audio segment and the corresponding first target language text.
[0054] Because the aforementioned steps align the first and second source language texts, audio segments with misaligned sounds and characters are removed, leaving only target audio segments with aligned sounds and characters. This generates simultaneous interpretation training data, improving its quality. It should be noted that the simultaneous interpretation training data generated by the present method can be used as weakly supervised training data for subsequent training of simultaneous interpretation models.
[0055] This application automatically generates simultaneous interpretation training data based on existing bilingual subtitled audio and video data, without the need for manual labeling, saving labor costs and improving the efficiency of training data acquisition.
[0056] In one possible implementation, after obtaining the source language aligned text and its corresponding target audio segment, the embodiment of the present application can further use the source language aligned text and its corresponding target audio segment as training data for the speech recognition model. When the set update conditions are met, this training data is used to update and train the speech recognition model used in step S110, thereby optimizing the performance of the speech recognition model and improving the recognition accuracy of the subsequent speech recognition model for the audio segment. By using the processed source language aligned text and its corresponding target audio segment as training data and continuously updating the speech recognition model, the recognition accuracy of the speech recognition model can be improved, thereby further improving the quality of the ultimately obtained simultaneous interpretation training data.
[0057] In some embodiments of the present application, some optional examples of source language alignment rules are introduced.
[0058] The source language alignment rules may include a first rule.
[0059] The process of aligning the second source language text with the first source language text according to the first rule to obtain the source language aligned text and selecting the target audio segment corresponding to the source language aligned text includes:
[0060] Step S200: Using the first source language text, the second source language text is corrected by matching similar characters to obtain a corrected second source language text.
[0061] Specifically, when the speech recognition model performs speech recognition on an audio clip, the recognized text may contain homophones but different characters, resulting in the inability to align the first source language text with the second source language text, thereby reducing the recall rate of the training data.
[0062] In this step, the first source language text can be used to perform similar character matching correction on the second source language text to obtain the second source language corrected text, thereby improving the quality of the second source language text recognized by the speech recognition model, thereby increasing the success rate of alignment between the second source language corrected text and the first source language text, and thus improving the recall rate of the training data.
[0063] In one possible implementation, an edit distance (such as Levenshtein distance or other types of edit distance) may be calculated for the first source language text and the second source language text, and elements may be aligned according to the edit distance to obtain the first source language aligned text and the second source language aligned text.
[0064] For each element in the aligned text of the second source language, determine the alignment element of the element in the aligned text of the first source language, and use the adjacent elements of the aligned element (such as the adjacent previous element, the adjacent next element, or the adjacent previous and next two elements) to perform similar character matching correction on the element to obtain the second source language corrected text. Exemplarily, if the adjacent element successfully matches the similar character of the element (that is, the adjacent element is sufficiently similar to the element in glyph to meet the matching requirements), the successfully matched adjacent element can be used to replace the element in the aligned text of the second source language to obtain the second source language corrected text. Among them, the process of similar character matching can be carried out by rule matching, such as pre-setting that the two characters have the same radicals to indicate that the similar character matching of the two characters is successful; in addition, a similar character matching model can be pre-trained to receive two input characters and predict the similar character matching results of the two characters.
[0065] Step S210: Calculate the similarity between the second source language corrected text and the first source language text. If the similarity threshold requirement is met, determine that the first rule is met, use the second source language corrected text as the source language aligned text, and use the audio segment corresponding to the first source language text as the target audio segment.
[0066] In the previous step, the second source language text was corrected to obtain the second source language corrected text. In this step, the similarity between the second source language corrected text and the first source language text is further calculated. If the similarity threshold is met, the current audio segment is considered acceptable. The second source language corrected text can be used as the source alignment text, and the audio segment corresponding to the first source language text (or the second source language text) can be used as the target audio segment.
[0067] The similarity threshold requirement refers to the similarity exceeding the set similarity threshold.
[0068] When calculating the similarity between the second source language revised text and the first source language text, an edit distance may be calculated for the second source language revised text and the first source language text, and elements may be aligned according to the edit distance, and a similarity value may be calculated based on the element alignment result.
[0069] This embodiment provides an example of applying the first rule:
[0070] First source language text: This land has nurtured five thousand years of glorious history and splendid civilization.
[0071] Second source language text: This land has nurtured 5,000 years of glorious history and profound heritage.
[0072] Calculate the edit distance between the first source language text and the second source language text, and align the elements according to the edit distance to obtain the first source language aligned text and the second source language aligned text:
[0073] The first source language aligned text:
[0074] [This] [piece of] [land] [that] [has] [bred] [a] [five-thousand-year] [splendid] [history] [and] [brilliant] [civilization].
[0075] The second source language aligned text:
[0076] [This] [piece of] [land] [that] [has] [bred] [a] [five-thousand-year] [glorious] [] [history] [and] [profound] [deposit] [].
[0077] It should be noted that when aligning elements according to the edit distance, the units such as characters, words, phrases, sentences, etc. can be used as elements for alignment. In this embodiment, only an example of aligning with characters as elements is shown. Among them, "[]" represents the inserted empty element.
[0078] Using the first source language text to perform correction by matching similar-shaped characters on the second source language text, the corrected second source language text is obtained: This piece of land that has bred a five-thousand-year splendid history and profound deposit. <s
[0079] In the above process of matching and correcting similar-shaped characters, taking [晖] in the second source language text as an example, its aligned element in the first source language text is "[]", and the adjacent elements before and after this element are: [年], [辉]. Through matching similar-shaped characters, it can be determined that [辉] and [晖] are successfully matched. Therefore, [辉] can be used to replace [晖].
[0080] Finally, calculate the similarity between the corrected second source language text and the first source language text, and judge that the similarity between the two does not exceed the similarity threshold, and it is determined that the first rule is not satisfied.
[0081] In a possible implementation, when it is determined that the first rule is not satisfied, the current audio segment and the corresponding first source language text and first target language text can be directly discarded as dirty data.
[0082] In another possible implementation, when it is determined that the first rule is not satisfied, other source language alignment rules can be further verified.
[0083] In this embodiment, by setting the first rule (source language alignment rule), introducing matching of similar-shaped words and combining with the first source language text to correct the recognition result of the audio segment by the speech recognition model, the recognition effect of the speech recognition model is improved, and the recall rate of the training data is increased.
[0084] In some embodiments of the present application, the source language alignment rule may include a second rule.
[0085] The process of aligning the second source language text with the first source language text according to the second rule to obtain the source language aligned text and selecting the target audio segment corresponding to the source language aligned text includes:
[0086] Step S300: When it is determined that the subtitles of the timestamps before and after the appearance of the bilingual subtitle information are misaligned, the first source language text of any audio segment is merged with the first source language text of the adjacent audio segment to obtain a source language merged text.
[0087] Specifically, for the collected bilingual subtitled audio and video data, there may be a situation where the subtitles of the previous and next timestamps are misaligned, such as the audio content of the current time period appears partially or completely in the subtitle text of the previous or next adjacent time period of the current time period. For example, the audio content of the t1-t2 time period appears in the subtitle text of the t2-t3 time period or the t0-t1 time period. Among them, in the conditions for determining whether there is a subtitle misalignment, the proportion of the audio content of the current time period that appears in the subtitle text of the previous or next adjacent time period of the current time period can be selected from 50%-100%, and the specific selection can be made by the staff according to business needs. The larger the ratio, the higher the quality requirements for the training data, and the smaller the ratio, the higher the recall rate of the training data.
[0088] In this step, if it is determined that the timestamps of the bilingual subtitles are misaligned, the first source language text of any audio segment can be merged with the first source language text of an adjacent audio segment (which can be the previous adjacent audio segment, the next adjacent audio segment, or both adjacent audio segments) to generate a merged source language text. The merged source language text is then used as the target for alignment and matching with the second source language text.
[0089] There are various ways to determine whether bilingual subtitle information exhibits subtitle misalignment between preceding and following timestamps. For example, if pre-acquired bilingual subtitle audio and video data is annotated with an indicator indicating whether subtitle misalignment exists, then the current bilingual subtitle audio and video data can be determined to exhibit subtitle misalignment based on the indicator. For another example, after segmenting the bilingual subtitle audio stream, alignment processing can be performed sequentially on each audio segment. During the alignment process, it can be detected whether the speech recognition result of the current audio segment exists in the subtitle text of the previous adjacent time period or the next adjacent time period. If so, it is determined that the bilingual subtitle information exhibits subtitle misalignment between preceding and following timestamps.
[0090] Step S310: Perform cross-segment fuzzy matching on the source language merged text and the second source language text, and calculate the matching similarity. If the similarity threshold requirement is met, it is determined that the second rule is met, and the second source language text is used as the source language alignment text, and the audio segment corresponding to the second source language text is used as the target audio segment.
[0091] Furthermore, if the similarity does not meet the similarity threshold requirement, it is determined that the second rule is not met.
[0092] Among them, when performing cross-segment fuzzy matching, a clustering-based fuzzy matching method can be used, and cluster outliers can be discarded to calculate the matching similarity between the source language merged text and the second source language text.
[0093] In a possible implementation, when it is determined that the second rule is not satisfied, the current audio segment and the corresponding first source language text and first target language text may be directly discarded as dirty data.
[0094] In another possible implementation, when it is determined that the second rule is not satisfied, other source language alignment rules may be further verified.
[0095] This embodiment provides an example of applying the second rule:
[0096] First source language text (fragment n): Writing the future with an unyielding spirit.
[0097] Source language combined text (fragments n-1, n, n+1): The people here are hardworking and intelligent, writing the future with an indomitable spirit. I love this land that has nurtured 5,000 years of glorious history and splendid civilization.
[0098] Second source language text (fragment n): I love this land that has nurtured five thousand years of glorious history and splendid civilization.
[0099] Perform cross-segment fuzzy matching on the source language merged text and the second source language text, calculate the matching similarity, determine whether the similarity threshold requirement is met, use the second source language text (segment n) as the source language alignment text, and use the audio segment corresponding to the second source language text (segment n) as the target audio segment.
[0100] In this embodiment, considering the phenomenon of misalignment of the subtitles before and after the timestamps in bilingual subtitle information, if the first source language text and the second source language text are directly matched to select the target audio segment, alignment failure due to the subtitle misalignment is likely to occur, a large amount of data is mistakenly discarded, and data utilization is too low. In this embodiment, when the phenomenon of misalignment of the subtitles before and after the timestamps in bilingual subtitle information is detected, the first source language text of any audio segment is merged with the first source language text of the adjacent audio segment to obtain a source language merged text, and a cross-segment fuzzy match is performed with the second source language text corresponding to the audio segment. This can greatly alleviate the problem of low data recall rate caused by subtitle misalignment and improve the utilization rate of bilingual subtitle audio and video data.
[0101] In some embodiments of the present application, the source language alignment rules may include a third rule. The third rule is an alignment rule formulated for the case where the source language belongs to a set special language. Taking into account the particularity of some languages, their speech recognition results may have alignment failures due to different punctuation or space division methods. To this end, the present application can formulate corresponding alignment rules for some set special languages. In this embodiment, the two cases of setting special languages including the first language and the second language are used as examples for explanation.
[0102] The first language is a language where the same word has different segmentation methods. For example, Korean is the first language. In Korean, the same word has different segmentation methods. When aligning the first and second source language texts, the alignment may fail due to differences in spaces between the sentences, resulting in a lower recall rate.
[0103] The second language is a language that uses spaces to separate sentences. For example, Thai is used as a second language. According to Thai grammar, sentences should be separated by spaces. However, the second source language text may be separated by words. As a result, when the second source language text is aligned with the first source language text, alignment may fail due to different spaces in the sentences, resulting in a low recall rate.
[0104] In a possible example, according to the third rule, the process of aligning the second source language text with the first source language text to obtain the source language aligned text and selecting the target audio segment corresponding to the source language aligned text includes:
[0105] When the source language belongs to the set first language, punctuation marks and spaces are removed from the first source language text and the second source language text, and then alignment and matching are performed. The matching similarity is calculated. If the similarity threshold requirement is met, the first source language text is used as the source language alignment text, and the audio segment corresponding to the first source language text is used as the target audio segment.
[0106] By removing punctuation marks and spaces from the first source language text and the second source language text, alignment failures caused by different segmentation methods can be avoided. In this embodiment, alignment matching is performed after removing punctuation marks and spaces. At this time, a higher similarity threshold can be set, such as setting the similarity threshold to 100%. That is, only when the first source language text and the second source language text after removing punctuation marks and spaces are accurately matched, is it determined that the third rule is met. At this time, the application chooses to trust the first source language text in the bilingual subtitle information as the source language alignment text, and uses the audio segment corresponding to the first source language text as the target audio segment.
[0107] Further, when it is determined that the similarity threshold requirement is not met, it is determined that the third rule is not met.
[0108] In a possible implementation, when it is determined that the second rule is not satisfied, the current audio segment and the corresponding first source language text and first target language text may be directly discarded as dirty data.
[0109] In another possible implementation, when it is determined that the second rule is not satisfied, other source language alignment rules may be further verified.
[0110] Taking Korean as the source language, an application example of the third rule is provided:
[0111] First source language text:
[0112] Second source language text:
[0113] The first source language text after removing punctuation and spaces:
[0114] The second source language text after removing punctuation and spaces:
[0115] As can be seen from the above example, after removing punctuation and spaces, the two are completely consistent and meet the similarity threshold requirement. The first source language text can be used as the source language alignment text, and the audio segment corresponding to the first source language text can be used as the target audio segment.
[0116] In the method of this embodiment, when the source language belongs to the set first language, punctuation marks and spaces are removed from the first source language text and the second source language text, and then alignment and matching are performed. This avoids alignment failure between the second source language text obtained by the speech recognition model and the first source language text due to the different positions of spaces in the sentences of the audio clip, thereby improving the recall rate of the training data.
[0117] In another possible example, according to the third rule, the process of aligning the second source language text with the first source language text to obtain the source language aligned text and selecting the target audio segment corresponding to the source language aligned text includes:
[0118] When the source language belongs to the set second language, the second source language text is re-spliced according to the space positions in the first source language text, and then aligned and matched with the first source language text, and the matching similarity is calculated. If the similarity threshold requirement is met, the first source language text is used as the source language alignment text, and the audio segment corresponding to the first source language text is used as the target audio segment.
[0119] This embodiment can avoid alignment failures caused by inappropriate space separation by re-splicing the second source language text according to the space positions in the first source language text and then aligning and matching it with the first source language text. In this embodiment, after re-splicing the second source language text according to the space positions in the first source language text and then aligning and matching it with the first source language text, a higher similarity threshold can be set at this time, such as setting the similarity threshold to 100%, that is, only when the re-spliced second source language text and the first source language text are exactly matched, is it determined that the third rule is met. At this time, this application chooses to trust the first source language text in the bilingual subtitle information as the source language alignment text, and uses the audio segment corresponding to the first source language text as the target audio segment.
[0120] Further, when it is determined that the similarity threshold requirement is not met, it is determined that the third rule is not met.
[0121] In a possible implementation, when it is determined that the second rule is not satisfied, the current audio segment and the corresponding first source language text and first target language text may be directly discarded as dirty data.
[0122] In another possible implementation, when it is determined that the second rule is not satisfied, other source language alignment rules may be further verified.
[0123] Taking Thai as the source language, an application example of the third rule is provided:
[0124] First source language text:
[0125] .
[0126] Second source language text:
[0127] .
[0128] After re-joining the second source language text according to the space positions in the first source language text: .
[0129] As can be seen from the above example, after the second source language text is re-spliced according to the space positions in the first source language text, it is completely consistent with the first source language text and meets the similarity threshold requirement. The first source language text can be used as the source language alignment text, and the audio segment corresponding to the first source language text can be used as the target audio segment.
[0130] In the method of this embodiment, when the source language belongs to the set second language, the second source language text is re-spliced according to the space positions in the first source language text and then aligned with the first source language text. This avoids the situation where the second source language text obtained by the speech recognition model fails to align with the first source language text due to space separation errors in the audio clip, thereby improving the recall rate of the training data.
[0131] The above embodiment provides three optional examples of source language alignment rules, namely the first rule, the second rule, and the third rule. The three rules can be used one by one, or any combination of two or all three rules can be used.
[0132] Combine Figure 3 , which illustrates an optional alignment processing flow when the source language alignment rules include the first rule, the second rule and the third rule.
[0133] For text 1 to be aligned and text 2 to be aligned (corresponding to the first source language text and the second source language text), first determine whether the first rule is met. If so, obtain the source language aligned text (using the second source language corrected text as the source language aligned text). If not, further determine whether there is a subtitle misalignment problem.
[0134] If there is no subtitle misalignment, the current audio segment, the first source language text, and the first target language text are treated as dirty data and discarded; if there is a subtitle misalignment, it can be further determined whether the second rule is met.
[0135] If the second rule is satisfied, the source language aligned text is obtained (the second source language text is used as the source language aligned text). If the second rule is not satisfied, it can be further determined whether the source language belongs to the set special language.
[0136] If it does not belong to the set special language, the current audio segment and the first source language text and the first target language text are treated as dirty data and discarded; if it belongs to the set special language, it is further determined whether the third rule is met.
[0137] If the third rule is met, the source language aligned text is obtained (when the source language belongs to the first or second language, the first source language text is used as the source language aligned text). If the third rule is not met, the current audio clip, the first source language text, and the first target language text are treated as dirty data and discarded.
[0138] The specific implementation process of the first, second and third rules can be referred to the relevant introduction in the previous article and will not be repeated here.
[0139] In some embodiments of the present application, the process of generating simultaneous interpretation training data based on the target audio segment and the first target language text corresponding to the target audio segment in step S130 of the aforementioned embodiment is described.
[0140] In a possible implementation, the target audio segment and its corresponding first target language text may be directly combined into simultaneous translation training data.
[0141] Since the target audio segment has already been aligned and filtered, the undesirable data, which is misaligned with the first source language text, has been removed. The first target language text corresponding to the speech segment has already been obtained in step S100. Therefore, after filtering the target audio segment, the first target language text corresponding to the target audio segment can be selected. The target audio segment and its corresponding first target language text form the simultaneous interpretation training data, allowing for rapid and relatively high-quality training data.
[0142] In another possible implementation, in order to further improve the quality of the training data, the first target language text can be further post-processed and aligned to obtain the target language aligned text. The simultaneous interpretation training data composed of the target audio clip and the target language aligned text can further eliminate the first target language text with errors and improve the quality of the final simultaneous interpretation training data.
[0143] The specific implementation process may include:
[0144] Step S400: Using the configured translation model, translate the source language aligned text corresponding to the target audio segment into the target language to obtain a second target language text corresponding to the target audio segment.
[0145] The translation model can employ various neural network models with translation capabilities. As discussed in the previous solution, source-aligned text, which is the result of aligning and correcting a second source-language text with the first source-language text, offers higher quality. On this basis, translating the source-aligned text can yield a high-quality second target-language text.
[0146] Step S410: Align the second target language text with the first target language text corresponding to the target audio segment according to the configured target language alignment rule to obtain the target language aligned text corresponding to the target audio segment.
[0147] Considering that translation models can also contain errors and other inaccuracies, this embodiment pre-configures alignment rules for the target language. Then, according to these alignment rules, the first target language text is used to align the second target language text, resulting in a target language-aligned text. The target language-aligned text is of higher quality than the first and second target language texts. Furthermore, if the second target language text, obtained by translating the source language-aligned text corresponding to a target audio segment through the translation model, fails to align with the first target language text (no corresponding target language-aligned text can be obtained for the target audio segment), the target audio segment can be determined to be dirty data and removed, thereby improving the quality of the resulting training data.
[0148] Step S420: The target audio segment and its corresponding aligned target language text form simultaneous translation training data.
[0149] In this embodiment, the translation model is introduced to translate the result of the first source language text screening (source language aligned text) to obtain the second target language text. The second target language text is then post-processed and aligned with the first target language text in the bilingual subtitle information, thereby secondary screening of available high-quality simultaneous translation training data.
[0150] In some embodiments of the present application, some optional examples of target language alignment rules are introduced.
[0151] The target language alignment rules may include a fourth rule.
[0152] According to the fourth rule, the process of aligning the second target language text with the first target language text to obtain the target language aligned text includes:
[0153] Step S500: Perform edit distance alignment on the first target language text and the second target language text, and calculate the similarity. If the similarity threshold requirement is met, it is determined that the fourth rule is met, and the second target language text is used as the target language aligned text; if the similarity threshold requirement is not met, it is determined that the fourth rule is not met.
[0154] In a possible implementation, when it is determined that the fourth rule is not satisfied, the current audio segment and the corresponding first source language text and first target language text may be directly discarded as dirty data.
[0155] In another possible implementation, when it is determined that the fourth rule is not satisfied, other target language alignment rules may be further verified.
[0156] In some embodiments of the present application, the target language alignment rules may include a fifth rule.
[0157] According to the fifth rule, the process of aligning the second target language text with the first target language text to obtain the target language aligned text includes:
[0158] Step S600: When it is determined that the subtitles of the timestamps before and after the bilingual subtitle information appear are misaligned, the first target language text of any audio segment is merged with the first target language text of the adjacent audio segment to obtain a target language merged text.
[0159] Specifically, for the collected bilingual subtitled audio and video data, there may be a situation where the subtitles of the previous and next timestamps are misaligned, such as the audio content of the current time period appears partially or completely in the subtitle text of the previous or next adjacent time period of the current time period. For example, the audio content of the t1-t2 time period appears in the subtitle text of the t2-t3 time period or the t0-t1 time period. Among them, in the conditions for determining whether there is a subtitle misalignment, the proportion of the audio content of the current time period that appears in the subtitle text of the previous or next adjacent time period of the current time period can be selected from 50%-100%, and the specific selection can be made by the staff according to business needs. The larger the ratio, the higher the quality requirements for the training data, and the smaller the ratio, the higher the recall rate of the training data.
[0160] In this step, if it is determined that the timestamps of the bilingual subtitles are misaligned, the first target language text of any audio segment can be merged with the first target language text of an adjacent audio segment (which can be the previous adjacent audio segment, the next adjacent audio segment, or both adjacent audio segments) to generate a merged target language text. The merged target language text is then used as the target for alignment and matching with the second target language text.
[0161] There are various ways to determine whether bilingual subtitle information exhibits subtitle misalignment between preceding and following timestamps. For example, if pre-acquired bilingual subtitle audio and video data is annotated with an indicator indicating whether subtitle misalignment exists, then the current bilingual subtitle audio and video data can be determined to exhibit subtitle misalignment based on the indicator. For another example, after segmenting the bilingual subtitle audio stream, alignment processing can be performed sequentially on each audio segment. During the alignment process, it can be detected whether the speech recognition result of the current audio segment exists in the subtitle text of the previous adjacent time period or the next adjacent time period. If so, it is determined that the bilingual subtitle information exhibits subtitle misalignment between preceding and following timestamps.
[0162] Step S610: Perform cross-segment fuzzy matching on the target language merged text and the second target language text, and calculate the matching similarity. If the similarity threshold requirement is met, it is determined that the second rule is met, and the second target language text is used as the target language aligned text.
[0163] Furthermore, if the similarity does not meet the similarity threshold requirement, it is determined that the fifth rule is not met.
[0164] In a possible implementation, when it is determined that the fifth rule is not satisfied, the current audio segment and the corresponding first source language text and first target language text may be directly discarded as dirty data.
[0165] In another possible implementation, when it is determined that the fifth rule is not satisfied, other target language alignment rules may be further verified.
[0166] In this embodiment, considering the misalignment of bilingual subtitles between the preceding and following timestamps, direct alignment of the first target language text and the second target language text can easily lead to alignment failure due to the misalignment, resulting in a large amount of data being erroneously discarded and low data utilization. When this embodiment detects misalignment of bilingual subtitles between the preceding and following timestamps, it merges the first target language text of any audio segment with the first target language text of an adjacent audio segment to produce a target language merged text. This text then performs a cross-segment fuzzy match with the second target language text corresponding to the audio segment. This significantly alleviates the low data recall rate caused by subtitle misalignment and improves the utilization of bilingual subtitle audio and video data.
[0167] In some embodiments of the present application, the target language alignment rules may include a sixth rule. The sixth rule is an alignment rule formulated for the case where the target language belongs to a set special language. Taking into account the particularity of some languages, the translation result of the target language obtained after translation by the translation model may fail to align with the first target language text due to different punctuation or space division methods. For this reason, the present application can formulate corresponding alignment rules for some set special languages. In this embodiment, the two cases of setting special languages including the first language and the second language are used as examples for explanation.
[0168] The first language is a language where the same word has different segmentation methods. For example, Korean is the first language. In Korean, the same word has different segmentation methods. When aligning the first target language text and the second target language text, different spaces in the sentences may cause alignment failure and reduce the recall rate.
[0169] The second language is a language that uses spaces to separate sentences. For example, let's take Thai as a second language. According to Thai grammar, sentences should be separated by spaces, but the second target language text may be separated by words. As a result, when aligning the second target language text with the first target language text, alignment may fail due to different spaces in the sentences, resulting in a low recall rate.
[0170] In a possible example, according to the sixth rule, the process of aligning the second target language text using the first target language text includes:
[0171] When the target language belongs to the set first language, punctuation marks and spaces are removed from the first target language text and the second target language text, and then alignment and matching are performed, and the matching similarity is calculated. If the similarity threshold requirement is met, the first target language text is used as the source language alignment text.
[0172] By removing punctuation and spaces from the first target language text and the second target language text, alignment failures caused by different segmentation methods can be avoided. In this embodiment, alignment matching is performed after removing punctuation and spaces. At this time, a higher similarity threshold can be set, such as setting the similarity threshold to 100%. In other words, only when the first target language text and the second target language text after removing punctuation and spaces are exactly matched is the sixth rule considered satisfied. In this case, the application chooses to trust the first target language text in the bilingual subtitle information and use it as the target language alignment text.
[0173] Further, when it is determined that the similarity threshold requirement is not met, it is determined that the sixth rule is not met.
[0174] In a possible implementation, when it is determined that the sixth rule is not satisfied, the current audio segment and the corresponding first source language text and first target language text may be directly discarded as dirty data.
[0175] In another possible implementation, when it is determined that the sixth rule is not satisfied, other target language alignment rules may be further verified.
[0176] In the method of this embodiment, when the target language belongs to the set first language, punctuation marks and spaces are removed from the first target language text and the second target language text, and then alignment matching is performed. This avoids alignment failure between the second target language text obtained by the translation model and the first target language text due to different space positions in the sentences, thereby improving the recall rate of the training data.
[0177] In another possible example, according to the sixth rule, the process of aligning the second target language text using the first target language text includes:
[0178] If the target language belongs to the set second language, the second target language text is re-spliced according to the space positions in the first target language text, and then aligned and matched with the first target language text, and the matching similarity is calculated. If the similarity threshold requirement is met, the first target language text is used as the target language alignment text.
[0179] This embodiment can avoid alignment failures caused by inappropriate space separation by re-splicing the second target language text according to the space positions in the first target language text and then aligning and matching it with the first target language text. In this embodiment, after re-splicing the second target language text according to the space positions in the first target language text and then aligning and matching it with the first source language text, a higher similarity threshold can be set at this time, such as setting the similarity threshold to 100%. That is, only when the re-spliced second target language text and the first target language text are exactly matched, is the sixth rule determined to be satisfied. At this time, this application chooses to trust the first target language text in the bilingual subtitle information and use it as the target language alignment text.
[0180] Further, when it is determined that the similarity threshold requirement is not met, it is determined that the sixth rule is not met.
[0181] In a possible implementation, when it is determined that the sixth rule is not satisfied, the current audio segment and the corresponding first source language text and first target language text may be directly discarded as dirty data.
[0182] In another possible implementation, when it is determined that the sixth rule is not satisfied, other target language alignment rules may be further verified.
[0183] In the method of this embodiment, when the target language is a set second language, the second target language text is re-spliced according to the space positions in the first target language text and then aligned with the first target language text. This avoids the situation where the second target language text obtained by the translation model fails to align with the first target language text due to space separation errors, thereby improving the recall rate of the training data.
[0184] The above embodiment provides three optional examples of target language alignment rules, namely the fourth rule, the fifth rule, and the sixth rule. The three rules can be used individually, or in any combination of two or all three rules.
[0185] Combine Figure 4 , which illustrates an optional alignment processing flow when the target language alignment rules include the fourth rule, the fifth rule, and the sixth rule.
[0186] For text 1 to be aligned and text 2 to be aligned (corresponding to the first target language text and the second target language text), first determine whether the fourth rule is met. If so, obtain the target language aligned text (take the second target language text as the target language aligned text). If not, further determine whether there is a subtitle misalignment problem.
[0187] If there is no subtitle misalignment, the current audio segment, the first source language text, and the first target language text are treated as dirty data and discarded; if there is a subtitle misalignment, it can be further determined whether the fifth rule is met.
[0188] If the fifth rule is satisfied, the target language aligned text is obtained (the second target language text is used as the target language aligned text). If the fifth rule is not satisfied, it can be further determined whether the target language belongs to the set special language.
[0189] If it does not belong to the set special language, the current audio segment and the first source language text and the first target language text are treated as dirty data and discarded; if it belongs to the set special language, it is further determined whether the sixth rule is met.
[0190] If the sixth rule is met, the target language aligned text is obtained (when the target language is the first or second language, the first target language text is used as the target language aligned text). If the sixth rule is not met, the current audio clip, the first source language text, and the first target language text are treated as dirty data and discarded.
[0191] Among them, the specific implementation process of the fourth, fifth and sixth rules can be referred to the relevant introduction in the previous article and will not be repeated here.
[0192] In one possible implementation, after obtaining the target language aligned text and its corresponding target audio segment, the embodiment of the present application can further use the target language aligned text and its corresponding target audio segment as training data for the translation model. When the set update conditions are met, this training data is used to update and train the translation model used in step S400, thereby optimizing the performance of the translation model and improving the translation accuracy of subsequent translation models for the source language aligned text. By using the processed target language aligned text and its corresponding target audio segment as training data and continuously updating the translation model, the translation accuracy of the translation model can be improved, thereby further improving the quality of the ultimately obtained simultaneous interpretation training data.
[0193] This embodiment of the present application can automatically generate simultaneous interpretation training data based on bilingual subtitled audio and video data. The bilingual subtitled audio and video data can include two types. Next, we will describe how to implement step S100 of the aforementioned embodiment, which obtains the first source language text and the first target language text corresponding to each audio segment based on the bilingual subtitle information in the bilingual subtitled audio and video data, for each type of bilingual subtitled audio and video data.
[0194] Reference Figure 5 As shown:
[0195] For audio and video data with bilingual external subtitle files (consisting of audio and video and their corresponding bilingual external subtitle files), this embodiment can obtain the first source language text and first target language text corresponding to each audio segment based on the bilingual external subtitle files. For example, timestamp information can be extracted from the bilingual external subtitle files, along with the source language text (as the first source language text) and target language text (as the first target language text) of the audio segments corresponding to adjacent timestamps. Furthermore, an audio stream can be extracted from the bilingual subtitled audio and video data and segmented into audio segments based on timestamps, resulting in one or more audio segments.
[0196] Reference Figure 6 As shown:
[0197] For video data with embedded bilingual subtitles, this embodiment extracts an audio stream from the video data and uses voice activity detection (VAD) to segment the audio stream into audio segments. Specifically, timestamp information is first obtained through VAD detection, and the audio stream is further segmented according to the timestamps to obtain audio segments. The audio segments can then be used as the targets for speech recognition in step S110 to obtain the second source language text corresponding to the audio segments.
[0198] This embodiment can also extract video frames within the duration corresponding to the audio segment from the video data according to the timestamp of each audio segment to obtain the video frame image.
[0199] The video frame image is divided into two or more sub-images, and the bilingual subtitles in each sub-image are respectively identified to obtain the source language text and the target language text contained in each sub-image.
[0200] When dividing a video frame image into sub-images, various dividing methods can be used, such as dividing it into several sub-images from left to right, or dividing it into several sub-images from top to bottom, etc.
[0201] Figure 6 In the example above, the video frame is evenly divided into three sub-images, defined as the upper, middle, and lower sub-images. Bilingual subtitles are recognized for each of the three sub-images, yielding the upper source and upper target texts contained in the upper sub-image, the middle source and target texts contained in the middle sub-image, and the lower source and target texts contained in the lower sub-image.
[0202] When recognizing bilingual subtitles for sub-images, an OCR model can be used. To improve the accuracy of subtitle recognition, a multilingual OCR large model can be used in this embodiment. The large model structure has the ability to recognize text in multiple languages.
[0203] The second source language text is aligned with the source language text of each subgraph by editing distance and similarity is calculated to determine a target subgraph that meets the similarity threshold requirement.
[0204] The video frame image may contain interference information. In this embodiment, by performing edit distance alignment between the second source language text and the source language text of each sub-image and calculating the similarity, the sub-images that do not meet the similarity threshold requirement can be eliminated, that is, the influence of the interference information can be eliminated.
[0205] The source language text contained in the target sub-image is merged to obtain the source language text contained in the video frame image, and the target language text contained in the target sub-image is merged to obtain the target language text contained in the video frame image.
[0206] Specifically, a video frame image contains several sub-images, and the target sub-images retained after the above-mentioned similarity threshold screening can be one or more. In this step, the source language texts contained in the retained target sub-images can be merged, and the result is used as the source language text contained in the video frame image. When merging the source language texts of each target sub-image, the position of each target sub-image in the video frame image can be referred to and merged in the set position order. Taking the case of dividing the video frame image into three sub-images of upper, middle and lower as an example, the source language texts contained in the target sub-images can be merged in sequence according to the order of upper, middle and lower during merging to obtain the source language text contained in the video frame image. Similarly, the target language texts contained in the target sub-images are also merged to obtain the target language text contained in the video frame image.
[0207] Furthermore, the source language texts contained in each video frame within the corresponding duration of the audio segment can be de-duplicated and merged to obtain the first source language text corresponding to the audio segment, and the target language texts contained in each video frame within the corresponding duration of the audio segment can be de-duplicated and merged to obtain the first target language text corresponding to the audio segment.
[0208] This embodiment targets video data with embedded bilingual subtitles. Considering that subtitles are not fixed in position and that video frames may contain other interfering text in addition to subtitle information, which can affect OCR recognition results, this embodiment segments the video frame image and feeds each segmented sub-image into the OCR model for recognition. The recognition results of each sub-image are aligned with the speech recognition results (the second source language text). Sub-images whose similarity meets a threshold are identified, and those whose similarity does not meet the threshold are removed. This reduces interfering information and improves the recognition accuracy of the first source language text and the first target language text.
[0209] In one possible implementation, the process of performing speech recognition on the audio clip in step S110 (that is, the process of the ASR model performing speech recognition on the audio clip) and the process of identifying the bilingual subtitles in each sub-image (that is, the process of the OCR model performing bilingual subtitle recognition on the sub-image) can be run in parallel. For example, the ASR model and the OCR model can be executed separately through different processes, thereby improving data processing efficiency and more efficiently obtaining simultaneous interpretation training data.
[0210] In one possible implementation, to further improve data processing efficiency, the video frame image can be compressed first and then divided into sub-images, thereby speeding up the processing speed of the OCR model. Of course, before compressing the video frame image, it is also possible to first determine whether the size of the video frame image exceeds the set size threshold. If it does, the video frame image is compressed. Otherwise, the video frame image can be directly divided into sub-images, thereby maximizing processing speed while ensuring image quality.
[0211] In some embodiments of the present application, Figure 5 As shown, a method for generating simultaneous translation training data is provided, which specifically includes the following steps:
[0212] For audio and video data with bilingual external subtitle files (consisting of audio and video and their corresponding bilingual external subtitle files), extract the audio stream from the audio and video, extract the timestamp information from the bilingual external subtitle files, and extract the source language text (as the first source language text) and target language text (as the first target language text) of the audio segments corresponding to adjacent timestamps.
[0213] The audio stream is segmented by timestamp to obtain several audio segments. The audio segments are then subjected to speech recognition using an ASR model (a large multilingual ASR model can be used) to obtain the corresponding second source language text.
[0214] The first source language text is used to align the second source language text, and a determination is made as to whether the source language alignment rules are satisfied. The source language alignment rules can be found in the previous embodiments and will not be described in detail here. If not, the current audio segment is discarded as dirty data. If satisfied, the source language aligned text and its corresponding target audio segment are obtained.
[0215] The source language aligned text is fed into the translation model to be translated into the target language to obtain the second target language text.
[0216] The first target language text is used to align the second target language text, and a determination is made as to whether the target language alignment rules are met. The target language alignment rules can be found in the previous embodiments and will not be described in detail here. If not, the current audio segment is discarded as dirty data. If so, the target language aligned text is obtained.
[0217] The target audio clips and target language aligned text obtained through the above process can be used as simultaneous translation training data.
[0218] The target audio clips and their corresponding source language aligned texts obtained through the above process can be fed into the ASR model as training data to update and train the ASR model to optimize the speech recognition effect of the ASR model, thereby improving the quality of the subsequent simultaneous translation training data.
[0219] The target audio clips and their corresponding target language aligned texts obtained through the above process can be fed into the translation model as training data to update and train the translation model to optimize its translation effect, thereby improving the quality of the subsequent simultaneous interpretation training data generated.
[0220] The method provided in this embodiment addresses the problem of low data recall due to timestamp and text misalignment in bilingual external subtitles. By introducing cross-segment fuzzy matching in the source language alignment rules, the low data recall caused by text and timestamp misalignment is addressed. By incorporating similar-word matching and combining it with the source language subtitle text, the ASR recognition results (second source language text) are corrected, improving the accuracy and recall of ASR recognition results, while also enhancing the quality of the translated data. Alignment rules are set for specific languages to adapt to the characteristics of each language, improving the accuracy and recall of ASR recognition results for these specific languages. A translation model is introduced to translate the source language aligned text to obtain a second target language text, which is then post-processed and aligned with the first target language text in the bilingual subtitle file. This allows for a secondary screening of available simultaneous interpretation training data, improving the quality of the training data. This embodiment further incorporates the screened and aligned target audio clips and source language aligned text into a multilingual ASR large model training to optimize the recognition results. Furthermore, the screened target audio clips and target language aligned text are used as translation data pairs and incorporated into the translation model training to optimize translation performance. It can further improve the quality of subsequent simultaneous translation training data.
[0221] In some embodiments of the present application, Figure 6 As shown in FIG, another method for generating simultaneous translation training data is provided, which specifically includes the following steps:
[0222] For video data with embedded bilingual subtitles, this embodiment extracts the audio stream from the video data, obtains timestamp information through VAD detection, and further segments the audio stream according to the timestamps to obtain audio segments. Speech recognition can be performed on the audio segments using an ASR model (a large multilingual ASR model can be used) to obtain the corresponding second source language text.
[0223] In addition, based on the timestamps obtained by VAD detection, video frames within the duration corresponding to adjacent timestamps can be intercepted from the video data to obtain video frame images.
[0224] The video frame image is evenly divided into three sub-images, defined as the upper, middle, and lower sub-images. An OCR model (a large multilingual OCR model can be used) is used to identify bilingual subtitles in each of the three sub-images. The results are: the upper source text and the upper target text in the upper sub-image; the middle source text and the middle target text in the sub-image; and the lower source text and the lower target text in the lower sub-image.
[0225] The ASR and OCR model processing can be executed in parallel to improve processing efficiency. Before using the OCR model to recognize sub-images, the video frame can be compressed and then segmented to further improve processing efficiency.
[0226] Furthermore, the second source language text is aligned with the source language text of each subgraph by edit distance and similarity is calculated, and target subgraphs that meet the similarity threshold requirement are selected.
[0227] The source language text contained in the target sub-images is merged to obtain the source language text contained in the video frame images. The target language text contained in the target sub-images is merged to obtain the target language text contained in the video frame images. The source language text contained in each video frame within the corresponding duration of the audio segment is deduplicated and merged to obtain the first source language text corresponding to the audio segment. The target language text contained in each video frame within the corresponding duration of the audio segment is deduplicated and merged to obtain the first target language text corresponding to the audio segment.
[0228] The first source language text is used to align the second source language text, and a determination is made as to whether the source language alignment rules are satisfied. The source language alignment rules can be found in the previous embodiments and will not be described in detail here. If not, the current audio segment is discarded as dirty data. If satisfied, the source language aligned text and its corresponding target audio segment are obtained.
[0229] The source language aligned text is fed into the translation model to be translated into the target language to obtain the second target language text.
[0230] The first target language text is used to align the second target language text, and a determination is made as to whether the target language alignment rules are met. The target language alignment rules can be found in the previous embodiments and will not be described in detail here. If not, the current audio segment is discarded as dirty data. If so, the target language aligned text is obtained.
[0231] The target audio clips and target language aligned text obtained through the above process can be used as simultaneous translation training data.
[0232] The target audio clips and their corresponding source language aligned texts obtained through the above process can be fed into the ASR model as training data to update and train the ASR model to optimize the speech recognition effect of the ASR model, thereby improving the quality of the subsequent simultaneous translation training data.
[0233] The target audio clips and their corresponding target language aligned texts obtained through the above process can be fed into the translation model as training data to update and train the translation model to optimize its translation effect, thereby improving the quality of the subsequent simultaneous interpretation training data generated.
[0234] The method provided in this embodiment addresses the problems of low processing efficiency for bilingual embedded audio and video data, uncertain embedded subtitle positioning, low recall rates due to ASR and OCR alignment issues, and low translation accuracy. First, to address low data processing efficiency, the ASR and OCR models are run in parallel, employing multiple processes to improve data processing efficiency. Second, the size of the OCR-processed image is appropriately compressed while ensuring optimal performance, improving OCR efficiency. Second, to address the issues of poor ASR and OCR alignment and low recall rates, multilingual ASR and OCR large models are introduced to improve recognition quality. Furthermore, similar character matching is introduced through source language alignment rules and combined with OCR results to correct ASR recognition results. To address the issue of subtitle positioning affecting OCR recognition results, the video frame is segmented and fed into the OCR model. Each segment is aligned with the ASR results to identify the segment with the highest similarity, thereby reducing interference and improving accuracy. Specific alignment rules are set for specific languages to adapt to their characteristics, improving the accuracy and recall of ASR recognition results for these languages. By introducing a translation model to translate the source language aligned text, a second target language text is obtained, which is then post-processed and aligned with the first target language text in the bilingual subtitle file, and the available simultaneous interpretation training data is screened again to improve the quality of the training data. This embodiment of the application further feeds the screened and aligned target audio clips and source language aligned text into a multilingual ASR large model training to optimize the multilingual ASR large model recognition results; at the same time, the screened target audio clips and target language aligned text are used as translation data pairs and added to the translation model training to optimize the translation effect. This can further improve the quality of the subsequently generated simultaneous interpretation training data.
[0235] The following describes a device for generating simultaneous translation training data provided in an embodiment of the present application. The device for generating simultaneous translation training data described below and the method for generating simultaneous translation training data described above can refer to each other.
[0236] See also Figure 7 , Figure 7 This is a structural diagram of a simultaneous translation training data generation device disclosed in an embodiment of the present application.
[0237] like Figure 7 As shown, the device may include:
[0238] A bilingual subtitle information processing unit 11 is configured to obtain a first source language text and a first target language text corresponding to each audio segment based on bilingual subtitle information in the bilingual subtitle audio and video data, wherein the audio segment is a segment of audio contained in the bilingual subtitle audio and video data;
[0239] a speech recognition unit 12, configured to perform speech recognition on the audio segment to obtain a second source language text corresponding to the audio segment;
[0240] a text alignment unit 13 configured to align the second source language text with the first source language text according to a configured source language alignment rule to obtain a source language aligned text and select a target audio segment corresponding to the source language aligned text;
[0241] The simultaneous interpretation training data generating unit 14 is configured to generate simultaneous interpretation training data based on the target audio segment and the first target language text corresponding to the target audio segment.
[0242] In one possible implementation, the source language alignment rule includes a first rule; and the text alignment unit aligns the second source language text using the first source language text according to the first rule to obtain a source language aligned text and selects a target audio segment corresponding to the source language aligned text, including:
[0243] Using the first source language text to perform similar character matching correction on the second source language text to obtain a second source language corrected text;
[0244] The similarity between the second source language corrected text and the first source language text is calculated. If a similarity threshold requirement is met, the second source language corrected text is used as the source language aligned text, and the audio segment corresponding to the first source language text is used as the target audio segment.
[0245] In one possible implementation, the text alignment unit uses the first source language text to perform similar character matching correction on the second source language text to obtain the corrected second source language text, including:
[0246] Calculating an edit distance between the first source language text and the second source language text, and aligning elements according to the edit distance to obtain a first source language aligned text and a second source language aligned text;
[0247] For each element in the second source language aligned text, the alignment element of the element in the first source language aligned text is determined, and the element is corrected by performing similar character matching using adjacent elements of the alignment element to obtain a second source language corrected text.
[0248] In one possible implementation, the source language alignment rule includes a second rule; and the text alignment unit aligns the second source language text using the first source language text according to the second rule to obtain a source language aligned text and selects a target audio segment corresponding to the source language aligned text, including:
[0249] When it is determined that the subtitles of the timestamps before and after the appearance of the bilingual subtitle information are misaligned, merging the first source language text of any audio segment with the first source language text of an adjacent audio segment to obtain a source language merged text;
[0250] Perform cross-segment fuzzy matching on the source language merged text and the second source language text, and calculate the matching similarity. If a similarity threshold requirement is met, use the second source language text as the source language aligned text, and use the audio segment corresponding to the second source language text as the target audio segment.
[0251] In one possible implementation, the source language alignment rule includes a third rule; and the text alignment unit aligns the second source language text using the first source language text according to the third rule to obtain a source language aligned text and selects a target audio segment corresponding to the source language aligned text, including:
[0252] When the source language belongs to a set first language, punctuation marks and spaces are removed from the first source language text and the second source language text, and then alignment and matching are performed, and matching similarity is calculated. If a similarity threshold requirement is met, the first source language text is used as the source language alignment text, and the audio segment corresponding to the first source language text is used as the target audio segment;
[0253] When the source language belongs to the set second language, the second source language text is re-spliced according to the space positions in the first source language text, and then aligned and matched with the first source language text, and the matching similarity is calculated. If the similarity threshold requirement is met, the first source language text is used as the source language alignment text, and the audio segment corresponding to the first source language text is used as the target audio segment.
[0254] In one possible implementation, the process of generating simultaneous interpretation training data by the simultaneous interpretation training data generating unit based on the target audio segment and the first target language text corresponding to the target audio segment includes:
[0255] The target audio segment and its corresponding first target language text are combined into simultaneous translation training data.
[0256] In another possible implementation, the process of generating simultaneous interpretation training data by the simultaneous interpretation training data generating unit based on the target audio segment and the first target language text corresponding to the target audio segment includes:
[0257] Using the configured translation model, translating the source language aligned text corresponding to the target audio segment into a target language to obtain a second target language text corresponding to the target audio segment;
[0258] Aligning the second target language text with the first target language text corresponding to the target audio segment according to a configured target language alignment rule to obtain a target language aligned text corresponding to the target audio segment;
[0259] The target audio segment and its corresponding aligned target language text constitute simultaneous translation training data.
[0260] In one possible implementation, the process of performing speech recognition on the audio segment is implemented by a configured speech recognition model, and the apparatus of the present application further includes:
[0261] The speech recognition model updating unit is used to use the source language aligned text and its corresponding target audio segment as training data for the speech recognition model to update the speech recognition model.
[0262] In one possible implementation, the apparatus of the present application further includes:
[0263] The translation model updating unit is configured to use the target audio segment and the corresponding target language aligned text as training data for the translation model to perform update training on the translation model.
[0264] In one possible implementation, the bilingual subtitle information processing unit obtains the first source language text and the first target language text corresponding to each audio segment based on the bilingual subtitle information in the bilingual subtitle audio and video data, including:
[0265] When the bilingual subtitled audio and video data is audio and video data with a bilingual external subtitle file, extracting timestamps, and first source language text and first target language text of audio segments corresponding to adjacent timestamps from the bilingual external subtitle file, and extracting an audio stream from the audio and video data;
[0266] The audio stream is divided into audio segments according to the timestamps.
[0267] In another possible implementation, the bilingual subtitle information processing unit obtains the first source language text and the first target language text corresponding to each audio segment based on the bilingual subtitle information in the bilingual subtitle audio and video data, including:
[0268] When the bilingual subtitled audio and video data is video data with embedded bilingual subtitles, extracting an audio stream from the video data, segmenting the audio stream into audio segments using voice activity detection (VAD), and intercepting video frames within a duration corresponding to the audio segments from the video data;
[0269] Dividing the video frame into two or more sub-images, respectively identifying bilingual subtitles in each sub-image, and obtaining the source language text and the target language text contained in each sub-image;
[0270] Performing edit distance alignment on the second source language text and the source language text of each subgraph, and calculating similarity, to determine a target subgraph that meets a similarity threshold requirement;
[0271] Merging the source language texts contained in the target subgraphs to obtain the source language texts contained in the video frame, and merging the target language texts contained in the target subgraphs to obtain the target language texts contained in the video frame;
[0272] The source language texts contained in each video frame within the corresponding duration of the audio segment are de-duplicated and merged to obtain a first source language text corresponding to the audio segment; and the target language texts contained in each video frame within the corresponding duration of the audio segment are de-duplicated and merged to obtain a first target language text corresponding to the audio segment.
[0273] In a possible implementation, the process of performing speech recognition on the audio segment by the speech recognition unit and the process of identifying the bilingual subtitles in each sub-graph by the bilingual subtitle information processing unit are run in parallel.
[0274] Each unit in the aforementioned simultaneous interpretation training data generation device may be implemented in whole or in part through software, hardware, or a combination thereof. Each of these units may be embedded in or independent of a processor within a computer device in the form of hardware, or may be stored in a computer device memory in the form of software, so that the processor can call and execute the corresponding operations of each of these units.
[0275] An electronic device is also provided in an embodiment of the present application. Figure 8 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to terminals such as mobile phones, tablet computers, translators, etc. Figure 8 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0276] like Figure 8As shown, the electronic device may include a processing device (e.g., a central processing unit, graphics processing unit, etc.) 1, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 2 or programs loaded from a storage device 8 into a random access memory (RAM) 3, thereby implementing the simultaneous interpretation training data generation method described in the aforementioned embodiments of this application. When the electronic device is powered on, the RAM 3 also stores various programs and data required for the operation of the electronic device. The processing device 1, ROM 2, and RAM 3 are interconnected via a bus 4. An input / output (I / O) interface 5 is also connected to the bus 4.
[0277] Typically, the following devices may be connected to the I / O interface 5: an input device 6 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 7 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 8 including, for example, a memory card, a hard disk, etc.; and a communication device 9. The communication device 9 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Figure 8 The electronic device is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0278] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the simultaneous interpretation training data generation methods provided in the embodiments of the present application.
[0279] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the simultaneous interpretation training data generation methods provided in the embodiments of the present application.
[0280] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0281] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0282] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0283] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0284] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referenced to each other.
Claims
1. A method for generating simultaneous translation training data, characterized in that: include: Based on bilingual subtitle information in the bilingual subtitle audio and video data, obtaining a first source language text and a first target language text corresponding to each audio segment, wherein the audio segment is a segment included in the audio of the bilingual subtitle audio and video data; Performing speech recognition on the audio segment to obtain a second source language text corresponding to the audio segment; Aligning the second source language text with the first source language text according to the configured source language alignment rule to obtain a source language aligned text and selecting a target audio segment corresponding to the source language aligned text; Simultaneous translation training data is generated based on the target audio segment and the first target language text corresponding to the target audio segment.
2. The method according to claim 1, characterized in that The source language alignment rule includes a first rule; and according to the first rule, the process of aligning the second source language text with the first source language text to obtain a source language aligned text and selecting a target audio segment corresponding to the source language aligned text includes: Using the first source language text to perform similar character matching correction on the second source language text to obtain a second source language corrected text; The similarity between the second source language corrected text and the first source language text is calculated. If a similarity threshold requirement is met, the second source language corrected text is used as the source language aligned text, and the audio segment corresponding to the first source language text is used as the target audio segment.
3. The method according to claim 2, characterized in that The process of performing similar character matching correction on the second source language text using the first source language text to obtain a second source language corrected text includes: Calculating an edit distance between the first source language text and the second source language text, and aligning elements according to the edit distance to obtain a first source language aligned text and a second source language aligned text; For each element in the second source language aligned text, the alignment element of the element in the first source language aligned text is determined, and the element is corrected by performing similar character matching using adjacent elements of the alignment element to obtain a second source language corrected text.
4. The method according to claim 2, characterized in that The source language alignment rule includes a second rule; and according to the second rule, the process of aligning the second source language text with the first source language text to obtain a source language aligned text and selecting a target audio segment corresponding to the source language aligned text includes: When it is determined that the subtitles of the timestamps before and after the appearance of the bilingual subtitle information are misaligned, merging the first source language text of any audio segment with the first source language text of an adjacent audio segment to obtain a source language merged text; Perform cross-segment fuzzy matching on the source language merged text and the second source language text, and calculate the matching similarity. If a similarity threshold requirement is met, use the second source language text as the source language aligned text, and use the audio segment corresponding to the second source language text as the target audio segment.
5. The method according to claim 2, characterized in that The source language alignment rule includes a third rule; and according to the third rule, the process of aligning the second source language text with the first source language text to obtain a source language aligned text and selecting a target audio segment corresponding to the source language aligned text includes: When the source language belongs to a set first language, punctuation marks and spaces are removed from the first source language text and the second source language text, and then alignment and matching are performed, and matching similarity is calculated. If a similarity threshold requirement is met, the first source language text is used as the source language alignment text, and the audio segment corresponding to the first source language text is used as the target audio segment; and / or, When the source language belongs to the set second language, the second source language text is re-spliced according to the space positions in the first source language text, and then aligned and matched with the first source language text, and the matching similarity is calculated. If the similarity threshold requirement is met, the first source language text is used as the source language alignment text, and the audio segment corresponding to the first source language text is used as the target audio segment.
6. The method according to any one of claims 1 to 5, characterized in that The process of generating simultaneous translation training data based on the target audio segment and the first target language text corresponding to the target audio segment includes: The target audio segment and its corresponding first target language text are combined into simultaneous translation training data.
7. The method according to any one of claims 1 to 5, characterized in that The process of generating simultaneous translation training data based on the target audio segment and the first target language text corresponding to the target audio segment includes: Using the configured translation model, translating the source language aligned text corresponding to the target audio segment into a target language to obtain a second target language text corresponding to the target audio segment; Aligning the second target language text with the first target language text corresponding to the target audio segment according to a configured target language alignment rule to obtain a target language aligned text corresponding to the target audio segment; The target audio segment and its corresponding aligned target language text constitute simultaneous translation training data.
8. The method according to any one of claims 1 to 5, characterized in that The process of obtaining a first source language text and a first target language text corresponding to each audio segment based on bilingual subtitle information in bilingual subtitle audio and video data includes: When the bilingual subtitled audio and video data is audio and video data with a bilingual external subtitle file, extracting timestamps, and first source language text and first target language text of audio segments corresponding to adjacent timestamps from the bilingual external subtitle file, and extracting an audio stream from the audio and video data; The audio stream is divided into audio segments according to the timestamps.
9. The method according to any one of claims 1 to 5, characterized in that The process of obtaining a first source language text and a first target language text corresponding to each audio segment based on bilingual subtitle information in bilingual subtitle audio and video data includes: When the bilingual subtitled audio and video data is video data with embedded bilingual subtitles, extracting an audio stream from the video data, segmenting the audio stream into audio segments using voice activity detection (VAD), and intercepting video frames within a duration corresponding to the audio segments from the video data; Dividing the video frame into two or more sub-images, respectively identifying bilingual subtitles in each sub-image, and obtaining the source language text and the target language text contained in each sub-image; Performing edit distance alignment on the second source language text and the source language text of each subgraph, and calculating similarity, to determine a target subgraph that meets a similarity threshold requirement; Merging the source language texts contained in the target subgraphs to obtain the source language texts contained in the video frame, and merging the target language texts contained in the target subgraphs to obtain the target language texts contained in the video frame; The source language texts contained in each video frame within the corresponding duration of the audio segment are de-duplicated and merged to obtain a first source language text corresponding to the audio segment; and the target language texts contained in each video frame within the corresponding duration of the audio segment are de-duplicated and merged to obtain a first target language text corresponding to the audio segment.
10. An electronic device, characterized in that: include: memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the method for generating simultaneous translation training data according to any one of claims 1 to 9.
11. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the method for generating simultaneous translation training data according to any one of claims 1 to 9 is implemented.
12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, each step of the method for generating simultaneous translation training data according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Method and device for acquiring speech recognition training data and computer equipment
CN117012179A
Sample audio data acquisition method, speech recognition method and related device
CN117894300A
Audio text pair acquisition method and device, electronic equipment and storage medium
CN117975934A
Data generation method and device
CN119864019A
Data processing method and device
CN119864032A