Simultaneous translation training data generation method, related equipment and program product
By utilizing bilingual subtitle audio and video data, combined with speech recognition and source language alignment rules, simultaneous interpretation training data is automatically generated, solving the problem of high manual annotation costs. This achieves efficient and low-cost simultaneous interpretation training data generation, improving data quality and speech recognition accuracy.
Patent Information
- Application Number
- CN202510984570.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Manual annotation of simultaneous interpreting training data is costly and inefficient, especially since it is difficult to find interpreters specializing in less commonly taught languages. Existing technologies are insufficient to efficiently generate high-quality simultaneous interpreting training data.
By utilizing bilingual subtitle audio and video data, simultaneous interpretation training data is generated through speech recognition and source language alignment rules. Dirty data with misaligned sounds and words is removed, and high-quality weakly supervised training data is automatically generated.
No manual annotation is required, reducing labor costs, improving the efficiency of training data acquisition, and enhancing the quality of training data and the recognition accuracy of speech recognition models.
Smart Images

Figure CN120493956B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of simultaneous interpretation, and more particularly to a simultaneous interpretation training data generation method, related equipment and program product. BACKGROUND
[0002] Simultaneous interpretation involves multiple scenarios, and is applied in various ways, and faces challenges of different languages, different accents and various expression methods. A simultaneous interpretation model needs a large amount of parallel corpus for training to improve the translation quality. The current simultaneous interpretation training data is generally manually annotated, which requires high labor cost and low efficiency. SUMMARY
[0003] In view of the above problems, the present application is proposed to provide a simultaneous interpretation training data generation method, related equipment and computer program product, so as to reduce the annotation cost of simultaneous interpretation training data and improve the data acquisition efficiency. The specific scheme is as follows:
[0004] In a first aspect, the present application provides a simultaneous interpretation training data generation method, comprising:
[0005] Based on the bilingual subtitle information in the bilingual subtitle audio-video data, the first source language text and the first target language text corresponding to each audio segment are obtained, the audio segment being a segment contained by the audio in the bilingual subtitle audio-video data;
[0006] The audio segment is subjected to speech recognition to obtain the second source language text corresponding to the audio segment;
[0007] According to the configured source language alignment rule, the second source language text is aligned using the first source language text to obtain the source language alignment text and filter the target audio segment corresponding to the source language alignment text;
[0008] Based on the target audio segment and the first target language text corresponding to the target audio segment, simultaneous interpretation training data is generated.
[0009] In a second aspect, the present application provides a simultaneous interpretation training data generation apparatus, comprising:
[0010] A bilingual subtitle information processing unit is configured to obtain the first source language text and the first target language text corresponding to each audio segment based on the bilingual subtitle information in the bilingual subtitle audio-video data, the audio segment being a segment contained by the audio in the bilingual subtitle audio-video data;
[0011] A speech recognition unit is configured to perform speech recognition on the audio segment to obtain the second source language text corresponding to the audio segment;
[0012] The text alignment unit is configured to align the second source language text with the first source language text according to a configured source language alignment rule, to obtain source language alignment text and to filter a target audio segment corresponding to the source language alignment text;
[0013] The simultaneous interpretation training data generation unit is configured to generate simultaneous interpretation translation training data based on the target audio segment and the first target language text corresponding to the target audio segment.
[0014] In a third aspect, the present application provides an electronic device, comprising a memory and a processor.
[0015] The memory is configured to store a program.
[0016] The processor is configured to execute the program to implement each step of the simultaneous interpretation translation training data generation method described in any one of the first aspects.
[0017] In a fourth aspect, the present application provides a readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements each step of the simultaneous interpretation translation training data generation method described in any one of the first aspects.
[0018] In a fifth aspect, the present application provides a computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements each step of the simultaneous interpretation translation training data generation method described in any one of the first aspects.
[0019] By employing the aforementioned technical solution, this application fully utilizes existing large-scale bilingual subtitle audio-visual data (such as audio-visual data including bilingual subtitle information collected at international conferences, press conferences, etc.) to generate weakly supervised simultaneous interpretation training data. This application does not directly extract audio and subtitle text from the bilingual subtitle audio-visual data to form training data. Instead, considering that the bilingual subtitle audio-visual data may contain some interference, such as audio and subtitles not being perfectly aligned, this application, based on obtaining the first source language text and the first target language text corresponding to each audio segment based on the bilingual subtitle information, obtains the second source language text corresponding to the audio segment through speech recognition. Then, according to the configured source language alignment rules, the first source language text is used to align the second source language text, obtaining the source language aligned text. The target audio segments corresponding to the source language aligned text are then selected. Through the above alignment process, the source language aligned text can be obtained, and the target audio segments can be selected. This process can eliminate dirty data with misaligned sounds and words in the bilingual subtitle audio-visual data. Generating weakly supervised simultaneous interpretation training data based on the target audio segments and their corresponding first target language text can improve the quality of the training data. This application automatically generates simultaneous interpretation training data based on existing bilingual subtitle audio and video data, eliminating the need for manual annotation, saving labor costs, and improving the efficiency of training data acquisition. Attached Figure Description
[0020] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0021] Figure 1 A schematic diagram of an implementation system architecture for the simultaneous interpretation training data generation method provided in this application embodiment;
[0022] Figure 2 This application provides a schematic flowchart of a method for generating training data for simultaneous interpreting.
[0023] Figure 3 This is a schematic diagram of a source language alignment rule processing flow provided in this embodiment;
[0024] Figure 4 This application provides a schematic diagram of a target language alignment rule processing flow.
[0025] Figure 5 This is a schematic flowchart of a method for generating training data for simultaneous interpretation, provided in an embodiment of this application, for audio and video with bilingual external subtitles.
[0026] Figure 6A same-translation training data generation method flowchart provided by the embodiment of the present application is provided for an embedded bilingual subtitle video;
[0027] Figure 7 An apparatus structure diagram of the same-translation training data generation device provided by the embodiment of the present application is provided.
[0028] Figure 8 An apparatus structure diagram of the electronic device provided by the embodiment of the present application is provided. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0030] It can be understood that, before using the technical solutions disclosed in the embodiments of the present application, the type, use range, use scenario, etc. of the personal information involved in the present application should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0031] The data involved in the technical solutions (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.
[0032] The annotation of the current same-translation training data is not only expensive but also time-consuming, especially for small languages, which are used by few people and it is time-consuming and laborious to find professional translators. There are a large amount of available audio and video data in the current international conferences, news conferences, etc., which contain both source language audio and bilingual subtitle information. The present application obtains high-quality weakly supervised same-translation training data by further processing such bilingual subtitle audio and video data, without manual annotation.
[0033] Among them, the bilingual subtitle audio and video data generally includes two categories, one is audio and video data with bilingual external subtitle files, and the other is video data with embedded bilingual subtitles. The audio and video data with bilingual external subtitle files contain timestamp and corresponding source language and target language subtitle text information in the bilingual external subtitle file. The video data with embedded bilingual subtitles does not have a separate subtitle file, and the bilingual subtitles are embedded in the video frame picture.
[0034] The present application provides a same-translation training data generation method, which can be applied to a system architecture as shown in Figure 1 The system can include a terminal 100 and a server 200. The server 200 can include one or more servers (such as a server cluster). Figure 1The server is taken as an example in the description.
[0035] The terminal 100 or the server 200 can be used alone to perform the simultaneous interpretation translation training data generation method provided in the embodiments of the present application. In addition, the terminal 100 and the server 200 can also be used cooperatively to perform the simultaneous interpretation translation training data generation method provided in the embodiments of the present application.
[0036] Next, the product form of the terminal 100 in the Figure 1
[0037] The terminal 100 in the embodiments of the present application can be a mobile phone, a tablet computer, a learning machine, a translation machine, a teaching large screen, a wearable device, a conference terminal, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiments of the present application do not make any limitation in this regard.
[0038] The embodiments of the present application provide a simultaneous interpretation translation training data generation method. The method is taken as an example in which the method is applied to a computer device, which can be specifically the terminal 100 in the Figure 1 or a system composed of the terminal 100 and the server 200. Referring to Figure 2 , the simultaneous interpretation translation training data generation method specifically includes the following steps:
[0039] In step S100, based on the bilingual subtitle information in the bilingual subtitle audio-video data, the first source language text and the first target language text corresponding to each audio segment are obtained.
[0040] The audio segment is a segment of audio contained in the bilingual subtitle audio-video data. The audio stream can be extracted from the bilingual subtitle audio-video data and divided into one or more audio segments. Each audio segment corresponds to a part of the bilingual subtitle information. In this step, the first source language text and the first target language text corresponding to the audio segment are obtained based on the bilingual subtitle information, that is, the first source language text and the first target language text are obtained by referring to the bilingual subtitle information.
[0041] The bilingual subtitle audio-video data generally includes two types, one type is audio-video data with bilingual external subtitle file, and the other type is video data with embedded bilingual subtitles. The audio-video data with bilingual external subtitle file contains timestamp and subtitle text information corresponding to source language and target language in the bilingual external subtitle file. The source language subtitle text and the target language subtitle text form bilingual subtitles. The video data with embedded bilingual subtitles does not have a separate subtitle file, and the bilingual subtitles are embedded in the video frame picture.
[0042] For the audio-video data with bilingual external subtitle file, the bilingual external subtitle file can be taken as bilingual subtitle information, and based on the bilingual subtitle information, the source language text (defined as first source language text) and the target language text (defined as first target language text) corresponding to each audio segment can be obtained.
[0043] For the video data with embedded bilingual subtitles, the embedded bilingual subtitles can be identified from the video frame image of the video data as bilingual subtitle information, and based on the bilingual subtitle information, the first source language text and the first target language text corresponding to each audio segment can be obtained.
[0044] Step S110, speech recognition is performed on the audio segment to obtain the second source language text corresponding to the audio segment.
[0045] In the above step, the first source language text and the first target language text corresponding to the audio segment are obtained by referring to the bilingual subtitle information in the bilingual subtitle audio-video data. In this step, speech recognition technology such as a speech recognition model can be used to perform speech recognition on the audio segment to obtain the recognition text as the second source language text corresponding to the audio segment.
[0046] The speech recognition model used in the speech recognition of the audio segment in this embodiment can be various types of speech recognition models. In order to improve the accuracy of the speech recognition result, a multi-language speech recognition large model can be used to perform speech recognition on the audio segment in this embodiment. The multi-language speech recognition large model refers to a large model that can understand and transcribe multiple language speech, which uses an end-to-end method to directly map from speech features to text sequences.
[0047] By using the multi-language speech recognition large model, multiple languages (from mainstream languages to small languages with scarce resources) are supported, greatly simplifying deployment and maintenance. In addition, speech representation knowledge is shared among different languages, and low-resource languages can improve performance with the help of high-resource language knowledge. The large model often reaches or exceeds single-language recognition models on mainstream languages with sufficient data; and significantly outperforms traditional speech recognition models on low-resource languages.
[0048] It should be noted that the above steps S100 and S110 can be executed in parallel or in any order, Figure 2Only one optional execution flow is illustrated.
[0049] Step S120: According to the configured source language alignment rules, align the second source language text with the first source language text to obtain the source language aligned text and filter the target audio segments corresponding to the source language aligned text.
[0050] Since bilingual subtitle audio and video data is weakly supervised, there may be some misalignment issues between sounds and words. Directly using each audio segment and its corresponding first target language text as training data for simultaneous interpretation can easily introduce dirty data and reduce the quality of the training data.
[0051] This application filters audio segments. Specifically, it can pre-configure alignment rules for the source language, and then align the second source language text with the first source language text according to these rules to obtain aligned source language text. The audio segments corresponding to the aligned source language text are then further filtered out as target audio segments. This effectively removes misaligned data between the first source language text and the audio segments.
[0052] Step S130: Generate simultaneous interpretation training data based on the target audio segment and the first target language text corresponding to the target audio segment.
[0053] Specifically, the target audio segment is the audio segment corresponding to the source language aligned text. The first target language text corresponding to the target audio segment can be obtained by referring to the bilingual subtitle information in the preceding steps. In this step, simultaneous interpretation training data can be generated based on the target audio segment and its corresponding first target language text.
[0054] Because the aforementioned steps involve aligning the first and second source language texts, audio segments with misaligned sounds and characters can be removed, leaving only the target audio segments with aligned sounds and characters. This generates simultaneous interpretation training data, improving the quality of the training data. It should be noted that the simultaneous interpretation training data generated by the method in this application can be used as weakly supervised training data for subsequent training of the simultaneous interpretation model.
[0055] This application automatically generates simultaneous interpretation training data based on existing bilingual subtitle audio and video data, eliminating the need for manual annotation, saving labor costs, and improving the efficiency of training data acquisition.
[0056] In a possible implementation, after obtaining the source language alignment text and the corresponding target audio segment, the embodiment of the present application can further use the source language alignment text and the corresponding target audio segment as training data of the speech recognition model. When a set update condition is reached, the speech recognition model used in step S110 is updated and trained using the training data, so as to optimize the performance of the speech recognition model and improve the recognition accuracy of the subsequent speech recognition model on the audio segment. By using the obtained source language alignment text and the corresponding target audio segment as training data to continuously update the speech recognition model, the recognition accuracy of the speech recognition model can be improved, thereby further improving the quality of the final simultaneous interpretation translation training data.
[0057] In some embodiments of the present application, some optional examples of source language alignment rules are introduced.
[0058] The source language alignment rule can include a first rule.
[0059] According to the first rule, the first source language text is used to align the second source language text to obtain the source language alignment text and filter the target audio segment corresponding to the source language alignment text.
[0060] In step S200, the first source language text is used to correct the second source language text by shape-based character matching to obtain a second source language corrected text.
[0061] Specifically, when the speech recognition model performs speech recognition on the audio segment, there can be a phenomenon that the recognized text has the same pronunciation but different characters, which leads to the failure of alignment between the first source language text and the second source language text and reduces the recall rate of the training data.
[0062] In this step, the first source language text can be used to correct the second source language text by shape-based character matching to obtain a second source language corrected text, which improves the quality of the second source language text recognized by the speech recognition model, thereby improving the alignment success rate of the second source language corrected text and the first source language text, and improving the recall rate of the training data.
[0063] In a possible implementation, the edit distance (for example, Levenshtein distance or other types of edit distance) can be calculated for the first source language text and the second source language text, and the elements are aligned according to the edit distance to obtain the first source language alignment text and the second source language alignment text.
[0064] For each element in the second source language alignment text, determine the alignment element of the element in the first source language alignment text, and perform a similar character matching correction on the element using adjacent elements of the alignment element (such as adjacent previous elements, adjacent subsequent elements, or adjacent previous and subsequent elements), to obtain a second source language corrected text. For example, if the adjacent elements successfully match the similar characters of the element (i.e., the adjacent elements are similar in character form to the element, meeting the matching requirements), the second source language corrected text can be obtained by replacing the element in the second source language alignment text with the adjacent elements that successfully match the similar characters. The process of similar character matching can be performed by rule matching, such as pre-setting that two characters match similar characters if they have the same component. In addition, a similar character matching model can be pre-trained to receive two input characters and predict the similar character matching result of the two characters.
[0065] In step S210, the similarity between the second source language corrected text and the first source language text is calculated, and if the similarity threshold requirement is met, it is determined that the first rule is met, the second source language corrected text is used as the source language alignment text, and the audio segment corresponding to the first source language text is used as the target audio segment.
[0066] In the previous step, the second source language corrected text is obtained by correcting the second source language text. In this step, the similarity between the second source language corrected text and the first source language text is further calculated, and if the similarity threshold requirement is met, it is considered that the current audio segment is available, and the second source language corrected text can be used as the source language alignment text, and the audio segment corresponding to the first source language text (or the second source language text) is used as the target audio segment.
[0067] The similarity threshold requirement refers to the similarity exceeding a set similarity threshold.
[0068] When calculating the similarity between the second source language corrected text and the first source language text, the edit distance between the second source language corrected text and the first source language text can be calculated, and the elements are aligned according to the edit distance, and the similarity value is calculated based on the element alignment result.
[0069] The present embodiment provides an application example of the first rule:
[0070] The first source language text is: This piece of land has nurtured five thousand years of glorious history and brilliant civilization.
[0071] The second source language text is: This piece of land has nurtured five thousand years of glorious history and brilliant civilization.
[0072] The edit distance between the first source language text and the second source language text is calculated, and the elements are aligned according to the edit distance, to obtain the first source language alignment text and the second source language alignment text:
[0073] First source language alignment text:
[0074] [This] [piece] [pregnant] [five] [thousand] [years] [glory] [history] [and] [brilliant] [land].
[0075] Second source language alignment text:
[0076] [This] [piece] [pregnant] [five] [thousand] [years] [glory] [history] [and] [brilliant] [land].
[0077] It should be noted that when aligning elements according to the edit distance, the elements can be aligned in units of words, phrases, sentences, etc. In this embodiment, only an example of aligning elements in words is shown. Wherein “[]” represents an inserted empty element.
[0078] The first source language text is used to correct the second source language text by near-homograph matching, and the second source language corrected text is obtained: This piece of land has nurtured five thousand years of brilliant history and profound roots.
[0079] The above near-homograph matching correction process takes [glory] in the second source language text as an example, and the aligned element in the first source language text is “[]”. The elements adjacent to the element are: [years], [glory]. Through near-homograph matching, it can be determined that [glory] and [glory] match successfully, so [glory] can be used to replace [glory].
[0080] Finally, the similarity between the second source language corrected text and the first source language text is calculated, and it is determined that the similarity does not exceed the similarity threshold, and it is determined that rule one is not satisfied.
[0081] In one possible implementation, when it is determined that the first rule is not satisfied, the current audio segment and the corresponding first source language text and first target language text can be directly discarded as dirty data.
[0082] In another possible implementation, when it is determined that the first rule is not satisfied, other source language alignment rules can be further verified.
[0083] In this embodiment, by setting the first rule (source language alignment rule), introducing near-homograph matching and combining the first source language text to correct the recognition result of the audio segment by the speech recognition model, the recognition effect of the speech recognition model is improved, and the recall rate of the training data is improved.
[0084] In some embodiments of the present application, the source language alignment rule can include a second rule.
[0085] According to the second rule, the first source language text is aligned with the second source language text to obtain a source language aligned text, and a process of screening a target audio segment corresponding to the source language aligned text includes:
[0086] In step S300, in a case where it is determined that the subtitle misplacement of the bilingual subtitle information occurs before and after the timestamp, the first source language text of any audio segment is merged with the first source language text of an adjacent audio segment to obtain a source language merged text.
[0087] Specifically, for the collected bilingual subtitle audio-video data, there can be a case of subtitle misplacement before and after the timestamp, for example, the audio content of the current time period is partially or entirely in the subtitle text of the previous adjacent or next adjacent time period. For example, the audio content of the t1-t2 time period is in the subtitle text of the t2-t3 time period or the t0-t1 time period. In the condition for determining whether the subtitle misplacement occurs, the proportion of the audio content of the current time period in the subtitle text of the previous adjacent or next adjacent time period can be selected from 50% to 100%, and can be selected by the staff according to the business requirements. When the proportion is larger, the quality requirement for the training data is higher, and when the proportion is smaller, the recall rate of the training data is higher.
[0088] In this step, in a case where it is determined that the subtitle misplacement of the bilingual subtitle occurs before and after the timestamp, the first source language text of any audio segment can be merged with the first source language text of an adjacent audio segment (which can be a previous adjacent audio segment, a next adjacent audio segment, or two adjacent audio segments) to obtain a source language merged text. The source language merged text is used as an object for aligning and matching with the second source language text.
[0089] The way of determining whether the subtitle misplacement of the bilingual subtitle information occurs before and after the timestamp can be various, for example, the pre-acquired bilingual subtitle audio-video data is labeled with an identifier indicating whether the subtitle misplacement occurs, and then whether the current bilingual subtitle audio-video data has the subtitle misplacement can be determined based on the identifier. For another example, after the bilingual subtitle audio stream is segmented, each audio segment can be sequentially aligned and processed, and when the alignment and processing is performed, it can be detected whether the speech recognition result of the current audio segment is in the subtitle text of the previous adjacent time period or the next adjacent time period, and if yes, it is determined that the subtitle misplacement of the bilingual subtitle information occurs before and after the timestamp.
[0090] In step S310, the source language merged text is cross-segmentally fuzzy matched with the second source language text, and a matching similarity is calculated. If the similarity threshold requirement is met, it is determined that the second rule is met, the second source language text is used as a source language aligned text, and the audio segment corresponding to the second source language text is used as a target audio segment.
[0091] Further, if the similarity does not satisfy the similarity threshold requirement, it is determined that the second rule is not satisfied.
[0092] In the cross-fragment fuzzy matching, a clustering-based fuzzy matching method can be used, and clustering outliers are discarded, and the matching similarity of the source language merged text and the second source language text is calculated.
[0093] In a possible implementation, when it is determined that the second rule is not satisfied, the current audio segment and the corresponding first source language text and first target language text can be directly discarded as dirty data.
[0094] In another possible implementation, when it is determined that the second rule is not satisfied, other source language alignment rules can be further verified.
[0095] The embodiment provides an application example of the second rule:
[0096] The first source language text (segment n): writes the future with indomitable spirit.
[0097] The source language merged text (segments n-1, n, n+1): the people here are hardworking and wise, writes the future with indomitable spirit, and I love this land that has nurtured a glorious history of five thousand years and a brilliant civilization.
[0098] The second source language text (segment n): I love this land that has nurtured a glorious history of five thousand years and a brilliant civilization.
[0099] The source language merged text and the second source language text are cross-fragment fuzzy matched, and the matching similarity is calculated, to determine that the similarity threshold requirement is satisfied, the second source language text (segment n) is taken as the source language alignment text, and the audio segment corresponding to the second source language text (segment n) is taken as the target audio segment.
[0100] In the embodiment, when the phenomenon of misalignment of subtitles before and after the timestamps of bilingual subtitle information occurs, if the first source language text and the second source language text are directly matched to screen the target audio segment, misalignment of subtitles is likely to cause alignment failure, a large amount of data is discarded incorrectly, and the data utilization rate is too low. In the embodiment, when the phenomenon of misalignment of subtitles before and after the timestamps of bilingual subtitle information is detected, the first source language text of any audio segment is merged with the first source language text of an adjacent audio segment to obtain a source language merged text, and cross-fragment fuzzy matching is performed between the source language merged text and the second source language text corresponding to the audio segment, which can greatly alleviate the problem of low data recall rate caused by misalignment of subtitles, and improve the utilization rate of bilingual subtitle audio and video data.
[0101] In some embodiments of the present application, the source language alignment rule can include a third rule. The third rule is an alignment rule formulated for the case where the source language belongs to a set of special languages. Considering the particularity of some languages, the speech recognition result thereof can fail to align due to different punctuation or space division methods. Therefore, the present application can formulate corresponding alignment rules for some set of special languages. The present embodiment takes the case where the set of special languages includes a first language and a second language as an example for illustration.
[0102] wherein the first language is a language with different segmentation methods for the same word, and Korean is taken as an example of the first language. The same word in Korean has different segmentation methods, and when the first source language text and the second source language text are aligned and matched, the alignment can fail and the recall rate can decrease due to different spaces in the sentence.
[0103] The second language is a language that separates sentences by spaces, and Thai is taken as an example of the second language. Thai grammar should separate sentences by spaces, but the second source language text can have words separated by spaces, which can cause the alignment to fail and the recall rate to decrease when the second source language text and the first source language text are aligned and matched due to different spaces in the sentence.
[0104] In one possible example, according to the third rule, the process of aligning the second source language text with the first source language text to obtain the source language alignment text and screening the target audio segment corresponding to the source language alignment text includes:
[0105] In the case where the source language belongs to the set of special languages, the punctuation marks and spaces in the first source language text and the second source language text are removed, and then aligned and matched, and the matching similarity is calculated. If the similarity threshold requirement is met, the first source language text is taken as the source language alignment text, and the audio segment corresponding to the first source language text is taken as the target audio segment.
[0106] By removing the punctuation marks and spaces in the first source language text and the second source language text, alignment failure due to different segmentation methods can be avoided. In the present embodiment, alignment is performed after removing the punctuation marks and spaces, and a higher similarity threshold can be set at this time, such as setting the similarity threshold to 100%. That is, only when the first source language text after removing the punctuation marks and spaces is exactly matched with the second source language text, it is determined that the third rule is met. At this time, the present application selects the first source language text in the bilingual subtitle information as the source language alignment text, and takes the audio segment corresponding to the first source language text as the target audio segment.
[0107] Further, when it is determined that the similarity threshold requirement is not met, it is determined that the third rule is not met.
[0108] In a possible implementation, when it is determined that the second rule is not satisfied, the current audio segment and the corresponding first source language text and first target language text can be directly discarded as dirty data.
[0109] In another possible implementation, when it is determined that the second rule is not satisfied, other source language alignment rules can be further verified.
[0110] Taking the source language as Korean as an example, an application example of the third rule is provided.
[0111] The first source language text is:
[0112] The second source language text is:
[0113] The first source language text after removing the punctuation and spaces is:
[0114] The second source language text after removing the punctuation and spaces is:
[0115] As can be seen from the above example, after removing the punctuation and spaces, the two are completely consistent, and the similarity threshold requirement is met, the first source language text can be used as the source language alignment text, and the audio segment corresponding to the first source language text can be used as the target audio segment.
[0116] In the embodiment, when the source language belongs to a set first language, punctuation marks and spaces in the first source language text and the second source language text are removed, and then alignment matching is performed, so that alignment failure caused by different space positions in the sentence between the second source language text obtained by recognizing the audio segment through the speech recognition model and the first source language text is avoided, and the recall rate of the training data is improved.
[0117] In another possible example, according to the third rule, the first source language text is used to align the second source language text to obtain the source language alignment text and filter the target audio segment corresponding to the source language alignment text, and the process includes:
[0118] In the case where the source language belongs to a set second language, the second source language text is re-spliced according to the space position in the first source language text, and then aligned with the first source language text, and the matching similarity is calculated, and if the similarity threshold requirement is met, the first source language text is used as the source language alignment text, and the audio segment corresponding to the first source language text is used as the target audio segment.
[0119] The embodiment can avoid alignment failure caused by improper space separation by re-splicing the second source language text according to the space position in the first source language text and then aligning and matching the first source language text. In the embodiment, after re-splicing the second source language text according to the space position in the first source language text and then aligning and matching the first source language text, a higher similarity threshold can be set, for example, the similarity threshold is set to 100%, that is, only when the re-spliced second source language text and the first source language text are exactly matched, it is determined that the third rule is met. At this time, the embodiment selects the first source language text in the bilingual subtitle information as the source language alignment text, and the audio segment corresponding to the first source language text as the target audio segment.
[0120] Further, when it is judged that the similarity threshold requirement is not met, it is determined that the third rule is not met.
[0121] In a possible implementation, when it is determined that the second rule is not met, the current audio segment and the corresponding first source language text and first target language text can be directly discarded as dirty data.
[0122] In another possible implementation, when it is determined that the second rule is not met, other source language alignment rules can be further verified.
[0123] Taking the source language Thai as an example, an application example of the third rule is provided:
[0124] The first source language text is:
[0125] .
[0126] The second source language text is:
[0127] .
[0128] After re-splicing the second source language text according to the space position in the first source language text: .
[0129] As can be seen from the above example, after re-splicing the second source language text according to the space position in the first source language text, the first source language text is completely consistent with the first source language text, the similarity threshold requirement is met, and the first source language text can be used as the source language alignment text, and the audio segment corresponding to the first source language text can be used as the target audio segment.
[0130] The method of the embodiment is used for re-splicing the second source language text according to the space position in the first source language text and then aligning and matching the first source language text, avoiding the alignment failure of the first source language text caused by the error of the second source language text due to the space separation when the audio segment is recognized by the speech recognition model, and thus improving the recall rate of the training data.
[0131] The foregoing embodiment provides three optional examples of the source language alignment rule, namely the first rule, the second rule and the third rule. The three rules can be used alternatively, or combined in any two or three.
[0132] In combination with Figure 3 which illustrates an optional alignment processing flow when the source language alignment rule includes the first rule, the second rule and the third rule.
[0133] For the text 1 and the text 2 to be aligned (corresponding to the first source language text and the second source language text), it is firstly judged whether the first rule is satisfied, if yes, the source language alignment text (the second source language correction text is taken as the source language alignment text) is obtained, if not, it is further judged whether there is a subtitle misplacement problem.
[0134] If there is no subtitle misplacement, the current audio segment and the first source language text and the first target language text are taken as dirty data and discarded; if there is a subtitle misplacement, it is further judged whether the second rule is satisfied.
[0135] If the second rule is satisfied, the source language alignment text (the second source language text is taken as the source language alignment text) is obtained, if the second rule is not satisfied, it is further judged whether the source language belongs to a special language.
[0136] If the source language does not belong to the special language, the current audio segment and the first source language text and the first target language text are taken as dirty data and discarded; if the source language belongs to the special language, it is further judged whether the third rule is satisfied.
[0137] If the third rule is satisfied, the source language alignment text (the first source language text is taken as the source language alignment text when the source language belongs to the first or second language) is obtained, if the third rule is not satisfied, the current audio segment and the first source language text and the first target language text are taken as dirty data and discarded.
[0138] The specific implementation process of the first rule, the second rule and the third rule can be referred to the foregoing relevant introduction, which will not be repeated here.
[0139] In some embodiments of the present application, the step S130 of the foregoing embodiments is based on the target audio segment and the first target language text corresponding to the target audio segment to generate the simultaneous interpretation translation training data.
[0140] In a possible implementation, the target audio segment and the first target language text corresponding to the target audio segment can be directly composed into the simultaneous interpretation translation training data.
[0141] Since the target audio segment is an audio segment after alignment screening, the dirty data in which the audio is not aligned with the first source language text is removed. The first target language text corresponding to the speech segment is obtained in the foregoing step S100, so that the first target language text corresponding to the target audio segment can be selected after the target audio segment is screened, and the simultaneous interpretation translation training data can be composed of the target audio segment and the first target language text corresponding to the target audio segment, so that relatively high-quality training data can be quickly obtained.
[0142] In another possible implementation, to further improve the quality of the training data, the first target language text can be further post-processed and aligned to obtain a target language aligned text, and the simultaneous interpretation translation training data can be composed of the target audio segment and the target language aligned text, so that the first target language text with errors can be further removed, and the quality of the finally obtained simultaneous interpretation translation training data can be improved.
[0143] The specific implementation process can include:
[0144] In step S400, a source language aligned text corresponding to the target audio segment is translated into a target language by using a configured translation model to obtain a second target language text corresponding to the target audio segment.
[0145] The translation model can be a neural network model with translation capability in multiple structures. As known from the foregoing scheme, the source language aligned text is a text obtained by aligning and correcting the second source language text with reference to the first source language text, and the quality of the source language aligned text is higher. On this basis, the second target language text with high quality can be obtained by translating the source language aligned text.
[0146] In step S410, a target language aligned text corresponding to the target audio segment is obtained by aligning the second target language text with the first target language text corresponding to the target audio segment according to a configured target language alignment rule.
[0147] In view of the inaccurate translation of the translation model, the alignment rules of the target language are preconfigured in the embodiment, and the second target language text is aligned with the first target language text according to the alignment rules to obtain the target language alignment text. The quality of the target language alignment text is higher than that of the first target language text and the second target language text. When the second target language text obtained by translating the source language alignment text corresponding to the target audio segment through the translation model fails to align with the first target language text (the target audio segment cannot obtain the corresponding target language alignment text), it can be determined that the target audio segment is dirty data, and the target audio segment is discarded, thereby improving the quality of the training data obtained finally.
[0148] Step S420: The simultaneous translation training data is composed of the target audio segment and the corresponding target language alignment text.
[0149] In the embodiment, the first source language text filtered result (source language alignment text) is introduced into the translation model for translation to obtain the second target language text, and the first target language text in the bilingual subtitle information is post-processed and aligned, and the available high-quality simultaneous translation training data is twice filtered.
[0150] In some embodiments of the application, some optional examples of the target language alignment rules are introduced.
[0151] The target language alignment rules can include a fourth rule.
[0152] According to the fourth rule, the second target language text is aligned with the first target language text to obtain the target language alignment text, and the process includes:
[0153] Step S500: The first target language text and the second target language text are aligned by the edit distance, and the similarity is calculated. If the similarity threshold requirement is met, it is determined that the fourth rule is met, and the second target language text is taken as the target language alignment text. If the similarity threshold requirement is not met, it is determined that the fourth rule is not met.
[0154] In a possible implementation, when it is determined that the fourth rule is not met, the current audio segment and the corresponding first source language text and first target language text can be directly discarded as dirty data.
[0155] In another possible implementation, when it is determined that the fourth rule is not met, the other target language alignment rules can be further verified.
[0156] In some embodiments of the application, the target language alignment rules can include a fifth rule.
[0157] According to the fifth rule, the process of aligning the second target language text with the first target language text to obtain the target language aligned text includes:
[0158] Step S600, in the case of determining that the bilingual subtitle information appears before and after the timestamp, the first target language text of any audio segment is merged with the first target language text of the adjacent audio segment to obtain the target language merged text.
[0159] Specifically, for the collected bilingual subtitle audio-video data, there can be a case of subtitle misplacement before and after the timestamp, such as the audio content of the current time period, part or all of which appears in the subtitle text of the previous adjacent or next adjacent time period of the current time period. For example, the audio content of the t1-t2 time period appears in the t2-t3 time period or the t0-t1 time period. Among the conditions for determining whether there is subtitle misplacement, the proportion of the audio content of the current time period appearing in the subtitle text of the previous adjacent or next adjacent time period of the current time period can be selected from 50%-100%, which can be selected by the staff according to business needs. When the proportion is larger, the quality requirement for the training data is higher, and when the proportion is smaller, the recall rate of the training data is higher.
[0160] In this step, in the case of determining that the bilingual subtitle appears before and after the timestamp, the first target language text of any audio segment can be merged with the first target language text of the adjacent audio segment (which can be the previous adjacent audio segment, the next adjacent audio segment, or the previous and next adjacent audio segments) to obtain the target language merged text. The target language merged text is used as the object for aligning and matching with the second target language text.
[0161] Among them, there can be many ways to determine whether the bilingual subtitle information appears before and after the timestamp, such as: the pre-acquired bilingual subtitle audio-video data is labeled with an identifier indicating whether there is subtitle misplacement, so based on the identifier, it can be determined whether the current bilingual subtitle audio-video data has subtitle misplacement. For example, after segmenting the bilingual subtitle audio stream, each audio segment can be sequentially aligned and processed. When aligning, it can be detected whether the speech recognition result of the current audio segment exists in the subtitle text of the previous adjacent time period or the next adjacent time period. If so, it is determined that the bilingual subtitle information appears before and after the timestamp.
[0162] Step S610, cross-segment fuzzy matching the target language merged text with the second target language text, and calculating the matching similarity. If the similarity threshold requirement is met, it is determined that the second rule is met, and the second target language text is used as the target language aligned text.
[0163] Further, if the similarity does not satisfy the similarity threshold requirement, it is determined that the fifth rule is not satisfied.
[0164] In a possible implementation, when it is determined that the fifth rule is not satisfied, the current audio segment and the corresponding first source language text and first target language text can be directly discarded as dirty data.
[0165] In another possible implementation, when it is determined that the fifth rule is not satisfied, other target language alignment rules can be further verified.
[0166] In this embodiment, considering that when the subtitle misplacement phenomenon of the timestamps before and after the bilingual subtitle information occurs, if the first target language text and the second target language text are directly aligned, the alignment can fail due to the subtitle misplacement, a large amount of data is discarded by mistake, and the data utilization rate is too low. In this embodiment, when the subtitle misplacement phenomenon of the timestamps before and after the bilingual subtitle information occurs, the first target language text of any audio segment is merged with the first target language text of the adjacent audio segment to obtain the target language merged text, and the cross-segment fuzzy matching is performed between the target language merged text and the second target language text corresponding to the audio segment, which can greatly alleviate the problem of low data recall rate caused by the subtitle misplacement, and improve the utilization rate of the bilingual subtitle audio and video data.
[0167] In some embodiments of the present application, the target language alignment rule can include a sixth rule. The sixth rule is an alignment rule formulated for the case where the target language belongs to a set special language. Considering the particularity of some languages, the translation result of the target language obtained after translation by the translation model can exist due to the different division modes of punctuation marks or spaces, resulting in alignment failure with the first target language text. Therefore, the present application can formulate corresponding alignment rules for some set special languages. In this embodiment, the case where the set special language includes two first languages and second languages is taken as an example for description.
[0168] The first language is a language with different segmentation modes for the same word, and Korean is taken as an example of the first language. The same word in Korean has different segmentation modes, and when the first target language text and the second target language text are aligned and matched, the alignment can fail and the recall rate can be reduced due to different spaces in the sentence.
[0169] The second language is a language that separates sentences by spaces, and Thai is taken as an example of the second language. Thai grammar should separate sentences by spaces, but the second target language text can exist in the case of separating by words, resulting in alignment failure and reduced recall rate when the second target language text and the first target language text are aligned and matched due to different spaces in the sentence.
[0170] In one possible example, according to the sixth rule, the process of aligning the second target language text with the first target language text includes:
[0171] In the case where the target language belongs to the set first language, the punctuation marks and spaces in the first target language text and the second target language text are removed, and then the alignment matching is performed, and the matching similarity is calculated. If the similarity threshold requirement is met, the first target language text is taken as the source language alignment text.
[0172] By removing the punctuation marks and spaces in the first target language text and the second target language text, the alignment failure caused by different segmentation methods can be avoided. In this embodiment, the alignment matching is performed after the punctuation marks and spaces are removed. At this time, a higher similarity threshold can be set, such as setting the similarity threshold to 100%. That is, only when the first target language text and the second target language text after removing the punctuation marks and spaces are exactly matched, it is determined that the sixth rule is met. At this time, the present application selects the first target language text in the bilingual subtitle information as the target language alignment text.
[0173] Further, when it is determined that the similarity threshold requirement is not met, it is determined that the sixth rule is not met.
[0174] In one possible implementation, when it is determined that the sixth rule is not met, the current audio segment and the corresponding first source language text and first target language text can be directly discarded as dirty data.
[0175] In another possible implementation, when it is determined that the sixth rule is not met, the other target language alignment rules can be further verified.
[0176] The method of the present embodiment removes the punctuation marks and spaces in the first target language text and the second target language text in the case where the target language belongs to the set first language, and then performs alignment matching, which avoids the alignment failure of the second target language text obtained by the translation model and the first target language text due to different space positions in the sentence, thereby improving the recall rate of the training data.
[0177] In another possible example, according to the sixth rule, the process of aligning the second target language text with the first target language text includes:
[0178] In the case where the target language belongs to the set second language, the second target language text is re-spliced according to the space position in the first target language text, and then the alignment matching is performed with the first target language text, and the matching similarity is calculated. If the similarity threshold requirement is met, the first target language text is taken as the target language alignment text.
[0179] The embodiment can avoid alignment failure caused by improper space separation by re-splicing the second target language text according to the space position in the first target language text and then aligning and matching the first target language text. In the embodiment, after re-splicing the second target language text according to the space position in the first target language text and then aligning and matching the first source language text, a higher similarity threshold can be set, for example, the similarity threshold is set to 100%, that is, only when the re-spliced second target language text and the first target language text are exactly matched, it is determined that the sixth rule is met. At this time, the embodiment selects the first target language text in the bilingual subtitle information as the target language alignment text.
[0180] Further, when it is judged that the similarity threshold requirement is not met, it is determined that the sixth rule is not met.
[0181] In a possible implementation, when it is determined that the sixth rule is not met, the current audio segment and the corresponding first source language text and first target language text can be directly discarded as dirty data.
[0182] In another possible implementation, when it is determined that the sixth rule is not met, other target language alignment rules can be further verified.
[0183] The embodiment method re-splices the second target language text according to the space position in the first target language text and then aligns and matches the first target language text when the target language belongs to a set second language, avoiding the case that the second target language text obtained by the translation model fails to align with the first target language text due to space separation errors, thereby improving the recall rate of the training data.
[0184] The foregoing embodiments provide three optional examples of target language alignment rules, namely the fourth rule, the fifth rule and the sixth rule. The three rules can be used individually, or any two can be combined or the three rules can be combined.
[0185] In combination Figure 4 which illustrates an optional alignment processing flow when the target language alignment rule includes the fourth rule, the fifth rule and the sixth rule.
[0186] For the to-be-aligned text 1 and the to-be-aligned text 2 (corresponding to the first target language text and the second target language text), it is first judged whether the fourth rule is met, if yes, the target language alignment text (the second target language text is taken as the target language alignment text) is obtained, if not, it can be further judged whether there is a subtitle misplacement problem.
[0187] If there is no subtitle misplacement, the current audio segment and the first source language text and the first target language text are discarded as dirty data; if there is subtitle misplacement, it can be further determined whether the fifth rule is satisfied.
[0188] If the fifth rule is satisfied, the target language alignment text is obtained (the second target language text is taken as the target language alignment text); if the fifth rule is not satisfied, it can be further determined whether the target language belongs to a set special language.
[0189] If the target language does not belong to the set special language, the current audio segment and the first source language text and the first target language text are discarded as dirty data; if the target language belongs to the set special language, it is further determined whether the sixth rule is satisfied.
[0190] If the sixth rule is satisfied, the target language alignment text is obtained (when the target language belongs to the first or second language, the first target language text is taken as the target language alignment text); if the sixth rule is not satisfied, the current audio segment and the first source language text and the first target language text are discarded as dirty data.
[0191] The specific implementation process of the fourth rule, the fifth rule and the sixth rule can be referred to the foregoing description, and will not be described here.
[0192] In a possible implementation, after obtaining the target language alignment text and the corresponding target audio segment, the target language alignment text and the corresponding target audio segment can be further used as training data of a translation model. When a set update condition is reached, the translation model used in step S400 is updated and trained by using the training data, so as to optimize the performance of the translation model and improve the translation accuracy of the translation model for the source language alignment text. By using the obtained target language alignment text and the corresponding target audio segment as training data, the translation model is constantly updated, so as to improve the translation accuracy of the translation model, thereby further improving the quality of the simultaneous interpretation translation training data obtained finally.
[0193] The embodiment of the present application can automatically generate simultaneous interpretation translation training data based on bilingual subtitle audio-video data. The bilingual subtitle audio-video data can include two types. Next, the implementation manner of step S100 in the foregoing embodiment, based on the bilingual subtitle information in the bilingual subtitle audio-video data, obtaining the first source language text and the first target language text corresponding to each audio segment, will be introduced respectively for the two types of bilingual subtitle audio-video data.
[0194] Referring to Figure 5
[0195] For audio and video data with bilingual subtitle file (composed of audio and video and its corresponding bilingual subtitle file), in this embodiment, the first source language text and the first target language text corresponding to each audio segment can be obtained based on the bilingual subtitle file. For example, the timestamp information is extracted from the bilingual subtitle file, and the source language text (as the first source language text) and the target language text (as the first target language text) of the audio segment corresponding to adjacent timestamps. Further, the audio stream can be extracted from the bilingual subtitle audio and video data, and divided into audio segments according to the timestamps, to obtain more than one audio segment.
[0196] Referring to Figure 6
[0197] For video data with embedded bilingual subtitles, in this embodiment, the audio stream is extracted from the video data, and the audio stream is divided into audio segments by using voice activity detection (VAD). Specifically, the timestamp information is obtained by VAD detection, and the audio stream is further divided according to the timestamps to obtain the audio segments. The audio segments can be used as the object of speech recognition in step S110 to obtain the second source language text corresponding to the audio segments.
[0198] In this embodiment, the video frames within the time length corresponding to each audio segment can also be intercepted from the video data according to the timestamp of each audio segment to obtain the video frame images.
[0199] The video frame images are divided into more than two subgraphs, and the bilingual subtitles in each subgraph are identified to obtain the source language text and the target language text contained in each subgraph.
[0200] When the video frame images are divided into subgraphs, various division methods can be used, such as dividing the video frame images into subgraphs from left to right, or dividing the video frame images into subgraphs from top to bottom, etc.
[0201] Figure 6 The video frame images are divided into three subgraphs, i.e., upper subgraph, middle subgraph and lower subgraph, as shown in the example in FIG. 8. The bilingual subtitles in the three subgraphs are identified to obtain the source language text and the target language text in the upper subgraph, the source language text and the target language text in the middle subgraph, and the source language text and the target language text in the lower subgraph.
[0202] When the subgraphs are identified for bilingual subtitles, an OCR model can be used. In order to improve the accuracy of subtitle recognition, in this embodiment, a multi-language OCR large model can be used, which uses a large model structure and has the ability to recognize multiple languages.
[0203] The second source language text is respectively aligned with the source language text of each subgraph by using the edit distance and the similarity is calculated to determine the target subgraph that meets the similarity threshold requirement.
[0204] The video frame image can contain interference information. In this embodiment, by respectively aligning the second source language text with the source language text of each subgraph by using the edit distance and calculating the similarity, the subgraph that does not meet the similarity threshold requirement can be removed, that is, the influence of the interference information is excluded.
[0205] The source language text contained in the target subgraph is merged to obtain the source language text contained in the video frame image, and the target language text contained in the target subgraph is merged to obtain the target language text contained in the video frame image.
[0206] Specifically, one video frame image contains several subgraphs, and the target subgraph retained after the above similarity threshold screening can be one or more. In this step, the source language text contained in each retained target subgraph can be merged, and the result is used as the source language text contained in the video frame image. When merging the source language text of each target subgraph, the position of each target subgraph in the video frame image can be referred to, and the merging is performed in the order of the set position. Taking the case of dividing the video frame image into three subgraphs of top, middle and bottom as an example, the source language text contained in the target subgraph can be sequentially merged in the order of top, middle and bottom to obtain the source language text contained in the video frame image. Similarly, the target language text contained in the target subgraph is also merged to obtain the target language text contained in the video frame image.
[0207] Further, the source language text contained in each video frame within the audio segment corresponding duration can be de-duplicated and merged to obtain the first source language text corresponding to the audio segment, and the target language text contained in each video frame within the audio segment corresponding duration can be de-duplicated and merged to obtain the first target language text corresponding to the audio segment.
[0208] This embodiment is aimed at video data with embedded bilingual subtitles. Considering that the position of the subtitles is not fixed and the video frame can contain other interference text in addition to the subtitle information, the OCR recognition result can be affected. Therefore, in this embodiment, the video frame image is divided, each divided subgraph is sent into the OCR model for recognition, the recognition result of each subgraph is respectively aligned with the speech recognition result (second source language text), the subgraph with a similarity meeting the threshold requirement is found, and the subgraph with a similarity not meeting the threshold requirement is removed, thereby reducing the interference information and improving the recognition accuracy of the first source language text and the first target language text.
[0209] In a possible implementation, the process of performing speech recognition on the audio segment (i.e., the process of performing speech recognition on the audio segment by the ASR model) in step S110 can be performed in parallel with the process of recognizing the bilingual subtitles in each subgraph (i.e., the process of recognizing the bilingual subtitles in the subgraph by the OCR model), such as performing the ASR model and the OCR model by different processes, so as to improve the data processing efficiency and obtain the simultaneous interpretation translation training data more efficiently.
[0210] In a possible implementation, to further improve the data processing efficiency, the video frame image can be compressed first, and then divided into subgraphs, so as to speed up the processing speed of the OCR model. Of course, before the video frame image is compressed, it can be first determined whether the size of the video frame image exceeds a set size threshold. If the size exceeds the set size threshold, the operation of compressing the video frame image is performed. Otherwise, the video frame image can be directly divided into subgraphs, so as to improve the processing speed as much as possible on the premise of ensuring the image quality.
[0211] In some embodiments of the present application, in combination with Figure 5 As shown in FIG. 1, a method for generating simultaneous interpretation translation training data is provided, which specifically includes the following steps:
[0212] For audio-video data with bilingual external subtitle files (composed of audio-video and corresponding bilingual external subtitle files), the audio stream is extracted from the audio-video, the timestamp information and the source language text (as the first source language text) and the target language text (as the first target language text) of the audio segment corresponding to adjacent timestamps are extracted from the bilingual external subtitle file.
[0213] The audio stream is divided according to the timestamps to obtain a plurality of audio segments. The audio segments are recognized by an ASR model (a multilingual ASR large model can be used) to obtain the second source language text corresponding to the audio segments.
[0214] The first source language text is aligned with the second source language text, and it is determined whether the source language alignment rule is satisfied. The source language alignment rule can refer to the related embodiments introduced in the foregoing description, which will not be described herein. If not, the current audio segment is discarded as dirty data. If yes, the source language alignment text and the corresponding target audio segment can be obtained.
[0215] The source language alignment text is input into a translation model to be translated into a target language to obtain the second target language text.
[0216] The first target language text is aligned with the second target language text, and it is determined whether a set target language alignment rule is met. The target language alignment rule can be referred to the related embodiments described above, and will not be described here again. If not met, the current audio segment is discarded as dirty data. If met, the target language alignment text can be obtained.
[0217] The target audio segment and the target language alignment text obtained by the above process can be used as the simultaneous interpretation translation training data.
[0218] For the target audio segment and the corresponding source language alignment text obtained by the above process, the ASR model can be updated and trained by using the target audio segment and the corresponding source language alignment text as the training data, so as to optimize the speech recognition effect of the ASR model, and further promote the quality of the subsequent generated simultaneous interpretation translation training data.
[0219] For the target audio segment and the corresponding target language alignment text obtained by the above process, the translation model can be updated and trained by using the target audio segment and the corresponding target language alignment text as the training data, so as to optimize the translation effect of the translation model, and further promote the quality of the subsequent generated simultaneous interpretation translation training data.
[0220] The method provided in the embodiment provides a solution to the problem of low data recall rate caused by misalignment of bilingual external subtitle timestamps and text. By introducing cross-segment fuzzy matching in the source language alignment rule, the problem of low data recall rate caused by misalignment of text and timestamps is solved. By introducing homophone matching and combining the source language subtitle text, the recognition result (second source language text) of the ASR is corrected, the accuracy of the ASR recognition result is improved, the recall rate is improved, and the quality of the translation data is improved. For special languages, corresponding alignment rules are set to adapt to the characteristics of different languages, improve the accuracy of the ASR recognition result for special languages, and improve the recall rate. By introducing a translation model to translate the source language alignment text, a second target language text is obtained, and the first target language text in the bilingual subtitle file is post-processed and aligned. The available simultaneous interpretation translation training data is screened twice to improve the quality of the training data. The embodiment of the application further sends the target audio segment and the source language alignment text after screening and alignment into a multilingual ASR large model for training, to optimize the recognition result of the multilingual ASR large model. At the same time, the target audio segment and the target language alignment text after screening are used as translation data pairs to join the translation model training, to optimize the translation effect. The quality of the subsequent generated simultaneous interpretation translation training data can be further promoted.
[0221] In some embodiments of the application, the method for generating simultaneous interpretation translation training data is combined with Figure 6 As shown in FIG. 6, another method for generating simultaneous interpretation translation training data is provided, which specifically includes the following steps:
[0222] For the video data with embedded bilingual subtitles, the embodiment extracts an audio stream from the video data, obtains timestamp information through VAD detection, further splits the audio stream according to the timestamp to obtain an audio segment. The audio segment can be subjected to speech recognition through an ASR model (a multilingual ASR large model can be used) to obtain second source language text corresponding to the audio segment.
[0223] In addition, video frames within a duration corresponding to adjacent timestamps can be intercepted from the video data according to the timestamp obtained through VAD detection to obtain video frame images.
[0224] The video frame images are uniformly divided into three subgraphs, i.e., upper, middle and lower subgraphs. The bilingual subtitles in the three subgraphs are recognized through an OCR model (a multilingual OCR large model can be used) to obtain source language text and target language text in the upper subgraph, source language text and target language text in the middle subgraph, and source language text and target language text in the lower subgraph.
[0225] The ASR model processing process and the OCR model processing process can be executed in parallel to improve processing efficiency. Before the subgraphs are recognized using the OCR model, the video frame images can be compressed and then divided to further improve processing efficiency.
[0226] Further, the second source language text is respectively aligned with the source language text of each subgraph in terms of edit distance and the similarity is calculated, and target subgraphs that meet the similarity threshold requirement are screened.
[0227] The source language text contained in the target subgraphs is merged to obtain source language text contained in the video frame images, and the target language text contained in the target subgraphs is merged to obtain target language text contained in the video frame images. The source language text contained in each video frame within the duration corresponding to the audio segment is de-duplicated and merged to obtain first source language text corresponding to the audio segment, and the target language text contained in each video frame within the duration corresponding to the audio segment is de-duplicated and merged to obtain first target language text corresponding to the audio segment.
[0228] The first source language text is used to align the second source language text, and it is determined whether the source language alignment rule is met, wherein the source language alignment rule can refer to the related embodiments described above, and will not be described here. If not, the current audio segment is discarded as dirty data, and if so, source language alignment text and corresponding target audio segments can be obtained.
[0229] The source language alignment text is input into a translation model to be translated into a target language to obtain second target language text.
[0230] The second target language text is aligned with the first target language text, and it is determined whether a set target language alignment rule is met. The target language alignment rule can be referred to the related embodiments described above, and will not be described here. If not met, the current audio segment is discarded as dirty data. If met, the target language alignment text can be obtained.
[0231] The target audio segment and the target language alignment text obtained by the above process can be used as simultaneous interpretation translation training data.
[0232] The target audio segment and the corresponding source language alignment text obtained by the above process can be used as training data to update and train the ASR model, so as to optimize the speech recognition effect of the ASR model, and further promote the quality of the subsequent generated simultaneous interpretation translation training data.
[0233] The target audio segment and the corresponding target language alignment text obtained by the above process can be used as training data to update and train the translation model, so as to optimize the translation effect of the translation model, and further promote the quality of the subsequent generated simultaneous interpretation translation training data.
[0234] The method provided by the embodiment provides a solution to the problems of low processing efficiency of bilingual embedded audio and video data, indefinite embedded subtitle position, low recall rate caused by the failure of ASR and OCR to align, and low accuracy of translated text. First, to solve the problem of low data processing efficiency, on the one hand, the ASR and OCR models are run in parallel, and multi-processes are used to improve the data processing efficiency; on the other hand, the size of the OCR processing image is appropriately compressed while ensuring the effect, to improve the efficiency of OCR. Second, to solve the problem of the failure of ASR and OCR to align and low recall rate, on the one hand, a multilingual ASR large model and a multilingual OCR large model are introduced to improve the recognition quality; on the other hand, the source language alignment rule is introduced to match similar characters and combine the results of OCR to correct the ASR recognition results. To solve the problem of the indefinite subtitle position affecting the recognition results of OCR, the video frame image is cut and sent into the OCR model, each part is aligned with the ASR result, and the part with the highest similarity is found to reduce the interference information and improve the accuracy. For special languages, corresponding alignment rules are set to adapt to the characteristics of different languages, improve the accuracy of ASR recognition results for special languages, and improve the recall rate. The source language alignment text is translated into a second target language text by introducing a translation model, and the first target language text in the bilingual subtitle file is post-processed and aligned to perform secondary screening on the available simultaneous interpretation translation training data, to improve the quality of the training data. The target audio segment and the source language alignment text after screening are further sent into a multilingual ASR large model for training, to optimize the recognition results of the multilingual ASR large model; at the same time, the target audio segment and the target language alignment text after screening are used as translation data pairs to join the training of the translation model, to optimize the translation effect. The quality of the subsequent generated simultaneous interpretation translation training data can be further improved.
[0235] The simultaneous interpretation translation training data generation device provided by the embodiment of the present application is described below. The simultaneous interpretation translation training data generation device described below can be correspondingly referred to the simultaneous interpretation translation training data generation method described above.
[0236] Referring to Figure 7 , Figure 7 A structure diagram of a simultaneous interpretation translation training data generation device disclosed by the embodiment of the present application is shown in FIG. 1.
[0237] As Figure 7 shown, the device can include:
[0238] The bilingual subtitle information processing unit 11 is configured to obtain a first source language text and a first target language text corresponding to each audio segment based on bilingual subtitle information in bilingual subtitle audio and video data, wherein the audio segment is a segment of audio contained in the bilingual subtitle audio and video data.
[0239] The speech recognition unit 12 is configured to perform speech recognition on the audio segment to obtain second source language text corresponding to the audio segment.
[0240] The text alignment unit 13 is configured to perform alignment on the second source language text according to a configured source language alignment rule and using the first source language text to obtain source language alignment text and screen target audio segments corresponding to the source language alignment text.
[0241] The simultaneous interpretation training data generation unit 14 is configured to generate simultaneous interpretation translation training data based on the target audio segment and the first target language text corresponding to the target audio segment.
[0242] In a possible implementation, the source language alignment rule includes a first rule, and the process of performing alignment on the second source language text according to the first rule and using the first source language text to obtain source language alignment text and screen target audio segments corresponding to the source language alignment text includes:
[0243] performing near-omophonic matching correction on the second source language text using the first source language text to obtain second source language corrected text;
[0244] calculating the similarity between the second source language corrected text and the first source language text, and if the similarity meets a similarity threshold requirement, regarding the second source language corrected text as the source language alignment text and regarding the audio segment corresponding to the first source language text as the target audio segment.
[0245] In a possible implementation, the process of performing near-omophonic matching correction on the second source language text using the first source language text to obtain second source language corrected text includes:
[0246] calculating the edit distance between the first source language text and the second source language text, and performing element alignment according to the edit distance to obtain first source language alignment text and second source language alignment text;
[0247] for each element in the second source language alignment text, determining an alignment element of the element in the first source language alignment text, and performing near-omophonic matching correction on the element using adjacent elements of the alignment element to obtain second source language corrected text.
[0248] In a possible implementation, the source language alignment rule includes a second rule, and the process of performing alignment on the second source language text according to the second rule and using the first source language text to obtain source language alignment text and screen target audio segments corresponding to the source language alignment text includes:
[0249] merge the first source language text of any audio segment with the first source language text of adjacent audio segments to obtain a source language merged text in a case of determining a subtitle misplacement of the bilingual subtitle information occurrence time stamp;
[0250] perform cross-segment fuzzy matching on the source language merged text and the second source language text, and calculate a matching similarity, and if a similarity threshold requirement is met, take the second source language text as the source language alignment text, and take an audio segment corresponding to the second source language text as the target audio segment.
[0251] In a possible implementation, the source language alignment rule includes a third rule; and the text alignment unit performs alignment on the second source language text by using the first source language text according to the third rule to obtain a source language alignment text and filters a target audio segment corresponding to the source language alignment text, including:
[0252] in a case where the source language belongs to a set first language, removing punctuation marks and spaces in the first source language text and the second source language text, and then performing alignment matching and calculating a matching similarity, and if a similarity threshold requirement is met, taking the first source language text as the source language alignment text and taking an audio segment corresponding to the first source language text as the target audio segment;
[0253] in a case where the source language belongs to a set second language, re-splicing the second source language text according to a space position in the first source language text, and then performing alignment matching with the first source language text and calculating a matching similarity, and if a similarity threshold requirement is met, taking the first source language text as the source language alignment text and taking an audio segment corresponding to the first source language text as the target audio segment.
[0254] In a possible implementation, the simultaneous interpretation training data generation unit generates simultaneous interpretation translation training data based on the target audio segment and the first target language text corresponding to the target audio segment, including:
[0255] The target audio segment and the first target language text corresponding thereto constitute simultaneous interpretation translation training data.
[0256] In another possible implementation, the simultaneous interpretation training data generation unit generates simultaneous interpretation translation training data based on the target audio segment and the first target language text corresponding to the target audio segment, including:
[0257] use a configured translation model to translate the source language alignment text corresponding to the target audio segment into a target language to obtain a second target language text corresponding to the target audio segment.
[0258] aligning the second target language text with the first target language text corresponding to the target audio segment according to the configured target language alignment rule, to obtain a target language alignment text corresponding to the target audio segment;
[0259] The simultaneous translation training data is composed of the target audio segment and the target language alignment text corresponding to the target audio segment.
[0260] In a possible implementation, the process of performing speech recognition on the audio segment is implemented by using a configured speech recognition model, and the device further includes:
[0261] The speech recognition model updating unit is configured to update and train the speech recognition model by using the source language alignment text and the target audio segment corresponding to the source language alignment text as training data of the speech recognition model.
[0262] In a possible implementation, the device further includes:
[0263] The translation model updating unit is configured to update and train the translation model by using the target audio segment and the target language alignment text corresponding to the target audio segment as training data of the translation model.
[0264] In a possible implementation, the process of obtaining, by the bilingual subtitle information processing unit, the first source language text and the first target language text corresponding to each audio segment based on the bilingual subtitle information in the bilingual subtitle audio and video data includes:
[0265] When the bilingual subtitle audio and video data is audio and video data with a bilingual external subtitle file, the bilingual subtitle information processing unit extracts, from the bilingual external subtitle file, a timestamp, the first source language text and the first target language text of an audio segment corresponding to adjacent timestamps, and extracts an audio stream from the audio and video data;
[0266] The audio stream is cut into audio segments according to the timestamp.
[0267] In another possible implementation, the process of obtaining, by the bilingual subtitle information processing unit, the first source language text and the first target language text corresponding to each audio segment based on the bilingual subtitle information in the bilingual subtitle audio and video data includes:
[0268] When the bilingual subtitle audio and video data is video data with embedded bilingual subtitles, the bilingual subtitle information processing unit extracts an audio stream from the video data, cuts the audio stream into audio segments by using voice activity detection (VAD), and intercepts video frames in a time period corresponding to the audio segment from the video data;
[0269] cutting the video frame into two or more subgraphs, respectively identifying the bilingual subtitles in each subgraph to obtain source language text and target language text contained in each subgraph;
[0270] aligning the second source language text with the source language text of each subgraph respectively according to the edit distance and calculating the similarity, and determining a target subgraph that meets the similarity threshold requirement;
[0271] merging the source language text contained in the target subgraph to obtain the source language text contained in the video frame, and merging the target language text contained in the target subgraph to obtain the target language text contained in the video frame;
[0272] performing deduplication and merging on the source language text contained in each video frame within the corresponding duration of the audio segment to obtain the first source language text corresponding to the audio segment, and performing deduplication and merging on the target language text contained in each video frame within the corresponding duration of the audio segment to obtain the first target language text corresponding to the audio segment.
[0273] In a possible implementation, the process of speech recognition performed by the speech recognition unit on the audio segment is performed in parallel with the process of identifying the bilingual subtitles in each subgraph by the bilingual subtitle information processing unit.
[0274] Each unit in the simultaneous interpretation translation training data generation apparatus can be implemented by software, hardware, or a combination thereof. Each unit can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so as to be invoked and executed by a processor to perform operations corresponding to each unit.
[0275] In the embodiments of the present application, an electronic device is also provided. Referring to Figure 8 , a structural diagram suitable for implementing the electronic device in the embodiments of the present application is shown. The electronic device in the embodiments of the present application can include but is not limited to a terminal such as a mobile phone, a tablet computer, a translation machine, and the like. Figure 8 The electronic device shown is only an example, and should not impose any limitation on the functions and use range of the embodiments of the present application.
[0276] As Figure 8As shown, the electronic device can include a processing device (e.g., a central processor, a graphics processor, etc.) 1 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 2 or programs loaded from a storage device 8 into a random access memory (RAM) 3 to implement the simultaneous interpretation translation training data generation method of the foregoing embodiments of the present application. In a state where the electronic device is powered on, various programs and data required for operation of the electronic device are also stored in the RAM 3. The processing device 1, the ROM 2, and the RAM 3 are connected to each other through a bus 4. An input / output (I / O) interface 5 is also connected to the bus 4.
[0277] Generally, the following devices can be connected to the I / O interface 5: input devices 6 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 7 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 8 including, for example, a memory card, a hard disk, etc.; and communication devices 9. The communication devices 9 can allow the electronic device to communicate wirelessly or wired with other devices to exchange data. Although Figure 8 The electronic device with various devices is shown, but it should be understood that it is not required to implement or have all the shown devices. More or less devices can alternatively be implemented or provided.
[0278] The embodiments of the present application also provide a computer program product including computer readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the simultaneous interpretation translation training data generation methods provided by the embodiments of the present application.
[0279] The embodiments of the present application also provide a computer readable storage medium carrying one or more computer programs, which, when executed by an electronic device, can cause the electronic device to implement any of the simultaneous interpretation translation training data generation methods provided by the embodiments of the present application.
[0280] In addition, it should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments. In addition, the connection relationship between the modules in the device embodiments provided by the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0281] Those skilled in the art can clearly understand, through the description of the foregoing embodiments, that the present application can be implemented by means of software and the necessary universal hardware, and of course can also be implemented by means of special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can be various, such as analog circuits, digital circuits, or special circuits. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions for causing a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0282] In the above embodiments, all or part can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part can be implemented in the form of a computer program product.
[0283] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, training device or data center to another through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0284] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The embodiments can be combined as needed, and the same or similar parts refer to each other.
Claims
1. A method for generating training data for simultaneous interpreting, characterized in that, The method comprises the following steps: Based on the bilingual subtitle information in the bilingual subtitle audio and video data, the first source language text and the first target language text corresponding to each audio segment are obtained, the audio segment being a segment contained in the audio in the bilingual subtitle audio and video data; Speech recognition is performed on the audio segment to obtain the second source language text corresponding to the audio segment; According to the configured source language alignment rule, the second source language text is aligned using the first source language text to obtain the source language alignment text and screen the target audio segment corresponding to the source language alignment text; Based on the target audio segment and the first target language text corresponding to the target audio segment, the simultaneous interpretation translation training data is generated; The process of generating the simultaneous interpretation translation training data based on the target audio segment and the first target language text corresponding to the target audio segment comprises: Using the configured translation model, the source language alignment text corresponding to the target audio segment is translated into the target language to obtain the second target language text corresponding to the target audio segment; According to the configured target language alignment rule, the second target language text is aligned using the first target language text corresponding to the target audio segment to obtain the target language alignment text corresponding to the target audio segment; The target audio segment and the target language alignment text corresponding thereto constitute the simultaneous interpretation translation training data.
2. The method of claim 1, wherein, The source language alignment rule comprises a first rule; according to the first rule, the second source language text is aligned using the first source language text to obtain the source language alignment text and screen the target audio segment corresponding to the source language alignment text, which comprises: The second source language text is corrected by shape similar word matching using the first source language text to obtain the second source language correction text; The similarity of the second source language correction text and the first source language text is calculated, and if the similarity threshold requirement is met, the second source language correction text is taken as the source language alignment text, and the audio segment corresponding to the first source language text is taken as the target audio segment.
3. The method of claim 2, wherein, The process of correcting the second source language text by shape similar word matching using the first source language text to obtain the second source language correction text comprises: The edit distance of the first source language text and the second source language text is calculated, and the elements are aligned according to the edit distance to obtain the first source language alignment text and the second source language alignment text; For each element in the second source language alignment text, the alignment element of the element in the first source language alignment text is determined, and the element is corrected by shape similar word matching using the adjacent elements of the alignment element to obtain the second source language correction text.
4. The method of claim 2, wherein, The source language alignment rule comprises a second rule; according to the second rule, the second source language text is aligned using the first source language text to obtain the source language alignment text and screen the target audio segment corresponding to the source language alignment text, which comprises: In a case where the subtitle misplacement of the bilingual subtitle information appearing timestamps is determined, the first source language text of any audio segment is merged with the first source language text of an adjacent audio segment to obtain a source language merged text; The source language merged text is cross-segment fuzzy matched with the second source language text, and a matching similarity is calculated. If a similarity threshold requirement is met, the second source language text is taken as the source language alignment text, and an audio segment corresponding to the second source language text is taken as the target audio segment.
5. The method of claim 2, wherein, The source language alignment rule includes a third rule. According to the third rule, the second source language text is aligned with the first source language text to obtain a source language alignment text and filter a target audio segment corresponding to the source language alignment text, including: In a case where the source language belongs to a set first language, punctuation marks and spaces in the first source language text and the second source language text are removed, and then alignment matching is performed, and a matching similarity is calculated. If a similarity threshold requirement is met, the first source language text is taken as the source language alignment text, and an audio segment corresponding to the first source language text is taken as the target audio segment; And / or, In a case where the source language belongs to a set second language, the second source language text is re-spliced according to a space position in the first source language text, and then alignment matching is performed with the first source language text, and a matching similarity is calculated. If a similarity threshold requirement is met, the first source language text is taken as the source language alignment text, and an audio segment corresponding to the first source language text is taken as the target audio segment.
6. The method according to any one of claims 1 to 5, characterized in that, Based on the target audio segment and the first target language text corresponding to the target audio segment, a process of generating a simultaneous interpretation translation training data includes: The target audio segment and the first target language text corresponding to the target audio segment are combined to form the simultaneous interpretation translation training data.
7. The method according to any one of claims 1 to 5, characterized in that, Based on bilingual subtitle information in bilingual subtitle audio and video data, a process of obtaining a first source language text and a first target language text corresponding to each audio segment includes: When the bilingual subtitle audio and video data is audio and video data with a bilingual external subtitle file, timestamps are extracted from the bilingual external subtitle file, and first source language texts and first target language texts of adjacent timestamp corresponding audio segments are extracted from the audio and video data; The audio stream is divided into audio segments according to the timestamps.
8. The method according to any one of claims 1 to 5, characterized in that, Based on bilingual subtitle information in bilingual subtitle audio and video data, a process of obtaining a first source language text and a first target language text corresponding to each audio segment includes: When the bilingual subtitle audio and video data is video data with embedded bilingual subtitles, an audio stream is extracted from the video data, the audio stream is divided into audio segments by using voice activity detection (VAD), and video frames within a time length corresponding to the audio segments are intercepted from the video data; The video frames are divided into two or more subgraphs, and bilingual subtitles in each subgraph are identified to obtain source language texts and target language texts contained in each subgraph; aligning and calculating the similarity between the second source language text and the source language text of each subgraph respectively to determine a target subgraph satisfying a similarity threshold requirement; merging the source language text contained in the target subgraph to obtain the source language text contained in the video frame, and merging the target language text contained in the target subgraph to obtain the target language text contained in the video frame; performing deduplication and merging on the source language text contained in each video frame within the audio segment corresponding time length to obtain the first source language text corresponding to the audio segment, and performing deduplication and merging on the target language text contained in each video frame within the audio segment corresponding time length to obtain the first target language text corresponding to the audio segment.
9. An electronic device, comprising: comprise: a memory and a processor; the memory is configured to store a program; the processor is configured to execute the program to implement each step of the simultaneous interpretation translation training data generation method according to any one of claims 1-8.
10. A readable storage medium, having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement each step of the simultaneous interpretation translation training data generation method according to any one of claims 1-8.
11. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement each step of the simultaneous interpretation translation training data generation method according to any one of claims 1-8.
Citation Information
Patent Citations
Sample audio data acquisition method, speech recognition method and related device
CN117894300A
Data generation method and device
CN119864019A