Simultaneous interpretation method, electronic device, and computer-readable storage medium

By using the target segmentation model in simultaneous translation to determine the text segmentation position and perform segmentation translation, the translation delay and accuracy problems are solved, and faster and more accurate simultaneous translation results are achieved.

CN119811419BActive Publication Date: 2025-07-11IFLYTEK CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510304418.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-10-23
Filing Date
2025-03-14
Publication Date
2025-07-11
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

In the existing simultaneous translation methods, the translation delay is long, making it difficult to achieve real-time translation, and the translation accuracy is insufficient.

Method used

The target segmentation model is used to determine the text segmentation position of the audio to be translated, and the audio to be translated is translated in segments according to the segmentation position, including the first text segmentation position and the second text segmentation position. The translation result before the first segmentation position is accurate, and the second segmentation position is the position where the target punctuation point is located.

Benefits of technology

It improves the timeliness and accuracy of simultaneous translations, reduces translation delays, and ensures the reliability of early translation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811419B_ABST
    Figure CN119811419B_ABST
Patent Text Reader

Abstract

The present application discloses a simultaneous translation method, an electronic device, and a computer-readable storage medium. The method includes: obtaining an audio to be translated; using a target segmentation model to determine the text segmentation positions of the text to be translated corresponding to the audio to be translated; wherein the text segmentation positions include a first text segmentation position and a second text segmentation position, the first text segmentation position is the position between a first sub-text to be translated and a second sub-text to be translated in the text to be translated, and the translation influence degree of the second sub-text to be translated on the first sub-text to be translated is lower than a first influence degree threshold, and the second text segmentation position is the position where a target punctuation mark is located in the text to be translated; segmenting and translating the audio to be translated according to the text segmentation positions. By the above method, the present application can improve the timeliness of simultaneous translation and reduce the delay of simultaneous translation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 2024114864667 and the invention title "Simultaneous Translation Method, Electronic Device and Computer Readable Storage Medium" filed on October 23, 2024, which is incorporated into this application in its entirety by reference. Technical Field

[0002] This application relates to the field of translation technology, and particularly to a simultaneous translation method, an electronic device and a computer readable storage medium. Background Art

[0003] In the field of simultaneous translation, cascaded translation methods and end-to-end translation methods are mostly adopted.

[0004] In the cascaded translation method, deciding when to output the translation result can only be done by combining strategies with the punctuation of the recognition result and the translation content, and the stability is too poor. In the end-to-end translation method, currently most output the result based on the punctuation of the translation result, that is, the translation result will only be output when commas, full stops, question marks, or exclamation marks are detected. Such a method inevitably brings translation latency, and it is very difficult to achieve real-time simultaneous translation. Moreover, as the audio continues to flow in and the audio length increases, the inference speed will also be affected, resulting in an increasingly slow simultaneous translation speed. Summary of the Invention

[0005] The main technical problem to be solved by this application is to provide a simultaneous translation method, an electronic device and a computer readable storage medium, which can improve the timeliness of simultaneous translation and reduce the latency of simultaneous translation.

[0006] To solve the above technical problem, one technical solution adopted by this application is: to provide a simultaneous translation method, which includes: obtaining the audio to be translated; using a target segmentation model to determine the text segmentation positions of the text to be translated corresponding to the audio to be translated; wherein, the text segmentation positions include a first text segmentation position and a second text segmentation position, the first text segmentation position is the position between the first sub-text to be translated and the second sub-text to be translated in the text to be translated, and the translation influence degree of the second sub-text to be translated on the first sub-text to be translated is lower than the first influence degree threshold, and the second text segmentation position is the position of the target punctuation in the text to be translated; segment and translate the audio to be translated according to the text segmentation positions.

[0007] To solve the above technical problem, another technical solution adopted by this application is: to provide an electronic device, which includes a memory and a processor, the memory stores program instructions, and the processor is used to execute the program instructions to implement the above simultaneous translation method.

[0008] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium for storing program instructions that can be executed to implement the above-mentioned simultaneous translation method.

[0009] In the above technical solution, after determining the text segmentation position of the text to be translated corresponding to the audio to be translated, the subsequent sub-text to be translated or the audio segment to be translated corresponding to the position before the text segmentation position is thrown out for translation. Therefore, the simultaneous translation result can be thrown out faster, improving the timeliness of simultaneous translation and reducing the delay of simultaneous translation. In addition, due to the appearance of new content after the text segmentation position, it basically has no impact on the translation result before the text segmentation position. Therefore, the translation of the sub-text to be translated or the audio segment to be translated corresponding to the position before the text segmentation position is accurate and reliable, improving the accuracy of simultaneous translation. Description of the Drawings

[0010] Figure 1 is a schematic flowchart of an embodiment of the simultaneous translation method provided by this application;

[0011] Figure 2 is a schematic structural diagram of an embodiment of the simultaneous translation device provided by this application;

[0012] Figure 3 is a schematic structural diagram of an embodiment of the electronic device provided by this application;

[0013] Figure 4 is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by this application. Detailed Embodiments

[0014] The following will combine the accompanying drawings of the specification to elaborate on the solutions of the embodiments of this application in detail.

[0015] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand this application.

[0016] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, the term "multiple" in this article means two or more than two. In addition, the term "at least one" in this article means any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0017] Please refer toFigure 1 , Figure 1 is a schematic flowchart of an embodiment of the simultaneous interpretation method provided by this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 1 the process sequence shown. As Figure 1 shown, this embodiment includes:

[0018] Step S11: Obtain the audio to be translated.

[0019] In this embodiment, the audio to be translated is obtained. Among them, the language of the audio to be translated is not limited. For example, the language of the audio to be translated is Chinese, English, French, local dialects (such as Zhejiang dialect, Beijing dialect, Jiangsu dialect, Anhui dialect, etc.). In addition, the translation language corresponding to the audio to be translated is not limited either. For example, the translation language is Chinese, English, French, local dialects (such as Zhejiang dialect, Beijing dialect, Jiangsu dialect, Anhui dialect, etc.).

[0020] In one embodiment, the audio to be translated can specifically be obtained from local storage or cloud storage. Of course, in other embodiments, a voice acquisition device can also be used to collect audio in real time as the audio to be translated, which is not limited here.

[0021] Step S12: Use the target segmentation model to determine the text segmentation positions of the text to be translated corresponding to the audio to be translated.

[0022] In this embodiment, the target segmentation model is used to determine the text segmentation positions of the text to be translated corresponding to the audio to be translated, where the text segmentation positions include a first text segmentation position and a second text segmentation position. The first text segmentation position is the position between the first sub-text to be translated and the second sub-text to be translated in the text to be translated, and the translation influence degree of the second sub-text to be translated on the first sub-text to be translated is lower than the first influence degree threshold. The second text segmentation position is the position of the target punctuation mark in the text to be translated.

[0023] Among them, the size of the first influence degree threshold is not limited and can be specifically set according to actual usage needs.

[0024] On the one hand, the first text segmentation position of the text to be translated corresponding to the audio to be translated will be determined. The first text segmentation position is the position between the first sub-text to be translated and the second sub-text to be translated of the text to be translated, and the translation influence degree of the second sub-text to be translated on the first sub-text to be translated is lower than the first influence degree threshold. That is to say, after translating the first sub-text to be translated before the first text segmentation position, the appearance of the second sub-text to be translated after the first text segmentation position, that is, the appearance of subsequent newly added content, has a translation influence degree on the translation result of the first sub-text to be translated lower than the first influence degree threshold, that is, basically there is no need to rewrite the translation result of the first sub-text to be translated; that is, it indicates that after determining the first text segmentation position of the text to be translated corresponding to the audio to be translated, throwing the sub-text to be translated or the audio segment to be translated corresponding to before the first text segmentation position for translation, the translation result is accurate and reliable.

[0025] After determining the first text segmentation position of the text to be translated corresponding to the audio to be translated, subsequently throwing the sub-text to be translated or the audio segment to be translated corresponding to before the first text segmentation position for translation, so that the simultaneous interpretation result can be thrown out faster, improving the timeliness of simultaneous interpretation and reducing the latency of simultaneous interpretation; in addition, due to the appearance of newly added content after the first text segmentation position, there is basically no influence on the translation result before the first text segmentation position, so the translation of the sub-text to be translated or the audio segment to be translated corresponding to before the first text segmentation position is accurate and reliable, improving the accuracy of simultaneous interpretation.

[0026] For example, taking the text to be translated of the audio to be translated as "Now I'm very hungry and eat rice" as an example: The first text segmentation position is the position between the first sub-text to be translated "Now" and the second sub-text to be translated "very hungry"; the translation result of the first sub-text to be translated "Now" is "now", and the translation result of the second sub-text to be translated "very hungry" is "its very hungry", and the second sub-text to be translated "very hungry" has basically no influence on the translation result of the first sub-text to be translated "Now"; so, when the first sub-text to be translated "Now" of the text to be translated is recognized in a streaming manner, throwing the first sub-text to be translated "Now" or the corresponding audio segment to be translated for translation, and then when the second sub-text to be translated "very hungry" of the text to be translated is recognized in a streaming manner, throwing the second sub-text to be translated "very hungry" or the corresponding audio segment to be translated for translation, compared with having to throw out for translation only after recognizing the complete text to be translated "Now I'm very hungry" before the comma, it improves the timeliness of simultaneous interpretation and reduces the latency of simultaneous interpretation.

[0027] For another example, taking the text to be translated in the audio to be translated as "I watch TV" as an example: If the first text segmentation position is between the first sub-text to be translated "I watch" and the second sub-text to be translated "TV"; the translation result of the first sub-text to be translated "I watch" may be "I look", and the translation result of the second sub-text to be translated "TV" is "TV". The second sub-text to be translated "TV" will cause the translation result of the first sub-text to be translated "I watch" to be rewritten as "I watch". Therefore, the second sub-text to be translated has a greater impact on the translation of the first sub-text to be translated, and the position between the first sub-text to be translated "I watch" and the second sub-text to be translated "TV" cannot be used as the first text segmentation position.

[0028] On the other hand, the second text segmentation position of the text to be translated corresponding to the audio to be translated will be determined. The second text segmentation position is the position of the target punctuation in the text to be translated. The text to be translated before the position of the target punctuation is a semantically complete text to be translated. The appearance of the text to be translated after the position of the target punctuation, that is, the appearance of subsequent newly added content, has basically no impact on the translation result before the position of the target punctuation; that is, it indicates that after determining the second text segmentation position of the text to be translated corresponding to the audio to be translated, the sub-text to be translated or the audio segment to be translated corresponding to before the second text segmentation position is thrown out for translation, and the translation result is accurate and reliable.

[0029] After determining the second text segmentation position of the text to be translated corresponding to the audio to be translated, the sub-text to be translated or the audio segment to be translated corresponding to before the second text segmentation position is subsequently thrown out for translation. Since the appearance of the newly added content after the second text segmentation position basically has no impact on the translation result before the second text segmentation position, the translation of the sub-text to be translated or the audio segment to be translated corresponding to before the second text segmentation position is accurate and reliable, which improves the accuracy of simultaneous interpretation.

[0030] In one embodiment, before using the target segmentation model to determine the text segmentation position of the text to be translated corresponding to the audio to be translated, the target audio feature of the audio to be translated will also be obtained, where the target audio feature is used to input the target segmentation model. That is to say, after obtaining the audio to be translated, the audio to be translated will be subjected to feature extraction to obtain the target audio feature; then, the target segmentation model is used to determine the text segmentation position of the text to be translated corresponding to the audio to be translated based on the target audio feature.

[0031] In a specific embodiment, the audio to be translated will be subjected to feature extraction to obtain Fbank features; then, the Fbank features will be subjected to feature extraction to obtain hidden layer features as the target audio features.

[0032] In one embodiment, a target segmentation model is used to determine the text segmentation positions of the text to be translated corresponding to the audio to be translated. Specifically, the target segmentation model is used to identify the text to be translated corresponding to the audio to be translated and obtain the text segmentation positions of the text to be translated. That is, the target segmentation model is used to identify the audio to be translated to obtain the text to be translated and the corresponding text segmentation positions.

[0033] In a specific embodiment, if the target segmentation model is used to identify the text to be translated including segmentation identifiers, then, at this time, to obtain the text segmentation positions of the text to be translated, specifically: the positions where the segmentation identifiers are located are used as the text segmentation positions of the text to be translated.

[0034] Of course, in other specific embodiments, what the target segmentation model identifies may also include two parts. One part is the text to be translated, and the other part is the text segmentation position. That is, the text segmentation position is not directly presented in the text to be translated by segmentation identifiers.

[0035] In one embodiment, the target segmentation model is a CTC segmentation model.

[0036] Step S13: Segment and translate the audio to be translated according to the text segmentation positions.

[0037] In this embodiment, the audio to be translated is segmented and translated according to the text segmentation positions. After determining the first text segmentation position of the text to be translated corresponding to the audio to be translated, the subsequent sub-text to be translated or audio segment to be translated corresponding to before the first text segmentation position is thrown out for translation. Therefore, the simultaneous translation results can be thrown out faster, improving the timeliness of simultaneous translation and reducing the latency of simultaneous translation. In addition, due to the appearance of new content after the first text segmentation position, it basically has no impact on the translation results before the first text segmentation position. Therefore, the translation of the sub-text to be translated or audio segment to be translated corresponding to before the first text segmentation position is accurate and reliable, improving the accuracy of simultaneous translation. After determining the second text segmentation position of the text to be translated corresponding to the audio to be translated, the subsequent sub-text to be translated or audio segment to be translated corresponding to before the second text segmentation position is thrown out for translation. Since the appearance of new content after the second text segmentation position basically has no impact on the translation results before the second text segmentation position, the translation of the sub-text to be translated or audio segment to be translated corresponding to before the second text segmentation position is accurate and reliable, improving the accuracy of simultaneous translation.

[0038] Therefore, segmenting and translating the audio to be translated according to the text segmentation positions, on the one hand, can throw out the simultaneous translation results faster, improving the timeliness of simultaneous translation and reducing the latency of simultaneous translation; on the other hand, it can improve the accuracy of simultaneous translation.

[0039] It should be noted that after determining the text segmentation position of the text to be translated corresponding to the audio to be translated using the target segmentation model, the sub-text to be translated before the text segmentation position can be directly thrown out for translation to achieve segmented translation of the audio to be translated. In the case of directly throwing out the sub-text to be translated for translation, it shows that this application is a cascaded translation method, that is, a speech recognition + text translation method. Of course, after determining the text segmentation position of the text to be translated corresponding to the audio to be translated using the target segmentation model, the audio segments to be thrown out can also be determined using the text segmentation position, and the thrown-out audio segments can be translated to achieve segmented translation of the audio to be translated. In the case of throwing out audio segments for translation, it shows that this application is an end-to-end translation method, that is, directly translating based on the audio rather than the speech recognition text of the audio.

[0040] In one embodiment, before determining the text segmentation position of the text to be translated corresponding to the audio to be translated using the target segmentation model, the target audio features of the audio to be translated are also obtained. At this time, segmented translation of the audio to be translated is performed according to the text segmentation position. Specifically: using the text segmentation position, the target audio features are divided into multiple feature segments, and each feature segment is translated using a translation model respectively.

[0041] In the case of dividing the target audio features into multiple feature segments using the text segmentation position defined in this application, the granularity of the thrown-out feature segments is smaller. Therefore, the feature segments can be thrown out faster, so that the simultaneous translation results can be thrown out faster, improving the timeliness of simultaneous translation and reducing the latency of simultaneous translation. In addition, due to the appearance of the newly added content after the text segmentation position defined in this application, it basically has no impact on the translation results before the text segmentation position. Therefore, the translation of the feature segments divided based on the text segmentation position is accurate and reliable, improving the accuracy of simultaneous translation.

[0042] In a specific embodiment, using the text segmentation position to divide the target audio features into multiple feature segments is specifically: using the text segmentation position to determine the audio segmentation position of the audio to be translated; using the audio segmentation position to divide the target audio features into multiple feature segments. That is to say, using the text segmentation position to determine the audio segmentation position of the audio to be translated, and using the audio segmentation position to segment the target audio features of the audio to be translated, so as to obtain multiple feature segments. Subsequently, direct translation based on the audio feature segments is an end-to-end translation, which can avoid some errors caused by speech recognition and improve the accuracy of simultaneous translation.

[0043] In a specific embodiment, the audio segmentation position includes the audio segmentation moment; using the text segmentation position to determine the audio segmentation position of the audio to be translated specifically includes: obtaining the target character in the text to be translated, where the target character is the character adjacent to and before the text segmentation position; and obtaining the target phoneme sequence of the audio to be translated; determining the target boundary of the target character in the target phoneme sequence; and obtaining the moment corresponding to the target boundary as the audio segmentation moment.

[0044] For example, taking the text to be translated of the audio to be translated as "Now I'm very hungry and want to eat rice." as an example: The text segmentation position a is between the sub-text to be translated "Now" and the sub-text to be translated "very hungry", so the character "Now" adjacent to and before the text segmentation position a is the target character α; the moment corresponding to the target boundary of the target character α in the target phoneme sequence is t1; therefore, the moment t1 is used as an audio segmentation moment. The text segmentation position b is the position where the target punctuation "," is located, so the character "hungry" adjacent to and before the text segmentation position b is the target character β; the moment corresponding to the target boundary of the target character β in the target phoneme sequence is t2; therefore, the moment t2 is used as an audio segmentation moment. The text segmentation position c is between the sub-text to be translated "eat" and the sub-text to be translated "rice", so the character "eat" adjacent to and before the text segmentation position c is the target character γ; the moment corresponding to the target boundary of the target character γ in the target phoneme sequence is t3; therefore, the moment t3 is used as an audio segmentation moment. The text segmentation position d is the position where the target punctuation "." is located, so the character "rice" adjacent to and before the text segmentation position d is the target character δ; the moment corresponding to the target boundary of the target character δ in the target phoneme sequence is t4; therefore, the moment t4 is used as an audio segmentation moment.

[0045] Therefore, when the audio to be translated is input in a streaming manner subsequently, the feature segment corresponding to the moment 0 - t1 will be cut out for translation, the feature segment corresponding to the moment t1 - t2 will be cut out for translation, the feature segment corresponding to the moment t2 - t3 will be cut out for translation, and the feature segment corresponding to the moment t3 - t4 will be cut out for translation.

[0046] In a specific embodiment, determining the target boundary of the target character in the target phoneme sequence specifically includes: matching the sub-phoneme sequence of the target character from the target phoneme sequence, and taking the boundary of the sub-phoneme sequence as the target boundary. That is to say, first match the phoneme sequence belonging to the target character, and take the boundary of the phoneme sequence belonging to the target character as the target boundary.

[0047] In one embodiment, translating each feature segment of the target audio feature is performed using a translation model, and determining the text segmentation position of the text to be translated corresponding to the audio to be translated is performed using a target segmentation model. The training steps of the translation model and the target segmentation model are specifically as follows: Feature extraction is performed on the sample audio to obtain sample audio features; the target segmentation model and the translation model are respectively trained using the sample audio features. That is to say, the training of the translation model and the target segmentation model shares the encoder.

[0048] Specifically, in the input layer, feature extraction is performed on the sample audio to obtain Fbank features; then the Fbank features are sent to the downsampling module and the conformer neural network to perform downsampling and encoding / decoding on the Fbank features, and a frame-level phoneme output, which is the hidden layer feature, is obtained. Then, the hidden layer feature of the phoneme output is sent to the target segmentation branch for training the target segmentation model. Specifically: The hidden layer feature of the phoneme output passes through a shallow convolutional neural network combined with softmax to obtain the output of the target segmentation model; and the hidden layer feature of the phoneme output is sent to the translation branch to train the translation model.

[0049] It should be noted that when training the target segmentation model and the translation model, the sample audio is segmented into sample segments according to the corresponding sample segmentation position, and after the sample segments are concatenated in a streaming manner, they are sent to the target segmentation model and the translation model for training. That is to say, one sentence will be converted into multiple sentences, and multiple segmentation results and translation results will be obtained.

[0050] In one embodiment, the target segmentation model is trained using the sample audio and the corresponding sample audio segmentation label. The steps for obtaining the sample audio segmentation label are specifically as follows: Obtain the sample audio text of the sample audio; respectively label the first sample segmentation position and the second sample segmentation position in the sample audio text to obtain the sample audio segmentation label of the sample audio; wherein, the first sample segmentation position is the position between the first sub-sample text and the second sub-sample text in the sample audio text, and the translation influence degree of the second sub-sample text on the first sub-sample text is lower than the second influence degree threshold, and the second sample segmentation position is the position where the sample punctuation in the sample audio text is located. Among them, the size of the second influence degree threshold is not limited and can be specifically set according to actual usage needs.

[0051] That is to say, the sample audio segmentation label used to train the target segmentation model is not constructed by manual annotation; compared with the method of constructing by manual annotation, the method for constructing the sample audio segmentation label provided in this application can, on the one hand, improve the accuracy of the constructed sample audio segmentation label; on the other hand, it can improve the efficiency of constructing the sample audio segmentation label and avoid excessive dependence on manual labor.

[0052] In addition, the sample audio segmentation label is obtained by annotating the first sample segmentation position and the second sample segmentation position in the sample audio text. The first sample segmentation position is the position between the first sub-sample text and the second sub-sample text in the sample audio text, and the translation influence degree of the second sub-sample text on the first sub-sample text is lower than the second influence degree threshold. The second sample segmentation position is the position where the sample punctuation in the sample audio text is located. Therefore, by using the sample audio and the corresponding sample audio segmentation label to train the target segmentation model, the trained and converged target segmentation model can determine the text segmentation position of the text to be translated corresponding to the audio to be translated. Subsequently, the sub-text to be translated or the audio segment to be translated corresponding to the position before the text segmentation position is thrown out for translation. Therefore, the simultaneous translation result can be thrown out faster, improving the timeliness of simultaneous translation and reducing the latency of simultaneous translation. In addition, due to the appearance of the newly added content after the text segmentation position, it basically has no impact on the translation result before the text segmentation position. Therefore, the translation of the sub-text to be translated or the audio segment to be translated corresponding to the position before the text segmentation position is accurate and reliable, improving the accuracy of simultaneous translation.

[0053] In a specific embodiment, the sample audio segmentation label is the first sample label text, and the first sample label text is formed by annotating the sample identifier in the sample audio text. The sample identifier includes the first identifier added at the first sample segmentation position and the second identifier replacing the sample punctuation at the first sample segmentation position.

[0054] Among them, the specific presentation forms of the first identifier and the second identifier are not limited. For example, the first identifier is " <c>”, and the second identifier is " <c> ”、" <ju> ”、" <wen>"or" <gan>” etc.

[0055] For example, taking the sample audio text of the sample audio as "Now extremely hungry" and the first identifier as " <c>"For example: The translation result of the first sub-sample text "now" is "now", and the translation result of the second sub-sample text "extremely hungry" is "its very hungry". The translation result of the second sub-sample text "extremely hungry" has basically no impact on the translation result of the first sub-sample text "now". Therefore, the position between the first sub-sample text "now" and the second sub-sample text "extremely hungry" is the first sample segmentation position. Thus, mark the first identifier at the position between the first sub-sample text "now" and the second sub-sample text "extremely hungry" <c>”, obtain the first sample label text "Now" <c>Very hungry."

[0056] In a specific implementation, the sample punctuation includes several punctuation types, and at least two punctuation types have different corresponding second identifiers. The presentation form of the second identifier corresponding to each punctuation type is not limited and can be specifically set according to actual use needs.

[0057] For example, sample punctuation marks include commas, periods, question marks, and exclamation marks; the second identifier corresponding to the comma is " <dou>”, and the second identifier corresponding to the full stop is " <ju>”, and the second identifier corresponding to the question mark is " <wen>”, the second identifier corresponding to the exclamation mark is " <gan>”.

[0058] In a specific embodiment, the sample punctuation includes a comma, and the second identifier corresponding to the comma and the first identifier are the same identifier.

[0059] For example, the first identifier is " <c>”, the second identifier corresponding to the comma is " <c>”.

[0060] For example, taking the sample audio text of the sample audio as "Then I was extremely hungry and ate rice.", the first identifier and the second identifier corresponding to the comma are " <c>”, and the second identifier corresponding to the full stop is " <ju>" is taken as an example: the translation result of the first sub-sample text "然后" is "afterwards", and the translation result of the second sub-sample text "特别饿" is "its veryhungry". The second sub-sample text "特别饿" has basically no effect on the translation result of the first sub-sample text "然后"; therefore, the position between the first sub-sample text "然后" and the second sub-sample text "特别饿" is the first sample segmentation position; therefore, the first identifier " <c>”. The translation result of the first sub-sample text "ate" is "eat", and the translation result of the second sub-sample text "rice" is "rice". The second sub-sample text "rice" has basically no influence on the translation result of the first sub-sample text "ate". Therefore, the position between the first sub-sample text "ate" and the second sub-sample text "rice" is the first sample segmentation position. Thus, the second identifier corresponding to the comma is marked at the position between the first sub-sample text "ate" and the second sub-sample text "rice" <c>”。At the position of the full stop in the sample audio text of the sample audio "Then I was extremely hungry and ate rice.", mark the second identifier corresponding to the full stop" <ju>”. Therefore, the sample audio segmentation label corresponding to the sample audio text of the sample audio "Then I was extremely hungry and ate rice." is "Then" <c>Especially hungry <c>Ate <c>Rice <ju>”。

[0061] In a specific embodiment, the first sample segmentation position of the labeled sample audio text is performed by a segmentation large model. The training steps of the segmentation large model include: using a number of first large models to respectively perform reference segmentation position annotation on the reference text to obtain first segmentation labels corresponding to each first large model; wherein, the reference segmentation position is the position between the first sub-reference text and the second sub-reference text in the reference text, and the translation influence degree of the second sub-reference text on the first sub-reference text is less than a third influence degree threshold; based on the first segmentation labels corresponding to each first large model, select at least one first large model with annotation quality meeting the quality requirements from the number of first large models as the second large model; fuse the first segmentation labels corresponding to each second large model to obtain a second segmentation label as the reference segmentation label of the reference text; use the reference text and the corresponding reference segmentation label to train the segmentation large model. Among them, the size of the third influence degree threshold is not limited and can be specifically set according to actual usage needs.

[0062] That is to say, use a number of currently open-source large models to perform reference segmentation position annotation on the reference text to obtain the first segmentation labels obtained by each large model's annotation; then, fuse the first segmentation labels of the large models with high annotation quality as the reference segmentation label of the reference text. The reference segmentation label of the reference text is obtained by fusing the first segmentation labels of the large models with high annotation quality, so the accuracy of the reference segmentation label of the reference text is high. Further, use the reference text and the corresponding reference segmentation label to train the segmentation large model, so that the trained and converged segmentation large model can accurately annotate the first sample segmentation position of the sample audio text of the sample audio.

[0063] In a specific embodiment, fusing the first segmentation labels corresponding to each second large model to obtain a second segmentation label includes: extracting the first common reference segmentation position from the first segmentation labels corresponding to each second large model and using the first common reference segmentation position as the second segmentation label. That is to say, extract the common reference segmentation position among the reference segmentation positions annotated by each large model with high annotation quality as the second segmentation label. The reference segmentation position annotated by each large model with high annotation quality must be accurate; using the reference segmentation position annotated by each large model with high annotation quality as the second segmentation label can improve the accuracy of the second segmentation label.

[0064] In a specific embodiment, the first segmentation label and the second segmentation label are the first label text, and the first label text is formed by adding a first identifier at the reference segmentation position corresponding to the reference text. That is to say, the first segmentation label and the second segmentation label are texts including the reference text and the first identifier formed by annotating the first identifier in the reference text.

[0065] In a specific embodiment, based on the first segmentation labels corresponding to each first large model, at least one first large model with annotation quality meeting the quality requirements is selected from several first large models as the second large model. Specifically: based on the first segmentation labels corresponding to each first large model, the annotation accuracy of each first large model is determined as the annotation quality of each first large model; from several first large models, the first large models with annotation quality meeting the quality requirements are selected as the second large model. That is to say, the annotation quality of each first large model is determined by the annotation accuracy of each first large model.

[0066] In a specific embodiment, the first large model meeting the quality requirements is: the first large model with annotation accuracy greater than that of a preset number of other large models. That is to say, the large models with top-ranked annotation accuracy are used as large models with high annotation quality.

[0067] Among them, the preset number is not limited. For example, the preset number is 2, 3, 5, etc. For example, taking the first large models ranked first and second in annotation accuracy as the second large model: by fusing the first segmentation labels of the first large models ranked first and second in annotation accuracy, the reference segmentation label of the reference text is obtained; therefore, the accuracy of the reference segmentation label of the reference text is high. Further, using the reference text and the corresponding reference segmentation label to train the segmentation large model, so that the trained and converged segmentation large model can accurately label the first sample segmentation position of the sample audio text of the sample audio.

[0068] In a specific embodiment, the segmentation large model is the first large model with the highest annotation quality. That is to say, the first large model with the highest annotation quality is used as the segmentation large model, and it is fine-tuned using the reference text and the corresponding reference segmentation label to make the segmentation large model converge, so that the converged segmentation large model can accurately label the first sample segmentation position of the sample audio text of the sample audio.

[0069] In a specific embodiment, after training the segmentation large model using the reference text and the corresponding reference segmentation label, the segmentation large model is used to label the reference segmentation position of the reference text to obtain a predicted segmentation label; the predicted segmentation label and the reference segmentation label are fused to obtain a new segmentation label; the new segmentation label is used as the new reference segmentation label, and the steps of training the segmentation large model using the reference text and the corresponding reference segmentation label and subsequent steps are repeated until the training converges.

[0070] That is to say, the reference text is annotated using the split large model after one fine-tuning to obtain predicted split labels; then, the reference split labels and the predicted split labels of the reference text are fused to serve as the new reference split labels of the reference text. The new reference split labels of the reference text are obtained by fusing the original reference split labels of the reference text and the predicted split labels of the split large model after one fine-tuning. Therefore, the new reference split labels of the reference text are highly accurate. Further, using the reference text and the new reference split labels to train the split large model enables the trained and converged split large model to accurately annotate the first sample split position of the sample audio text of the sample audio.

[0071] In a specific embodiment, the predicted split labels and the reference split labels are fused to obtain new split labels, specifically: from the predicted split labels and the reference split labels, the second common reference split positions are extracted, and the second common reference split positions are used as the new split labels. That is to say, the common reference split positions in the predicted split labels and the original reference split labels of the reference text are extracted as the new reference split labels. The reference split positions that exist in both the reference split labels and the predicted split labels of the reference text must be accurate; using the reference split positions that exist in both the reference split labels and the predicted split labels of the reference text as the new reference split labels of the reference text can improve the accuracy of the obtained reference split labels.

[0072] In a specific embodiment, the steps for obtaining the reference text are specifically: obtaining a number of sample texts; screening out the sample texts with correct reference punctuation from the number of sample texts as the reference text. That is to say, a large number of sample texts will be obtained, and the sample texts with correct punctuation will be used as the reference text.

[0073] In a specific embodiment, screening out the sample texts with correct reference punctuation from the number of sample texts as the reference text is specifically: for each sample text, accuracy scores are given to each reference punctuation in the sample text to obtain the accuracy scores of each reference punctuation in the sample text; the accuracy scores of each reference punctuation in the sample text are summed to obtain a total score; in response to the total score being greater than the first threshold, the sample text is used as the reference text. Among them, the size of the first threshold is not limited and can be specifically set according to actual usage needs. That is to say, the sample texts are screened by scoring the accuracy of the punctuation of each sample text.

[0074] In a specific embodiment, sample texts with correct reference punctuation are selected from a number of sample texts as reference texts. Specifically, for each sample text, a number of predicted punctuations of the sample text are generated using a punctuation generation model; the differences between the number of predicted punctuations and a number of reference punctuations in the sample text are compared; in response to the degree of difference between the number of predicted punctuations and the number of reference punctuations in the sample text being less than a second threshold, the sample text is used as a reference text. Among them, the size of the second threshold is not limited and can be specifically set according to actual usage needs. That is to say, the sample texts are screened by simulating the generation of punctuation marks using a punctuation generation model.

[0075] In an embodiment, the sample translation labels corresponding to the sample audio are generated using the sample audio segmentation labels. The steps for obtaining the sample translation labels are specifically as follows: obtaining the sub-sample texts between each sample segmentation position; translating each sub-sample text to obtain the sample translation results of each sub-sample text; for each sample segmentation position, combining the sample translation results of the sub-sample texts before the sample segmentation position to obtain the sample translation result corresponding to the sample segmentation position; among them, the sample translation results corresponding to each sample segmentation position constitute the sample translation labels.

[0076] That is to say, the sample translation labels corresponding to the sample audio are not constructed by manual annotation; the method for constructing the sample translation labels provided in this application can, on the one hand, improve the accuracy of the constructed sample translation labels compared with the method constructed by manual annotation; on the other hand, it can improve the efficiency of constructing the sample translation labels and avoid over-reliance on manual labor.

[0077] In addition, the sample translation results corresponding to each sample segmentation position are obtained by combining the sample translation results of the sub-sample texts before the sample segmentation position, and are incremental translation results obtained by fixing the historical translation results at the same time. Further, the translation model is trained using the sample audio and the corresponding sample translation labels, so that the translation model can learn to translate the sample audio segments in combination with the historical translation results to improve the translation performance.

[0078] For example, taking the sample audio text of the sample audio as "Then I'm extremely hungry." and its corresponding sample audio segmentation label as "Then" <c>Very hungry <ju>” For example, for the sample segmentation position" <c>”, and its corresponding sub-sample text is "then", and the sample translation result of the sub-sample text "then" is "afterwards"; therefore, the sample segmentation position" <c>The corresponding sample translation result is "afterwards", and the sample translation result "afterwards" constitutes a sample translation tag. For the sample segmentation position <ju>", the corresponding sub-sample text is "特别饿", and the sample translation result corresponding to the sub-sample text "特别饿" is "its very hungry"; the sample segmentation position is " <ju>The sample translation result corresponding to " is obtained by combining the sample translation result "afterwards" of the sub-sample text "然后" and the translation result "its veryhungry" of the sub-sample text "特别饿". That is, the sample segmentation position " <ju>”The corresponding sample translation result is "afterwards its veryhungry”. The sample translation result "afterwards its very hungry” constitutes a sample translation tag.

[0079] Please refer to Figure 2 , Figure 2 FIG. Figure 2 is a schematic structural diagram of an embodiment of the simultaneous translation device provided by the present application. The simultaneous translation device 20 includes an acquisition module 21, a determination module 22, and a translation module 23. The acquisition module 21 is configured to acquire an audio to be translated. The determination module 22 is configured to determine a text segmentation position of the text to be translated corresponding to the audio to be translated by using a target segmentation model. Wherein, the text segmentation position includes a first text segmentation position and a second text segmentation position. The first text segmentation position is the position between a first sub-text to be translated and a second sub-text to be translated in the text to be translated, and the translation influence degree of the second sub-text to be translated on the first sub-text to be translated is lower than a first influence degree threshold. The second text segmentation position is the position of a target punctuation mark in the text to be translated. The translation module 23 is configured to perform segmented translation on the audio to be translated according to the text segmentation position.

[0080] Wherein, the acquisition module 21 is further configured to acquire a target audio feature of the audio to be translated before determining the text segmentation position of the text to be translated corresponding to the audio to be translated by using the target segmentation model. Wherein, the target audio feature is used for inputting into the target segmentation model; and / or, the determination module 22 is configured to determine the text segmentation position of the text to be translated corresponding to the audio to be translated by using the target segmentation model, including: identifying the text to be translated corresponding to the audio to be translated by using the target segmentation model, and obtaining the text segmentation position of the text to be translated.

[0081] Wherein, the acquisition module 21 is configured to acquire a target audio feature of the audio to be translated after determining the text segmentation position of the text to be translated corresponding to the audio to be translated. The translation module 23 is configured to perform segmented translation on the audio to be translated according to the text segmentation position, including: determining an audio segmentation position of the audio to be translated by using the text segmentation position; dividing the target audio feature into multiple feature segments by using the audio segmentation position; and respectively translating each feature segment by using a translation model.

[0082] Wherein, the above audio segmentation position includes an audio segmentation moment. The translation module 23 is configured to determine the audio segmentation position of the audio to be translated by using the text segmentation position, including: acquiring a target character in the text to be translated. Wherein, the target character is a character located before and adjacent to the text segmentation position; and, acquiring a target phoneme sequence of the audio to be translated; determining a target boundary of the target character in the target phoneme sequence; and acquiring the moment corresponding to the target boundary as the audio segmentation moment.

[0083] Among them, the above-mentioned target segmentation model is trained using sample audio and corresponding sample audio segmentation labels. The simultaneous translation device 20 further includes a tagging module 24. The tagging module 24 is used for the step of obtaining sample audio segmentation labels, including: obtaining the sample audio text of the sample audio; respectively marking the first sample segmentation position and the second sample segmentation position in the sample audio text to obtain the sample audio segmentation label of the sample audio; among them, the first sample segmentation position is the position between the first sub-sample text and the second sub-sample text in the sample audio text, and the translation influence degree of the second sub-sample text on the first sub-sample text is lower than the second influence degree threshold, and the second sample segmentation position is the position where the sample punctuation in the sample audio text is located.

[0084] Among them, the above-mentioned sample audio segmentation label is the first sample label text, and the first sample label text is formed by setting sample identifiers in the sample audio text. The sample identifiers include the first identifier added at the first sample segmentation position and the second identifier replacing the sample punctuation at the second sample segmentation position.

[0085] Among them, the marking of the first sample segmentation position of the sample audio text is performed by a segmentation large model. The simultaneous translation device 20 further includes a training module 25. The training module 25 is used for the training of the segmentation large model, including: using a number of first large models to respectively perform reference segmentation position marking on the reference text to obtain first segmentation labels corresponding to each first large model; among them, the reference segmentation position is the position between the first sub-reference text and the second sub-reference text in the reference text, and the translation influence degree of the second sub-reference text on the first sub-reference text is less than the third influence degree threshold; based on the first segmentation labels corresponding to each first large model, select at least one first large model whose marking quality meets the quality requirements from the number of first large models as the second large model; fuse the first segmentation labels corresponding to each second large model to obtain a second segmentation label as the reference segmentation label of the reference text; use the reference text and the corresponding reference segmentation label to train the segmentation large model.

[0086] Among them, the above-mentioned first segmentation label and second segmentation label are the first label text, and the first label text is formed by adding a first identifier at the reference segmentation position marked correspondingly in the reference text; and / or, the training module 25 is used for fusing the first segmentation labels corresponding to each second large model to obtain a second segmentation label, including: extracting the first common reference segmentation position from the first segmentation labels corresponding to each second large model and using the first common reference segmentation position as the second segmentation label.

[0087] Among them, the training module 25 is used to select at least one first large model with annotation quality meeting the quality requirements from several first large models based on the first segmentation labels corresponding to each first large model as the second large model, including: determining the annotation accuracy of each first large model based on the first segmentation labels corresponding to each first large model as the annotation quality of each first large model; selecting the first large model with annotation quality meeting the quality requirements from several first large models as the second large model; where the first large model meeting the quality requirements is: the first large model with annotation accuracy greater than the annotation accuracy of a preset number of other large models; and / or, the segmentation large model is the first large model with the highest annotation quality.

[0088] Among them, the training module 25 is also used for obtaining reference texts, including: obtaining several sample texts; for each sample text, scoring the accuracy of each reference punctuation in the sample text to obtain the accuracy scores of each reference punctuation in the sample text; summing up the accuracy scores of each reference punctuation in the sample text to obtain the total score; in response to the total score being greater than the first threshold, taking the sample text as the reference text; or, for each sample text, using a punctuation generation model to generate several predicted punctuations of the sample text; comparing the differences between the several predicted punctuations and the several reference punctuations in the sample text; in response to the degree of difference between the several predicted punctuations and the several reference punctuations in the sample text being less than the second threshold, taking the sample text as the reference text.

[0089] Among them, after training the segmentation large model using the reference text and the corresponding reference segmentation labels, the training module 25 is used to perform reference segmentation position annotation on the reference text using the segmentation large model to obtain predicted segmentation labels; fusing the predicted segmentation labels and the reference segmentation labels to obtain new segmentation labels; taking the new segmentation labels as the new reference segmentation labels, and re-performing the steps of training the segmentation large model using the reference text and the corresponding reference segmentation labels and subsequent steps until the training converges.

[0090] Please refer to Figure 3 , Figure 3 FIG. is a schematic structural diagram of an embodiment of an electronic device provided by this application. The electronic device 30 includes a memory 31 and a processor 32 that are coupled to each other. The processor 32 is used to execute program instructions stored in the memory 31 to implement the steps of any of the above simultaneous translation method embodiments. In a specific implementation scenario, the electronic device 30 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 30 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited here.

[0091] Specifically, the processor 32 is used to control itself and the memory 31 to implement the steps of any of the above simultaneous translation method embodiments. The processor 32 can also be referred to as a CPU (Central Processing Unit). The processor 32 may be an integrated circuit chip with signal processing capabilities. The processor 32 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Additionally, the processor 32 can be implemented jointly by integrated circuit chips.

[0092] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by the present application. The computer-readable storage medium 40 of the embodiment of the present application stores program instructions 41, and when the program instructions 41 are executed, the methods provided by any of the simultaneous translation method embodiments of the present application and any non-conflicting combinations are implemented. Among them, the program instructions 41 can form a program file and be stored in the above computer-readable storage medium 40 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) can execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned computer-readable storage medium 40 includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, or a tablet.

[0093] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0094] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.< / ju> < / ju> < / ju> < / c> < / c> < / ju> < / c> < / ju> < / c> < / c> < / c> < / ju> < / c> < / c> < / ju> < / c> < / c> < / c> < / gan> < / wen> < / ju> < / dou> < / c> < / c> < / c> < / gan> < / wen> < / ju> < / c> < / c>

Claims

1. A simultaneous interpretation method, characterized in that, The method includes: Obtain the audio to be translated; Use the target segmentation model to determine the text segmentation positions of the text to be translated corresponding to the audio to be translated; wherein, the text segmentation positions include a first text segmentation position and a second text segmentation position, the first text segmentation position is the position between a first sub-text to be translated and a second sub-text to be translated in the text to be translated, and the translation influence degree of the second sub-text to be translated on the first sub-text to be translated is lower than a first influence degree threshold, and the second text segmentation position is the position of the target punctuation in the text to be translated; Segment and translate the audio to be translated according to the text segmentation positions; Wherein, the target segmentation model is trained using sample audio and corresponding sample audio segmentation labels, and the sample audio segmentation labels are obtained by labeling a first sample segmentation position and a second sample segmentation position of the sample audio text of the sample audio. The first sample segmentation position of the sample audio text is labeled by a segmentation large model, and the training steps of the segmentation large model include: Use a number of first large models to respectively label the reference segmentation positions of the reference text to obtain first segmentation labels corresponding to each of the first large models; wherein, the reference segmentation position is the position between a first sub-reference text and a second sub-reference text in the reference text, and the translation influence degree of the second sub-reference text on the first sub-reference text is less than a third influence degree threshold; Based on the first segmentation labels corresponding to each of the first large models, select at least one first large model with labeling quality meeting the quality requirements from the number of first large models as a second large model; Fuse the first segmentation labels corresponding to each of the second large models to obtain a second segmentation label as the reference segmentation label of the reference text; Use the reference text and the corresponding reference segmentation label to train the segmentation large model; Use the trained segmentation large model to label the reference segmentation positions of the reference text to obtain predicted segmentation labels; Fuse the predicted segmentation labels and the reference segmentation labels to obtain new segmentation labels; Use the new segmentation labels as the new reference segmentation labels, and re-execute the steps of using the reference text and the corresponding reference segmentation labels to train the segmentation large model and subsequent steps until the training converges.

2. The method according to claim 1, wherein Before using the target segmentation model to determine the text segmentation positions of the text to be translated corresponding to the audio to be translated, the method further includes: Obtain the target audio features of the audio to be translated; wherein, the target audio features are used to input the target segmentation model; And / or, using the target segmentation model to determine the text segmentation positions of the text to be translated corresponding to the audio to be translated includes: Use the target segmentation model to identify the text to be translated corresponding to the audio to be translated and obtain the text segmentation positions of the text to be translated.

3. The method according to claim 1, wherein After using the target segmentation model to determine the text segmentation positions of the text to be translated corresponding to the audio to be translated, the method further includes: Obtain the target audio features of the audio to be translated; Performing segmented translation on the audio to be translated according to the text segmentation position includes: Determining the audio segmentation position of the audio to be translated by using the text segmentation position; Dividing the target audio features into multiple feature segments by using the audio segmentation position; Translating each of the feature segments by using a translation model.

4. The method according to claim 3, wherein The audio segmentation position includes an audio segmentation moment; and determining the audio segmentation position of the audio to be translated by using the text segmentation position includes: Obtaining a target character in the text to be translated; wherein the target character is a character located before and adjacent to the text segmentation position; And obtaining a target phoneme sequence of the audio to be translated; Determining a target boundary of the target character in the target phoneme sequence; Obtaining the moment corresponding to the target boundary as the audio segmentation moment.

5. The method according to claim 1, wherein The step of obtaining the sample audio segmentation label includes: Obtaining a sample audio text of the sample audio; Respectively labeling a first sample segmentation position and a second sample segmentation position in the sample audio text to obtain the sample audio segmentation label of the sample audio; wherein the first sample segmentation position is the position between a first sub-sample text and a second sub-sample text in the sample audio text, and the influence degree of the second sub-sample text on the translation of the first sub-sample text is lower than a second influence degree threshold, and the second sample segmentation position is the position where a sample punctuation mark is located in the sample audio text.

6. The method according to claim 5, wherein The sample audio segmentation label is a first sample label text, and the first sample label text is formed by setting sample identifiers in the sample audio text, and the sample identifiers include a first identifier added at the first sample segmentation position and a second identifier replacing the sample punctuation mark at the second sample segmentation position.

7. The method according to claim 1, characterized in that, The first segmentation label and the second segmentation label are a first label text, and the first label text is formed by adding a first identifier at a reference segmentation position correspondingly marked in the reference text; And / or, fusing the first segmentation labels corresponding to each of the second large models to obtain a second segmentation label, including: Extracting a first common reference segmentation position from the first segmentation labels corresponding to each of the second large models, and using the first common reference segmentation position as the second segmentation label.

8. The method according to claim 1, wherein Selecting at least one first large model with a labeling quality meeting the quality requirement from the several first large models based on the first segmentation labels corresponding to each of the first large models as the second large model, including: Determining the labeling accuracy of each of the first large models based on the first segmentation labels corresponding to each of the first large models as the labeling quality of each of the first large models; Selecting a first large model with a labeling quality meeting the quality requirement from the several first large models as the second large model; Among them, the first large model that meets the quality requirement is: the first large model whose annotation accuracy is greater than that of a preset number of other large models; and / or, the segmentation large model is the first large model with the highest annotation quality.

9. The method according to claim 1, wherein The steps for obtaining the reference text include: Obtain a number of sample texts; For each of the sample texts, score the accuracy of each reference punctuation in the sample text to obtain the accuracy scores of each reference punctuation in the sample text; Sum up the accuracy scores of each reference punctuation in the sample text to obtain a total score; In response to the total score being greater than the first threshold, use the sample text as the reference text; Alternatively, for each of the sample texts, use a punctuation generation model to generate a number of predicted punctuations for the sample text; Compare the differences between the number of predicted punctuations and the number of reference punctuations in the sample text; In response to the degree of difference between the number of predicted punctuations and the number of reference punctuations in the sample text being less than the second threshold, use the sample text as the reference text.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores program instructions, and the processor is configured to execute the program instructions to implement the method according to any one of claims 1-9.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program instructions, and the program instructions can be executed to implement the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Speech translation method and device, computer equipment and storage medium

    CN111310481A

  • Speech translation method and device and storage medium

    CN113591495A

  • Voice evaluation method and device, equipment and storage medium

    CN117542346A

  • Data labeling method, electronic equipment and computer readable storage medium

    CN118013273A

  • System and method for translating real-time speech using segmentation based on conjunction locations

    US20150134320A1