Simultaneous Interpretation Quality Evaluation Method and Related Devices, Equipment and Storage Media
By dividing the simultaneous text into subtexts and analyzing the brushing data of the sub-voice, and integrating the scores to evaluate the quality of simultaneous transmission, the problem of evaluation accuracy in streaming simultaneous transmission is solved, and higher evaluation accuracy is achieved.
Patent Information
- Application Number
- CN202411858505.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-12-17
AI Technical Summary
It is difficult for the prior art to accurately evaluate the quality of simultaneous transmission in streaming simultaneous transmission application scenarios, especially in large-scale model simultaneous transmission.
By dividing the simultaneous text into several subtexts, and obtaining the sub-voice brushing data corresponding to each subtext, analyzing the simultaneous quality score of the sub-voice, and fusing the quality scores of each sub-voice to obtain the simultaneous quality score of the target voice.
The evaluation granularity is refined, the accuracy of simultaneous transmission quality evaluation is improved, and the simultaneous transmission quality can be effectively measured in the streaming simultaneous transmission scenario.
Smart Images

Figure CN119312818B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and particularly to a simultaneous interpretation quality evaluation method and related devices, equipment, and storage media. Background Art
[0002] Benefiting from the continuous development of machine learning, applying large models to simultaneous interpretation has achieved considerable progress. Different from traditional simultaneous interpretation, large model simultaneous interpretation requires directly generating corresponding translation results in a streaming manner while the user inputs speech, without going through recognition and then translation. Therefore, it is particularly important to evaluate the quality of simultaneous interpretation end-to-end.
[0003] However, existing evaluation methods are usually applicable to traditional simultaneous interpretation. If applied to large model simultaneous interpretation, it will be difficult to truly reflect the quality of simultaneous interpretation. In view of this, how to improve the accuracy of simultaneous interpretation quality evaluation in the application scenario of streaming simultaneous interpretation has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem to be solved by this application is to provide a simultaneous interpretation quality evaluation method and related devices, equipment, and storage media, which can improve the accuracy of simultaneous interpretation quality evaluation in the application scenario of streaming simultaneous interpretation.
[0005] To solve the above technical problem, a first aspect of this application provides a simultaneous interpretation quality evaluation method, including: segmenting the simultaneous interpretation text of the target speech to obtain a number of sub-texts; obtaining the typing data of the sub-speech corresponding to the sub-text in the target speech; wherein, the typing data of the sub-speech includes: a number of texts of the sub-speech from the first appearance of words to gradual correction until finally translated into the sub-text during simultaneous interpretation; analyzing to obtain the simultaneous interpretation quality score of the sub-speech based on the typing data of the sub-speech; and fusing to obtain a target score representing the simultaneous interpretation quality of the target speech based on the simultaneous interpretation quality scores of each sub-speech.
[0006] To solve the above technical problem, a second aspect of this application provides a simultaneous interpretation quality evaluation device, including: a text segmentation module, a data acquisition module, a quality analysis module, and a score fusion module. The text segmentation module is used to segment the simultaneous interpretation text of the target speech to obtain a number of sub-texts; the data acquisition module is used to obtain the typing data of the sub-speech corresponding to the sub-text in the target speech; wherein, the typing data of the sub-speech includes: a number of texts of the sub-speech from the first appearance of words to gradual correction until finally translated into the sub-text during simultaneous interpretation; the quality analysis module is used to analyze to obtain the simultaneous interpretation quality score of the sub-speech based on the typing data of the sub-speech; and the score fusion module is used to fuse to obtain a target score representing the simultaneous interpretation quality of the target speech based on the simultaneous interpretation quality scores of each sub-speech.
[0007] To solve the above technical problems, a third aspect of the present application provides an electronic device, which at least includes a memory and a processor coupled to each other. The memory stores at least program instructions, and the processor is configured to execute the program instructions to implement the simultaneous interpretation quality evaluation method in the first aspect above.
[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor. The program instructions are used to implement the simultaneous interpretation quality evaluation method in the first aspect above.
[0009] In the above solution, the simultaneous interpretation text of the target speech is segmented to obtain several sub-texts, and the brushing data of the sub-speech corresponding to the sub-text in the target speech is obtained. The brushing data of the sub-speech includes: several texts of the sub-speech from the first appearance of words to gradual correction until finally translated into the sub-text during the simultaneous interpretation process. Then, based on the brushing data of the sub-speech, the simultaneous interpretation quality score of the sub-speech is analyzed, and further based on the simultaneous interpretation quality scores of each sub-speech, a target score representing the simultaneous interpretation quality of the target speech is fused. Therefore, on the one hand, by dividing the simultaneous interpretation text into several sub-texts, and then separately evaluating the simultaneous interpretation quality of each sub-text corresponding to the sub-speech and finally fusing the scores, compared with the overall quality evaluation of the target speech and its simultaneous interpretation text, the evaluation granularity can be further refined, which helps to improve the accuracy of the simultaneous interpretation quality evaluation to a certain extent. On the other hand, when evaluating each sub-speech, since the brushing data of the sub-speech is combined, and the brushing data includes several texts of the sub-speech from the first appearance of words to gradual correction until finally translated into the sub-text during the simultaneous interpretation process, the simultaneous interpretation process can be concerned during the simultaneous interpretation quality evaluation of the sub-speech, which helps to measure the simultaneous interpretation quality during the brushing process in the application scenario of streaming simultaneous interpretation. Therefore, the accuracy of the simultaneous interpretation quality evaluation can be improved in the application scenario of streaming simultaneous interpretation. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a schematic flowchart of an embodiment of the simultaneous interpretation quality evaluation method of the present application;
[0011] Figure 2a is a schematic diagram of an embodiment of the brushing data;
[0012] Figure 2b is a schematic diagram of an embodiment of the simultaneous interpretation time sequence;
[0013] Figure 2c is a schematic diagram of another embodiment of the brushing data;
[0014] Figure 2d is a schematic diagram of the process of an embodiment of the simultaneous interpretation quality evaluation of the present application;
[0015] Figure 3It is a schematic framework diagram of an embodiment of the simultaneous interpretation quality evaluation device of the present application;
[0016] Figure 4 It is a schematic framework diagram of an embodiment of the electronic device of the present application;
[0017] Figure 5 It is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Specific embodiments
[0018] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.
[0019] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0020] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the fragment " / " in this article generally represents an "or" relationship between the preceding and following associated objects. In addition, "multiple" in this article means two or more than two.
[0021] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the simultaneous interpretation quality evaluation method of the present application. Specifically, it may include the following steps:
[0022] Step S11: Segment the simultaneous interpretation text based on the target speech to obtain a number of sub-texts.
[0023] In one implementation scenario, the simultaneous interpretation text can be obtained by a simultaneous interpretation large model for the target speech. It should be noted that the simultaneous interpretation large model can include but is not limited to: open-source large models such as LLAMA and Bloom, and can also include but is not limited to large models obtained by fine-tuning the parameters of open-source large models with a specific corpus, or can also include but is not limited to custom large models. The specific source of the simultaneous interpretation large model is not limited here.
[0024] In one implementation scenario, before segmenting the simultaneous interpretation text based on the target speech, the simultaneous interpretation text can also be preprocessed first, such as including but not limited to: removing redundant spaces, special characters, etc., to ensure the text quality.
[0025] In an implementation scenario, the simultaneous interpretation text can be segmented based on punctuation marks to obtain several sub-texts. For example, the simultaneous interpretation text can be segmented based on punctuation marks such as semicolons, commas, and periods to obtain several sub-texts. Exemplarily, taking the simultaneous interpretation text "Today's meeting discussed the project progress and put forward constructive opinions; meanwhile, the manager made a summary and arranged the next step plan" as an example, when segmenting based on punctuation marks, the following sub-texts can be obtained: "Today's meeting discussed the project progress", "Put forward constructive opinions", "Meanwhile", "The manager made a summary", "And arranged the next step plan". Of course, the above example is only one possible example of text segmentation in the actual application process, and other possible situations will not be exemplified one by one here.
[0026] In another implementation scenario, different from the foregoing implementation manner, the simultaneous interpretation text can also be segmented based on semantics to obtain several sub-texts. For example, the simultaneous interpretation text can be segmented based on a semantic segmentation model to obtain several sub-texts. Exemplarily, still taking the simultaneous interpretation text "Today's meeting discussed the project progress and put forward constructive opinions; meanwhile, the manager made a summary and arranged the next step plan" as an example, when segmenting based on semantics, the following sub-texts can be obtained: "Today's meeting discussed the project progress", "Put forward constructive opinions", "Meanwhile, the manager made a summary", "And arranged the next step plan". Of course, the above example is only one possible example of text segmentation in the actual application process, and other possible situations will not be exemplified one by one here.
[0027] In yet another implementation scenario, different from the foregoing implementation manners, the simultaneous interpretation text may also be segmented based on punctuation marks to obtain a number of first sub-texts, and the simultaneous interpretation text may be segmented based on semantics to obtain a number of second sub-texts, and then the number of first sub-texts and the number of second sub-texts are integrated to obtain a number of sub-texts. For example, during the integration process, if it is found that any first sub-text belongs to a part of a certain second sub-text, then the first sub-text may follow the segmentation manner of the second sub-text to which it belongs. Conversely, if it is found that any second sub-text belongs to a part of a certain first sub-text, then the second sub-text may follow the segmentation manner of the first sub-text to which it belongs, so as to achieve the integration of the two segmentation manners. Still taking the simultaneous interpretation text "Today's meeting discussed the project progress and put forward constructive opinions; at the same time, the manager made a summary and arranged the next step plan" as an example, since the first sub-text "at the same time" belongs to a part of the second sub-text "at the same time, the manager made a summary", the first sub-text "at the same time" may follow the segmentation manner of the second sub-text "at the same time, the manager made a summary". In addition, on the basis of the above integration, the adjacent sub-texts of each sub-text may be further examined. If there is logical coherence between the adjacent sub-texts and it, the adjacent sub-texts may be further segmented until they belong to the same sub-text as this sub-text. Still taking the foregoing example as an example, after the above integration, the following sub-texts may be obtained: "Today's meeting discussed the project progress", "put forward constructive opinions", "at the same time, the manager made a summary", "and arranged the next step plan". By examining the logical coherence between the adjacent sub-texts, the sub-text "Today's meeting discussed the project progress" and the sub-text "put forward constructive opinions" may be integrated into a sub-text "Today's meeting discussed the project progress and put forward constructive opinions", and the sub-text "at the same time, the manager made a summary" and the sub-text "and arranged the next step plan" may be integrated into a sub-text "at the same time, the manager made a summary and arranged the next step plan". Of course, the above example is only a possible example of sub-text integration, and other possible situations are not listed one by one here. The above method, which segments the simultaneous interpretation text based on punctuation marks to obtain a number of first sub-texts, segments the simultaneous interpretation text based on semantics to obtain a number of second sub-texts, and then integrates the number of first sub-texts and the number of second sub-texts to obtain a number of sub-texts, can combine the grammatical structure and semantic information to more accurately reflect the logical relationship and meaning of the sentence, making the result more natural and coherent.
[0028] Step S12: Obtain the brushing data of the sub-voice corresponding to the sub-text in the target voice.
[0029] In the embodiments of the present disclosure, the brushing data of the sub-voice includes: a number of texts of the sub-voice from the first appearance of characters to gradual correction until finally translated into the sub-text during the simultaneous interpretation process. Please combine Figure 2a , Figure 2aIt is a schematic diagram of an embodiment of the brush character data. As Figure 2a shown, the gray text indicates that since the speaker has not finished a sentence, the system directly translates according to the word, and the subsequent results may change due to factors such as grammar and context. The black text indicates the definite translation result given by the system after comprehensively considering the context and grammar, and this result will not change subsequently. Of course, Figure 2a what is shown is only a possible example of the brush character data, and the specific content of the brush character data is not limited here. Still taking the aforementioned example as an example, the brush character data corresponding to the sub-voice of the sub-text "Today's meeting discussed the project progress and put forward constructive opinions" can be obtained, and the brush character data corresponding to the sub-voice of the sub-text "At the same time, the manager made a summary and arranged the next plan" can be obtained, so as to facilitate the subsequent simultaneous interpretation quality analysis of each sub-voice respectively.
[0030] Step S13: Based on the brush character data of the sub-voice, analyze and obtain the simultaneous interpretation quality score of the sub-voice.
[0031] In the embodiments of the present disclosure, the simultaneous interpretation quality can be evaluated by several evaluation indicators, and the several evaluation indicators can specifically include at least one of: the character output response time, the result response time, the brush character ratio, and the jump degree. The evaluation indicators are not limited here. It should be noted that the character output response time represents the time interval from when the user starts to pronounce (such as the first word or phrase) to when the translation system first outputs the translation result. As a possible example, the character output response time can be the primary indicator for the user to perceive the speed of the translation system; the result response time represents the time from when the user stops speaking (such as the last word of a sentence or statement) to when the translation system completes the translation of this segment. As a possible example, the result response time directly affects the completion speed of the translation. Especially in continuous conversations, too long a result response time will cause the interruption of the conversation rhythm and affect the fluency of communication; the brush character ratio represents the ratio of the number of times the intermediate results (i.e., several texts in the brush character data) are output to the number of words (terms) of the entire text result of the semantic clause when seeing the entire final result of a complete semantic clause; the jump degree represents the average value of the change in the number of words jumped each time for the intermediate results (i.e., several texts in the brush character data) when seeing the entire final result of a complete semantic clause. For the specific calculation processes of the above evaluation indicators, please refer to the following relevant descriptions respectively, and will not be elaborated here for the time being.
[0032] In an implementation scenario, when the simultaneous interpretation quality is used as an evaluation indicator with the character output response time, the time when the speaker first makes a sound in the sub-voice can be obtained as the first voice time, and the time when the first character appears in the brush character data of the sub-voice can be obtained as the first response time. Then, based on the time difference between the first voice time and the first response time, the simultaneous interpretation quality score of the sub-voice regarding the character output response time can be obtained. Please refer to Figure 2b , Figure 2bIt is a schematic diagram of an embodiment of simultaneous interpretation timing. As Figure 2b shown, the time when the speaker first speaks in sub-speech 1 is t 11 , and the time when the first character appears in the character brushing data of sub-speech 1 is t' 11 , the time when the speaker first speaks in sub-speech 2 is t 21 , and the time when the first character appears in the character brushing data of sub-speech 1 is t' 21 , the time when the speaker first speaks in sub-speech n-1 is t (n-1)1 , and the time when the first character appears in the character brushing data of sub-speech 1 is t' (n-1)1 , the time when the speaker first speaks in sub-speech n is t n1 , and the time when the first character appears in the character brushing data of sub-speech 1 is t' n1 . On this basis, for sub-speech 1, the simultaneous interpretation quality score t' 11 -t 11 regarding the character output response time can be obtained. For sub-speech 2, the simultaneous interpretation quality score t' 21 - t 21 regarding the character output response time can be obtained. For sub-speech n-1, the simultaneous interpretation quality score t' (n-1)1 - t (n-1)1 regarding the character output response time can be obtained. For sub-speech n, the simultaneous interpretation quality score t' n1 - t n1 regarding the character output response time can be obtained. Of course, the above examples are only possible examples of the first speech time, the first response time, and the simultaneous interpretation quality score regarding the character output response time, and other possible situations are not exemplified one by one here. In the above manner, when the simultaneous interpretation quality is evaluated by the character output response time, the time when the speaker first speaks in the sub-speech is obtained as the first speech time, and the time when the first character appears in the character brushing data of the sub-speech is obtained as the first response time. Then, based on the time difference between the first speech time and the first response time, the simultaneous interpretation quality score of the sub-speech regarding the character output response time is obtained, which can evaluate as accurately as possible the translation speed perceived by the user for each sub-speech during the simultaneous interpretation process.
[0033] In an implementation scenario, when the simultaneous interpretation quality is evaluated by the result response time, the time when the speaker finally speaks in the sub-speech can be obtained as the second speech time, and the time when the sub-text appears in the character brushing data of the sub-speech can be obtained as the second response time. Then, based on the time difference between the second speech time and the second response time, the simultaneous interpretation quality score of the sub-speech regarding the result response time is obtained. Please refer to Figure 2b , Figure 2b It is a schematic diagram of an embodiment of simultaneous interpretation timing. As Figure 2b shown, the time when the speaker finally speaks in sub-speech 1 is t 12, the occurrence time of the sub - text in the brushing data of the sub - voice 1 is t' 12 , the last speaking time of the speaker in the sub - voice 2 is t 22 , the occurrence time of the sub - text in the brushing data of the sub - voice 2 is t' 22 , the last speaking time of the speaker in the sub - voice n - 1 is t (n-1)2 , the occurrence time of the sub - text in the brushing data of the sub - voice n - 1 is t' (n-1)2 , the last speaking time of the speaker in the sub - voice n is t n2 , the occurrence time of the sub - text in the brushing data of the sub - voice n is t' n2 . On this basis, for the sub - voice 1, the simultaneous interpretation quality score t' regarding the result response time can be obtained 12 - t 12 , for the sub - voice 2, the simultaneous interpretation quality score t' regarding the result response time can be obtained 22 - t 22 , for the sub - voice n - 1, the simultaneous interpretation quality score t' regarding the result response time can be obtained (n-1)2 - t (n-1)2 , for the sub - voice n, the simultaneous interpretation quality score t' regarding the result response time can be obtained n2 - t n2 . Of course, the above example is only a possible example of the second voice time, the second response time, and the simultaneous interpretation quality score regarding the result response time. Other possible situations are not exemplified one by one here. In the above - mentioned way, when the simultaneous interpretation quality is evaluated by the result response time, the last speaking time of the speaker in the sub - voice is obtained as the second voice time, and the occurrence time of the sub - text in the brushing data of the sub - voice is obtained as the second response time. Then, based on the time difference between the second voice time and the second response time, the simultaneous interpretation quality score of the sub - voice regarding the result response time is obtained, which can evaluate the completion speed and fluency of the translation of each sub - voice in the simultaneous interpretation process as accurately as possible.
[0034] In an implementation scenario, when the simultaneous interpretation quality is evaluated by the brushing ratio, the total number of several texts in the brushing data of the sub - voice can be obtained, and the total length of the sub - text in the brushing data of the sub - voice can be obtained. Then, based on the ratio of the total number to the total length, the simultaneous interpretation quality score of the sub - voice regarding the brushing ratio is obtained. Please refer to Figure 2c , Figure 2c is a schematic diagram of another embodiment of the brushing data. As Figure 2c shown, for the sub - text "I am Chinese.", the brushing data of its corresponding sub - voice in the simultaneous interpretation process can include n texts (i.e., S 1, S 2 … , Sn-1 , S n ). That is to say, in this example, the total number of several texts in the brushing data of the sub-voice can be denoted as n, and since the sub-text in the brushing data of the sub-voice is "I am Chinese.", its total length can be denoted as . On this basis, the ratio of the two can be obtained to get the simultaneous interpretation quality score of the sub-voice regarding the brushing ratio, that is . Of course, the above example is only a possible example of the simultaneous interpretation quality score regarding the brushing ratio, and other possible situations will not be exemplified one by one here. It should be noted that the "brushing ratio" in the embodiments of the present disclosure can intuitively reflect the jump frequency of the translation result when the simultaneous interpretation function translates an audio segment. For example, if the brushing ratio is 0.8, it means that on average, each character needs to jump 0.8 times to correct it to the final result. That is to say, when 100 characters are translated, the number of intermediate result outputs is 80 times. In the above manner, when the simultaneous interpretation quality takes the brushing ratio as an evaluation index, the total number of several texts in the brushing data of the sub-voice is obtained, and the total length of the sub-text in the brushing data of the sub-voice is obtained, and then based on the ratio of the total number to the total length, the simultaneous interpretation quality score of the sub-voice regarding the brushing ratio is obtained, which can as accurately as possible reflect the jump frequency of each sub-voice during the simultaneous interpretation process.
[0035] In an implementation scenario, when the simultaneous interpretation quality takes the jump degree as an evaluation index, several texts in the brushing data of the sub-voice can be respectively selected as the current text, and the previous text located at the current text can be selected as the reference text, and then based on the respective total lengths of the current text and the reference text, the first length is obtained, and the length of the longest common prefix of the current text and the reference text is obtained as the second length, so that the jump degree of the current text can be obtained based on the first length and the second length, and then based on the jump degrees of each text in the brushing data of the sub-voice, the simultaneous interpretation quality score of the sub-voice regarding the jump degree can be obtained. It should be noted that the jump degree can be used to reflect how many characters change each time the intermediate result jumps. For example, if the jump degree is 2.6, it means that on average, 2.6 characters change each time the intermediate result changes. In the above manner, by performing jump analysis on each group of adjacent texts and determining the simultaneous interpretation quality score of the sub-voice regarding the jump degree by comprehensively analyzing the jumps of each group of adjacent texts, the jump degree of each sub-voice during the simultaneous interpretation process can be analyzed as accurately as possible.
[0036] In a specific implementation scenario, specifically, the larger value of the respective total lengths of the current text and the reference text can be obtained as the first length, and based on the length difference between the first length and the second length, the jump degree of the current text can be obtained, and then based on the average of the jump degrees of each text, the simultaneous interpretation quality score of the sub-voice regarding the jump degree can be obtained.
[0037] In a specific implementation scenario, please continue to refer to Figure 2c to select the first text S Taking 1 as the current text as an example, since the previous text is empty, the larger value among the total lengths of the current text and the reference text, that is, the first length is the total length of the current text itself, which can be denoted as L1 = Max(l1, 0). In addition, since the previous text is empty, there is no common prefix between the current text and the reference text, so the length of the longest common prefix is 0, that is, the second length is 0. Therefore, for the first text S 1, its jump degree can be denoted as L1. In contrast, taking 2 as the current text as an example to select the second text S , since the total length of its previous text S 1 is l1, the larger value among the total lengths of the current text and the reference text, that is, the first length can be denoted as L2 = Max(l2, l1). In addition, the length of the longest common prefix between the current text and the reference text, that is, the second length can be denoted as L'2 = len(LCP(S2, S1)), where LCP represents the longest common prefix and len(LCP) represents the length of the longest common prefix. Therefore, for the second text S 2, its jump degree can be denoted as L2 - L'2. In the case of selecting other texts as the current text, it can be deduced by analogy, and no more examples will be given here.
[0038] In a specific implementation scenario, still taking the Figure 2c shown brush character data as an example, the simultaneous interpretation quality score of the sub-voice regarding the jump degree can be expressed as [L1 + (L2 - L'2) + (L3 - L'3)… + (L n-1 - L' n-1 ) + (L n - L' n )] / n. Of course, the above example is only a possible example regarding the jump degree of the sub-voice when taking the Figure 2c shown brush character data as an example, and no more examples will be given for other possible situations here.
[0039] It should be noted that the above examples are possible examples for obtaining the simultaneous interpretation quality score of the sub-voice when taking the character output response time, result response time, brush character ratio, and jump degree as evaluation indicators respectively, and it does not limit the use of other evaluation indicators accordingly. No more examples will be given for the specific process of obtaining the simultaneous interpretation quality score when using other simultaneous interpretation indicators here.
[0040] Step S14: Based on the simultaneous interpretation quality scores of each sub-voice, fuse to obtain a target score representing the simultaneous interpretation quality of the target voice.
[0041] Specifically, after obtaining the simultaneous interpretation quality scores of each sub-speech, the simultaneous interpretation quality scores of the target speech for the corresponding evaluation index can be obtained by fusing the simultaneous interpretation quality scores of each sub-speech for the same evaluation index. Exemplarily, taking the word output response time as the evaluation index, the simultaneous interpretation quality scores of each sub-speech for the word output response time can be averaged to obtain the simultaneous interpretation quality score of the target speech for the word output response time:
[0042] =[( - ) + ( - ) +... + ( - ) + ( - )] / n (1)
[0043] In the above formula (1), represents the simultaneous interpretation quality score of the target speech for the word output response time. In addition, for the specific meanings of other parameters in formula (1), reference can be made to the foregoing relevant descriptions and will not be elaborated here. Or, taking the result response time as the evaluation index, the simultaneous interpretation quality scores of each sub-speech for the result response time can be averaged to obtain the simultaneous interpretation quality score of the target speech for the result response time:
[0044] =[( - ) + ( - ) +... + ( - ) + ( - )] / n (2)
[0045] In the above formula (2), represents the simultaneous interpretation quality score of the target speech for the result response time. In addition, for the specific meanings of other parameters in formula (2), reference can be made to the foregoing relevant descriptions and will not be elaborated here. Or, taking the character brushing ratio and the jump degree as the evaluation indexes, the simultaneous interpretation quality scores of the target speech for the character brushing ratio and the jump degree can be obtained respectively by referring to the calculation methods of the foregoing word output response time and result response time (such as averaging), which will not be elaborated here. It should be noted that if the performance is excellent in terms of the word output response time and the result response time, and the character brushing ratio and the jump degree are low, it indicates that the simultaneous interpretation has the ability of fast and stable translation output, which can bring users a low-latency and smooth usage experience. Please refer to Table 1 in combination. Table 1 is a schematic table of an embodiment of the reference range of user perception indexes.
[0046] Table 1 Schematic Table of the Reference Range of User Perception Indexes in an Embodiment
[0047]
[0048] As shown in Table 1, if the character output response time is within 1 second, it can be considered to be at an optimized level, and if it is within 2 seconds, it can be considered to be at a qualified level; if the result response time is within 1 second, it can be considered to be at an optimized level, and if it is within 2 seconds, it can be considered to be at a qualified level; if the jump degree is between 1 and 2, it can be considered that the output is stable; if the character brushing ratio is between 0.5 and 1, it can be considered that the jump frequency is moderate. Specifically, reference can be made to Table 1, which will not be elaborated here. On this basis, the target scores representing the simultaneous interpretation quality of the target speech can be obtained by further synthesizing the simultaneous interpretation quality scores of the target speech with respect to various evaluation indexes. Exemplarily, the simultaneous interpretation quality scores can be overall quantified according to the weights of the various evaluation indexes to obtain the target scores representing the simultaneous interpretation quality of the target speech. Specifically, the simultaneous interpretation quality scores of the various evaluation indexes can be first uniformly scored (e.g., unified within 0 to 10 points) according to Table 1 or a reference table similar to Table 1, and then the unified simultaneous interpretation quality scores are weighted using the weights of the various evaluation indexes to obtain the target scores representing the simultaneous interpretation quality of the target speech.
[0049] In an implementation scenario, please refer to Figure 2d , Figure 2d which is a schematic diagram of the process of an embodiment of the simultaneous interpretation quality evaluation of the present application. As Figure 2d shown, after the simultaneous interpretation text of the target speech is segmented based on punctuation and semantics, a number of sub-texts can be obtained. For each sub-text, the character brushing data of the corresponding sub-speech in the target speech can be obtained, and the character brushing data includes a number of texts of the sub-speech from the first character output to gradual correction until finally translated into the sub-text during the simultaneous interpretation process. Then, based on the character brushing data of the sub-speech, the simultaneous interpretation quality score of the sub-speech can be analyzed. On this basis, the target scores representing the simultaneous interpretation quality of the target speech can be fused based on the simultaneous interpretation quality scores of the respective sub-speeches. It should be noted that although Figure 2d the number of texts in the character brushing data of each sub-speech is expressed as "text 1, text 2,..., text m", it does not thereby limit that the character brushing data of each sub-speech has the same number of texts. In actual application processes, the number of texts in the character brushing data of each sub-speech can also be completely different; or, the number of texts in the character brushing data of each sub-speech can also be not completely the same, that is, the character brushing data of some sub-speeches has the same number of texts, while the character brushing data of some sub-speeches has different numbers of texts. The number of texts in the character brushing data of each sub-speech is not limited here.
[0050] Based on the simultaneous interpretation text of the target voice, the above solution is segmented to obtain several sub-texts, and the brushing data of the sub-voice corresponding to the sub-text in the target voice is obtained. The brushing data of the sub-voice includes: several texts of the sub-voice from the first appearance of words to gradual correction until finally translated into the sub-text during the simultaneous interpretation process. Then, based on the brushing data of the sub-voice, the simultaneous interpretation quality score of the sub-voice is analyzed, and further, based on the simultaneous interpretation quality scores of each sub-voice, the target score representing the simultaneous interpretation quality of the target voice is fused. Therefore, on the one hand, by dividing the simultaneous interpretation text into several sub-texts, and then separately evaluating the simultaneous interpretation quality of each sub-text corresponding to the sub-voice and finally performing score fusion, compared with the overall quality evaluation by combining the target voice and its simultaneous interpretation text, the evaluation granularity can be further refined, which helps to improve the accuracy of the simultaneous interpretation quality evaluation to a certain extent. On the other hand, when evaluating each sub-voice, since the brushing data of the sub-voice is combined, and the brushing data includes several texts of the sub-voice from the first appearance of words to gradual correction until finally translated into the sub-text during the simultaneous interpretation process, the simultaneous interpretation process can be concerned during the simultaneous interpretation quality evaluation of the sub-voice, which helps to measure the simultaneous interpretation quality during the brushing process in the application scenario of streaming simultaneous interpretation. Therefore, the accuracy of the simultaneous interpretation quality evaluation can be improved in the application scenario of streaming simultaneous interpretation.
[0051] Please refer to Figure 3 , Figure 3 which is a framework schematic diagram of an embodiment of the simultaneous interpretation quality evaluation device of the present application. The simultaneous interpretation quality evaluation device 30 includes: a text segmentation module 31, a data acquisition module 32, a quality analysis module 33, and a score fusion module 34. The text segmentation module 31 is configured to segment based on the simultaneous interpretation text of the target voice to obtain several sub-texts; the data acquisition module 32 is configured to acquire the brushing data of the sub-voice corresponding to the sub-text in the target voice; wherein, the brushing data of the sub-voice includes: several texts of the sub-voice from the first appearance of words to gradual correction until finally translated into the sub-text during the simultaneous interpretation process; the quality analysis module 33 is configured to analyze and obtain the simultaneous interpretation quality score of the sub-voice based on the brushing data of the sub-voice; the score fusion module 34 is configured to fuse based on the simultaneous interpretation quality scores of each sub-voice to obtain the target score representing the simultaneous interpretation quality of the target voice.
[0052] In the above solution, the simultaneous interpretation quality evaluation device 30 segments the simultaneous interpretation text based on the target speech to obtain a number of sub-texts, and obtains the typing data of the sub-speech corresponding to the sub-text in the target speech. The typing data of the sub-speech includes: a number of texts of the sub-speech from the first appearance of words to gradual correction until finally translated into the sub-text during simultaneous interpretation. Then, based on the typing data of the sub-speech, the simultaneous interpretation quality score of the sub-speech is analyzed, and further, based on the simultaneous interpretation quality scores of each sub-speech, the target score representing the simultaneous interpretation quality of the target speech is fused. Therefore, on the one hand, by dividing the simultaneous interpretation text into a number of sub-texts, and then separately evaluating the simultaneous interpretation quality of each sub-text corresponding to the sub-speech and finally fusing the scores, compared with the overall quality evaluation by combining the target speech and its simultaneous interpretation text, the evaluation granularity can be further refined, which helps to improve the accuracy of the simultaneous interpretation quality evaluation to a certain extent. On the other hand, when evaluating each sub-speech, since the typing data of the sub-speech is combined, and the typing data includes a number of texts of the sub-speech from the first appearance of words to gradual correction until finally translated into the sub-text during simultaneous interpretation, the simultaneous interpretation process can be concerned during the simultaneous interpretation quality evaluation of the sub-speech, which helps to measure the simultaneous interpretation quality during the typing process in the application scenario of streaming simultaneous interpretation. Therefore, the accuracy of the simultaneous interpretation quality evaluation can be improved in the application scenario of streaming simultaneous interpretation.
[0053] In some disclosed embodiments, the text segmentation module 31 includes a first segmentation sub-module for segmenting the simultaneous interpretation text based on punctuation marks to obtain a number of first sub-texts; the text segmentation module 31 includes a second segmentation sub-module for segmenting the simultaneous interpretation text based on semantics to obtain a number of second sub-texts; the text segmentation module 31 includes a segmentation integration sub-module for integrating based on the number of first sub-texts and the number of second sub-texts to obtain a number of sub-texts.
[0054] In some disclosed embodiments, the quality analysis module 33 includes a word appearance response evaluation sub-module for obtaining the time of the first voice of the speaker in the sub-speech as the first voice time and obtaining the time of the first word appearance in the typing data of the sub-speech as the first response time when the simultaneous interpretation quality is evaluated by the word appearance response time; the quality analysis module 33 includes a word appearance quality evaluation sub-module for obtaining the simultaneous interpretation quality score of the sub-speech regarding the word appearance response time based on the time difference between the first voice time and the first response time.
[0055] In some disclosed embodiments, the quality analysis module 33 includes a result response evaluation sub-module, which is used to obtain the time when the speaker of the sub-speech makes the last utterance as the second speech time, and obtain the appearance time of the sub-text in the word brushing data of the sub-speech as the second response time when the simultaneous interpretation quality is evaluated by the result response time; the quality analysis module 33 includes a result quality evaluation sub-module, which is used to obtain the simultaneous interpretation quality score of the sub-speech with respect to the result response time based on the time difference between the second speech time and the second response time.
[0056] In some disclosed embodiments, the quality analysis module 33 includes a text information acquisition sub-module, which is used to obtain the total number of several texts in the word brushing data of the sub-speech and obtain the total length of the sub-text in the word brushing data of the sub-speech when the simultaneous interpretation quality is evaluated by the word brushing ratio; the quality analysis module 33 includes a word brushing quality evaluation sub-module, which is used to obtain the simultaneous interpretation quality score of the sub-speech with respect to the word brushing ratio based on the ratio of the total number to the total length.
[0057] In some disclosed embodiments, the quality analysis module 33 includes a text selection sub-module, which is used to respectively select several texts in the word brushing data of the sub-speech as the current text and select the previous text located before the current text as the reference text when the simultaneous interpretation quality is evaluated by the jump degree; the quality analysis module 33 includes a length measurement sub-module, which is used to obtain the first length based on the total lengths of the current text and the reference text respectively, and obtain the length of the longest common prefix of the current text and the reference text as the second length; the quality analysis module 33 includes a jump degree measurement sub-module, which is used to obtain the jump degree of the current text based on the first length and the second length; the quality analysis module 33 includes a jump evaluation sub-module, which is used to fuse the jump degrees of each text in the word brushing data of the sub-speech to obtain the simultaneous interpretation quality score of the sub-speech with respect to the jump degree.
[0058] In some disclosed embodiments, the length measurement sub-module is specifically used to obtain the larger value of the total lengths of the current text and the reference text respectively as the first length; the jump degree measurement sub-module is specifically used to obtain the jump degree of the current text based on the length difference between the first length and the second length; the jump evaluation sub-module is specifically used to average the jump degrees of each text to obtain the simultaneous interpretation quality score of the sub-speech with respect to the jump degree.
[0059] In some disclosed embodiments, the scoring fusion module 34 includes a scoring fusion sub-module, which is used to fuse the simultaneous interpretation quality scores of each sub-speech with respect to the same evaluation index to obtain the simultaneous interpretation quality score of the target speech with respect to the corresponding evaluation index; the scoring fusion module 34 includes a scoring synthesis sub-module, which is used to synthesize the simultaneous interpretation quality scores of the target speech with respect to various evaluation indexes to obtain the target score representing the simultaneous interpretation quality of the target speech.
[0060] In some disclosed embodiments, the quality of simultaneous interpretation is evaluated using several evaluation metrics, and the several evaluation metrics include at least one of: character output response time, result response time, character brushing ratio, and jump degree; and / or, the simultaneous interpretation text is obtained by a simultaneous interpretation large model for simultaneous interpretation of the target speech.
[0061] Please refer to Figure 4 , Figure 4 FIG. is a schematic framework diagram of an embodiment of an electronic device according to the present application. The electronic device 40 at least includes a memory 41 and a processor 42 that are coupled to each other. At least program instructions are stored in the memory 41, and the processor 42 is configured to execute the program instructions to implement the steps in any of the above embodiments of the simultaneous interpretation quality evaluation method. Specifically, reference can be made to the foregoing disclosed embodiments, which will not be elaborated herein. As a possible example, the electronic device 40 may include, but is not limited to, a server, etc., and the specific type of the electronic device 40 is not limited herein.
[0062] Specifically, the processor 42 is configured to control itself and the memory 41 to implement the steps in any of the above embodiments of the simultaneous interpretation quality evaluation method. The processor 42 may also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 42 may be implemented jointly by integrated circuit chips.
[0063] In the above solution, the electronic device 40 segments the simultaneous interpretation text based on the target voice to obtain a number of sub-texts, and obtains the brushing data of the sub-voices corresponding to the sub-texts in the target voice. The brushing data of the sub-voices includes: a number of texts of the sub-voices from the first appearance of words to gradual correction until finally translated into the sub-texts during simultaneous interpretation. Then, based on the brushing data of the sub-voices, the simultaneous interpretation quality score of the sub-voices is analyzed, and further based on the simultaneous interpretation quality scores of each sub-voice, a target score representing the simultaneous interpretation quality of the target voice is fused. Therefore, on the one hand, by dividing the simultaneous interpretation text into a number of sub-texts, and then separately evaluating the simultaneous interpretation quality of each sub-text corresponding to the sub-voice and finally performing score fusion, compared with the overall quality evaluation by combining the target voice and its simultaneous interpretation text, the evaluation granularity can be further refined, which helps to improve the accuracy of the simultaneous interpretation quality evaluation to a certain extent. On the other hand, when evaluating each sub-voice, since the brushing data of the sub-voice is combined, and the brushing data includes a number of texts of the sub-voice from the first appearance of words to gradual correction until finally translated into the sub-text during simultaneous interpretation, the simultaneous interpretation process can be concerned during the simultaneous interpretation quality evaluation of the sub-voice, which helps to measure the simultaneous interpretation quality during the brushing process in the application scenario of streaming simultaneous interpretation. Therefore, the accuracy of the simultaneous interpretation quality evaluation can be improved in the application scenario of streaming simultaneous interpretation.
[0064] Please refer to Figure 5 , Figure 5 which is a framework schematic diagram of an embodiment of the computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be run by a processor, and the program instructions 51 are used to implement the steps in any of the above embodiments of the simultaneous interpretation quality evaluation method.
[0065] In the above solution, the computer-readable storage medium 50 segments based on the simultaneous interpretation text of the target voice to obtain a number of sub-texts, acquires the brushing data of the sub-voices corresponding to the sub-texts in the target voice, and the brushing data of the sub-voices includes: a number of texts of the sub-voices from the first appearance of words to gradual correction until finally translated into the sub-texts during simultaneous interpretation. Then, based on the brushing data of the sub-voices, the simultaneous interpretation quality score of the sub-voices is analyzed, and further, based on the simultaneous interpretation quality scores of each sub-voice, a target score representing the simultaneous interpretation quality of the target voice is fused. Therefore, on the one hand, by dividing the simultaneous interpretation text into a number of sub-texts, and then separately evaluating the simultaneous interpretation quality of each sub-voice corresponding to each sub-text and finally fusing the scores, compared with the overall quality evaluation by combining the target voice and its simultaneous interpretation text, the evaluation granularity can be further refined, which helps to improve the accuracy of the simultaneous interpretation quality evaluation to a certain extent. On the other hand, when evaluating each sub-voice, since the brushing data of the sub-voice is combined, and the brushing data includes a number of texts of the sub-voice from the first appearance of words to gradual correction until finally translated into the sub-text during simultaneous interpretation, the simultaneous interpretation process can be concerned during the simultaneous interpretation quality evaluation of the sub-voice, which helps to measure the simultaneous interpretation quality during the brushing process in the application scenario of streaming simultaneous interpretation. Therefore, the accuracy of the simultaneous interpretation quality evaluation can be improved in the application scenario of streaming simultaneous interpretation.
[0066] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0067] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated in this article.
[0068] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0069] The unit described as a separate component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0070] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0071] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0072] If the technical solution of the present application involves personal information, before the product applying the technical solution of the present application processes personal information, it has clearly informed the personal information processing rules and obtained the individual's independent consent. If the technical solution of the present application involves sensitive personal information, before the product applying the technical solution of the present application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirement of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection range has been entered and personal information will be collected. If an individual voluntarily enters the collection range, it is regarded as consenting to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
Claims
1. A method for evaluating simultaneous interpretation quality, characterized in that: include: The simultaneous interpretation text of the target speech is segmented to obtain a plurality of subtexts; wherein the simultaneous interpretation text is segmented based on punctuation to obtain a plurality of first subtexts, and the simultaneous interpretation text is segmented based on semantics to obtain a plurality of second subtexts, and the plurality of first subtexts and the plurality of second subtexts are integrated to obtain the plurality of subtexts, and in the integration process, if any of the first subtexts is a partial text of the second subtext, the corresponding first subtext follows the segmentation method of the second subtext to which it belongs, and if any of the second subtexts is a partial text of the first subtext, the corresponding second subtext follows the segmentation method of the first subtext to which it belongs; Acquire the brush word data of the sub-speech corresponding to the sub-text in the target speech; wherein the brush word data of the sub-speech includes: a plurality of texts of the sub-speech from the first word generation to the gradual correction and finally translation into the sub-text in the simultaneous interpretation process; Based on the brush word data of the sub-speech, analyzing and obtaining the simultaneous interpretation quality score of the sub-speech; Based on the simultaneous interpretation quality scores of the sub-speech, a target score representing the simultaneous interpretation quality of the target speech is obtained by fusion; wherein, In the case where the simultaneous interpretation quality is evaluated by the brush word ratio, the simultaneous interpretation quality score of the sub-speech is obtained by analyzing the brush word data of the sub-speech, including: obtaining the total number of the plurality of texts in the brush word data of the sub-speech, and obtaining the total length of the sub-texts in the brush word data of the sub-speech; obtaining the simultaneous interpretation quality score of the sub-speech with respect to the brush word ratio based on the ratio of the total number to the total length; In the case where the simultaneous interpretation quality is evaluated by the jump degree, the simultaneous interpretation quality score of the sub-speech is obtained by analyzing the brush word data of the sub-speech, including: selecting the several texts in the brush word data of the sub-speech as the current text, and selecting the previous text located before the current text as the reference text; obtaining the larger value of the total length of the current text and the reference text as the first length, and obtaining the longest common prefix length of the current text and the reference text as the second length; obtaining the jump degree of the current text based on the length difference between the first length and the second length; and obtaining the simultaneous interpretation quality score of the sub-speech with respect to the jump degree by averaging the jump degrees of each of the texts in the brush word data of the sub-speech.
2. The method according to claim 1, characterized in that In the case where the simultaneous interpretation quality is evaluated by the word output response time, the simultaneous interpretation quality score of the sub-speech is obtained by analyzing the word output data based on the sub-speech, including: The time when the speaker first utters a sound in the sub-speech is obtained as the first speech time, and the time when the first word is produced in the brush word data of the sub-speech is obtained as the first response time; Based on the time difference between the first speech time and the first response time, a simultaneous interpretation quality score of the sub-speech with respect to the word output response time is obtained.
3. The method according to claim 1, characterized in that In the case where the simultaneous interpretation quality is evaluated by the result response time, the simultaneous interpretation quality score of the sub-speech is obtained by analyzing the brush word data based on the sub-speech, including: Acquire the last utterance time of the speaker in the sub-speech as the second speech time, and acquire the appearance time of the sub-text in the brush word data of the sub-speech as the second response time; Based on the time difference between the second speech time and the second response time, a simultaneous interpretation quality score of the sub-speech with respect to the result response time is obtained.
4. The method according to claim 1, characterized in that: The simultaneous interpretation quality scores based on the sub-speech are integrated to obtain a target score representing the simultaneous interpretation quality of the target speech, including: Based on the simultaneous interpretation quality scores of the sub-speech with respect to the same evaluation index, the simultaneous interpretation quality score of the target speech with respect to the corresponding evaluation index is obtained by fusing; The simultaneous interpretation quality scores of the target speech with respect to various evaluation indicators are integrated to obtain a target score representing the simultaneous interpretation quality of the target speech.
5. The method according to any one of claims 1 to 4, characterized in that: The simultaneous interpretation quality is evaluated by a number of evaluation indicators, and the several evaluation indicators also include: at least one of the word response time and the result response time; And / or, the simultaneous interpretation text is obtained by simultaneously interpreting the target speech by a large simultaneous interpretation model.
6. A simultaneous interpretation quality evaluation device, characterized in that: include: A text segmentation module, for segmenting a simultaneous interpretation text based on the target speech to obtain a plurality of subtexts; wherein the simultaneous interpretation text is segmented based on punctuation to obtain a plurality of first subtexts, and the simultaneous interpretation text is segmented based on semantics to obtain a plurality of second subtexts, and the plurality of first subtexts and the plurality of second subtexts are integrated to obtain the plurality of subtexts, and during the integration process, if any of the first subtexts is a partial text in the second subtext, the corresponding first subtext follows the segmentation method of the second subtext to which it belongs, and if any of the second subtexts is a partial text in the first subtext, the corresponding second subtext follows the segmentation method of the first subtext to which it belongs; A data acquisition module is used to acquire the brush word data of the sub-speech corresponding to the sub-text in the target speech; wherein the brush word data of the sub-speech includes: a plurality of texts of the sub-speech from the first word generation to the gradual correction and finally translation into the sub-text during the simultaneous interpretation process; A quality analysis module, used for analyzing the word-brushing data of the sub-speech to obtain a simultaneous interpretation quality score of the sub-speech; A scoring fusion module is used to obtain a target score representing the simultaneous interpretation quality of the target speech based on the simultaneous interpretation quality scores of each of the sub-speech; wherein, In the case where the simultaneous interpretation quality is evaluated by the brush word ratio, the simultaneous interpretation quality score of the sub-speech is obtained by analyzing the brush word data of the sub-speech, including: obtaining the total number of the plurality of texts in the brush word data of the sub-speech, and obtaining the total length of the sub-texts in the brush word data of the sub-speech; obtaining the simultaneous interpretation quality score of the sub-speech with respect to the brush word ratio based on the ratio of the total number to the total length; In the case where the simultaneous interpretation quality is evaluated by the jump degree, the simultaneous interpretation quality score of the sub-speech is obtained by analyzing the brush word data of the sub-speech, including: selecting the several texts in the brush word data of the sub-speech as the current text, and selecting the previous text located before the current text as the reference text; obtaining the larger value of the total length of the current text and the reference text as the first length, and obtaining the longest common prefix length of the current text and the reference text as the second length; obtaining the jump degree of the current text based on the length difference between the first length and the second length; and obtaining the simultaneous interpretation quality score of the sub-speech with respect to the jump degree by averaging the jump degrees of each of the texts in the brush word data of the sub-speech.
7. An electronic device, characterized in that: The method comprises at least a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the simultaneous interpretation quality evaluation method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the simultaneous interpretation quality evaluation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Machine simultaneous interpretation system and method, test method and device and related equipment
CN116935853A
Reading evaluation method and device, electronic equipment and storage medium
CN117935863A