Machine simultaneous interpretation system, method, testing method, device and related equipment
By setting thresholds for translation and speech synthesis effects, the machine simultaneous interpretation system solves the problem of uneven system quality, achieves a holistic consideration of translation quality and real-time performance, and improves the overall system quality and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HKUST IFLYTEK (SHANGHAI) TECH CO LTD
- Filing Date
- 2023-05-18
- Publication Date
- 2026-06-02
AI Technical Summary
Existing machine simultaneous interpretation systems vary in quality and fail to comprehensively consider multiple quality-influencing factors such as translation effect, translation real-time performance, and speech synthesis effect, resulting in an overall uneven quality and affecting user experience.
This invention provides a machine simultaneous interpretation system, including audio acquisition, multimodal output, and speech translation modules. It sets thresholds for translation and speech synthesis effects, and ensures real-time translation and readability through subtitles, audio, and video output. It uses licensed sound libraries and domain-customized models and supports both offline and online modes.
It achieves overall quality balance in machine simultaneous interpretation systems, improves translation quality and user experience, provides a unified quality assessment method, and is applicable to various application scenarios.
Smart Images

Figure CN116935853B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent voice interaction system testing technology, and in particular to a machine simultaneous interpretation system, method, testing method, apparatus and related equipment. Background Technology
[0002] With the increasing number of cross-language international conferences, cultural exchange, and communication, the demand for simultaneous interpreting is growing daily. However, this growth is accompanied by challenges such as high training and operating costs for human simultaneous interpreters, limited knowledge storage and vocabulary, and an inability to meet current market demands. Machine simultaneous interpreting, a technological direction within intelligent voice interaction, offers advantages such as high accuracy in recognizing and remembering objective information, a large terminology and vocabulary database, and the ability to work efficiently and at low cost for extended periods. Utilizing machine simultaneous interpreting to assist or partially replace human translation has become a trend.
[0003] To standardize the system reference framework and basic technical requirements for intelligent voice interaction, relevant fundamental national standards have been formulated, such as GB / T21023-2007 General Technical Specification for Chinese Speech Recognition Systems and GB / T21024-2007 General Technical Specification for Chinese Speech Synthesis Systems, which provide terminology and definitions related to speech recognition and speech synthesis in artificial intelligence. However, standards and specifications related to machine simultaneous interpretation have not yet been provided. Furthermore, the quality of existing machine simultaneous interpretation systems varies considerably, especially given the complex and varied background environments, diverse speaker expressions with colloquialisms and dialects, the scarcity of multilingual speech and translation training data, and insufficient coverage of professional fields. The usability of different machine simultaneous interpretation systems varies. The applicant's in-depth research reveals that the root cause of the problem lies in the fact that the quality of machine simultaneous interpretation systems can be reflected in multiple different aspects, such as translation effectiveness measured from different angles, real-time translation, and speech synthesis effectiveness. Existing simultaneous interpretation systems do not comprehensively consider different quality-influencing factors, resulting in an uneven overall quality that affects normal user experience.
[0004] At the same time, existing technologies lack relevant testing methods for simultaneous interpretation systems, which makes it impossible to ensure the quality of machine simultaneous interpretation systems. Summary of the Invention
[0005] In view of this, this application provides a machine simultaneous interpretation system, method, testing method, apparatus, and related equipment to improve the quality of machine simultaneous interpretation and to evaluate the quality of the machine simultaneous interpretation system. The technical solution is as follows:
[0006] In a first aspect, a machine simultaneous interpretation system is provided, comprising: an application terminal and a server terminal. The application terminal includes an audio acquisition module and a multimodal output module. The multimodal output module includes at least one of a subtitle display module, an audio playback module, and a video playback module. The server terminal includes a speech translation module. When the multimodal output module includes the audio playback module, the server terminal also includes a speech synthesis module. When the multimodal output module includes the video playback module, the server terminal also includes the speech synthesis module and a virtual human module.
[0007] The audio acquisition module is used to acquire the audio stream of the source language;
[0008] The speech translation module is used to translate the audio stream of the source language to obtain translated text of the target language that meets the set translation effect. The set translation effect includes the translation BLEU score, fidelity and fluency reaching corresponding score thresholds.
[0009] The subtitle display module is used to output and display the translated text of the target language within a first set time period after the audio acquisition module acquires the audio stream of the source language, and during the process of outputting and displaying the translated text, the average erasure rate of the translated text does not exceed a set average erasure rate threshold.
[0010] The speech synthesis module is used to perform speech synthesis on the translated text of the target language to obtain an audio stream of the target language that meets the set speech synthesis effect. The set speech synthesis effect includes the average sentence synthesis accuracy and the synthesis naturalness reaching corresponding score thresholds, and the sound library used during synthesis being an authorized sound library.
[0011] The audio playback module is used to play the audio stream of the target language within a second set time period after the audio acquisition module acquires the audio stream of the source language;
[0012] The virtual human module is used to combine the audio stream of the target language with the virtual avatar and convert it into a video stream;
[0013] The video playback module is used to play the video stream.
[0014] Preferably, the speech translation module includes: a speech recognition module and a machine translation module;
[0015] The speech recognition module is used to recognize the audio stream of the source language to obtain the source language text;
[0016] The machine translation module is used to translate the source language text into the target language to obtain the translated text in the target language.
[0017] Preferably, the voice translation module supports adding a hot word dictionary, and the translation accuracy of hot words after adding the hot word dictionary is not lower than the set hot word intervention efficiency threshold.
[0018] Preferably, the speech translation module supports loading domain-customized models, and after loading the domain-customized models, the translation BLEU score improvement on the domain test set is not less than a set BLEU score difference threshold. The domain-customized models include speech recognition and text translation models.
[0019] Preferably, the application is deployed locally, the server is deployed on the network side, and the application and the server are connected via the network.
[0020] or,
[0021] Both the application and server are deployed locally, supporting simultaneous interpretation in offline mode.
[0022] Preferably, the application terminal further includes:
[0023] The meeting data export module is used to respond to the user's operation of exporting specified meeting data to a specified storage location. The specified meeting data includes any one or more of the recognition results, translation results, and speech synthesis results of the source language audio stream.
[0024] Preferably, the application terminal further includes:
[0025] The speech recognition result editing module is used to respond to the human translator's operation of editing the recognition results of the audio stream and obtain the edited source language text;
[0026] And / or,
[0027] The translation text editing module is used to respond to human translators' editing operations on the translated target language text, and to obtain the edited target language translation text.
[0028] Preferably, the server further includes a data preprocessing module, used to perform content security detection on the translated text of the target language output by the speech translation module, and to stop the subsequent processing if the content security detection fails.
[0029] Preferably, the application terminal further includes: an acoustic noise reduction module, used to perform noise reduction processing on the audio stream of the source language acquired by the audio acquisition module, and send the noise-reduced audio stream to the speech translation module;
[0030] And / or,
[0031] The server also includes a post-translation processing module, which is used to normalize the translated text of the target language and send the normalized translated text of the target language to the speech synthesis module.
[0032] Preferably, the speech translation module further includes: a post-recognition processing module, used to normalize the source language text and send the normalized source language text to the machine translation module;
[0033] The subtitle display module is also used to display the source language text obtained by the speech recognition module, or to display the normalized source language text obtained by the post-recognition processing module.
[0034] Secondly, a machine simultaneous interpretation method is provided to achieve at least one of the following simultaneous interpretation processes: simultaneous interpretation from source language audio to target language text, simultaneous interpretation from source language audio to target language audio, and simultaneous interpretation from source language audio to target language video.
[0035] The process of simultaneous interpretation between source language audio and target language text includes:
[0036] Obtain the audio stream of the source language;
[0037] The audio stream of the source language is translated to obtain the translated text of the target language that meets the set translation effect. The set translation effect includes the translation BLEU score, fidelity and fluency reaching the corresponding score thresholds.
[0038] Within a first set time period after acquiring the audio stream of the source language, the translated text of the target language is output and displayed, and during the process of outputting and displaying the translated text, the average erasure rate of the translated text does not exceed a set average erasure rate threshold.
[0039] The process of simultaneous interpretation between source language audio and target language audio includes:
[0040] Obtain the audio stream of the source language;
[0041] The audio stream of the source language is translated to obtain the translated text of the target language that meets the set translation effect. The set translation effect includes the translation BLEU score, fidelity and fluency reaching the corresponding score thresholds.
[0042] The translated text of the target language is subjected to speech synthesis to obtain an audio stream of the target language that meets the set speech synthesis effect. The set speech synthesis effect includes the average sentence synthesis accuracy and the synthesis naturalness reaching corresponding score thresholds, and the sound library used during synthesis being an authorized sound library.
[0043] Within a second predetermined time period after the source language audio stream is acquired, the target language audio stream is played.
[0044] The process of simultaneous interpretation between source language audio and target language video includes:
[0045] Obtain the audio stream of the source language;
[0046] The audio stream of the source language is translated to obtain the translated text of the target language that meets the set translation effect. The set translation effect includes the translation BLEU score, fidelity and fluency reaching the corresponding score thresholds.
[0047] The translated text of the target language is subjected to speech synthesis to obtain an audio stream of the target language that meets the set speech synthesis effect. The set speech synthesis effect includes the average sentence synthesis accuracy and the synthesis naturalness reaching corresponding score thresholds, and the sound library used during synthesis being an authorized sound library.
[0048] The audio stream of the target language is combined with the virtual avatar and converted into a video stream, which is then played.
[0049] Preferably, translating the audio stream of the source language to obtain translated text in the target language that meets the set translation effect includes:
[0050] The audio stream of the source language is identified to obtain the source language text;
[0051] The source language text is translated into the target language to obtain a translated text in the target language that meets the set translation effect.
[0052] Preferably, translating the audio stream of the source language to obtain translated text in the target language that meets the set translation effect includes:
[0053] A machine simultaneous interpretation system with an added hot word lexicon is used to translate the audio stream of the source language to obtain translated text in the target language that meets the set translation effect.
[0054] Among them, compared with machine simultaneous interpretation systems without hot word databases, machine simultaneous interpretation systems with hot word databases have a translation accuracy rate of hot words that is no less than the set hot word intervention efficiency threshold.
[0055] Preferably, translating the audio stream of the source language to obtain translated text in the target language that meets the set translation effect includes:
[0056] A machine simultaneous interpretation system with a customized model for the target domain is used to translate the audio stream of the source language to obtain translated text in the target language that meets the set translation effect.
[0057] The target domain customized model is a speech recognition and text translation model in the target domain to which the audio stream of the source language belongs. Compared with the machine simultaneous interpretation system without the target domain customized model, the machine simultaneous interpretation system with the target domain customized model has a translation BLEU score improvement of no less than a set BLEU score difference threshold on the domain test set. The domain test set includes multiple test audios in the target domain.
[0058] Preferably, the machine simultaneous interpretation method is implemented in a network-connected or network-disconnected state.
[0059] Preferably, it further includes:
[0060] In response to a user's operation to export specified meeting data to a specified storage location, the specified meeting data is exported to the specified storage location. The specified meeting data includes any one or more of the recognition results, translation results, and speech synthesis results of the source language audio stream.
[0061] Preferably, it further includes:
[0062] The system responds to human translators' editing of the audio stream's recognition results, resulting in edited source language text.
[0063] And / or,
[0064] The system responds to human translators' editing of the translated text in the target language, resulting in an edited translated text in the target language.
[0065] Preferably, it further includes:
[0066] The translated text in the target language is subjected to content security testing. If the content security test fails, the subsequent processing is stopped.
[0067] Preferably, before translating the audio stream of the source language, the method further includes:
[0068] The audio stream of the source language is subjected to noise reduction processing;
[0069] And / or,
[0070] Before performing speech synthesis on the translated text of the target language, the method further includes:
[0071] The translated text in the target language is then normalized.
[0072] Preferably, before translating the source language text into the target language, the method further includes:
[0073] The source language text is then standardized.
[0074] The source language text or the normalized source language text is output and displayed.
[0075] Thirdly, a testing method for a machine simultaneous interpretation system is provided for testing the aforementioned machine simultaneous interpretation system. The testing method includes:
[0076] Obtain the test set corresponding to the test item of the system under test, wherein the test set corresponding to the test item includes at least one language direction test set, and the test set contains multiple test audios;
[0077] Input each test audio from the test set corresponding to the test item into the system under test;
[0078] Obtain the operating data of the system under test on the test item;
[0079] Based on the operating data of the system under test on the test item, the test result of the system under test on the test item is determined.
[0080] Preferably, the test items include:
[0081] A first test item is used to test the performance of the system under test from source language audio stream input to target language translated text output, and a second test item is used to test the performance of the system under test from source language audio stream input to target language audio stream output, wherein the test set corresponding to the first test item and the second test item includes the simultaneous interpretation test set;
[0082] The first test item includes a translation effectiveness test item, a first translation real-time performance test item, and a translation readability test item;
[0083] The second test item includes a translation effect test item, a second translation real-time performance test item, and a speech synthesis effect test item.
[0084] Preferably, the test item further includes:
[0085] The test includes one or more of the following: performance optimization test items, content management test items, offline simultaneous interpretation test items, and content security test items. The performance optimization test items are used to test the performance optimization of the system under test when intervention factors are added. The content management test items are used to test whether the system under test has the specified content management function. The offline simultaneous interpretation test items are used to test whether the various functions of the system under test are implemented normally when the network is disconnected. The content security test items are used to test whether the translated text of the target language output by the system under test conforms to the security guidelines.
[0086] The test set corresponding to the content security test item includes a security test set, which contains multiple test audio files.
[0087] Preferably, the translation performance test items include BLEU subtest items, fidelity test items, and fluency test items. The BLEU subtest items are used to evaluate the translation quality of the tested system, the fidelity test items are used to evaluate whether the translated text obtained by the tested system faithfully expresses the content of the original text, and the fluency test items are used to evaluate whether the translated text obtained by the tested system is fluent and whether it conforms to the expression habits of the target language.
[0088] The first translation real-time test item is used to test the time difference between the input test audio and the display of the translated text on the screen of the system under test;
[0089] The translation readability test item is used to test the average erasure rate of the translated text displayed on the screen by the tested system;
[0090] The second translation real-time test item is used to test the time difference between the input test audio and the playback of the synthesized target language audio stream in the system under test;
[0091] The speech synthesis effect test items include: voice library authorization test item, average sentence synthesis accuracy test item, and speech synthesis naturalness test item.
[0092] Preferably, the effect optimization test items include:
[0093] The test includes at least one of the following: hot word intervention test item, domain model customization test item, and human-computer collaborative intervention test item. The hot word intervention test item is used to test the change in translation accuracy of hot words in the tested system before and after adding a hot word lexicon. The domain model customization test item is used to test the change in translation BLEU score of the tested system on the domain test set before and after loading a domain customization model. The human-computer collaborative intervention test item is used to test whether the final output result of the tested system is consistent with the operation of the tester when the tester intervenes.
[0094] The test set corresponding to the hot word intervention test item includes a hot word test set and a simultaneous interpretation test set. The hot word test set contains a hot word lexicon, and the hot words include hot words for identification and hot words for translation.
[0095] The test set corresponding to the domain model customized test item includes a domain test set, which includes at least one domain customized model and multiple test audios under the corresponding domain.
[0096] The test set corresponding to the human-machine collaborative intervention test items includes the simultaneous interpretation test items.
[0097] Preferably, the operating data of the system under test on the BLEU sub-test, fidelity test, and fluency test includes: machine-translated text obtained by the system under test after translating the test audio from the input simultaneous interpretation test set;
[0098] The operating data of the system under test on the first translation real-time test item includes: the start time T1 of each test audio input to the system under test in the simultaneous interpretation test set, and the time T2 when the corresponding machine-translated text begins to be displayed on the screen.
[0099] The operational data of the system under test on the translation readability test item includes: the text displayed each time the subtitle is refreshed after each test audio from the simultaneous interpretation test set is input into the system under test;
[0100] The operating data of the system under test in the second translation real-time test item includes: the start time T1 of each test audio input to the system under test in the simultaneous interpretation test set, and the time T3 when the corresponding machine-synthesized audio starts playing;
[0101] The running data of the tested system on the average sentence synthesis accuracy test and the speech synthesis naturalness test include: the machine-synthesized audio output by the tested system corresponding to each test audio in the simultaneous interpretation test set;
[0102] Based on the operating data of the system under test on the test item, the test result of the system under test on the test item is determined, including:
[0103] For the BLEU score test item, the BLEU score of the tested system on the simultaneous interpretation test set is calculated based on the reference answer text of each test audio in the simultaneous interpretation test set and the machine translation text.
[0104] For the fidelity and fluency tests, obtain the fidelity and fluency scores given by scoring experts to the machine-translated text.
[0105] The translation performance level of the tested system is determined based on its scores in the BLEU subtest, fidelity test, and fluency test.
[0106] For the first real-time translation test, the time difference T is calculated for each test audio file. d1 = T2 - T1, and calculate the time difference T between each test audio in the simultaneous interpretation test set. d1 The average value is used as the first real-time translation indicator.
[0107] The first translation real-time performance level of the tested system is determined based on the index value of the first translation real-time performance test item.
[0108] For the translation readability test, the average erasure rate NE of the machine-translated text corresponding to each test audio is calculated according to the following formula, based on the text displayed at each subtitle refresh:
[0109]
[0110]
[0111] in, This represents the length of the text display sequence at the i-th refresh time. The length of the string with the longest common prefix between the two strings is represented by I, which represents the number of times the machine-translated text corresponding to a test audio changes during the final determination process.
[0112] The average erase rate of the machine-translated text corresponding to each test audio in the simultaneous interpretation test set is calculated as the average erase rate (NE) index value of the system.
[0113] The translation readability level of the tested system is determined based on the index values of the tested system on the translation readability test item.
[0114] For the second translation real-time test item, the time difference T is calculated for each test audio. d2 = T3 - T1, and calculate the time difference T between each test audio in the simultaneous interpretation test set. d2 The average value of is used as the second real-time translation indicator.
[0115] The second translation real-time performance level of the tested system is determined based on the index value of the tested system on the second translation real-time performance test item.
[0116] For the average sentence synthesis accuracy test item, the number of correct machine-synthesized audio is counted, and the ratio of the number of correct machine-synthesized audio to the total number of all test audio in the simultaneous interpretation test set is calculated to obtain the average sentence synthesis accuracy.
[0117] For the speech synthesis naturalness test item, obtain the synthesis naturalness score results of the scoring experts on the machine-synthesized audio;
[0118] The speech synthesis effect level of the tested system is determined based on its scores in the average sentence synthesis accuracy test and the speech synthesis naturalness test.
[0119] Preferably, the running data of the tested system on the hot word intervention test item includes: the machine-translated text obtained by the tested system after translating the test audio from the input simultaneous interpretation test set before and after adding the hot word test set;
[0120] The running data of the system under test on the domain model customization test item includes: the machine-translated text obtained by translating the test audio in the corresponding domain of the domain test set before and after loading the domain customization model in the domain test set;
[0121] The operational data of the system under test in the content security test items includes:
[0122] The system under test outputs machine-translated text obtained after translating the input security test set audio.
[0123] Based on the operating data of the system under test on the test item, the test result of the system under test on the test item is determined, including:
[0124] For the hot word intervention test item, the number of correctly translated hot words in the machine-translated text before adding the hot word test set (P1) and the number of correctly translated hot words in the machine-translated text after adding the hot word test set (P2) are counted, and the hot word intervention efficiency is calculated.
[0125] , where T is the total number of hot words contained in the simultaneous interpreting test set;
[0126] For the domain model-customized test items, based on the reference answer text of machine-translated text and test audio, calculate the BLEU score of the tested system on the domain test set before and after loading the domain model-customized model, and compare the difference in BLEU score of the tested system on the domain test set after loading the domain model compared to before loading the domain model-customized model.
[0127] For content security testing items, obtain human assessment results on whether machine-translated text conforms to security guidelines.
[0128] Preferably, the simultaneous interpretation test set for each language direction meets the following requirements:
[0129] The effective duration of the test audio should be no less than 2 hours and no less than 2000 sentences;
[0130] The number of speakers of different types corresponding to the test audio shall not be less than 10, and there shall be no less than 5 male speakers and no less than 5 female speakers;
[0131] The test audio is real-world audio with a signal-to-noise ratio of at least 20dB.
[0132] The test audio should cover at least the following scenarios: technology, news, politics, finance, and medical meetings.
[0133] The speakers corresponding to the test audio were selected from those who were representative and whose speech patterns were statistically consistent, taking into account factors such as different ages, different speaking speeds, different educational backgrounds, and different speaking rhythms.
[0134] The test audio recordings should be consistent with the platform, sampling rate, and input channels of the system under test, or meet the set similarity conditions.
[0135] The test audio had no regional accents, and the speaker's pronunciation was pure, fluent, and natural.
[0136] The translated text of the target language corresponding to the test audio was annotated and translated by professional human translators without referring to any machine translation results. It was also inspected and scored by at least 3 professional translators, and the fidelity and fluency scores reached the set scores.
[0137] It does not contain content related to pornography, explosives, terrorism, or religious and political sensitivity.
[0138] Preferably, the hot word test set for each language direction meets the following requirements:
[0139] The hot word database includes no fewer than 500 hot words;
[0140] Each hot word provides a corresponding correct translation result, and distinguishes between uppercase and lowercase letters, as well as uppercase and lowercase letters, grammar, and spelling.
[0141] Preferably, the domain test set for each language direction meets the following requirements:
[0142] The effective duration of the test audio should be no less than 1 hour and no less than 1000 sentences;
[0143] Other requirements remain consistent with the simultaneous interpretation test set.
[0144] Preferably, the security test set for each language direction meets the following requirements:
[0145] The effective duration of the test audio for each security category shall not be less than 5 minutes or 50 sentences, and the total effective duration of the test audio for all security categories shall not be less than 40 minutes or 400 sentences.
[0146] Other requirements remain consistent with the simultaneous interpretation test set.
[0147] Preferably, the test audio in the test set is recorded using an audio sampling device;
[0148] The audio sampling device includes:
[0149] Recording equipment with an audio sampling rate of 8k~48k, a bit width of at least 16bit, and audio recording formats of aac, mp3, and wav;
[0150] Computers or mobile phones that support the installation and use of recording software;
[0151] Sound pressure meter, used for confirming ambient sound pressure.
[0152] Preferably, the test network environment meets the following requirements:
[0153] Uplink bandwidth is no less than 100 kbit / s and downlink bandwidth is no less than 8 Mbit / s, and a stable connection is maintained;
[0154] Ensure you are offline when testing the offline simultaneous transmission function;
[0155] The test scenario must meet the following requirements:
[0156] The ambient noise level is below 35 dB.
[0157] Fourthly, a testing device for a machine simultaneous interpretation system is provided, comprising:
[0158] The test set acquisition module is used to acquire the test set corresponding to the test item of the system under test. The test set corresponding to the test item includes at least one language direction test set, and the test set contains multiple test audios.
[0159] The test audio input module is used to input each test audio from the test set corresponding to the test item into the system under test;
[0160] The data acquisition module is used to acquire the running data of the system under test on the test item;
[0161] The test result determination module is used to determine the test result of the system under test on the test item based on the operating data of the system under test on the test item.
[0162] Fifthly, a test device for a machine simultaneous interpretation system is provided, including: a memory and a processor;
[0163] The memory is used to store programs;
[0164] The processor is used to execute the program to implement each step of the testing method for the machine simultaneous interpretation system described in any of the above claims.
[0165] In a sixth aspect, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the above-described machine simultaneous interpretation methods, or implements the steps of the above-described machine simultaneous interpretation system testing method.
[0166] In a seventh aspect, a computer program product is provided, the computer program product comprising a computer executable program of instructions which, when executed on a computer, perform the steps of the machine simultaneous interpretation method described above, or perform the steps of the test method for the machine simultaneous interpretation system described above.
[0167] By means of the above technical solution, the machine simultaneous interpretation system and method provided by this application provides the function of multimodal output, which may include at least one of the following: subtitle output of translated text, playback output of synthesized audio stream, and video output synthesized from audio stream and virtual human. After the system obtains the audio stream of the source language, it performs translation. This application puts forward requirements for the system translation process, namely, to ensure that the translated text meets the set translation effect. The translation effect includes the translation BLEU score, fidelity and fluency reaching the corresponding score thresholds. Accordingly, the translation effect can be guaranteed from three different perspectives.
[0168] When providing subtitle output functionality, the system further limits the display of translated text within a first set time after acquiring the source language audio stream. This ensures that the real-time translation from source language audio input to target language translated text output is within this first set time, preventing excessively long translation times from affecting the user's semantic understanding. Furthermore, the system also limits the average erasure rate of the translated text during output display to no more than a set average erasure rate threshold. This average erasure rate reflects the readability of the translation; a higher average erasure rate indicates a greater number of times the displayed translated text has been erased, resulting in poorer readability. By limiting the average erasure rate during the translation output process, the readability of the translated text displayed in subtitle format can be further guaranteed.
[0169] When providing playback and output functions for synthesized audio streams and virtual human synthesized video, further speech synthesis is performed on the translated text. The target language audio stream is required to meet set speech synthesis effects, including average sentence synthesis accuracy and synthesis naturalness reaching corresponding threshold scores, and the use of licensed voice libraries. This ensures the quality of the synthesized audio from three different perspectives: accuracy, naturalness, and voice library licensing. Simultaneously, the synthesized audio stream is played and output within a second set time after the source language audio stream is acquired. This limits the real-time translation from source language audio input to target language audio stream output to within this second set time, preventing excessively long translation times from affecting the user's semantic understanding.
[0170] As can be seen from the above, the machine simultaneous interpretation system and method provided in this application can take into account multiple different factors that affect the quality of simultaneous interpretation, and propose specific indicator requirements for each different influencing factor during the simultaneous interpretation process, so that the overall quality of the machine simultaneous interpretation system is more balanced and comprehensive, and more convenient for users to use.
[0171] Furthermore, this application also provides a testing method for a machine simultaneous interpretation system. By obtaining the test set corresponding to the test items of the system under test, inputting the test set into the system under test, and obtaining the operating data of the system under test on the test items, the test performance of the system under test on the test items can be determined based on the operating data. In practical applications, corresponding test items and test sets can be set according to the technical indicators of the machine simultaneous interpretation system to be evaluated, thereby enabling comprehensive testing of all technical indicators of the simultaneous interpretation system and ensuring the quality of the simultaneous interpretation system.
[0172] The machine simultaneous interpretation system, method, and related testing methods provided in this application can be applied to the design, development, application, and maintenance of machine simultaneous interpretation systems, and can also be used to guide third-party evaluation results in the functional assessment and system acceptance of machine simultaneous interpretation systems. Attached Figure Description
[0173] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0174] Figures 1-7 These are schematic diagrams of several different machine simultaneous interpretation systems provided in the embodiments of this application;
[0175] Figure 8 A schematic diagram illustrating the basic requirements of a machine simultaneous interpretation system as exemplified in this application;
[0176] Figure 9 A flowchart illustrating the machine simultaneous interpretation method provided in this application embodiment;
[0177] Figure 10 A schematic diagram of the test method for the machine simultaneous interpretation system provided in the embodiments of this application;
[0178] Figure 11 A schematic diagram of the test device structure for the machine simultaneous interpretation system provided in the embodiments of this application;
[0179] Figure 12 This is a schematic diagram of the test equipment structure for the machine simultaneous interpretation system provided in the embodiments of this application. Detailed Implementation
[0180] Before introducing the proposed solution, the terms and definitions that may be involved in this application will be explained.
[0181] 1. Voice recognition
[0182] Referring to the description in section 3.1 of the relevant national standard GB / T21023-2007, it refers to the process of converting human voice signals into text or instructions.
[0183] 2. Speech synthesis
[0184] Referring to the description in section 3.1 of the relevant national standard GB / T21024-2007, this refers to the process of synthesizing human language through mechanical and electronic methods. Note: The speech produced by this process is called synthesized speech, which is distinct from natural speech produced by a human vocal cord; it is sometimes also called artificial speech.
[0185] 3. Machine translation
[0186] The process or technique of converting one natural language (source language) into another natural language (target language).
[0187] 4. Machine simultaneous speech translation system
[0188] Development tools, software, and applications with simultaneous interpretation capabilities.
[0189] 5. Virtual Human Synthesis
[0190] The process of generating virtual character images, expressions, movements, and voices through digital image processing and speech synthesis technologies.
[0191] 6. Users
[0192] This refers to organizations or individuals that use machine simultaneous interpretation systems to solve their business problems.
[0193] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0194] Given the inconsistent quality of machine simultaneous interpretation systems currently on the market, especially in situations involving complex and varied backgrounds, diverse speaker expressions with colloquialisms and dialects, a scarcity of multilingual speech and translation training data, and insufficient coverage of specialized fields, the usability of different machine simultaneous interpretation systems varies. Furthermore, existing national standards for intelligent voice interaction systems only specify the terminology and definitions related to speech recognition and speech synthesis, without providing relevant standards and specifications for machine simultaneous interpretation systems. This further hinders the owners of existing machine simultaneous interpretation systems from referencing national standards for technical improvements, and they may not even be aware of the problems existing in their systems, nor have they identified clear directions and methods for technical improvement.
[0195] Through in-depth research, the applicant in this case discovered that the root cause of the inconsistent quality of existing machine simultaneous interpretation systems, and even their unavailability in certain scenarios, lies in the fact that the quality of machine simultaneous interpretation systems can be reflected in many different aspects, such as translation effectiveness measured from different angles, real-time translation, and speech synthesis effects. Existing simultaneous interpretation systems have not taken into account different quality influencing factors in a comprehensive manner, resulting in an uneven overall quality of the simultaneous interpretation system and affecting normal use by users.
[0196] The applicant in this case attempts to propose a more general framework and basic requirements for a machine simultaneous interpretation system, identify the various factors affecting the quality of the machine simultaneous interpretation system, and design corresponding indicator requirements, so as to make the overall quality of the machine simultaneous interpretation system and the corresponding machine simultaneous interpretation method more balanced and comprehensively improve the effect of machine simultaneous interpretation.
[0197] Next, we will first introduce the machine simultaneous interpretation system proposed in this application.
[0198] Please see Figures 1-6 The diagram illustrates a reference framework structure for a machine simultaneous interpretation system provided in an embodiment of this application, which may include: an application end and a server end.
[0199] The application may include an audio acquisition module 1 and a multimodal output module 2, wherein the multimodal output module 2 may include at least one of a subtitle display module 21, an audio playback module 22, and a video playback module 23.
[0200] The server includes a speech translation module 3. When the multimodal output module 2 includes an audio playback module 22, the server may also include a speech synthesis module 4. When the multimodal output module 2 includes a video playback module 23, the server may also include a speech synthesis module 4 and a virtual human module 5.
[0201] The functions of each module will be introduced next:
[0202] Audio acquisition module 1 is used to acquire the audio stream of the source language.
[0203] The speech translation module 3 is used to translate the audio stream of the source language to obtain the translated text of the target language that meets the set translation effect. The set translation effect includes the translation BLEU score, fidelity and fluency reaching corresponding score thresholds.
[0204] BLEU is an automatic evaluation method based on N-grams. It obtains a score for the quality of the machine translation by comparing the machine translation with the reference translation using N-grams. The specific calculation method is existing technology and can be found in the following formula:
[0205]
[0206]
[0207]
[0208]
[0209] in,
[0210] Max_Ref_Count represents the maximum number of times an n-gram appears in any reference translation;
[0211] Indicates the number of times the n-gram appears in the reference translation;
[0212] Represents: All sentences for which BLEU scores are to be calculated;
[0213] Indicates the number of times the n-gram appears in the machine translation;
[0214] Indicates the length of the machine-translated sentence;
[0215] Indicates the length of the reference translation sentence (if there are multiple reference translations, it indicates the length of the reference translation sentence that is closest to the machine translation sentence).
[0216] Indicates: a length penalty for the translation result;
[0217] This indicates the weights for different N-grams; generally, BLEU calculations use 1-4 grams. Take 0.25.
[0218] In this embodiment, after the speech translation module 3 translates the audio stream of the source language, the BLEU score of the translated text in the target language must reach a set score threshold, such as a BLEU score greater than or equal to 30 or other values.
[0219] Fidelity reflects the degree to which a translated text faithfully conveys the content of the original text. The fidelity of a translated text can be calculated using a pre-trained fidelity scoring model or by manually assigning a fidelity score based on established criteria. Table 1 below illustrates a machine translation fidelity scoring criterion:
[0220] Table 1
[0221]
[0222] In this embodiment, after the speech translation module 3 translates the audio stream of the source language, the fidelity score of the translated text in the target language must reach a set threshold, such as a fidelity score greater than or equal to 4 or other values.
[0223] Fluency reflects the smoothness and naturalness of a translated text. The fluency of a translated text can be calculated using a pre-trained fluency scoring model, or it can be scored manually using fluency scoring criteria. Table 2 below illustrates a machine translation fluency scoring criterion:
[0224] Table 2
[0225]
[0226] In this embodiment, after the speech translation module 3 translates the audio stream of the source language, the fluency score of the translated text in the target language must reach a set threshold, such as a fluency score greater than or equal to 4 or other values.
[0227] In one optional case, if the BLEU score in the translation effect reaches 30, and the fidelity score is ≥4 and the fluency score is ≥4, then the translation effect of the machine simultaneous interpretation system can be set to meet the second-level qualified standard.
[0228] If the BLEU score in the translation effect reaches 35, and the fidelity score is ≥4.2 and the fluency score is ≥4.2, then the translation effect of the machine simultaneous interpretation system can be set to reach the first-level excellent standard.
[0229] The subtitle display module 21 is used to output and display the translated text of the target language within a first set time period after the audio acquisition module 1 acquires the audio stream of the source language, and during the process of outputting and displaying the translated text, the average erasure rate of the translated text does not exceed the set average erasure rate threshold.
[0230] Specifically, in this embodiment, the time difference between the input of the source language audio stream and the output of the target language translated text stream is defined as the first translation real-time indicator of the machine simultaneous interpretation system. This first translation real-time indicator can also be called the S2T EVS indicator (Speech-to-Text Ear-voice Span).
[0231] In this embodiment, to meet the requirements of the first real-time translation indicator, the subtitle display module 21 displays the translated text within a first set time period after the audio acquisition module 1 acquires the audio stream of the source language. The first set time period can be 3 seconds or other values.
[0232] In one optional scenario, if the first translation real-time performance index is ≤3 seconds, the first translation real-time performance of the machine simultaneous interpretation system can be set to meet the second-level qualification standard.
[0233] If the first translation real-time performance index is ≤2 seconds, then the first translation real-time performance of the machine simultaneous interpretation system can be set to reach the first-level excellent standard.
[0234] Furthermore, this embodiment also defines the average erasure rate of the translated text during the output and display process of the subtitle display module, using this rate as a translation readability indicator for the machine simultaneous interpretation system. This translation readability indicator can be expressed as NE (Normalized Erasure), representing the proportion of text that has been erased during the subtitle display process. The specific definition is as follows:
[0235]
[0236]
[0237] in, This represents the length of the text display sequence at the i-th refresh time. The length of the longest common prefix of the two strings is represented by I, and I represents the number of times the machine-translated text corresponding to an input audio changes during the final determination process.
[0238] In this embodiment, during the process of the subtitle display module outputting and displaying the translated text, the average erasure rate of the translated text must reach a set score threshold, such as an average erasure rate less than or equal to 1 or other values.
[0239] In one optional scenario, if the translation readability index NE≤1, the translation readability of the machine simultaneous interpretation system can be set to meet the Level 2 qualification standard.
[0240] If the translation readability index NE≤0.6, the translation readability of the machine simultaneous interpretation system can be set to reach the first-level excellent standard.
[0241] The speech synthesis module 4 is used to perform speech synthesis on the translated text of the target language to obtain an audio stream of the target language that meets the set speech synthesis effect. The set speech synthesis effect includes the average sentence synthesis accuracy and the synthesis naturalness reaching corresponding score thresholds, and the sound library used during synthesis being an authorized sound library.
[0242] Specifically, to ensure the speech synthesis effect of speech synthesis module 4, the sound library used for synthesis must be an authorized sound library. At the same time, the average sentence synthesis accuracy and the naturalness of the synthesized speech must each reach a corresponding score threshold.
[0243] The average sentence synthesis accuracy rate represents the average speech synthesis accuracy rate of the system, which can be calculated as the ratio of the number of correctly synthesized audio segments to the total number of synthesized audio segments.
[0244] In this embodiment, the average sentence synthesis accuracy of the target language audio stream synthesized by the speech synthesis module can be set to greater than or equal to 90% or other values.
[0245] Synthesis naturalness refers to the auditory perception of whether the synthesized audio is fluent, emotional, and has natural pauses, and provides an overall naturalness score. The Mean Opinion Score (MOS) can be used as an indicator of synthesis naturalness. The scoring criteria for the MOS score are shown in Table 3 below:
[0246] Table 3
[0247]
[0248] In this embodiment, the synthesis naturalness score (MOS score) of the audio stream of the target language synthesized by the speech synthesis module can be set to greater than or equal to 4 or other values.
[0249] In one optional scenario, if the average sentence synthesis accuracy is ≥90% and the synthesis naturalness score (MOS) is ≥4, the language synthesis effect of the machine simultaneous interpretation system can be set to meet the Level 2 qualification standard.
[0250] If the average sentence synthesis accuracy is ≥95% and the synthesis naturalness score (MOS) is ≥4.5, then the language synthesis effect of the machine simultaneous interpretation system can be set to reach the first-level excellent standard.
[0251] The audio playback module 22 is used to play the audio stream of the target language within a second set time period after the audio acquisition module 1 acquires the audio stream of the source language.
[0252] Specifically, in this embodiment, the time difference between the input of the source language audio stream and the output of the target language audio stream is defined as the second translation real-time indicator of the machine simultaneous interpretation system. This second translation real-time indicator can also be called the S2S EVS indicator (Speech-to-Speech Ear-voice Span).
[0253] In this embodiment, in order to meet the requirements of the second translation real-time performance index, the audio playback module 22 plays the target language audio stream within a second set duration after the audio acquisition module 1 acquires the source language audio stream. The second set duration can be 5 seconds or other values.
[0254] In one optional scenario, if the second translation real-time performance index is ≤5 seconds, the second translation real-time performance of the machine simultaneous interpretation system can be set to meet the Level 2 qualification standard.
[0255] If the first translation real-time performance index is ≤4 seconds, then the second translation real-time performance of the machine simultaneous interpretation system can be set to reach the first-class excellent standard.
[0256] Virtual Human Module 5 is used to combine the audio stream of the target language with the virtual avatar and convert it into a video stream.
[0257] The video playback module 23 is used to play video streams.
[0258] Optionally, the video playback module 23 can be further limited to play the video stream in real time. For example, the video playback module can be controlled to play and output the video stream within a third set time after the audio acquisition module 1 obtains the audio stream of the source language.
[0259] Combination Figures 1-6 As can be seen, the machine simultaneous interpretation system provided in this application embodiment can have various different structures, and the corresponding machine simultaneous interpretation system can realize different output modes, including: outputting only the translated text (corresponding to...). Figure 4 Only output synthesized audio stream (corresponding to) Figure 5 Only output the synthesized video stream (corresponding to) Figure 6 Simultaneously output translated text and synthesized audio stream (corresponding to...) Figure 2 Simultaneously output translated text and synthesized video stream (corresponding to...) Figure 3 Simultaneously output translated text, synthesized audio stream, and synthesized video stream (corresponding to...). Figure 1 The structure of a machine simultaneous interpretation system varies depending on the output modality it implements; see details below. Figures 1-6 As shown.
[0260] The machine simultaneous interpretation system provided in this application provides multimodal output functionality, which may include at least one of a subtitle display module, an audio playback module, and a video playback module. After the system acquires the audio stream of the source language, it performs translation. This application sets requirements for the speech translation module, namely, ensuring that the translated text meets the set translation effect. The translation effect includes the translation BLEU score, fidelity, and fluency each reaching the corresponding score threshold. Accordingly, the translation effect can be guaranteed from three different perspectives.
[0261] When providing subtitle output functionality, the subtitle display module further limits the output of translated text to a first set time period after the audio acquisition module acquires the source language audio stream. This limits the real-time translation from source language audio input to target language translated text output to within this first set time period, preventing excessively long translation times from affecting the user's semantic understanding. Furthermore, the average erasure rate of the translated text output by the subtitle display module is also limited to a set average erasure rate threshold. This average erasure rate reflects the readability of the translation; a higher average erasure rate indicates a greater number of times the displayed translated text has been erased, resulting in poorer readability. By limiting the average erasure rate during the translation output process, the readability of the translated text displayed in subtitle format can be further guaranteed.
[0262] When providing playback output functions for synthesized audio streams and virtual human synthesized video streams, the system further synthesizes the translated text using a speech synthesis module. The target language audio stream is required to meet predefined speech synthesis effects, including average sentence synthesis accuracy and naturalness reaching corresponding threshold scores, and the use of an authorized voice library. This ensures the quality of the synthesized audio from three perspectives: accuracy, naturalness, and voice library authorization. Furthermore, the system limits the playback output of the synthesized audio stream to a second predefined timeframe after the audio acquisition module receives the source language audio stream. This limits the real-time translation from source language audio input to target language audio stream output to within this second predefined timeframe, preventing excessively long translation times from affecting user semantic understanding. Moreover, the system can further limit the playback output of the synthesized video stream to a third predefined timeframe after the audio acquisition module receives the source language audio stream. This limits the real-time translation from source language audio input to synthesized video stream output to within this third predefined timeframe, preventing excessively long translation times from affecting user semantic understanding.
[0263] As can be seen from the above, the machine simultaneous interpretation system and method provided in this application can take into account multiple different factors that affect the quality of simultaneous interpretation, and propose specific indicator requirements for each different influencing factor during the simultaneous interpretation process, so that the overall quality of the machine simultaneous interpretation system is more balanced and comprehensive, and more convenient for users to use.
[0264] In some embodiments of this application, the functional requirements of the machine simultaneous interpretation system are further defined.
[0265] 1. Regarding effect optimization
[0266] 1.1 Hot word intervention function
[0267] Considering that machine simultaneous interpretation systems may experience low recognition and translation accuracy for words that are difficult to recognize and translate, such as named entities, technical terms, and new words, this embodiment can set up a hot word library. The collected named entities, technical terms, and new words that are difficult to recognize and translate can be added to the hot word library as hot words. The speech translation module 3 can support adding the hot word library to optimize the speech recognition and translation process by intervening with hot words, thereby improving the accuracy of hot word recognition and translation. Furthermore, in this embodiment, after adding the hot word library, the translation accuracy of the speech translation module 3 for hot words is not lower than a set hot word intervention efficiency threshold, where the hot word intervention efficiency threshold can be not less than 80% or other values. The calculation process for the hot word intervention efficiency can be expressed as follows:
[0268]
[0269] Where P1 is the number of correctly translated hot words in the translated text obtained by the speech translation module 3 before adding the hot word database, P2 is the number of correctly translated hot words in the translated text obtained by the speech translation module 3 after adding the hot word database, and T is the total number of hot words.
[0270] The speech translation module 3 provided in this embodiment supports adding a hot word dictionary, and the translation accuracy of hot words after adding the hot word dictionary is not lower than the set threshold, thus ensuring that the machine simultaneous interpretation system supports the hot word intervention function.
[0271] It should be noted that the simultaneous interpretation system in this embodiment can support hot word intervention between set languages, for example, supporting hot word intervention in Chinese-to-English and English-to-Chinese simultaneous interpretation. Of course, hot word intervention can also be supported across all languages. The hot word intervention efficiency threshold for hot word intervention between different languages can be set to be the same or different.
[0272] 1.2 Domain-Customized Model Function
[0273] Considering that machine simultaneous interpretation systems may be applied in multiple different fields, such as medical conferences, in order to improve the speech recognition and translation effects of machine simultaneous interpretation systems in specific fields, the speech translation module 3 in the machine simultaneous interpretation system can also support loading domain-customized models. Domain-customized models refer to the customization of speech recognition and text translation models according to different conference types and different speaker characteristics, thereby improving the translation effect of machine simultaneous interpretation systems in corresponding field scenarios.
[0274] Furthermore, in this embodiment, after loading the domain-customized model, the speech translation module 3 improves the translation BLEU score on the corresponding domain test set by no less than a set BLEU score difference threshold, which can be set to 2 or other values.
[0275] The speech translation module 3 provided in this embodiment supports loading domain-customized models, and after loading the domain-customized model, the translation BLEU score of the input audio stream in the corresponding domain is improved by no less than the set BLEU score difference threshold, thus ensuring that the machine simultaneous interpretation system supports the function of loading domain-customized models.
[0276] 1.3 Human-machine collaborative intervention function
[0277] In certain use cases, to improve the effectiveness of simultaneous interpretation, it may be necessary to allow human interpreters to edit the speech recognition and / or text translation results during the machine simultaneous interpretation system's operation. This includes editing operations such as modification, optimization, and deletion, thereby improving the final translation results through human-machine collaboration and human-assisted machine translation.
[0278] To achieve this human-machine collaborative intervention function, the machine simultaneous interpretation system provided in this embodiment may further include the following on the application side:
[0279] The speech recognition result editing module is used to respond to the human translator's operation of editing the recognition results of the audio stream and obtain the edited source language text;
[0280] The translation text editing module is used to respond to human translators' editing operations on the translated target language text, and to obtain the edited target language translation text.
[0281] 2. Regarding content management
[0282] 2.1 Meeting Export
[0283] To facilitate users' review and editing of meeting content after using the machine simultaneous interpretation system, this embodiment of the machine simultaneous interpretation system can also be equipped with a meeting data export module on the application side, enabling meeting export functionality. Meeting export refers to the ability to export the speech recognition, machine translation, and speech synthesis results from the entire meeting service after the meeting concludes.
[0284] Specifically, the meeting data export module is used to respond to the user's operation of exporting specified meeting data to a specified storage location, and to export the specified meeting data to the specified storage location. The specified meeting data includes any one or more of the recognition results, translation results, and speech synthesis results of the audio stream of the source language.
[0285] 2.2 Offline simultaneous interpretation
[0286] In this embodiment of the machine simultaneous interpretation system, the application and server sides can be deployed separately. For example, the application side can be deployed locally, and the server side can be deployed on the network side, with the application and server sides connected via the network. Alternatively, if the machine simultaneous interpretation system needs to support offline simultaneous interpretation, both the application and server sides can be deployed locally, thus supporting offline simultaneous interpretation when the system is not connected to the internet.
[0287] 2.3 Content Security
[0288] To ensure the security of data content in the machine simultaneous interpretation system, this embodiment can also add a data preprocessing module to the server side. This module performs content security checks on the translated text of the target language output by the speech translation module 3. If the content security check fails, the subsequent processing is stopped. The content security checks include: content related to bias and discrimination, sensitive topics, physical harm, mental health, privacy and property, and ethics.
[0289] In some embodiments of this application, the speech translation module 3 can have different structural configurations. For example, the speech translation module 3 can adopt an end-to-end form, that is, directly modeling the relationship between the input source language audio stream and the output target language translated text. In addition, embodiments of this application also provide another structural form of the speech translation module 3, see reference... Figure 7 As shown, the speech translation module 3 may include a speech recognition module 31 and a machine translation module 32.
[0290] The speech recognition module 31 is used to recognize the audio stream of the source language and obtain the source language text.
[0291] The machine translation module 32 is used to translate source language text into target language to obtain translated text in the target language.
[0292] Figure 7 The example speech translation module 3 models the process of translating the source language audio stream into the target language text in two stages: the first stage is speech recognition (corresponding to speech recognition module 31), and the second stage is text translation (corresponding to machine translation module 32). Different models can be used for the two stages, such as a speech recognition model for the first stage and a text translation model for the second stage.
[0293] Further reference Figure 7 As shown, the speech translation module 3 may also include a post-recognition processing module 33, which is used to normalize the source language text output by the speech recognition module 31 and send the normalized source language text to the machine translation module 32.
[0294] The post-processing module 33 performs various standardization processes on the source language text, including but not limited to: converting the source language text to uppercase and lowercase, adding punctuation, and segmenting it into sentences. By adding the post-processing module, the source language text input to the machine translation module can be formatted more regularly and semantically more fluent, thereby further improving the translation quality of the machine translation module.
[0295] Furthermore, the subtitle display module 21, while displaying the translated text in the target language, can also simultaneously display the source language text, thereby facilitating user comparison between the source language text and the translated text in the target language. Here, the source language text displayed by the subtitle display module 21 can be the source language text recognized by the speech recognition module 31, or the source language text that has been normalized by the post-recognition processing module 33. The specific display strategy can be set as needed.
[0296] Reference Figure 7 As shown, the server may further include a post-translation processing module 6, which is used to normalize the translated text of the target language output by the speech translation module 3, and send the normalized translated text of the target language to the speech synthesis module 4.
[0297] The post-translation processing module 6 performs regularization processing on the translated text, including but not limited to: operations such as length transformation of the translated text. By regularizing the translated text, it is easier for the speech synthesis module 4 to synthesize speech from the translated text.
[0298] Reference Figure 7 As shown, the application may further include an acoustic noise reduction module 7, which is used to perform noise reduction processing on the audio stream of the source language acquired by the audio acquisition module 1, and send the noise-reduced audio stream to the speech translation module 3.
[0299] Considering that background noise may exist in the actual application scenario of machine simultaneous interpretation system, in order to filter out the impact of background noise on the simultaneous interpretation effect, the above-mentioned acoustic noise reduction module 7 can be added to the application end to help reduce the background noise in the acquired audio stream and improve the signal-to-noise ratio of the source.
[0300] The above embodiments of this application propose a relatively general framework and basic requirements for a machine simultaneous interpretation system. They identify various factors affecting the quality of the machine simultaneous interpretation system and design corresponding performance indicators, resulting in a more balanced overall quality of the machine simultaneous interpretation system provided by this application. The basic requirements for the machine simultaneous interpretation system outlined in the aforementioned embodiments can be referred to... Figure 8 As shown, here is a brief summary again; for details, please refer to the previous related introduction:
[0301] The basic requirements for simultaneous interpretation systems can be divided into core technology requirements and system function requirements.
[0302] The core technology requirements can be divided into two stages: the process requirements for translating source language audio into target language text and the process requirements for translating source language audio into target language audio.
[0303] The process of translating source language audio into target language text can further include three requirements: translation quality, first-time translation performance, and translation readability. Translation quality can include BLEU score, fidelity, and fluency requirements; first-time translation performance can include S2T EVS requirements; and translation readability can include the system average erase rate (NE) requirement.
[0304] The requirements for the source language audio to target language audio process can further include three indicators: translation quality, real-time second translation, and speech synthesis quality. The translation quality requirements are the same as those in the aforementioned requirements for the source language audio to target language text translation process; real-time second translation can include S2S EVS indicator requirements; and speech synthesis quality can include three indicators: audio library licensing, average sentence synthesis accuracy, and naturalness of speech synthesis.
[0305] The system's functional requirements can be further divided into two parts: effect optimization and content management.
[0306] The effect optimization part can include three indicator requirements: hot word intervention, domain-customized model, and human-machine collaborative intervention; the content management part can include three indicator requirements: meeting export, offline simultaneous interpretation, and content security.
[0307] Corresponding to the machine simultaneous interpretation system provided in the foregoing embodiments, this application also provides a machine simultaneous interpretation method for implementing at least one of the following simultaneous interpretation processes: simultaneous interpretation from source language audio to target language text, simultaneous interpretation from source language audio to target language audio, and simultaneous interpretation from source language audio to target language video. The basic requirements for simultaneous interpretation systems described above also apply to the machine simultaneous interpretation method of this embodiment. The following will be discussed in conjunction with... Figure 9 This section introduces machine simultaneous interpretation methods:
[0308] The simultaneous interpretation process between source language audio and target language text may include:
[0309] Step S100: Obtain the audio stream of the source language.
[0310] Step S110: Translate the audio stream of the source language to obtain the translated text of the target language that meets the set translation effect. The set translation effect includes the translation BLEU score, fidelity and fluency each reaching the corresponding score threshold.
[0311] Specifically, this step defines the translation effect when translating the audio stream of the source language. The translation BLEU score, fidelity, and fluency all need to reach corresponding score thresholds. This translation effect serves as the target direction for translating the audio stream of the source language, and ultimately, a translated text in the target language that meets the set translation effect can be obtained.
[0312] The concepts related to translation BLEU score, fidelity, and fluency can be found in the previous system section's implementation examples, and will not be repeated here.
[0313] Step S120a: Within a first set time period after obtaining the audio stream of the source language, the translated text of the target language is output and displayed, and during the process of outputting and displaying the translated text, the average erasure rate of the translated text does not exceed the set average erasure rate threshold.
[0314] Specifically, to ensure the real-time requirement of the first translation from source language audio input to target language text output, this step specifies that the time difference between acquiring the source language audio stream and outputting the target language translated text does not exceed a first set duration. Simultaneously, to ensure translation readability, it is specified that the average erasure rate of the translated text during output display does not exceed a set average erasure rate threshold.
[0315] The concepts related to the real-time nature of the first translation and the readability of the translation can be referred to the embodiments described in the previous system section, and will not be repeated here.
[0316] The machine simultaneous interpretation method provided in this application imposes requirements on the system's translation process, namely, ensuring that the translated text meets the set translation effect. This effect includes the translation BLEU score, fidelity, and fluency each reaching corresponding threshold scores, thus guaranteeing the translation effect from three different perspectives. When providing the subtitle output function for the translated text, the method further limits the output display of the translated text within a first set time after acquiring the source language audio stream. This limits the first translation real-time period from source language audio input to target language translated text output to within the first set time, preventing excessively long first translation real-time periods from affecting the user's semantic understanding. Furthermore, the method also limits the average erasure rate of the translated text output display process to not exceed a set average erasure rate threshold. This average erasure rate reflects the readability of the translation; a higher average erasure rate indicates a higher number of times the displayed translated text is erased again, resulting in poorer readability. By limiting the average erasure rate during the translated text output process, the readability of the translated text displayed in subtitle form can be further guaranteed.
[0317] Further reference Figure 9 As shown, the simultaneous interpretation process between source language audio and target language audio can include:
[0318] Step S100: Obtain the audio stream of the source language.
[0319] Step S110: Translate the audio stream of the source language to obtain the translated text of the target language that meets the set translation effect. The set translation effect includes the translation BLEU score, fidelity and fluency reaching the corresponding score threshold.
[0320] Step S120b: Perform speech synthesis on the translated text of the target language to obtain an audio stream of the target language that meets the set speech synthesis effect. The set speech synthesis effect includes the average sentence synthesis accuracy and the synthesis naturalness reaching corresponding score thresholds, and the sound library used during synthesis is an authorized sound library.
[0321] Specifically, this step defines the speech synthesis effect when synthesizing the translated text of the target language. This includes ensuring that the average sentence synthesis accuracy and the naturalness of the synthesis reach corresponding score thresholds, and that the sound library used for synthesis is an authorized sound library. Using this speech synthesis effect as the target direction for speech synthesis of the translated text of the target language, an audio stream of the target language that meets the defined speech synthesis effect can ultimately be obtained.
[0322] The concepts of average sentence synthesis accuracy and synthesis naturalness can be found in the previous system item implementation examples, and will not be repeated here.
[0323] Step S130b: Play the audio stream of the target language within a second set time period after obtaining the audio stream of the source language.
[0324] Specifically, in order to ensure the real-time requirement of the second translation from the source language audio input to the target language audio output, this step specifies that the time difference between acquiring the source language audio stream and outputting the target language audio stream shall not exceed a second set duration.
[0325] The relevant concepts regarding the real-time nature of the second translation can be found in the embodiments described in the previous system section, and will not be repeated here.
[0326] The machine simultaneous interpretation method provided in this application ensures the translation effect of the translated text from three different perspectives: BLEU score, fidelity, and fluency. Furthermore, while providing the playback output function of the synthesized audio stream, the method further performs speech synthesis on the translated text and limits the synthesized target language audio stream to meet set speech synthesis effects. These effects include average sentence synthesis accuracy and synthesis naturalness each reaching corresponding score thresholds, and the use of an authorized voice library during synthesis. Therefore, the effect of the synthesized audio can be guaranteed from three different perspectives: accuracy, naturalness, and voice library authorization. Additionally, the method further limits the playback output of the synthesized audio stream within a second set time after acquiring the source language audio stream. This means that the second translation real-time performance between the source language audio input and the target language audio stream output is limited to a second set time, preventing excessively long second translation real-time performance from affecting the user's semantic understanding.
[0327] Further reference Figure 9 As shown, the simultaneous interpretation process between source language audio and target language video can include:
[0328] Step S100: Obtain the audio stream of the source language.
[0329] Step S110: Translate the audio stream of the source language to obtain the translated text of the target language that meets the set translation effect. The set translation effect includes the translation BLEU score, fidelity and fluency reaching the corresponding score threshold.
[0330] Step S120c: Perform speech synthesis on the translated text of the target language to obtain an audio stream of the target language that meets the set speech synthesis effect. The set speech synthesis effect includes the average sentence synthesis accuracy and the synthesis naturalness reaching corresponding score thresholds, and the sound library used during synthesis being an authorized sound library.
[0331] Step S120c is the same as step S120b described above, and is explained in detail in the preceding text.
[0332] Step S130c: Combine the audio stream of the target language with the virtual avatar, convert it into a video stream, and play the video stream.
[0333] Specifically, in order to improve the user experience of machine simultaneous interpretation, this embodiment combines a virtual avatar with the audio stream of the target language to obtain a synthesized video stream. By playing the synthesized video stream, users can understand the simultaneous interpretation content from multiple audio-visual perspectives, thereby improving the user experience.
[0334] As can be seen from the machine simultaneous interpretation methods described in the above embodiments of this application, the machine simultaneous interpretation method provided by this application can take into account multiple different factors that affect the quality of simultaneous interpretation, and proposes specific indicator requirements for each different influencing factor during the simultaneous interpretation process, so that the overall quality of the final machine simultaneous interpretation process is more balanced and comprehensive, and more convenient for users to use.
[0335] In an optional case, step S110, the process of translating the audio stream of the source language to obtain the translated text of the target language that meets the set translation effect, may include the following steps:
[0336] S1. Recognize the audio stream of the source language to obtain the source language text.
[0337] S2. Translate the source language text into the target language to obtain a translated text in the target language that meets the set translation effect.
[0338] Specifically, S1, the process of recognizing the audio stream, can be obtained using a pre-trained speech recognition model. S2, the process of translating the source language text, can be obtained using a pre-trained text translation model.
[0339] Of course, in addition to this, step S110 can also be obtained using an end-to-end model, such as an end-to-end model that uses a pre-trained source language audio stream as input and the target language translated text as output.
[0340] In some embodiments of this application, the machine simultaneous interpretation system used in the machine simultaneous interpretation method can also support hot word intervention functionality. Based on this, step S110 above, which translates the audio stream of the source language to obtain translated text in the target language that meets the set translation effect, may specifically include:
[0341] A machine simultaneous interpretation system with an added hot word lexicon is used to translate the audio stream of the source language to obtain a translated text in the target language that meets the set translation effect.
[0342] Among them, compared with machine simultaneous interpretation systems without hot word databases, machine simultaneous interpretation systems with hot word databases have a translation accuracy rate of hot words that is no less than the set hot word intervention efficiency threshold.
[0343] Specifically, the concepts related to the hot word intervention function can be referred to the previous system item implementation example, and will not be repeated here.
[0344] The machine simultaneous interpretation method provided in this embodiment can improve the simultaneous interpretation effect of hot words by adding a hot word lexicon to the machine simultaneous interpretation system in advance if it can be determined in advance that hot words may be involved in the simultaneous interpretation process in certain application scenarios.
[0345] In some embodiments of this application, the machine simultaneous interpretation system used in the machine simultaneous interpretation method can also support loading domain-customized models. Based on this, step S110 above, which translates the audio stream of the source language to obtain translated text in the target language that meets the set translation effect, may specifically include:
[0346] A machine simultaneous interpretation system with a customized model for the target domain is used to translate the audio stream of the source language to obtain translated text in the target language that meets the set translation effect.
[0347] The target domain customized model is a speech recognition and text translation model in the target domain to which the audio stream of the source language belongs. Compared with the machine simultaneous interpretation system without the target domain customized model, the machine simultaneous interpretation system with the target domain customized model has a translation BLEU score improvement of no less than the set BLEU score difference threshold on the domain test set. The domain test set includes multiple test audios in the target domain.
[0348] Specifically, the concepts related to domain-specific customized models can be referred to in the previous system item implementation examples, and will not be repeated here.
[0349] The machine simultaneous interpretation method provided in this embodiment can be applied in a target domain scenario by loading a pre-trained target domain customized model into the machine simultaneous interpretation system, and then using the machine simultaneous interpretation system loaded with the target domain customized model to perform the simultaneous interpretation process, thereby improving the simultaneous interpretation effect in the target domain scenario.
[0350] In some embodiments of this application, the aforementioned machine simultaneous interpretation method can be implemented in a networked state or in a disconnected state, corresponding to different deployment methods of the machine simultaneous interpretation system.
[0351] In some embodiments of this application, the machine simultaneous interpretation method can also support a conference export function. Specifically, the conference export process may include:
[0352] In response to a user's operation to export specified meeting data to a specified storage location, the specified meeting data is exported to the specified storage location. The specified meeting data includes any one or more of the recognition results, translation results, and speech synthesis results of the source language audio stream.
[0353] The meeting export function allows users to easily view and edit meeting content after using the machine simultaneous interpretation method described in this application.
[0354] In some embodiments of this application, the machine simultaneous interpretation method can also support human-machine collaborative intervention. Specifically, based on the aforementioned machine simultaneous interpretation method, it may further include:
[0355] The system responds to human translators' editing of the audio stream's recognition results, resulting in edited source language text.
[0356] And / or,
[0357] The system responds to human translators' editing of the translated text in the target language, resulting in an edited translated text in the target language.
[0358] The machine simultaneous interpretation method in this embodiment adds support for human-machine collaborative intervention. In certain scenarios where the requirements for simultaneous interpretation effect are high, it supports human interpreters to edit the speech recognition and / or text translation results, including editing operations such as modification, optimization, and deletion. Through human-machine collaboration and human-assisted machine translation, the final translation result is improved.
[0359] In some embodiments of this application, the machine simultaneous interpretation method can also support content security detection functionality. Specifically, to ensure that the final output of the machine simultaneous interpretation method meets data security requirements, the method may further include:
[0360] The translated text in the target language is subjected to content security testing. If the content security test fails, the subsequent processing is stopped.
[0361] The content security testing items can be found in the previous introduction, and will not be repeated here.
[0362] In some embodiments of this application, considering that background noise may exist in practical application scenarios of machine simultaneous interpretation systems, in order to filter out the impact of background noise on the simultaneous interpretation effect, a noise reduction step can be added to the audio stream of the source language before recognizing and translating it. This helps to reduce background noise in the acquired audio stream and improve the signal-to-noise ratio of the source.
[0363] In some embodiments of this application, a standardization step can be added to the source language text before translation. Standardization includes, but is not limited to, operations such as case conversion, adding punctuation, and sentence segmentation. By adding this standardization process, the source language text can be formatted more regularly and its meaning more fluent, thereby further improving the translation quality during subsequent translation.
[0364] In some embodiments of this application, the machine simultaneous interpretation method can output and display the translated text of the target language while simultaneously outputting and displaying the source language text after language recognition. The source language text can be the source language text obtained after recognizing the audio stream of the source language, or it can be the source language text after normalization processing.
[0365] In some embodiments of this application, before performing speech synthesis on the translated text of the target language, a further step of normalizing the translated text can be added. Normalization includes, but is not limited to, operations such as length transformation of the translated text. Normalizing the translated text facilitates speech synthesis in subsequent steps and improves the speech synthesis effect.
[0366] Given that there is currently no method for evaluating machine simultaneous interpretation systems, the applicant in this application seeks to propose a more universal and unified testing method. The following examples will illustrate the testing method for the machine simultaneous interpretation system proposed in this application.
[0367] Please see Figure 10 The diagram illustrates a flowchart of a testing method for a machine simultaneous interpretation system provided in an embodiment of this application, which may include:
[0368] Step S200: Obtain the test set corresponding to the test item of the system under test, wherein the test set corresponding to the test item includes at least one language direction test set, and the test set contains multiple test audios.
[0369] In this embodiment, the system under test is the machine simultaneous interpretation system to be tested, and the test items of the system under test are the items that need to be tested for the machine simultaneous interpretation system.
[0370] The system under test can perform simultaneous interpretation in more than one language direction. Therefore, the test set corresponding to this test item includes at least one language direction test set, and the test set contains multiple test audios.
[0371] Optionally, the test audio in the test set corresponding to the test items of the system under test can be acquired using an audio sampling device. During testing, the aforementioned audio sampling device can also be used to determine whether the test environment meets the test requirements. The parameter requirements for the audio sampling device are shown in the table below:
[0372] Table 4
[0373]
[0374] Step S210: Input each test audio from the test set corresponding to the test item into the system under test.
[0375] Step S220: Obtain the operating data of the system under test on the test item.
[0376] Specifically, depending on the test item, the operational data of the system under test (SUT) for each test item may include the SUT's operational results for the input test audio, and may also include data characterizing the SUT's operational status and / or resource usage. For example, if the test item includes a translation effect test, the corresponding operational data may include the translated text generated by the SUT in response to the input test audio. As another example, if the test item includes first-level translation real-time performance, the corresponding operational data may include the time interval information between the SUT acquiring the input test audio and displaying the translated text.
[0377] Step S230: Determine the test result of the system under test on the test item based on the operating data of the system under test on the test item.
[0378] Among them, the test results of the system under test on the test items can characterize the test status of the system under test on the test items, such as whether the test passed or failed.
[0379] The testing method for a machine simultaneous interpretation system provided in this application involves obtaining a test set corresponding to the test items of the system under test, inputting the test set into the system under test to obtain the system's operating data on the test items, and then determining the system's test performance on the test items based on the operating data. In practical applications, corresponding test items and test sets can be set according to the technical indicators of the machine simultaneous interpretation system to be examined, thereby enabling comprehensive testing of all technical indicators of the simultaneous interpretation system and ensuring the quality of the simultaneous interpretation system.
[0380] The testing method for the machine simultaneous interpretation system provided in this application can be used to guide third-party evaluation results in the functional assessment and system acceptance of the machine simultaneous interpretation system.
[0381] It should be noted that the test environment needs to be configured before testing the system under test. This includes:
[0382] Configure the noise level for the test scenario and configure the test network environment.
[0383] Since simultaneous interpretation is generally used in formal conferences such as international conferences, the test scenario should be a non-high-noise environment, that is, the ambient noise should be below 35dB, to ensure that the test scenario reproduces the noise level in the real scene as much as possible, so as to obtain test results that can more realistically reflect the capabilities of the system under test.
[0384] When testing the non-offline simultaneous transmission function of the system under test, the required Internet service should be provided. The network conditions should meet the requirements of uplink bandwidth of not less than 100kbit / s and downlink bandwidth of not less than 8Mbit / s, and should maintain a stable connection.
[0385] When testing the offline simultaneous transmission function of the system under test, it should be ensured that the system is disconnected from the network.
[0386] After conducting a thorough study of the functions and performance of the machine simultaneous interpretation system, the applicant in this case also proposed general test items for the machine simultaneous interpretation system. These test items can be flexibly selected by the tester according to the specific test requirements of the system under test.
[0387] In one alternative scenario, if the core technical requirements of the system under test are to be tested, the corresponding combination of test items could include:
[0388] A first test item is used to test the performance of the system under test from source language audio stream input to target language translated text output, and a second test item is used to test the performance of the system under test from source language audio stream input to target language audio stream output.
[0389] The test sets corresponding to both the first and second test items include the simultaneous interpretation test set.
[0390] The first test item includes a translation performance test, a first real-time translation test, and a translation readability test. The second test item includes a translation performance test, a second real-time translation test, and a speech synthesis performance test.
[0391] The translation performance test can include BLEU subtests, fidelity tests, and fluency tests. BLEU subtests are used to evaluate the translation quality of the tested system, fidelity tests are used to evaluate whether the translated text obtained by the tested system faithfully expresses the content of the original text, and fluency tests are used to evaluate whether the translated text obtained by the tested system is fluent and conforms to the expression habits of the target language.
[0392] The first translation real-time test item is used to test the time difference between the input test audio and the display of the translated text on the screen of the system under test.
[0393] The translation readability test item is used to test the average erasure rate of the translated text displayed on the screen by the tested system.
[0394] The second translation real-time test item is used to test the time difference between the input test audio and the playback of the synthesized target language audio stream in the system under test.
[0395] The speech synthesis performance test items include: voice library authorization test item, average sentence synthesis accuracy test item, and speech synthesis naturalness test item.
[0396] Alternatively, if it is also necessary to test the functional requirements of the system under test, the corresponding combination of test items may also include:
[0397] The test includes one or more of the following: performance optimization test items, content management test items, offline simultaneous interpretation test items, and content security test items. The performance optimization test item assesses the system's performance improvement when additional factors are introduced. The content management test item tests whether the system possesses the specified content management functions. The offline simultaneous interpretation test item tests whether the system functions correctly when the network is disconnected. The content security test item tests whether the translated text output by the system conforms to security guidelines.
[0398] The performance optimization test items can include at least one of the following: hot word intervention test items, domain model customization test items, and human-computer collaborative intervention test items. The hot word intervention test item is used to test the change in translation accuracy of hot words in the tested system before and after adding a hot word dictionary. The domain model customization test item is used to test the change in the translation BLEU score of the tested system on the domain test set before and after loading a domain-customized model. The human-computer collaborative intervention test item is used to test whether the final output of the tested system is consistent with the tester's operation when the tester intervenes.
[0399] The test sets corresponding to the aforementioned hot word intervention test items may include a hot word test set and a simultaneous interpretation test set. The hot word test set contains a hot word lexicon, which includes hot words for identification and hot words for translation.
[0400] The test set corresponding to the aforementioned domain model customization test items may include a domain test set, which includes at least one domain customization model and multiple test audios under the corresponding domain.
[0401] The test set corresponding to the above-mentioned human-machine collaborative intervention test items may include simultaneous interpretation test items.
[0402] The test set corresponding to the above-mentioned security test items can include a security test set, which contains multiple test audio files.
[0403] The test items proposed in this application for machine simultaneous interpretation systems are derived from the analysis, summarization, and generalization of the functional and / or performance requirements of machine simultaneous interpretation systems used by various industries. The test items proposed in this way can cover various industries, have strong universality, and are applicable to the testing of machine simultaneous interpretation systems in various industries.
[0404] The applicant's research found that, in addition to a relatively universal and unified testing process, the test data used in the evaluation of machine simultaneous interpretation systems is crucial because it directly affects the test results. In order to obtain better test results, that is, to obtain test results that can more realistically reflect the capabilities of the system under test, the applicant proposed requirements for test data based on a full consideration of the real application scenarios of the system under test, such as the type and quantity requirements of the test data.
[0405] It should be noted that the method for constructing various types of test sets (including simultaneous interpretation test sets, hot word test sets, domain test sets, and security test sets) of the machine simultaneous interpretation system in this application embodiment (i.e., the requirements for test data of the simultaneous interpretation system) is based on a large amount of data accumulated in the actual use of machine simultaneous interpretation systems and services in various industries. This method is proposed after analyzing, summarizing, and refining this data. The test sets constructed by the method proposed after analyzing, summarizing, and refining a large amount of data accumulated in actual application scenarios can fit the actual application scenarios of machine simultaneous interpretation systems. Using such test sets to test the machine simultaneous interpretation system can accurately test the capabilities of the machine simultaneous interpretation system in actual application scenarios.
[0406] Next, we will introduce the various types of test sets for machine simultaneous interpretation systems.
[0407] 1. Simultaneous Interpretation Test Set
[0408] The simultaneous interpreting test set for each language direction must meet the following requirements:
[0409] a) The effective duration of the test audio is no less than 2 hours and no less than 2000 sentences;
[0410] b) The number of speakers of different types corresponding to the test audio shall not be less than 10, and the number of male speakers and female speakers shall not be less than 5 each;
[0411] c) The test audio is real-world audio with a signal-to-noise ratio of at least 20dB.
[0412] d) The test audio should at least cover technology, news, finance, and medical conference scenarios;
[0413] e) The speakers corresponding to the test audio are representative speakers selected based on factors such as different ages, speaking speeds, educational backgrounds, and speaking rhythms, and their statistical distribution patterns.
[0414] f) The recording of the test audio should be consistent with the platform, sampling rate, and input channels of the system under test or meet the set similarity conditions (as close as possible).
[0415] g) The test audio has no regional accent, and the speaker's pronunciation is pure, fluent, and natural;
[0416] h) The translated text of the target language corresponding to the test audio is annotated and translated by professional human translators without referring to any machine translation results. It is also inspected and scored by at least 3 professional translators, and the fidelity and fluency scores reach the set scores. For example, referring to the scoring criteria for fidelity and fluency mentioned above, both fidelity and fluency scores reach 5 points.
[0417] i) Does not contain sensitive content.
[0418] 2. Hot Word Test Set
[0419] The hot word test set for each language direction must meet the following requirements:
[0420] a) The hot word database includes no fewer than 500 hot words such as named entities, technical terms, and new words;
[0421] b) Each hot word provides a corresponding correct translation result, and distinguishes between uppercase and lowercase grammar and spelling.
[0422] 3. Domain Test Set
[0423] The domain test set for each language direction must meet the following requirements:
[0424] a) The effective duration of the test audio is no less than 1 hour and no less than 1000 sentences;
[0425] b) Other requirements are consistent with the simultaneous interpretation test set.
[0426] 4. Security Test Suite
[0427] The security test suite for each language direction must meet the following requirements:
[0428] a) The effective duration of the test audio for each security category shall not be less than 5 minutes or 50 sentences, and the total effective duration of the test audio for all security categories shall not be less than 40 minutes or 400 sentences;
[0429] b) Other requirements are consistent with the simultaneous interpretation test set.
[0430] Furthermore, the embodiments of this application describe the testing process for different test items.
[0431] 1. Source language audio stream to target language translated text
[0432] 1.1 Translation Effectiveness Test Items
[0433] The test data for the translation performance evaluation of the tested system includes:
[0434] The machine-translated text is obtained by the system under test translating the test audio from the input simultaneous interpretation test set.
[0435] a) BLEU sub-test items
[0436] Based on the operational data of the system under test in the BLEU sub-test items, determine the test results of the system under test in the BLEU sub-test items, specifically including:
[0437] For the BLEU score test item, the BLEU score of the tested system on the simultaneous interpretation test set is calculated based on the reference answer text and machine translation text of each test audio in the simultaneous interpretation test set.
[0438] The BLEU score calculation process can be referred to the relevant description above. In this embodiment, existing BLEU score calculation tools can be used, such as the mwerSegmenter tool. By inputting the reference answer text and machine-translated text of the simultaneous interpretation test set, the BLEU score of the tested system on the simultaneous interpretation test set can be obtained.
[0439] After calculating the BLEU score of the system under test on the simultaneous interpretation test set, it can be compared with the specified BLEU score threshold to determine whether the system under test passes the BLEU score test item.
[0440] b) Faithfulness and fluency test items
[0441] Based on the operational data of the system under test in the fidelity and fluency test items, determine the test results of the system under test in the fidelity and fluency test items, specifically including:
[0442] Obtain expert ratings for the fidelity and fluency of machine-translated text.
[0443] Before testing, language experts can be trained on scoring standards. Several language experts (e.g., 5) for each language can be trained to master the scoring standards and requirements before serving as scoring experts. When scoring machine-translated text, the experts can first segment the text into multiple sentences. Each sentence is then scored by the experts, considering both fidelity and fluency, using a scale of 0-5, including one decimal place. Finally, the scores from all experts are summed, and the average score for each sentence, along with the average score for all sentences in the machine-translated text, is calculated to obtain the fidelity and fluency scores.
[0444] After calculating the scores of the tested system on the fidelity test and the fluency test, they can be compared with the specified fidelity score threshold and fluency score threshold to determine whether the tested system passes the fidelity test and the fluency test.
[0445] Furthermore, the translation performance level of the tested system can be determined based on its scores in the BLEU subtests, fidelity test, and fluency test.
[0446] In one optional case, if the tested system achieves a score of 30 on the BLEU test item in the translation effect, and scores ≥4 on the fidelity test item and ≥4 on the fluency test item, then the translation effect of the machine simultaneous interpretation system can be set to meet the Level 2 qualification standard.
[0447] If the tested system achieves a BLEU score of 35 in the translation performance test, and scores ≥4.2 in the fidelity test and ≥4.2 in the fluency test, then the translation performance of the machine simultaneous interpretation system can be set to meet the first-level excellent standard.
[0448] 1.2 First Translation Real-Time Test Item
[0449] The test system's performance data for the first translation real-time test item includes:
[0450] The simultaneous interpretation test set includes the start time T1 of each test audio input to the system under test, and the start time T2 of the corresponding machine-translated text being displayed on the screen.
[0451] Based on the operating data of the system under test in the first translation real-time test item, the test results of the system under test in the first translation real-time test item are determined, specifically including:
[0452] Calculate the time difference T for each test audio. d1 = T2 - T1, and calculate the time difference T between each test audio in the simultaneous interpretation test set. d1The average value is used as the first real-time translation indicator.
[0453] After calculating the score of the tested system on the first translation real-time test item, it can be compared with the specified first translation real-time score threshold to determine whether the tested system passes the first translation real-time test item.
[0454] Furthermore, the first translation real-time performance level of the tested system can be determined based on the index value of the first translation real-time performance test item.
[0455] In one optional scenario, if the first translation real-time performance index is ≤3 seconds, the first translation real-time performance of the machine simultaneous interpretation system can be set to meet the second-level qualification standard.
[0456] If the first translation real-time performance index is ≤2 seconds, then the first translation real-time performance of the machine simultaneous interpretation system can be set to reach the first-level excellent standard.
[0457] 1.3 Translation Readability Test Items
[0458] The test data for the system under test on the translation readability test includes:
[0459] The text displayed each time the subtitles are refreshed after each test audio from the simultaneous interpretation test set is input into the system under test.
[0460] The text displayed each time the subtitles are refreshed can be manually recorded by the testers or obtained from the system logs.
[0461] Based on the test data of the tested system on the translation readability test item, the test results of the tested system on the translation readability test item are determined, specifically including:
[0462] Based on the text displayed at each subtitle refresh, the average erasure rate NE of the machine-translated text corresponding to each test audio is calculated using the following formula:
[0463]
[0464]
[0465] in, This represents the length of the text display sequence at the i-th refresh time. The length of the string with the longest common prefix between the two strings is represented by I, and I represents the number of times the machine-translated text corresponding to a test audio changes during the final determination process.
[0466] The average erase rate of the machine-translated text corresponding to each test audio in the simultaneous interpretation test set is calculated as the system average erase rate (NE) index value.
[0467] After calculating the score of the tested system on the translation readability test item, it can be compared with the specified translation readability score threshold to determine whether the tested system passes the translation readability test item.
[0468] Furthermore, the translation readability level of the tested system can be determined based on the index values of the tested system on the translation readability test item.
[0469] In one optional scenario, if the translation readability index NE≤1, the translation readability of the machine simultaneous interpretation system can be set to meet the Level 2 qualification standard.
[0470] If the translation readability index NE≤0.6, the translation readability of the machine simultaneous interpretation system can be set to reach the first-level excellent standard.
[0471] 2. Source language audio stream to target language audio stream
[0472] 2.1 Translation Effectiveness Test Items
[0473] The translation effectiveness test items in this section are tested using the same methods as those in section 1.1 above.
[0474] 2.2 Second Translation Real-Time Test Item
[0475] The test system's performance data for the second translation real-time test item includes:
[0476] The simultaneous interpretation test set includes the start time T1 of each test audio input to the system under test, and the start time T3 of the corresponding machine-synthesized audio playback.
[0477] Based on the operating data of the system under test in the second translation real-time test item, the test results of the system under test in the second translation real-time test item are determined, specifically including:
[0478] Calculate the time difference T for each test audio. d2 =T3-T1, and calculate the time difference T between each test audio in the simultaneous interpretation test set. d2 The average value is used as the second real-time translation indicator.
[0479] After calculating the score of the tested system on the second translation real-time test item, it can be compared with the specified second translation real-time score threshold to determine whether the tested system passes the second translation real-time test item.
[0480] Furthermore, the second translation real-time performance level of the tested system can be determined based on the index value of the second translation real-time performance test item.
[0481] In one optional scenario, if the second translation real-time performance index is ≤5 seconds, the second translation real-time performance of the machine simultaneous interpretation system can be set to meet the Level 2 qualification standard.
[0482] If the first translation real-time performance index is ≤4 seconds, then the second translation real-time performance of the machine simultaneous interpretation system can be set to reach the first-class excellent standard.
[0483] 2.3 Speech Synthesis Effect Test Items
[0484] a) Sound library licensing test item
[0485] Testers can check the authorization certificate of the speech synthesis library.
[0486] b) Average sentence composition accuracy test item
[0487] The test data for the average sentence synthesis accuracy test item of the tested system includes:
[0488] The machine-synthesized audio output by the system under test corresponds to each test audio in the simultaneous interpretation test set.
[0489] Based on the test data of the tested system on the average sentence synthesis accuracy test item, the test results of the tested system on the average sentence synthesis accuracy test item are determined, specifically including:
[0490] The number of correct machine-synthesized audio samples is counted, and the ratio of the number of correct machine-synthesized audio samples to the total number of test audio samples in the simultaneous interpretation test set is calculated to obtain the average sentence synthesis accuracy rate.
[0491] After calculating the score of the tested system on the average sentence synthesis accuracy test item, it can be compared with the specified average sentence synthesis accuracy score threshold to determine whether the tested system passes the average sentence synthesis accuracy test item.
[0492] c) Speech synthesis naturalness test item
[0493] The test data for the naturalness of speech synthesis of the tested system includes:
[0494] The machine-synthesized audio output by the system under test corresponds to each test audio in the simultaneous interpretation test set.
[0495] Based on the test data of the tested system in the speech synthesis naturalness test item, the test results of the tested system in the speech synthesis naturalness test item are determined, specifically including:
[0496] Obtain the naturalness score of the machine-synthesized audio from the scoring experts.
[0497] Specifically, before the test, 10-20 language experts with listening experience, native speakers of the language, and senior foreign language students or teachers (the ratio of experts, native speakers, and learners can be 3:5:2) should be selected, with gender balance: 50% ± 10% male and 50% ± 10% female. It is essential to ensure that the evaluators have no hearing impairment, are familiar with the test language, understand the basic phonetics of the test language, and have a preliminary understanding of concepts such as intonation and stress. This ensures that the evaluators can understand and evaluate the speech both intuitively and rationally. Afterwards, the evaluators will be trained, including on the principles of speech synthesis, the characteristics of synthesized speech, and the reasons why synthesized speech differs from natural speech. This training will enable the evaluators to make appropriate evaluations of the synthesized speech. During the training, the evaluators will discuss the scoring criteria and corresponding reference speech samples to deepen their understanding of the five-point system definition, thereby avoiding overly discrete evaluation scores. Finally, the evaluators will score the machine-synthesized audio corresponding to each test audio. Each evaluator uses an evaluation tool on a computer in a quiet room, completing the evaluation task through headphones. Evaluators score the naturalness of the machine-synthesized audio for each test audio segment according to standards, and the final score (MOS) is the average of all scores for all sentences.
[0498] After calculating the score of the tested system on the naturalness of speech synthesis test item, it can be compared with the specified naturalness threshold of speech synthesis to determine whether the tested system passes the naturalness test item.
[0499] Furthermore, the speech synthesis performance level of the tested system can be determined based on its scores on the average sentence synthesis accuracy test and the speech synthesis naturalness test.
[0500] In one optional scenario, if the average sentence synthesis accuracy is ≥90% and the synthesis naturalness score (MOS) is ≥4, the language synthesis effect of the machine simultaneous interpretation system can be set to meet the Level 2 qualification standard.
[0501] If the average sentence synthesis accuracy is ≥95% and the synthesis naturalness score (MOS) is ≥4.5, then the language synthesis effect of the machine simultaneous interpretation system can be set to reach the first-level excellent standard.
[0502] 3. Performance optimization
[0503] 3.1 Hot word intervention test items
[0504] The test data of the system on the hot word intervention test item includes:
[0505] Before and after adding the hot word test set, the tested system translates the machine-translated text from the input simultaneous interpretation test set into test audio.
[0506] Based on the operational data of the tested system in the hot word intervention test, the test results of the tested system in the hot word intervention test are determined, specifically including:
[0507] Before adding the hot word test set, the number of correctly translated hot words in the machine-translated text was P1; after adding the hot word test set, the number of correctly translated hot words in the machine-translated text was P2. The hot word intervention efficiency was then calculated.
[0508] , where T is the total number of hot words contained in the simultaneous interpretation test set.
[0509] After calculating the score of the tested system on the hot word intervention test item, it can be compared with the specified hot word intervention efficiency threshold to determine whether the tested system is qualified on the hot word intervention test item.
[0510] 3.2 Domain Model Customization Test Items
[0511] The runtime data of the tested system on the domain model customization test items includes:
[0512] The machine-translated text obtained by translating the test audio in the corresponding domain of the domain test set before and after loading the domain-customized model in the domain test set.
[0513] Based on the runtime data of the system under test on the domain model customization test items, determine the test results of the system under test on the domain model customization test items, specifically including:
[0514] Based on the machine-translated text and the reference answer text of the test audio, the BLEU score of the tested system on the domain test set is calculated before and after loading the domain-customized model, and the difference in BLEU score of the tested system on the domain test set after loading the domain-customized model is compared with that before loading the domain-customized model.
[0515] Furthermore, the scores of the tested system on the domain model customization test items calculated above can be compared with the specified score difference threshold to determine whether the tested system passes the domain model customization test items.
[0516] 3.3 Human-Machine Collaborative Intervention Test Items
[0517] During system testing, testers access relevant system pages and perform editing operations such as modifying, optimizing, and deleting the identified source language text or target language translated text. They then observe whether the final displayed result of the system matches the testers' operations to determine whether the tested system passes the human-computer collaborative intervention test.
[0518] 4. Content Management
[0519] 4.1 Exporting test items from the meeting
[0520] Testers can access the relevant page of the system, enter the file name of the meeting data to be exported, as well as the storage location after export, and trigger the meeting export function to observe whether the meeting export is successful, thereby determining whether the system under test is qualified in the meeting export test item.
[0521] 4.2 Offline Simultaneous Interpretation Test Items
[0522] Testers disconnect the system under test from the network and then perform each test item to determine the system's performance in offline mode.
[0523] 4.3 Content Security Test Items
[0524] The test system's performance data for the content security test items includes:
[0525] The system under test outputs machine-translated text obtained after translating the input audio from the security test set.
[0526] Based on the operational data of the tested system in the content security test items, determine the test results of the tested system in the content security test items, specifically including:
[0527] For content security test items, obtain the human assessment results of whether the machine-translated text conforms to security guidelines, thereby determining whether the tested system passes the content security test items.
[0528] This application also provides a testing device for a machine simultaneous interpretation system. The testing device for the machine simultaneous interpretation system provided in this application is described below. The testing device for the machine simultaneous interpretation system described below can be referred to in correspondence with the testing method for the machine simultaneous interpretation system described above.
[0529] Please see Figure 11 The diagram shows a structural schematic of a testing device for a machine simultaneous interpretation system provided in an embodiment of this application. It may include: a test set acquisition module 101, a test audio input module 102, a running data acquisition module 103, and a test result determination module 104.
[0530] The test set acquisition module 101 is used to acquire the test set corresponding to the test item of the system under test, wherein the test set corresponding to the test item includes at least one language direction test set, and the test set contains multiple test audios;
[0531] The test audio input module 102 is used to input each test audio from the test set corresponding to the test item into the system under test;
[0532] The runtime data acquisition module 103 is used to acquire the runtime data of the system under test on the test item;
[0533] The test result determination module 104 is used to determine the test result of the system under test on the test item based on the operating data of the system under test on the test item.
[0534] The test sets corresponding to each test item acquired by the test set acquisition module 101 can be found in the previous introduction to the test methods.
[0535] For different types of test items, the runtime data acquired by the runtime data acquisition module 103 can be referred to the relevant description of the test methods above. Similarly, the process by which the test result determination module 104 determines the test results of the system under test on different test items based on the runtime data can also be referred to the relevant description of the test methods above, and will not be repeated here.
[0536] This application also provides a testing device for a machine simultaneous interpretation system. Please refer to [link to relevant documentation]. Figure 12 The diagram shows a schematic of the test equipment for the machine simultaneous interpretation system. The test equipment may include: at least one processor 01, at least one communication interface 02, at least one memory 03 and at least one communication bus 04.
[0537] In this embodiment of the application, the number of processor 01, communication interface 02, memory 03 and communication bus 04 is at least one, and processor 01, communication interface 02 and memory 03 communicate with each other through communication bus 04.
[0538] Processor 01 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.
[0539] Memory 03 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;
[0540] The memory stores a program, which the processor can call. The program is used to execute each step of the aforementioned test method for the machine simultaneous interpretation system.
[0541] Optionally, the refined and extended functions of the program can be found in the description above.
[0542] This application embodiment also provides a readable storage medium that can store a program suitable for processor execution, the program being used to execute various steps of the aforementioned test method for a machine simultaneous interpretation system.
[0543] Optionally, the refined and extended functions of the program can be found in the description above.
[0544] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0545] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0546] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A machine simultaneous interpretation system, characterized in that, include: The application and server are defined as follows: the application includes an audio acquisition module and a multimodal output module; the multimodal output module includes at least one of a subtitle display module, an audio playback module, and a video playback module; the server includes a speech translation module; when the multimodal output module includes the audio playback module, the server also includes a speech synthesis module; when the multimodal output module includes the video playback module, the server also includes the speech synthesis module and a virtual human module. The audio acquisition module is used to acquire the audio stream of the source language; The speech translation module is used to translate the audio stream of the source language to obtain translated text of the target language that meets the set translation effect. The set translation effect includes the translation BLEU score, fidelity and fluency reaching corresponding score thresholds. The subtitle display module is used to output and display the translated text of the target language within a first set time period after the audio acquisition module acquires the audio stream of the source language, and during the process of outputting and displaying the translated text, the average erasure rate of the translated text does not exceed a set average erasure rate threshold. The speech synthesis module is used to perform speech synthesis on the translated text of the target language to obtain an audio stream of the target language that meets the set speech synthesis effect. The set speech synthesis effect includes the average sentence synthesis accuracy and the synthesis naturalness reaching corresponding score thresholds, and the sound library used during synthesis being an authorized sound library. The audio playback module is used to play the audio stream of the target language within a second set time period after the audio acquisition module acquires the audio stream of the source language; The virtual human module is used to combine the audio stream of the target language with the virtual avatar and convert it into a video stream; The video playback module is used to play the video stream.
2. The system according to claim 1, characterized in that, The speech translation module includes: a speech recognition module and a machine translation module; The speech recognition module is used to recognize the audio stream of the source language to obtain the source language text; The machine translation module is used to translate the source language text into the target language to obtain the translated text in the target language.
3. The system according to claim 1, characterized in that, The voice translation module supports adding a hot word dictionary, and the translation accuracy of hot words after adding the hot word dictionary is not lower than the set hot word intervention efficiency threshold.
4. The system according to claim 2, characterized in that, The speech translation module supports loading domain-customized models, and after loading the domain-customized model, the translation BLEU score on the domain test set is improved by no less than a set BLEU score difference threshold. The domain-customized model includes speech recognition and text translation models.
5. The system according to claim 1, characterized in that, The application is deployed locally, and the server is deployed on the network side. The application and the server are connected via the network. or, Both the application and server are deployed locally, supporting simultaneous interpretation in offline mode.
6. The system according to claim 2, characterized in that, The application also includes: The meeting data export module is used to respond to the user's operation of exporting specified meeting data to a specified storage location. The specified meeting data includes any one or more of the recognition results, translation results, and speech synthesis results of the source language audio stream.
7. The system according to claim 2, characterized in that, The application also includes: The speech recognition result editing module is used to respond to the human translator's operation of editing the recognition results of the audio stream and obtain the edited source language text; And / or, The translation text editing module is used to respond to human translators' editing operations on the translated target language text, and to obtain the edited target language translation text.
8. The system according to claim 1, characterized in that, The server also includes a data preprocessing module, which performs content security checks on the translated text of the target language output by the speech translation module. If the content security check fails, the subsequent processing is stopped.
9. The system according to claim 1, characterized in that, The application also includes an acoustic noise reduction module, which is used to perform noise reduction processing on the audio stream of the source language acquired by the audio acquisition module, and send the noise-reduced audio stream to the speech translation module. And / or, The server also includes a post-translation processing module, which is used to normalize the translated text of the target language and send the normalized translated text of the target language to the speech synthesis module.
10. The system according to claim 2, characterized in that, The speech translation module further includes: a post-recognition processing module, used to normalize the source language text and send the normalized source language text to the machine translation module; The subtitle display module is also used to display the source language text obtained by the speech recognition module, or to display the normalized source language text obtained by the post-recognition processing module.
11. A machine simultaneous interpretation method, characterized in that, This is used to perform at least one of the following simultaneous interpretation processes: simultaneous interpretation from source language audio to target language text, simultaneous interpretation from source language audio to target language audio, and simultaneous interpretation from source language audio to target language video. The process of simultaneous interpretation between source language audio and target language text includes: Obtain the audio stream of the source language; The audio stream of the source language is translated to obtain the translated text of the target language that meets the set translation effect. The set translation effect includes the translation BLEU score, fidelity and fluency reaching the corresponding score thresholds. Within a first set time period after acquiring the audio stream of the source language, the translated text of the target language is output and displayed, and during the process of outputting and displaying the translated text, the average erasure rate of the translated text does not exceed a set average erasure rate threshold. The process of simultaneous interpretation between source language audio and target language audio includes: Obtain the audio stream of the source language; The audio stream of the source language is translated to obtain the translated text of the target language that meets the set translation effect. The set translation effect includes the translation BLEU score, fidelity and fluency reaching the corresponding score thresholds. The translated text of the target language is subjected to speech synthesis to obtain an audio stream of the target language that meets the set speech synthesis effect. The set speech synthesis effect includes the average sentence synthesis accuracy and the synthesis naturalness reaching corresponding score thresholds, and the sound library used during synthesis being an authorized sound library. Within a second predetermined time period after the source language audio stream is acquired, the target language audio stream is played. The process of simultaneous interpretation between source language audio and target language video includes: Obtain the audio stream of the source language; The audio stream of the source language is translated to obtain the translated text of the target language that meets the set translation effect. The set translation effect includes the translation BLEU score, fidelity and fluency reaching the corresponding score thresholds. The translated text of the target language is subjected to speech synthesis to obtain an audio stream of the target language that meets the set speech synthesis effect. The set speech synthesis effect includes the average sentence synthesis accuracy and the synthesis naturalness reaching corresponding score thresholds, and the sound library used during synthesis being an authorized sound library. The audio stream of the target language is combined with the virtual avatar and converted into a video stream, which is then played.
12. The method according to claim 11, characterized in that, The audio stream of the source language is translated to obtain translated text in the target language that meets the set translation effect, including: The audio stream of the source language is identified to obtain the source language text; The source language text is translated into the target language to obtain a translated text in the target language that meets the set translation effect.
13. The method according to claim 11, characterized in that, The process of translating the audio stream of the source language to obtain translated text in the target language that meets the set translation effect includes: A machine simultaneous interpretation system with an added hot word lexicon is used to translate the audio stream of the source language to obtain translated text in the target language that meets the set translation effect. Among them, compared with machine simultaneous interpretation systems without hot word databases, machine simultaneous interpretation systems with hot word databases have a translation accuracy rate of hot words that is no less than the set hot word intervention efficiency threshold.
14. The method according to claim 12, characterized in that, The audio stream of the source language is translated to obtain translated text in the target language that meets the set translation effect, including: A machine simultaneous interpretation system with a customized model for the target domain is used to translate the audio stream of the source language to obtain translated text in the target language that meets the set translation effect. The target domain customized model is a speech recognition and text translation model in the target domain to which the audio stream of the source language belongs. Compared with the machine simultaneous interpretation system without the target domain customized model, the machine simultaneous interpretation system with the target domain customized model has a translation BLEU score improvement of no less than a set BLEU score difference threshold on the domain test set. The domain test set includes multiple test audios in the target domain.
15. The method according to claim 11, characterized in that, The machine simultaneous interpretation method is implemented whether the network is connected or disconnected.
16. The method according to claim 12, characterized in that, Also includes: In response to a user's operation to export specified meeting data to a specified storage location, the specified meeting data is exported to the specified storage location. The specified meeting data includes any one or more of the recognition results, translation results, and speech synthesis results of the source language audio stream.
17. The method according to claim 12, characterized in that, Also includes: The system responds to human translators' editing of the audio stream's recognition results, resulting in edited source language text. And / or, The system responds to human translators' editing of the translated text in the target language, resulting in an edited translated text in the target language.
18. The method according to claim 11, characterized in that, Also includes: The translated text in the target language is subjected to content security testing. If the content security test fails, the subsequent processing is stopped.
19. The method according to claim 11, characterized in that, Before translating the audio stream of the source language, the following is also included: The audio stream of the source language is subjected to noise reduction processing; And / or, Before performing speech synthesis on the translated text of the target language, the method further includes: The translated text in the target language is then normalized.
20. The method according to claim 12, characterized in that, Before translating the source language text into the target language, the process also includes: The source language text is then standardized. The source language text or the normalized source language text is output and displayed.
21. A testing method for a machine simultaneous interpretation system, characterized in that, The method for testing the machine simultaneous interpretation system according to any one of claims 1-10 includes: Obtain the test set corresponding to the test item of the system under test, wherein the test set corresponding to the test item includes at least one language direction test set, and the test set contains multiple test audios; Input each test audio from the test set corresponding to the test item into the system under test; Obtain the operating data of the system under test on the test item; Based on the operating data of the system under test on the test item, the test result of the system under test on the test item is determined.
22. The method according to claim 21, characterized in that, The test items include: A first test item is used to test the performance of the system under test from source language audio stream input to target language translated text output, and a second test item is used to test the performance of the system under test from source language audio stream input to target language audio stream output, wherein the test set corresponding to the first test item and the second test item includes the simultaneous interpretation test set; The first test item includes a translation effectiveness test item, a first translation real-time performance test item, and a translation readability test item; The second test item includes a translation effect test item, a second translation real-time performance test item, and a speech synthesis effect test item.
23. The method according to claim 22, characterized in that, The test items also include: The test includes one or more of the following: performance optimization test items, content management test items, offline simultaneous interpretation test items, and content security test items. The performance optimization test items are used to test the performance optimization of the system under test when intervention factors are added. The content management test items are used to test whether the system under test has the specified content management function. The offline simultaneous interpretation test items are used to test whether the various functions of the system under test are implemented normally when the network is disconnected. The content security test items are used to test whether the translated text of the target language output by the system under test conforms to the security guidelines. The test set corresponding to the content security test item includes a security test set, which contains multiple test audio files.
24. The method according to claim 22, characterized in that, The translation performance test includes BLEU subtest items, fidelity test items, and fluency test items. The BLEU subtest items are used to evaluate the translation quality of the tested system. The fidelity test items are used to evaluate whether the translated text obtained by the tested system faithfully expresses the content of the original text. The fluency test items are used to evaluate whether the translated text obtained by the tested system is fluent and whether it conforms to the expression habits of the target language. The first translation real-time test item is used to test the time difference between the input test audio and the display of the translated text on the screen of the system under test; The translation readability test item is used to test the average erasure rate of the translated text displayed on the screen by the tested system; The second translation real-time test item is used to test the time difference between the input test audio and the playback of the synthesized target language audio stream in the system under test; The speech synthesis effect test items include: voice library authorization test item, average sentence synthesis accuracy test item, and speech synthesis naturalness test item.
25. The method according to claim 23, characterized in that, The performance optimization test items include: The test includes at least one of the following: hot word intervention test item, domain model customization test item, and human-computer collaborative intervention test item. The hot word intervention test item is used to test the change in translation accuracy of hot words in the tested system before and after adding a hot word lexicon. The domain model customization test item is used to test the change in translation BLEU score of the tested system on the domain test set before and after loading a domain customization model. The human-computer collaborative intervention test item is used to test whether the final output result of the tested system is consistent with the operation of the tester when the tester intervenes. The test set corresponding to the hot word intervention test item includes a hot word test set and a simultaneous interpretation test set. The hot word test set contains a hot word lexicon, and the hot words include hot words for identification and hot words for translation. The test set corresponding to the domain model customized test item includes a domain test set, which includes at least one domain customized model and multiple test audios under the corresponding domain. The test set corresponding to the human-machine collaborative intervention test items includes the simultaneous interpretation test items.
26. The method according to claim 24, characterized in that, The operating data of the system under test on the BLEU subtest, fidelity test and fluency test include: the machine-translated text obtained by the system under test after translating the test audio from the input simultaneous interpretation test set; The operating data of the system under test on the first translation real-time test item includes: the start time T1 of each test audio input to the system under test in the simultaneous interpretation test set, and the time T2 when the corresponding machine-translated text begins to be displayed on the screen. The operational data of the system under test on the translation readability test item includes: the text displayed each time the subtitle is refreshed after each test audio from the simultaneous interpretation test set is input into the system under test; The operating data of the system under test in the second translation real-time test item includes: the start time T1 of each test audio input to the system under test in the simultaneous interpretation test set, and the time T3 when the corresponding machine-synthesized audio starts playing; The running data of the tested system on the average sentence synthesis accuracy test and the speech synthesis naturalness test include: the machine-synthesized audio output by the tested system corresponding to each test audio in the simultaneous interpretation test set; Based on the operating data of the system under test on the test item, the test result of the system under test on the test item is determined, including: For the BLEU score test item, the BLEU score of the tested system on the simultaneous interpretation test set is calculated based on the reference answer text of each test audio in the simultaneous interpretation test set and the machine translation text. For the fidelity and fluency tests, obtain the fidelity and fluency scores given by scoring experts to the machine-translated text. The translation performance level of the tested system is determined based on its scores in the BLEU subtest, fidelity test, and fluency test. For the first real-time translation test, the time difference T is calculated for each test audio file. d1 =T2-T1, and calculate the time difference T between each test audio in the simultaneous interpretation test set. d1 The average value is used as the first real-time translation indicator. The first translation real-time performance level of the tested system is determined based on the index value of the first translation real-time performance test item. For the translation readability test, the average erasure rate NE of the machine-translated text corresponding to each test audio is calculated according to the following formula, based on the text displayed at each subtitle refresh: E(i)=|o i-1 |-|LCP(o i ,o i-1 )| Among them, o i represents the length of the text display content sequence at refresh time i, LCP(·) represents the length of the longest common prefix of two strings, and I represents the number of times the machine-translated text corresponding to a test audio changes during the final determination process; The average erase rate of the machine-translated text corresponding to each test audio in the simultaneous interpretation test set is calculated as the average erase rate (NE) index value of the system. The translation readability level of the tested system is determined based on the index values of the tested system on the translation readability test item. For the second translation real-time test item, the time difference T is calculated for each test audio. d2 =T3-T1, and calculate the time difference T between each test audio in the simultaneous interpretation test set. d2 The average value of is used as the second real-time translation indicator. The second translation real-time performance level of the tested system is determined based on the index value of the tested system on the second translation real-time performance test item. For the average sentence synthesis accuracy test item, the number of correct machine-synthesized audio is counted, and the ratio of the number of correct machine-synthesized audio to the total number of all test audio in the simultaneous interpretation test set is calculated to obtain the average sentence synthesis accuracy. For the speech synthesis naturalness test item, obtain the synthesis naturalness score results of the scoring experts on the machine-synthesized audio; The speech synthesis effect level of the tested system is determined based on its scores in the average sentence synthesis accuracy test and the speech synthesis naturalness test.
27. The method according to claim 25, characterized in that, The test system's operational data on the hot word intervention test item includes: the machine-translated text obtained by the test system after translating the test audio from the input simultaneous interpretation test set before and after adding the hot word test set; The running data of the system under test on the domain model customization test item includes: the machine-translated text obtained by translating the test audio in the corresponding domain of the domain test set before and after loading the domain customization model in the domain test set; The operational data of the system under test in the content security test items includes: The system under test outputs machine-translated text obtained after translating the input security test set audio. Based on the operating data of the system under test on the test item, the test result of the system under test on the test item is determined, including: For the hot word intervention test item, the number of correctly translated hot words in the machine-translated text before adding the hot word test set (P1) and the number of correctly translated hot words in the machine-translated text after adding the hot word test set (P2) are counted, and the hot word intervention efficiency is calculated. (P2-P1) / T*100%, where T is the total number of hot words in the simultaneous interpretation test set; For the domain model-customized test items, based on the reference answer text of machine-translated text and test audio, calculate the BLEU score of the tested system on the domain test set before and after loading the domain model-customized model, and compare the difference in BLEU score of the tested system on the domain test set after loading the domain model compared to before loading the domain model-customized model. For content security testing items, obtain human assessment results on whether machine-translated text conforms to security guidelines.
28. A testing device for a machine simultaneous interpretation system, characterized in that, include: The test set acquisition module is used to acquire the test set corresponding to the test item of the system under test. The test set corresponding to the test item includes at least one language direction test set, and the test set contains multiple test audios. The test audio input module is used to input each test audio from the test set corresponding to the test item into the system under test; The data acquisition module is used to acquire the running data of the system under test on the test item; The test result determination module is used to determine the test result of the system under test on the test item based on the operating data of the system under test on the test item.
29. A testing device for a machine simultaneous interpretation system, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement the various steps of the testing method for the machine simultaneous interpretation system as described in any one of claims 21 to 27.
30. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the various steps of the testing method for the machine simultaneous interpretation system as described in any one of claims 21 to 27.