A voice relay system testing method and related device

By employing a full-link, multi-scenario testing method for voice simultaneous interpretation systems, the performance and collaborative operation of speech recognition, language conversion, and speech synthesis modules are evaluated in a tiered manner. This addresses the issue of fragmented testing dimensions in the integrated architecture and improves the system's real-time performance and fluency.

CN120913540BActive Publication Date: 2026-02-06IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511422493.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-02-06
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

Existing testing methods for simultaneous voice interpretation systems cannot effectively cover the integrated architecture, making it difficult to guarantee real-time performance, fluency, and reliability. In particular, in multi-scenario and multi-module collaborative work, there are problems of fragmented testing dimensions and a lack of indicator system.

Method used

This paper presents a full-link, multi-scenario testing method for voice simultaneous interpretation systems. By acquiring test audio sets and classifying them into user experience, system function, and core module dimensions, the performance and collaborative work of speech recognition, language conversion, and speech synthesis modules are evaluated respectively, and multiple test indicators are used for quantitative evaluation.

Benefits of technology

It achieves end-to-end full-link coverage and problem localization, and can accurately identify and solve performance and collaborative problems of the voice simultaneous interpretation system at all levels, thereby improving the real-time performance and smoothness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913540B_ABST
    Figure CN120913540B_ABST
Patent Text Reader

Abstract

The application discloses a voice simultaneous interpretation system test method and related device, relates to the technical field of system test, and the voice simultaneous interpretation system test method comprises the following steps: obtaining a test audio set, inputting test audio in the test audio set into a voice simultaneous interpretation system for processing, determining test results of the voice simultaneous interpretation system in user experience dimensions, system function dimensions and core module dimensions according to processing condition data of the voice simultaneous interpretation system on the test audio in the test audio set, the user experience dimensions focus on the perception of the user on the end-to-end voice simultaneous interpretation system, the system function dimensions focus on the collaborative working condition among the modules in the voice simultaneous interpretation system, and the core module dimensions focus on the performance of each independent module in the voice simultaneous interpretation system. The voice simultaneous interpretation system test method disclosed by the application classifies the complex system test into three clear levels, problems can be found through the test of the three levels, the problem source can be accurately located, end-to-end full-link coverage and verification are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of system testing technology, and in particular to a method and related apparatus for testing a voice simultaneous interpretation system. Background Technology

[0002] With the development of voice simultaneous interpretation technology, its application scenarios have expanded from professional conferences to daily communication. Traditional voice simultaneous interpretation systems have a modular, serial architecture, which decomposes the complex "voice-to-voice" task into three independent, serial sub-tasks, which are handled by dedicated modules respectively. In essence, traditional voice simultaneous interpretation systems are a distributed architecture.

[0003] To reduce network latency and achieve true streaming processing, integrated architecture voice interpretation systems have emerged. These systems deeply optimize and integrate the three core modules—speech recognition, language conversion (translation), and speech synthesis—and run them in a unified software framework or service platform. This architecture reduces data transmission and processing latency in intermediate links, significantly improving the real-time performance and fluency of voice interpretation.

[0004] To ensure the real-time performance, fluency, and reliability of simultaneous voice interpretation, testing of the system is necessary. Current testing methods are designed for traditional simultaneous voice interpretation systems, which have significant limitations for the integrated architecture described above. Therefore, a testing method specifically for the integrated architecture is urgently needed to support system updates and guarantee the real-time performance, fluency, and reliability of simultaneous voice interpretation. Summary of the Invention

[0005] In view of this, this application provides a testing method and related apparatus for a voice simultaneous interpretation system, which provides a testing method for a voice simultaneous interpretation system with an integrated architecture, and the technical solution is as follows:

[0006] The first aspect of this application provides a method for testing a voice simultaneous interpretation system, including:

[0007] Obtain the test audio set;

[0008] The test audio from the test audio set is input into the target speech interpretation system for processing;

[0009] Based on the processing data of the target speech interpretation system for the test audio in the test audio set, the test results of the target speech interpretation system under several test dimensions are determined.

[0010] The several test dimensions include a user experience dimension, a system function dimension, and a core module dimension. The user experience dimension focuses on user perception of the target voice simultaneous interpretation system end-to-end. The system function dimension focuses on the collaborative work of modules in the target voice simultaneous interpretation system. The core module dimension focuses on the performance of each independent module in the target voice simultaneous interpretation system.

[0011] In a possible implementation, the test audio set is obtained by:

[0012] The constructed basic corpus set, interference corpus set, and professional field corpus set are obtained, wherein the basic corpus set includes several basic audios of general fields, the basic audios in the basic corpus set cover multiple languages, multiple speech speeds, and multiple types of scenes, the interference corpus set includes multiple interference audios, and the professional field corpus set includes several professional field audios.

[0013] According to the basic audios in the basic corpus set and the interference audios in the interference corpus set, several interference-containing audios are constructed as general field test audios, and the professional field audios in the professional field corpus set are taken as professional field test audios, to obtain a test audio set including the general field test audios and the professional field test audios.

[0014] In a possible implementation, each of the test dimensions includes several test items.

[0015] The test results of the target voice simultaneous interpretation system in the several test dimensions are determined according to the processing condition data of the target voice simultaneous interpretation system for the test audios in the test audio set, and the test results include:

[0016] For each of the test dimensions, data related to each test item in the test dimension is obtained from the processing condition data of the target voice simultaneous interpretation system for the test audios in the test audio set.

[0017] According to the data related to each test item in the test dimension, a test index of each test item in the test dimension is determined.

[0018] According to the test index of each test item in the test dimension, a test result of the target voice simultaneous interpretation system on each test item in the test dimension is determined.

[0019] In a possible implementation, the test items in the user experience dimension include some or all of the following test items: a real-time test item, a fluency test item, and a subjective experience test item.

[0020] The test items under the system function dimension include some or all of the following test items: a cross-language conversion accuracy test item and a multi-module coordination performance test item. The cross-language conversion accuracy test item evaluates whether the meaning expressed by the test audio is accurately and completely converted and delivered into target language audio through cooperation of the modules. The multi-module coordination performance test item evaluates the fluency, stability and timeliness of the cooperation of the modules.

[0021] The test items under the core module dimension include some or all of the following test items: a speech recognition module performance test item, a language conversion module performance test item and a speech synthesis module performance test item.

[0022] In a possible implementation, the test indicators of the real-time test item under the user experience dimension include some or all of the following indicators: a first response delay and a tail response delay. The first response delay is a time difference between a time when test audio starts to be input and a time when a first valid phoneme of target language audio output by the target speech simultaneous interpretation system, and the tail response delay is a time difference between a time when the input of the test audio ends and a time when a last phoneme of the target language audio output by the target speech simultaneous interpretation system.

[0023] The test indicators of the fluency test item under the user experience dimension include some or all of the following indicators: a non-punctuation position stuttering frequency, a longest duration of single stuttering and a translation stuttering number.

[0024] The test indicators of the subjective experience test item under the user experience dimension include some or all of the following indicators: an adaptation degree of a simultaneous interpretation speed of the target speech simultaneous interpretation system to a scene, intelligibility and naturalness of the final output audio of the target speech simultaneous interpretation system.

[0025] In a possible implementation, according to data related to the real-time test item under the user experience dimension, the first response delay and the tail response delay of the real-time test item are determined, including:

[0026] For each piece of test audio in the set of test audios, according to a time when the test audio starts to be input into the target speech simultaneous interpretation system, a time when the target speech simultaneous interpretation system outputs a first valid phoneme of target language audio for the test audio, an input end time of the test audio and a time when the target speech simultaneous interpretation system outputs a last phoneme of target language audio for the test audio, the first response delay and the tail response delay corresponding to the test audio are determined.

[0027] a mean of the first response delays corresponding to each of the test audios in the test audio set is calculated to obtain a first first response delay of the target speech simultaneous interpretation system; and / or, the first response delays corresponding to each of the test audios in the test audio set are sorted in ascending order to obtain a first response delay of the target position as a second first response delay of the target speech simultaneous interpretation system; and / or, a mean of the first response delays corresponding to each of the test audios in the short sentence scenario is calculated to obtain a short sentence first response delay of the target speech simultaneous interpretation system; and / or, a mean of the first response delays corresponding to each of the test audios in the long sentence scenario is calculated to obtain a long sentence first response delay of the target speech simultaneous interpretation system.

[0028] a mean of the last response delays corresponding to each of the test audios in the test audio set is calculated to obtain a last response delay of the target speech simultaneous interpretation system.

[0029] In a possible implementation, the test indicators of the cross-language conversion accuracy test item in the system function dimension include part or all of the following indicators: a word-level accuracy indicator of cross-language conversion, a sentence-level semantic integrity indicator of cross-language conversion, and a context adaptability indicator of cross-language conversion.

[0030] The test indicators of the multi-module coordination performance test item in the system function dimension include part or all of the following indicators: an indicator representing data accumulation among modules in the target speech simultaneous interpretation system, and an indicator representing synchronization among modules in the target speech simultaneous interpretation system.

[0031] In a possible implementation, the word-level accuracy indicator includes a word-level translation accuracy rate, the sentence-level semantic integrity indicator includes a proportion of sentences that reserve core semantic roles, and the context adaptability indicator of cross-language conversion includes a correct conversion rate of pronouns and referring words in a multi-round dialogue.

[0032] The indicator representing data accumulation among modules in the target speech simultaneous interpretation system includes an accumulation amount of data among modules and / or an accumulation duration of data among modules, and the indicator representing synchronization among modules in the target speech simultaneous interpretation system includes a time difference between an output of a speech recognition module and an input of a speech synthesis module.

[0033] In a possible implementation, the test indicators of the speech recognition module performance test item in the core module dimension include part or all of the following indicators: a word-level speech recognition accuracy rate, a sentence-level speech recognition accuracy rate, and an endpoint detection accuracy rate.

[0034] The test indicators of the language conversion module performance test item in the core module dimension include part or all of the following indicators: a translation delay, a conversion accuracy rate for professional terms, a conversion naturalness for specified sentence patterns, and an overlap degree with a standard translation.

[0035] The test indicators of the speech synthesis module performance test item under the core module dimension include part or all of the following indicators: mean opinion score of synthesized audio, synthesized multi-sound character accuracy, synthesized text regularization accuracy, and indicators representing audio synthesis real-time performance.

[0036] In a possible implementation, the speech simultaneous interpretation system testing method further includes:

[0037] According to the test indicators of the test items under the test dimensions, an index distribution heat map of each scenario, each speech speed, or each language is generated.

[0038] And / or, according to the test indicators of the test items under the test dimensions, a performance correlation curve of different modules is drawn.

[0039] In a possible implementation, the speech simultaneous interpretation system testing method further includes:

[0040] According to the test results of the target speech simultaneous interpretation system under the test dimensions, problems existing in the target speech simultaneous interpretation system are determined.

[0041] The problems existing in the target speech simultaneous interpretation system are output; and / or, according to the problems existing in the target speech simultaneous interpretation system, a problem use case library is constructed, so that the target speech simultaneous interpretation system is tested based on the problem use case library after being optimized.

[0042] The second aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:

[0043] The memory is used to store a computer program;

[0044] The processor is used to execute the computer program, so that the electronic device can implement the steps of any one of the speech simultaneous interpretation system testing methods.

[0045] The third aspect of the present application is a computer storage medium, which carries one or more computer programs, when the one or more computer programs are executed by an electronic device, the electronic device can implement the steps of any one of the speech simultaneous interpretation system testing methods.

[0046] The fourth aspect of the present application provides a computer program product, comprising computer readable instructions, when the computer readable instructions run on an electronic device, the electronic device can implement the steps of any one of the speech simultaneous interpretation system testing methods.

[0047] By means of the technical solutions, the voice simultaneous interpretation system testing method provided by the application first acquires a test audio set, then inputs test audio in the test audio set into a target voice simultaneous interpretation system for processing, and finally determines the test result of the target voice simultaneous interpretation system in the user experience dimension, the system function dimension and the core module dimension according to the processing condition data of the target voice simultaneous interpretation system for the test audio in the test audio set. The voice simultaneous interpretation system testing method provided by the application classifies the complex system testing into three clear levels (user experience -> system function -> core module), on the one hand, the three-level testing forms a positioning funnel from top to bottom and layer by layer, that is, through the three-level testing, the problem can be found and the root cause of the problem can be accurately located, on the other hand, the three-level testing realizes end-to-end full-link coverage and verification. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only constitute the embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of the provided drawings.

[0049] Figure 1 The flowchart of the voice simultaneous interpretation system testing method provided by the embodiment of the application;

[0050] Figure 2 The schematic diagram of the test dimension, the test item under the test dimension and the test index of the test item of the voice simultaneous interpretation system testing method provided by the embodiment of the application;

[0051] Figure 3 The structural schematic diagram of the voice simultaneous interpretation system testing device provided by the embodiment of the application. DETAILED DESCRIPTION

[0052] The embodiments of the application will be described below in conjunction with the drawings in the embodiments of the application. The terms used in the embodiment part of the application are only used to explain the specific embodiments of the application, and are not intended to limit the application.

[0053] The embodiments of the application will be described below in conjunction with the drawings in the embodiments of the application. The embodiments of the application provided by the embodiments of the application are also applicable to similar technical problems as the technology develops and new scenarios appear.

[0054] The terms "first", "second", and the like in the description and in the claims of the present application and above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the terms so used are interchangeable under appropriate circumstances and are merely employed to distinguish one object from another. Furthermore, the terms "comprise", "comprising", "have", "having", "include", "including", "contain", "containing", and any variations thereof are intended to cover a non-exclusive inclusion, such that processes, methods, systems, products, or devices that comprise, have, include, contain, or are otherwise including a series of elements do not include only those elements but can include other elements not expressly listed or inherent to such processes, methods, systems, products, or devices.

[0055] The current voice simultaneous interpretation system test method has problems such as fragmented test dimensions (such as only verifying the voice recognition accuracy or the synthesis naturalness), incomplete scene coverage, missing quantitative index system, and the like. In view of the problems of the current voice simultaneous interpretation system test method, the present application provides a systematic test method covering the whole link, multiple scenes, and being quantifiable. Next, the voice simultaneous interpretation system test method provided by the present application will be introduced through the following embodiments.

[0056] Please refer to Figure 1 , which shows a flowchart of the voice simultaneous interpretation system test method provided by the embodiment of the present application. The method can include:

[0057] Step S101: Obtain a test audio set.

[0058] Among the test audio set, there are multiple test audios. The test audios in the test audio set are audios input to the voice simultaneous interpretation system under test for processing.

[0059] In one possible implementation, a basic corpus set, an interference corpus set, and a professional field corpus set can be constructed. Then, the test audio set is obtained based on the basic corpus set, the interference corpus set, and the professional field corpus set.

[0060] Among the basic corpus set, there are several basic corpora. Each basic corpus can include several basic audios in general fields without interference, and related information of the basic audios (such as the standard source language text, the standard target language text, the standard target language audio, and the like corresponding to the basic audios (i.e. source language audios)). Preferably, the basic audios in the basic corpus set cover multiple languages (such as Chinese, English, Japanese, Korean, and the like), multiple speech speeds (such as slow, normal, fast, extremely fast, variable speed, and the like), and multiple types of scenes (such as conference, interview, speech, telephone, live broadcast, daily conversation, academic report, news broadcast, and the like).

[0061] The interference corpus set includes multiple interference corpora, such as background noise, overlapping speech, and abnormal input. The background noise can include some or all of the following: office noise (such as keyboard sound, air conditioner sound, etc.), street noise (such as vehicle sound, horn, crowd noise, construction noise, etc.), restaurant noise (such as cutlery collision sound, dense human voice, background music, etc.), conference room echo, etc. The overlapping speech refers to speech of multiple people speaking at the same time (such as 2-5 people speaking at the same time). The abnormal input can include some or all of the following: silence (such as 30s of silence), super-long speech (such as 1 hour of super-long speech), and mixed dialect and standard Chinese.

[0062] The professional field corpus set includes a plurality of professional field corpora. Each professional field corpus can include professional field audio and related information of the professional field audio (such as standard source language text, standard target language text, and standard target language audio corresponding to the professional field audio). Preferably, the professional field corpora in the professional field corpus set cover multiple vertical fields, such as the medical field, the legal field, the manufacturing field, the financial field, and the technology / IT field. The professional field audio is audio containing professional terms of a vertical field. For example, the medical field contains medical terms such as "myocardial infarction", "targeted drug", and "coronary angiography"; the legal field contains legal terms such as "joint and several liability", "statute of limitations", and "presumption of innocence"; the manufacturing field contains manufacturing terms such as "numerical control lathe", "quenching", and "injection molding"; the financial field contains financial terms such as "profit rate", "hedge fund", and "options futures"; and the technology / IT field contains IT terms such as "application programming interface", "deep learning", and "blockchain".

[0063] The test audio set can be obtained by constructing a plurality of interference-containing audios (the interference corpus in the interference corpus subset can be superimposed on the base audio in the base corpus set to obtain the interference-containing audios, so as to simulate the audio in a real scene) as general field test audios and the professional field audio in the professional field corpus set as a professional field test audio.

[0064] Step S102: inputting the test audio in the test audio set into the target speech simultaneous interpretation system for processing.

[0065] The target speech simultaneous interpretation system is a speech simultaneous interpretation system to be tested. The target speech simultaneous interpretation system includes a speech recognition module, a language conversion module (i.e., a translation module), and a speech synthesis module.

[0066] The target speech simultaneous interpretation system inputs each test audio in the test audio set, a speech recognition module in the target speech simultaneous interpretation system performs speech recognition on the test audio (i.e., source language audio), the recognition text output by the speech recognition module is input into a language conversion module (i.e., translation module) of the speech simultaneous interpretation system, the language conversion module converts the input recognition text (source language text) into target language text, the target language text output by the language conversion module is input into a speech synthesis module of the speech simultaneous interpretation system for speech synthesis, the speech synthesis module synthesizes the audio corresponding to the target language text, and outputs the target language audio.

[0067] Step S103: determining the test results of the target speech simultaneous interpretation system in several test dimensions according to the processing condition data of the target speech simultaneous interpretation system for the test audio in the test audio set.

[0068] The processing condition data of the target speech simultaneous interpretation system for the test audio in the test audio set includes data involved in the processing of the target speech simultaneous interpretation system for the test audio in the test audio set, such as processing time, processing result, etc.

[0069] The several test dimensions can include a user experience dimension, a system function dimension, and a core module dimension, i.e., the target speech simultaneous interpretation system is tested in the user experience dimension, the system function dimension, and the core module dimension.

[0070] The user experience dimension focuses on the user's perception of the end-to-end speech simultaneous interpretation system, and the test in the user experience dimension is used to simulate real users to externally evaluate the overall use of the speech simultaneous interpretation system. The user's perception of the end-to-end speech simultaneous interpretation system refers to the user's comprehensive understanding and feeling of the overall processing process and the final output result of the speech simultaneous interpretation system. It should be noted that the test in the user experience dimension does not care how the speech simultaneous interpretation system works internally, but from the user's perspective, it cares whether the speech simultaneous interpretation system helps the user to accurately, easily, and timely understand the content of the source language audio.

[0071] The system function dimension focuses on the collaborative working condition of the internal modules (speech recognition module, language conversion module, speech synthesis module) of the speech simultaneous interpretation system, and the test in the system function dimension is used to verify the correctness and stability of the collaborative working of the internal modules of the speech simultaneous interpretation system.

[0072] The core module dimension focuses on the performance of the independent modules (speech recognition module, language conversion module, speech synthesis module) in the speech simultaneous interpretation system, i.e., the test in the core module dimension is used to verify the performance of each module itself in the speech simultaneous interpretation system.

[0073] The voice simultaneous interpretation system test method provided in the embodiments of the present application first acquires a test audio set, then inputs test audio in the test audio set into a target voice simultaneous interpretation system for processing, and finally determines test results of the target voice simultaneous interpretation system in user experience dimensions, system function dimensions, and core module dimensions according to processing condition data of the test audio in the test audio set by the target voice simultaneous interpretation system. The voice simultaneous interpretation system test method provided in the embodiments of the present application classifies complex system testing into three clear levels (user experience -> system function -> core module), on the one hand, the three levels of testing form a positioning funnel from top to bottom and layer by layer, that is, through the three levels of testing, problems can be found and the root cause of the problems can be accurately located, on the other hand, through the three levels of testing, end-to-end full-link coverage and verification are realized.

[0074] In some embodiments of the present application, the specific implementation process of "step S103: determining test results of the target voice simultaneous interpretation system in a plurality of test dimensions according to the processing condition data of the test audio in the test audio set by the target voice simultaneous interpretation system" is introduced.

[0075] The process of determining test results of the target voice simultaneous interpretation system in a plurality of test dimensions according to the processing condition data of the test audio in the test audio set by the target voice simultaneous interpretation system can include:

[0076] Step a1, for each test dimension, obtaining data related to each test item in the test dimension from the processing condition data of the test audio in the test audio set by the target voice simultaneous interpretation system.

[0077] The processing condition data of the test audio in the test audio set by the target voice simultaneous interpretation system can include, but is not limited to, part or all of the following data: the time (such as the start input time and the end input time) when the test audio in the test audio set is input into the target voice simultaneous interpretation system, the time (such as the time when the first valid syllable is output and the time when the last syllable is output) when the target voice simultaneous interpretation system outputs target language audio for the test audio (i.e. source language audio), the time when the speech recognition module outputs recognized text, the time when the text output by the language conversion module is input into the speech synthesis module, the time when the speech synthesis module outputs synthesized audio, the processing results (such as recognized text output by the speech recognition result, translated text output by the language conversion result, and synthesized audio output by the speech synthesis result) of each module of the target voice simultaneous interpretation system, and the data accumulation situation between modules in the target voice simultaneous interpretation system, and the like.

[0078] The present application sets a plurality of test items for each test dimension. The test item in a test dimension refers to the item or content to be tested in the test dimension.

[0079] Step a2, determining the test index of each test item under the test dimension according to the data related to each test item under the test dimension.

[0080] The test index of any test item is used to measure and evaluate the performance of the target speech simultaneous interpretation system on the test item.

[0081] Step a3, determining the test result of the target speech simultaneous interpretation system on each test item under the test dimension according to the test index of each test item under the test dimension.

[0082] Specifically, each test index of each test item under each test dimension corresponds to a preset reference threshold, and the test result corresponding to each test index of each test item under each test dimension can be determined by comparing each test index of each test item under each test dimension with the corresponding reference threshold.

[0083] Please refer to Figure 2 , which shows examples of test items and test indexes of test items under the user experience dimension, the system function dimension and the core module dimension.

[0084] The test items under the user experience dimension can include part or all (preferably all) of the following test items: real-time test item, fluency test item, subjective experience test item.

[0085] Among them, the real-time test item under the user experience dimension is used to evaluate the subjective perception of the whole target speech simultaneous interpretation system on the user's waiting or time lag, the fluency test item under the user experience dimension is used to evaluate the continuity and naturalness of the output audio of the target speech simultaneous interpretation system, that is, whether the user perceives the simultaneous interpretation process to be smooth and uninterrupted, and whether it conforms to the rhythm and habit of human speech, and the subjective experience test item under the user experience dimension evaluates the comprehensive sensory quality and acceptability of the output audio of the target speech simultaneous interpretation system, that is, the effect of the output audio of the target speech simultaneous interpretation system on the perception and cognition of people.

[0086] The test index of the real-time test item under the user experience dimension can include part or all of the following indexes: first response delay, tail response delay.

[0087] Among them, the first response delay is the time difference between the start of the test audio (i.e. the source language audio) input to the first valid syllable of the target language audio output by the speech simultaneous interpretation system, and the tail response delay is the time difference between the end of the test audio (i.e. the source language audio) input to the last syllable of the target language audio output by the speech simultaneous interpretation system.

[0088] For each test audio in the test audio set, a first response delay corresponding to the test audio can be determined according to a time at which the test audio starts to be input into the target speech simultaneous interpretation system and a time at which the target speech simultaneous interpretation system outputs a first valid phoneme of the target language audio corresponding to the test audio, and a last response delay corresponding to the test audio can be determined according to a time at which the input of the test audio ends and a time at which the target speech simultaneous interpretation system outputs a last phoneme of the target language audio corresponding to the test audio.

[0089] After obtaining the first response delays corresponding to the test audios in the test audio set respectively, a first response delay of the target speech simultaneous interpretation system on the entire test audio set can be calculated. Specifically, a mean of the first response delays corresponding to the test audios in the test audio set respectively can be calculated as a first first response delay of the target speech simultaneous interpretation system on the entire test audio set, and / or, the first response delays corresponding to the test audios in the test audio set respectively can be sorted in ascending order, and a first response delay at a target position (such as a 95% position) can be taken as a second first response delay of the target speech simultaneous interpretation system on the entire test audio set, and / or, a mean of the first response delays corresponding to the test audios in the short sentence scenario respectively can be calculated as a short sentence first response delay of the target speech simultaneous interpretation system on the entire test audio set, and / or, a mean of the first response delays corresponding to the test audios in the long sentence scenario respectively can be calculated as a long sentence first response delay of the target speech simultaneous interpretation system on the entire test audio set. Exemplarily, a short sentence can be a sentence with less than or equal to 5 words, and a long sentence can be a sentence with more than 20 words.

[0090] After obtaining the last response delays corresponding to the test audios in the test audio set respectively, a last response delay of the target speech simultaneous interpretation system on the entire test audio set can be determined. Specifically, a mean of the last response delays corresponding to the test audios in the test audio set respectively can be calculated as the last response delay of the target speech simultaneous interpretation system on the entire test audio set.

[0091] The following table shows the definition, calculation method, unit and reference threshold of the first response delay and the last response delay of the target speech simultaneous interpretation system on the entire test audio set.

[0092] Table 1. Test index examples of real-time test items under user experience dimensions

[0093]

[0094] It should be noted that after obtaining the test index, the test result corresponding to the test index can be determined according to the test index and the reference threshold corresponding to the test index. Taking the first first response delay as an example, if the first first response delay is less than or equal to 500 ms, the test result is determined to be “excellent”, if the first first response delay is greater than 500 ms and less than or equal to 800 ms, the test result of the target speech simultaneous interpretation system is determined to be “acceptable”, and other indexes are the same.

[0095] The test indicators of the fluency test item under the user experience dimension can include some or all of the following indicators: non-punctuation position stuttering frequency, single stuttering maximum duration, and translation stuttering times.

[0096] The non-punctuation position stuttering frequency is the number of stuttering at non-punctuation positions per T minutes (such as 10 minutes), the single stuttering maximum duration is the longest stuttering time among all stuttering times corresponding to all stuttering, and the translation stuttering times is the number of times of translation stuttering.

[0097] The stuttering of the voice simultaneous interpretation system includes broadcast stuttering and translation stuttering. The translation stuttering is stuttering caused by the language conversion module (i.e., the translation module) in the voice simultaneous interpretation system failing to keep up with the processing speed and failing to provide translated text to the speech synthesis module in time, and the broadcast stuttering is stuttering caused by the speech synthesis module itself in the process of synthesizing speech after receiving the translated text.

[0098] In order to evaluate the fluency of the entire target voice simultaneous interpretation system, indicators representing stuttering of the target voice simultaneous interpretation system can be determined, such as non-punctuation position stuttering frequency, single stuttering maximum duration, and translation stuttering times.

[0099] During the test, the stuttering can be identified by analyzing the waveform of the output audio of the target voice simultaneous interpretation system (for example, at non-punctuation positions, if the silence duration is greater than 300 ms, it can be determined that stuttering is identified), and then the non-punctuation position stuttering frequency, single stuttering maximum duration, etc. can be counted.

[0100] The definition, calculation method, unit, and reference threshold of the test indicators of the fluency test item under the user experience dimension are shown in the following table.

[0101] Table 2 Example of test indicators of fluency test item under user experience dimension

[0102]

[0103] It should be noted that after obtaining the test indicators, the test results corresponding to the test indicators can be determined according to the test indicators and the reference thresholds corresponding to the test indicators. Taking the non-punctuation position stuttering frequency as an example, if the non-punctuation position stuttering frequency is less than or equal to 1 time, the test result is determined to be “excellent”, if the non-punctuation position stuttering frequency is greater than 1 time and less than or equal to 3 times, the test result is determined to be “acceptable”, and other indicators are the same.

[0104] The test indicators of the subjective experience test item under the user experience dimension can include some or all of the following indicators: the adaptation of the simultaneous interpretation speed of the voice simultaneous interpretation system to the scene, the intelligibility and naturalness of the final output audio of the voice simultaneous interpretation system.

[0105] The speed of the simultaneous interpretation system, the adaptability to the scene, the intelligibility and naturalness of the final output audio of the simultaneous interpretation system are all subjective indexes given by multiple evaluators (such as 20 evaluators, including 5 professional interpreters).

[0106] After the multiple evaluators compare the test audio (i.e. the source language audio) and the target language audio output by the target voice simultaneous interpretation system for the test audio, the intelligibility (whether the semantic meaning can be accurately understood) and the naturalness (whether it conforms to the natural speaking habit) of the target language audio are evaluated and scored (for example, a 1-5 scoring system can be used). It should be noted that when scoring, multiple evaluators score the same target language audio (the audio output by the target voice simultaneous interpretation system), and the average score of the multiple evaluators for the same target language audio is taken as the score of the target language audio. After obtaining the scores of the target language audio corresponding to each test audio in the test audio set, the average score of the target language audio corresponding to each test audio can be obtained.

[0107] The multiple evaluators can also score whether the simultaneous interpretation speed of the target voice simultaneous interpretation system for the test audio matches the scene (a 1-5 scoring system can be used) to obtain the adaptability score of the simultaneous interpretation speed of the voice simultaneous interpretation system to the scene.

[0108] As shown in FIG. 8, the test items under the system function dimension can include part or all of the following test items: cross-language conversion accuracy test item, multi-module collaborative performance test item. Figure 2

[0109] Among them, the cross-language conversion accuracy test item under the system function dimension evaluates whether the meaning expressed by the test audio is completely and correctly converted and transmitted to the target language audio through the cooperation of each module in the target voice simultaneous interpretation system, and the multi-module collaborative performance test item under the system function dimension evaluates the fluency, stability and timeliness of the collaborative work of each module in the target voice simultaneous interpretation system.

[0110] The test indicators of the cross-language conversion accuracy test item under the system function dimension can include part or all of the following indicators: word-level accuracy indicator of cross-language conversion, sentence-level semantic integrity indicator of cross-language conversion, context adaptability indicator of cross-language conversion.

[0111] ​The word-level accuracy index of cross-language conversion can be the word-level translation accuracy from the "recognition-conversion" link, the sentence-level semantic integrity index of cross-language conversion can be the proportion of the original sentence core semantic roles (agent, patient, etc.) included in the audio output by the speech simultaneous interpretation system, and the context adaptability index of cross-language conversion can be the correct conversion rate of pronouns (such as he, she, it) and referring words (this, that, etc.) in multi-round dialogue. The test index of the cross-language conversion accuracy test item can be determined by analyzing the translation text output by the language conversion module and the synthesized audio output by the speech synthesis module, and combining the corresponding standard target language text and standard target language audio. The following table shows the definition, calculation method, unit and reference threshold of the sentence-level semantic integrity index of cross-language conversion.

[0112] Table 3 Sentence-level semantic integrity index of cross-language conversion

[0113]

[0114] The test index of the multi-module collaborative performance test item under the system function dimension can include some or all of the following indexes: an index representing the data backlog between modules in the speech simultaneous interpretation system, and an index representing the synchronization between modules in the speech simultaneous interpretation system.

[0115] The index representing the data backlog between modules in the speech simultaneous interpretation system can include the backlog amount of data between modules and the backlog duration of data between modules, and the index representing the synchronization between modules in the speech simultaneous interpretation system can include the mean of the time difference between the output of the speech recognition module and the input of the speech synthesis module (the mean of the time difference corresponding to each test audio in the test audio set, and the time difference corresponding to a test audio is the time difference between the input of the translation text of the recognized text of the test audio into the speech synthesis module and the output of the recognized text of the test audio by the speech recognition module).

[0116] When testing the target speech simultaneous interpretation system on the multi-module collaborative performance test item, backlog testing and synchronization testing can be performed. Backlog testing refers to testing the backlog between modules in the target speech simultaneous interpretation system under a high concurrency scenario (such as 100 simultaneous simultaneous interpretation). Synchronization testing refers to testing the synchronization of modules in the target speech simultaneous interpretation system.

[0117] The following table shows the definition, calculation method, unit and reference threshold of the test index of the multi-module collaborative performance test item under the system function dimension.

[0118] Table 4 Example of test index of multi-module collaborative performance test item under system function dimension

[0119]

[0120] The test items under the core module dimension can include some or all of the following test items: speech recognition module performance test items, language conversion module performance test items, and speech synthesis module performance test items.

[0121] The speech recognition module performance test items are used to test the performance of the speech recognition module, the language conversion module performance test items are used to test the performance of the language conversion module, and the speech synthesis module performance test items are used to test the performance of the speech synthesis module.

[0122] The test indicators of the speech recognition module performance test items under the core module dimension can include some or all of the following indicators: speech recognition word accuracy, speech recognition sentence accuracy, and endpoint detection accuracy. The definitions, calculation methods, units, and reference thresholds of the speech recognition word accuracy and the speech recognition sentence accuracy are shown in the following table.

[0123] Table 5: Examples of test indicators of speech recognition module performance test items under the core module dimension

[0124]

[0125] The reference word number is the number of words in the standard text of the test audio.

[0126] The test indicators of the language conversion module performance test items under the core module dimension include some or all of the following indicators: translation delay, translation (or conversion) accuracy for professional terms, conversion naturalness for specified sentence patterns (such as complex sentence patterns), and overlap degree with standard translation (i.e., translation BLEU score).

[0127] The definitions, calculation methods, units, and reference thresholds of the translation BLEU score and the translation accuracy for professional terms are shown in the following table.

[0128] Table 6: Examples of test indicators of language conversion module performance test items under the core module dimension

[0129]

[0130] The test indexes of the performance test items of the speech synthesis module under the core module dimension include some or all of the following indexes: mean opinion score of the synthesized audio, synthesis accuracy rate for multi-syllable words, synthesis text regularization accuracy rate, and indexes reflecting the real-time performance of the audio synthesis. The indexes reflecting the real-time performance of the audio synthesis can include throughput, first word delay, etc. Among them, the synthesis accuracy rate for multi-syllable words and the synthesis text regularization accuracy rate can be determined by analyzing the synthesized audio output by the speech synthesis module, and the indexes reflecting the real-time performance of the audio synthesis can be determined by analyzing the time when the translation text of the recognition result of the test audio is input into the speech synthesis module and the time when the synthesized audio is output by the speech synthesis module. The following table shows the definition, calculation method, unit, and reference threshold of some test indexes of the performance test items of the speech synthesis module under the core module dimension.

[0131] Table 7 Examples of test indexes of performance test items of speech synthesis module under core module dimension

[0132]

[0133] Among them, text regularization is a preprocessing step of speech synthesis, which refers to the process of converting “non-standard” written text into “standard” readable text that conforms to pronunciation rules. In simple terms, text regularization is the process of converting numbers, symbols, abbreviations, dates, currencies, etc. in the text into complete words or phrases that can be read correctly.

[0134] In some embodiments of the present application, the speech simultaneous interpretation system testing method can further include: generating an index distribution heat map (for example, a delay distribution heat map of each scene) of each scene, each speech speed, or each language according to the test indexes of the test items under the test dimensions; and / or drawing a performance correlation curve of different modules (for example, a correlation curve of recognition accuracy and synthesis naturalness) according to the test indexes of the test items under the test dimensions, so that the tester can intuitively understand the performance of the speech simultaneous interpretation system.

[0135] In some embodiments of the present application, the speech simultaneous interpretation system testing method can further include: determining the problems existing in the speech simultaneous interpretation system according to the test results of the speech simultaneous interpretation system under the test dimensions (the test results corresponding to the test indexes of the test items under the test dimensions); and outputting the problems existing in the speech simultaneous interpretation system.

[0136] When analyzing the test results under the test dimensions (the test results corresponding to the test indexes of the test items under the test dimensions), the test indexes of the test items under the test dimensions, and the processing data of the speech simultaneous interpretation system for the test audio in the test audio set can be analyzed.

[0137] Optionally, when the problems existing in the output speech simultaneous interpretation system are output, the problems can be ranked according to the degree of influence on the user experience, and then topK problems (K problems with the greatest influence on the user experience) are output.

[0138] In some embodiments of the present application, the speech simultaneous interpretation system testing method can further include: constructing a problem use case library according to the problems existing in the speech simultaneous interpretation system, so as to test the optimized speech simultaneous interpretation system based on the problem use case library after the speech simultaneous interpretation system is optimized.

[0139] Among them, the problem use case library can include the problems existing in the speech simultaneous interpretation system and the test data corresponding to the problems existing in the speech simultaneous interpretation system. After the speech simultaneous interpretation system is optimized, the test data in the problem use case library can be used to test the optimized speech simultaneous interpretation system.

[0140] When the optimized speech simultaneous interpretation system is tested, 100% coverage testing of the problem use cases in the problem use case library can be performed to ensure that the repair rate is ≥ 99%. Among them, the repair rate is the ratio of the number of problem use cases passing the test to the total number of problem use cases.

[0141] In addition, the speech simultaneous interpretation system testing method provided by the present application can be used to test other speech simultaneous interpretation systems under the same test environment, and then the differences between the test indicators of the target speech simultaneous interpretation system and the test indicators of other speech simultaneous interpretation systems can be compared, and then the optimization direction of the target speech simultaneous interpretation system can be determined according to the index difference. Optionally, a differential analysis report of the test indicators of the target speech simultaneous interpretation system and the test indicators of other speech simultaneous interpretation systems can also be output.

[0142] The present application also provides a speech simultaneous interpretation system testing device, as shown in Figure 3 The speech simultaneous interpretation system testing device can include a data acquisition unit 301, a data input unit 302, and a data processing unit 303.

[0143] The data acquisition unit 301 is configured to acquire a test audio set.

[0144] The data input unit 302 is configured to input the test audio in the test audio set into the target speech simultaneous interpretation system for processing.

[0145] The data processing unit 303 is configured to determine the test results of the target speech simultaneous interpretation system in several test dimensions according to the processing data of the test audio in the test audio set by the target speech simultaneous interpretation system.

[0146] The several test dimensions include a user experience dimension, a system function dimension, and a core module dimension. The user experience dimension focuses on user perception of the target voice simultaneous interpretation system end-to-end. The system function dimension focuses on the collaborative work of modules in the target voice simultaneous interpretation system. The core module dimension focuses on the performance of each independent module in the target voice simultaneous interpretation system.

[0147] In a possible implementation, the process in which the data acquisition unit 301 acquires the test audio set can include:

[0148] The constructed basic corpus set, interference corpus set, and professional field corpus set are acquired, where the basic corpus set includes basic audios of several general fields, the basic audios in the basic corpus set cover multiple languages, multiple speech speeds, and multiple types of scenes, the interference corpus set includes multiple interference audios, and the professional field corpus set includes several professional field audios.

[0149] According to the basic audios in the basic corpus set and the interference audios in the interference corpus set, several interference-containing audios are constructed as general field test audios. The professional field audios in the professional field corpus set are taken as professional field test audios, to obtain a test audio set including the general field test audios and the professional field test audios.

[0150] In a possible implementation, each test dimension includes several test items.

[0151] According to the processing condition data of the target voice simultaneous interpretation system on the test audios in the test audio set, test results of the target voice simultaneous interpretation system in the several test dimensions are determined, including:

[0152] For each test dimension, data related to each test item in the test dimension is obtained from the processing condition data of the target voice simultaneous interpretation system on the test audios in the test audio set.

[0153] According to the data related to each test item in the test dimension, a test index of each test item in the test dimension is determined.

[0154] According to the test index of each test item in the test dimension, a test result of the target voice simultaneous interpretation system on each test item in the test dimension is determined.

[0155] In a possible implementation, the test items in the user experience dimension include some or all of the following test items: real-time performance test item, fluency test item, and subjective experience test item.

[0156] The test items under the system function dimension include part or all of the following test items: a cross-language conversion accuracy test item and a multi-module coordination performance test item. The cross-language conversion accuracy test item evaluates whether the meaning expressed by the test audio is accurately and completely converted and delivered into the target language audio through the cooperation of the modules. The multi-module coordination performance test item evaluates the fluency, stability and timeliness of the cooperation of the modules.

[0157] The test items under the core module dimension include part or all of the following test items: a speech recognition module performance test item, a language conversion module performance test item and a speech synthesis module performance test item.

[0158] In a possible implementation, the test indicators of the real-time test item under the user experience dimension include part or all of the following indicators: a first response delay and a tail response delay. The first response delay is a time difference between the start of input of the test audio and the output of the first valid syllable of the target language audio by the target speech simultaneous interpretation system. The tail response delay is a time difference between the end of input of the test audio and the output of the last syllable of the target language audio by the target speech simultaneous interpretation system.

[0159] The test indicators of the fluency test item under the user experience dimension include part or all of the following indicators: a non-punctuation position stuttering frequency, a single stuttering longest duration and a translation stuttering number.

[0160] The test indicators of the subjective experience test item under the user experience dimension include part or all of the following indicators: an adaptation degree of the simultaneous interpretation speed of the target speech simultaneous interpretation system to the scene, an intelligibility and naturalness of the final output audio of the target speech simultaneous interpretation system.

[0161] In a possible implementation, the first response delay and the tail response delay of the real-time test item are determined according to data related to the real-time test item under the user experience dimension, and the determination includes the following steps.

[0162] For each test audio in the test audio set, the first response delay and the tail response delay corresponding to the test audio are determined according to the time at which the test audio is input into the target speech simultaneous interpretation system, the time at which the target speech simultaneous interpretation system outputs the first valid syllable of the target language audio for the test audio, the end time of input of the test audio, and the time at which the target speech simultaneous interpretation system outputs the last syllable of the target language audio for the test audio.

[0163] The mean of the first response delay corresponding to each test audio in the test audio set is calculated to obtain the first first response delay of the target speech simultaneous interpretation system; and / or, the first response delay corresponding to each test audio in the test audio set is sorted in ascending order to obtain the first response delay of the target position as the second first response delay of the target speech simultaneous interpretation system; and / or, the mean of the first response delay corresponding to each test audio in the short sentence scenario is calculated to obtain the short sentence first response delay of the target speech simultaneous interpretation system; and / or, the mean of the first response delay corresponding to each test audio in the long sentence scenario is calculated to obtain the long sentence first response delay of the target speech simultaneous interpretation system.

[0164] The mean of the first response delay corresponding to each test audio in the test audio set is calculated to obtain the first response delay of the target speech simultaneous interpretation system; and / or, the first response delay corresponding to each test audio in the test audio set is sorted in ascending order to obtain the first response delay of the target position as the second first response delay of the target speech simultaneous interpretation system; and / or, the mean of the first response delay corresponding to each test audio in the short sentence scenario is calculated to obtain the short sentence first response delay of the target speech simultaneous interpretation system; and / or, the mean of the first response delay corresponding to each test audio in the long sentence scenario is calculated to obtain the long sentence first response delay of the target speech simultaneous interpretation system.

[0165] In a possible implementation, the test indicators of the cross-language conversion accuracy test item in the system function dimension include part or all of the following indicators: a word-level accuracy indicator of cross-language conversion, a sentence-level semantic integrity indicator of cross-language conversion, and a context adaptability indicator of cross-language conversion.

[0166] The test indicators of the multi-module collaborative performance test item in the system function dimension include part or all of the following indicators: an indicator representing the data backlog between modules in the target speech simultaneous interpretation system, and an indicator representing the synchronization between modules in the target speech simultaneous interpretation system.

[0167] In a possible implementation, the word-level accuracy indicator includes a word-level translation accuracy, the sentence-level semantic integrity indicator includes a proportion of sentences that retain core semantic roles, and the context adaptability indicator of cross-language conversion includes a correct conversion rate of pronouns and referring words in a multi-turn dialogue.

[0168] The indicator representing the data backlog between modules in the target speech simultaneous interpretation system includes a backlog amount of data between modules and / or a backlog duration of data between modules, and the indicator representing the synchronization between modules in the target speech simultaneous interpretation system includes a time difference between the output of the speech recognition module and the input of the speech synthesis module.

[0169] In a possible implementation, the test indicators of the speech recognition module performance test item in the core module dimension include part or all of the following indicators: a speech recognition word accuracy, a speech recognition sentence accuracy, and an endpoint detection accuracy.

[0170] The test indicators of the language conversion module performance test item in the core module dimension include part or all of the following indicators: a translation delay, a conversion accuracy for professional terms, a conversion naturalness for specified sentence patterns, and an overlap degree with a standard translation.

[0171] The test indexes of the performance test items of the speech synthesis module under the core module dimension include some or all of the following indexes: mean opinion score of synthesized audio, synthesized multi-sound character accuracy, synthesized text regularization accuracy, and an index representing audio synthesis real-time performance.

[0172] In a possible implementation, the speech simultaneous interpretation system testing apparatus can further include an index distribution thermodynamic map generation unit and / or a performance correlation curve drawing unit.

[0173] The index distribution thermodynamic map generation unit is configured to generate an index distribution thermodynamic map of each scenario, each speech speed, or each language according to the test indexes of the test items under the test dimensions.

[0174] The performance correlation curve drawing unit is configured to draw a performance correlation curve of different modules according to the test indexes of the test items under the test dimensions.

[0175] In a possible implementation, the speech simultaneous interpretation system testing apparatus can further include a problem determination module, a problem output unit, and / or a problem use case library construction unit.

[0176] The problem determination unit is configured to determine a problem existing in the target speech simultaneous interpretation system according to the test results of the target speech simultaneous interpretation system under the test dimensions.

[0177] The problem output unit is configured to output the problem existing in the target speech simultaneous interpretation system.

[0178] The problem use case library construction unit is configured to construct a problem use case library according to the problem existing in the target speech simultaneous interpretation system, so as to test the optimized target speech simultaneous interpretation system based on the problem use case library.

[0179] The embodiments of the present application also provide an electronic device including at least one processor and a memory connected to the processor, wherein:

[0180] The memory is configured to store a computer program.

[0181] The processor is configured to execute the computer program, so that the electronic device can implement the steps of the speech simultaneous interpretation system testing method provided by the above embodiments.

[0182] The present disclosure also provides a computer readable storage medium having computer program instructions stored thereon, the program instructions being executed by a processor to implement the steps of the speech simultaneous interpretation system testing method provided by the above embodiments.

[0183] The embodiment of the present application further provides a computer program product comprising computer readable instructions which, when executed on an electronic device, cause the electronic device to implement the steps of the voice transmission system testing method provided by the above embodiment.

[0184] In addition, it should be noted that the above-described apparatus embodiments are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. In addition, the connection relationship between the modules in the apparatus embodiments provided by the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.

[0185] Through the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and the necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions to make a computer device (which can be a personal computer, training device, or network device, etc.) execute the methods described in various embodiments of the present application.

[0186] In the above embodiments, all or part can be realized by software, hardware, firmware or any combination thereof. When realized by software, it can be realized in the form of a computer program product in whole or in part.

[0187] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as a training device, a data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

Claims

1. A method for testing a voice relay system, characterized in that, The method comprises: acquiring a test audio set; inputting test audio in the test audio set into a target voice simultaneous interpretation system for processing; determining test results of the target voice simultaneous interpretation system in a plurality of test dimensions according to processing data of the target voice simultaneous interpretation system on test audio in the test audio set; wherein the plurality of test dimensions comprise a user experience dimension, a system function dimension, and a core module dimension, the user experience dimension focuses on user perception of an end-to-end target voice simultaneous interpretation system, the system function dimension focuses on the collaborative work of modules in the target voice simultaneous interpretation system, and the core module dimension focuses on the performance of each independent module in the target voice simultaneous interpretation system; the test items in the user experience dimension comprise some or all of the following test items: real-time performance test item, fluency test item, and subjective experience test item; and the test indicators of the subjective experience test item in the user experience dimension comprise some or all of the following indicators: adaptation of simultaneous interpretation speed of the target voice simultaneous interpretation system to a scene, intelligibility, and naturalness of final output audio of the target voice simultaneous interpretation system; the test items in the core module dimension comprise some or all of the following test items: voice recognition module performance test item, language conversion module performance test item, and voice synthesis module performance test item; and the test indicators of the language conversion module performance test item in the core module dimension comprise some or all of the following indicators: translation delay, conversion accuracy for professional terms, conversion naturalness for specified sentence patterns, and overlap with standard translation.

2. The voice over system test method of claim 1, wherein, The acquiring of the test audio set comprises: acquiring a constructed basic corpus set, an interference corpus set, and a professional field corpus set, wherein the basic corpus set comprises a plurality of basic audios in general fields, the basic audios in the basic corpus set cover multiple languages, multiple speech speeds, and multiple types of scenes, the interference corpus set comprises a plurality of interference audios, and the professional field corpus set comprises a plurality of professional field audios; constructing a plurality of interference-containing audios as general field test audios according to the basic audios in the basic corpus set and the interference audios in the interference corpus set, and taking the professional field audios in the professional field corpus set as professional field test audios to obtain a test audio set comprising the general field test audios and the professional field test audios.

3. The voice over system test method of claim 1, wherein, Each of the test dimensions comprises a plurality of test items; The determining of the test results of the target voice simultaneous interpretation system in a plurality of test dimensions according to processing data of the target voice simultaneous interpretation system on test audio in the test audio set comprises: for each of the test dimensions, obtaining data related to each test item in the test dimension from the processing data of the target voice simultaneous interpretation system on test audio in the test audio set; determining test indicators of each test item in the test dimension according to the data related to each test item in the test dimension; determining test results of each test item in the test dimension of the target voice simultaneous interpretation system according to the test indicators of each test item in the test dimension.

4. The voice over system test method of claim 1, wherein, The test items under the system function dimension include some or all of the following test items: a cross-language conversion accuracy test item and a multi-module coordination performance test item. The cross-language conversion accuracy test item evaluates whether the meaning expressed by the test audio is accurately and completely converted and delivered into the target language audio through the cooperation of the modules. The multi-module coordination performance test item evaluates the fluency, stability and timeliness of the cooperation of the modules.

5. The voice over system test method of claim 1, wherein, The test indicators of the real-time test item under the user experience dimension include some or all of the following indicators: a first response delay and a tail response delay. The first response delay is the time difference between the start of the input of the test audio and the output of the first valid syllable of the target language audio by the target speech simultaneous interpretation system. The tail response delay is the time difference between the end of the input of the test audio and the output of the last syllable of the target language audio by the target speech simultaneous interpretation system. The test indicators of the fluency test item under the user experience dimension include some or all of the following indicators: a non-punctuation position stuttering frequency, a single stuttering maximum duration and a translation stuttering number.

6. The voice over system test method of claim 5, wherein, According to the data related to the real-time test item under the user experience dimension, the first response delay and the tail response delay of the real-time test item are determined, including: For each test audio in the test audio set, the first response delay and the tail response delay corresponding to the test audio are determined according to the time when the test audio is input into the target speech simultaneous interpretation system, the time when the first valid syllable of the target language audio is output by the target speech simultaneous interpretation system for the test audio, the end time of the input of the test audio and the time when the last syllable of the target language audio is output by the target speech simultaneous interpretation system for the test audio; The mean value of the first response delays corresponding to each test audio in the test audio set is calculated to obtain a first first response delay of the target speech simultaneous interpretation system. And / or, the first response delays corresponding to each test audio in the test audio set are sorted in ascending order to obtain the first response delay at the target position as a second first response delay of the target speech simultaneous interpretation system. And / or, the mean value of the first response delays corresponding to each test audio in a short sentence scenario is calculated to obtain a short sentence first response delay of the target speech simultaneous interpretation system. And / or, the mean value of the first response delays corresponding to each test audio in a long sentence scenario is calculated to obtain a long sentence first response delay of the target speech simultaneous interpretation system. The mean value of the tail response delays corresponding to each test audio in the test audio set is calculated to obtain a tail response delay of the target speech simultaneous interpretation system.

7. The voice over system test method of claim 4, wherein, The test indicators of the cross-language conversion accuracy test item under the system function dimension include some or all of the following indicators: a word-level accuracy indicator of cross-language conversion, a sentence-level semantic integrity indicator of cross-language conversion and a context adaptability indicator of cross-language conversion. The test indicators of the multi-module coordination performance test item under the system function dimension include some or all of the following indicators: an indicator representing the data backlog among the modules in the target speech simultaneous interpretation system and an indicator representing the synchronization among the modules in the target speech simultaneous interpretation system.

8. The voice over system test method of claim 7, wherein, The word-level accuracy indicator includes a word-level translation accuracy rate, the sentence-level semantic integrity indicator includes a proportion of sentences that retain core semantic roles, and the contextual adaptability indicator of cross-language conversion includes a correct conversion rate of pronouns and referents in multi-turn dialogues. The indicators representing the data backlog between modules in the target voice simultaneous interpretation system include the amount of data backlog and / or the duration of data backlog between modules, and the indicators representing the synchronization between modules in the target voice simultaneous interpretation system include the time difference between the output of the speech recognition module and the input of the speech synthesis module.

9. The voice over system test method of claim 1, wherein, The test indicators of the performance test items of the speech recognition module under the core module dimension include some or all of the following indicators: word-level speech recognition accuracy, sentence-level speech recognition accuracy, and endpoint detection accuracy. The test indicators of the performance test items of the speech synthesis module under the core module dimension include some or all of the following indicators: mean opinion score of synthesized audio, synthesized multi-sound character accuracy, synthesized text regularization accuracy, and indicators representing the real-time performance of audio synthesis.

10. The voice over system test method of claim 3, wherein, Further comprising: generating an index distribution heat map for each scenario, speed, or language based on the test indicators of the test items under the test dimensions; and / or drawing a performance correlation curve of different modules based on the test indicators of the test items under the test dimensions.

11. The voice over system test method of claim 3, wherein, Further comprising: determining the problems of the target voice simultaneous interpretation system based on the test results of the target voice simultaneous interpretation system under the test dimensions; outputting the problems of the target voice simultaneous interpretation system; and / or, constructing a problem use case library based on the problems of the target voice simultaneous interpretation system, so as to test the optimized target voice simultaneous interpretation system based on the problem use case library.

12. An electronic device, comprising: The electronic device comprises at least one processor and a memory connected to the processor, wherein: the memory is used to store a computer program; the processor is used to execute the computer program, so that the electronic device can implement the steps of the voice simultaneous interpretation system testing method according to any one of claims 1-11.

13. A computer storage medium, characterized in that The storage medium carries one or more computer programs, which can enable the electronic device to implement the steps of the voice simultaneous interpretation system testing method according to any one of claims 1-11 when the one or more computer programs are executed by the electronic device.

14. A computer program product, characterised in that, The computer readable instructions enable the electronic device to implement the steps of the voice simultaneous interpretation system testing method according to any one of claims 1-11 when the computer readable instructions are run on the electronic device.

Citation Information

Patent Citations

  • Machine simultaneous interpretation system and method, test method and device and related equipment

    CN116935853A

  • Voice transcription system evaluation method and device, related equipment and program product

    CN120071903A