Multilingual testing method and system for voice interaction system
Patent Information
- Application Number
- CN202611231101.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]但是,上述说出测试语音的过程和判断响应结果是否正确的过程,均由测试人员执行,耗费了较大的人力成本和时间成本,导致多语言测试的效率较低
[0034]本申请实施例提供了一种对语音交互系统的多语言测试方案,能够在车辆座舱内播放属于多种目标测试语言的多个测试语音,分别获取每个测试语音对应的第一响应数据和第二响应数据,第一响应数据由语音交互系统确定,第二响应数据通过人工智能模型确定,每个测试语音对应的第一响应数据与第二响应数据的对比结果能够反映语音交互系统对每个测试语音的响应过程是否正确,因此基于每个测试语音对应的第一响应数据与第二响应数据的对比结果,确定语音交互系统的多语言测试结果。测试过程中无需测试人员说出属于多种目标测试语言的测试语音,也无需测试人员人工判断语音交互系统的响应数据是否正确,降低了人力成本和时间成本,提高了多语言测试的效率。
Smart Images

Figure CN122799809A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technology, and in particular to a multilingual testing method and system for voice interaction systems. Background Technology
[0002] A voice interaction system is an in-vehicle system that uses voice as the interaction medium, enabling drivers to control the vehicle through natural language, thus improving ease of interaction and driving safety. To fully test the cross-language capabilities of a voice interaction system, multilingual testing is necessary.
[0003] Currently, manual testing is commonly used, where testers speak test voices in multiple test languages. Each time the voice interaction system collects the test voices spoken by the testers, it generates response data based on the test voices. The testers then determine the multilingual test results of the voice interaction system by judging whether the response data corresponding to each test voice is correct.
[0004] However, the process of speaking the test audio and judging whether the response is correct is performed by the testers, which consumes a lot of manpower and time, resulting in low efficiency of multilingual testing. Summary of the Invention
[0005] This application provides a multilingual testing method and system for a voice interaction system, reducing labor and time costs and improving the efficiency of multilingual testing. The technical solution is as follows: According to one aspect of the embodiments of this application, a multilingual testing method for a voice interaction system is provided, the method comprising: Multiple test voices are played inside the vehicle cabin, and the multiple test voices belong to various target test languages to be tested; The first response data corresponding to each test voice is obtained; the first response data is obtained by the voice interaction system responding to the test voice. By using an artificial intelligence model, each test speech is responded to separately, and second response data corresponding to each test speech is obtained; Based on the comparison results of the first response data and the second response data corresponding to each test voice, the multilingual test result of the voice interaction system is determined.
[0006] In one possible implementation, the playing of multiple test voices within the vehicle cabin, the multiple test voices belonging to various target test languages to be tested, including: Inside the vehicle cabin, test audio for each of the multiple target test languages is played sequentially according to their order; or, Inside the vehicle cabin, multiple test voices are played in parallel each time, and each of the multiple test voices corresponds one-to-one with the multiple target test languages.
[0007] In one possible implementation, acquiring the first response data corresponding to each test voice statement includes: Receive combined response data sent by the voice interaction system, wherein the combined response data is generated by the voice interaction system based on multiple test voices played in parallel; From the combined response data, extract the first response data corresponding to each test voice.
[0008] In one possible implementation, extracting the first response data corresponding to each test speech from the combined response data includes: The combined response data is split to obtain multiple first response data; Determine the content correlation degree between each test speech and each first response data; Based on the content correlation between each test speech and each first response data, the plurality of first response data are respectively assigned to the plurality of test speech, so that the plurality of first response data corresponds one-to-one with the plurality of test speech.
[0009] In one possible implementation, the method is performed by a multilingual testing system, which includes multiple robotic arms and multiple speakers, each speaker located on one robotic arm; The process involves playing multiple test voice recordings within the vehicle cabin. These multiple test voice recordings belong to various target test languages to be tested, including: Identify multiple target test locations to be tested; each target test location is located within the vehicle's cabin. Move each robotic arm to its corresponding target test position; Each speaker plays a test audio message belonging to each target test language.
[0010] In one possible implementation, the playing of multiple test voices within the vehicle cabin, the multiple test voices belonging to various target test languages to be tested, including: For each target test language, a target test corpus corresponding to the target test language is determined. The target test corpus includes test texts belonging to the target test language and belonging to multiple interaction types. Obtain test text belonging to the target interaction type to be tested from the target test corpus; The test text is converted into the corresponding test speech.
[0011] In one possible implementation, the step of responding to each test speech using an artificial intelligence model to obtain second response data corresponding to each test speech includes: For each test speech, speech recognition is performed on the test speech to obtain the first recognized text corresponding to the test speech; The artificial intelligence model responds to the first identified text to obtain a first response operation that matches the first identified text. Obtain the first response voice that matches the first response operation.
[0012] In one possible implementation, the second response data includes at least one of the first identified text, the first response operation, and the first response speech, and the first response data includes at least one of the second identified text, the second response operation, and the second response speech; The determination of the multilingual test result of the voice interaction system based on the comparison result between the first response data and the second response data corresponding to each test voice includes at least one of the following: Based on the comparison results of the first recognized text and the second recognized text corresponding to each test speech, the speech recognition accuracy of the voice interaction system is determined; Based on the comparison results of the first response operation and the second response operation corresponding to each test voice, the operation decision accuracy of the voice interaction system is determined; Based on the comparison results between the first response voice and the second response voice corresponding to each test voice, the voice quality parameters of the voice interaction system are determined.
[0013] In one possible implementation, the multilingual test results include the average response latency of the voice interaction system; the method further includes: Determine the response delay corresponding to each test voice, wherein the response delay is the time interval between the time when the test voice is played and the time when the response voice of the voice interaction system is acquired, or the response delay is the difference between the time interval between the time when the test voice is played and the time when the response voice of the voice interaction system is acquired and a preset time interval. The average response delay of the voice interaction system is determined based on the response delay corresponding to each test voice.
[0014] According to another aspect of the embodiments of this application, a multilingual testing system is provided, the multilingual testing system comprising: a testing unit and a control unit; The test unit is used to play multiple test voices in the vehicle cabin, and the multiple test voices belong to multiple target test languages to be tested. The control unit is used to acquire first response data corresponding to each test voice; the first response data is obtained by the voice interaction system responding to the test voice. The control unit is also used to respond to each test voice through an artificial intelligence model to obtain second response data corresponding to each test voice. The control unit is further configured to determine the multilingual test result of the voice interaction system based on the comparison result of the first response data and the second response data corresponding to each test voice.
[0015] In one possible implementation, the testing unit is configured to sequentially play test audio belonging to each of the multiple target test languages in the vehicle cabin, according to the order of the multiple target test languages; or, in the vehicle cabin, to play multiple test audios in parallel each time, wherein the multiple test audios correspond one-to-one with the multiple target test languages.
[0016] In one possible implementation, the control unit is configured to receive combined response data sent by the voice interaction system, the combined response data being generated by the voice interaction system based on multiple test voices played in parallel; and to extract the first response data corresponding to each test voice from the combined response data.
[0017] In one possible implementation, the control unit is configured to split the combined response data to obtain multiple first response data; determine the content correlation degree between each test speech and each first response data; and, based on the content correlation degree between each test speech and each first response data, assign the multiple first response data to the multiple test speeches, so that the multiple first response data corresponds one-to-one with the multiple test speeches.
[0018] In one possible implementation, the test unit includes multiple robotic arms and multiple speakers, with each speaker located on one robotic arm; The control unit is also used to determine multiple target test locations to be tested; each target test location is located inside the vehicle cabin; The testing unit is used to move each robotic arm to its corresponding target testing position. The control unit is also used to play test voice belonging to each target test language through each speaker.
[0019] In one possible implementation, the control unit is further configured to, for each target test language, determine a target test corpus corresponding to the target test language, the target test corpus including test text belonging to the target test language and belonging to multiple interaction types; obtain test text belonging to the target interaction type to be tested from the target test corpus; and convert the test text into the corresponding test speech.
[0020] In one possible implementation, the control unit is configured to perform speech recognition on each test speech to obtain a first recognized text corresponding to the test speech; respond to the first recognized text through the artificial intelligence model to obtain a first response operation matching the first recognized text; and obtain a first response speech matching the first response operation.
[0021] In one possible implementation, the second response data includes at least one of the first identified text, the first response operation, and the first response speech, and the first response data includes at least one of the second identified text, the second response operation, and the second response speech; The control unit is configured to perform at least one of the following steps: Based on the comparison results of the first recognized text and the second recognized text corresponding to each test speech, the speech recognition accuracy of the voice interaction system is determined; Based on the comparison results of the first response operation and the second response operation corresponding to each test voice, the operation decision accuracy of the voice interaction system is determined; Based on the comparison results between the first response voice and the second response voice corresponding to each test voice, the voice quality parameters of the voice interaction system are determined.
[0022] In one possible implementation, the multilingual test results include the average response latency of the voice interaction system; the control unit is further configured to determine the response latency corresponding to each test voice, wherein the response latency is the time interval between the completion time of the test voice playback and the time of acquiring the response voice of the voice interaction system, or the response latency is the difference between the time interval between the completion time of the test voice playback and the time of acquiring the response voice of the voice interaction system and a preset time interval; and the average response latency of the voice interaction system is determined based on the response latency corresponding to each test voice.
[0023] According to another aspect of the embodiments of this application, a multilingual testing device for a voice interaction system is provided, the multilingual testing device for the voice interaction system comprising: The voice playback module is used to play multiple test voices in the vehicle cabin, wherein the multiple test voices belong to various target test languages to be tested; The data acquisition module is used to acquire the first response data corresponding to each test voice; the first response data is obtained by the voice interaction system responding to the test voice. The voice response module is used to respond to each test voice through an artificial intelligence model to obtain the second response data corresponding to each test voice. The result determination module is used to determine the multilingual test result of the voice interaction system based on the comparison result between the first response data and the second response data corresponding to each test voice.
[0024] In one possible implementation, the voice playback module is configured to play test voices belonging to each of the multiple target test languages sequentially in the vehicle cabin, according to the order of the multiple target test languages; or, in the vehicle cabin, multiple test voices are played in parallel each time, and the multiple test voices correspond one-to-one with the multiple target test languages.
[0025] In one possible implementation, the data acquisition module is configured to receive combined response data sent by the voice interaction system, the combined response data being generated by the voice interaction system based on multiple test voices played in parallel; and to extract the first response data corresponding to each test voice from the combined response data.
[0026] In one possible implementation, the data acquisition module is configured to split the combined response data to obtain multiple first response data; determine the content correlation between each test speech and each first response data; and, based on the content correlation between each test speech and each first response data, assign the multiple first response data to the multiple test speech, so that the multiple first response data corresponds one-to-one with the multiple test speech.
[0027] In one possible implementation, the multilingual testing device for the voice interaction system includes multiple robotic arms and multiple speakers, with each speaker located on one robotic arm; The voice playback module is used to determine multiple target test locations to be tested; each target test location is located inside the vehicle cabin; each robotic arm is moved to its corresponding target test location; and test voice belonging to each target test language is played through each speaker.
[0028] In one possible implementation, the device further includes: The speech determination module is used to determine the target test corpus corresponding to each target test language, wherein the target test corpus includes test texts belonging to the target test language and belonging to multiple interaction types; obtain test texts belonging to the target interaction type to be tested from the target test corpus; and convert the test texts into corresponding test speech.
[0029] In one possible implementation, the voice response module is configured to perform speech recognition on each test speech to obtain a first recognized text corresponding to the test speech; respond to the first recognized text through the artificial intelligence model to obtain a first response operation matching the first recognized text; and obtain a first response speech matching the first response operation.
[0030] In one possible implementation, the second response data includes at least one of the first identified text, the first response operation, and the first response speech, and the first response data includes at least one of the second identified text, the second response operation, and the second response speech; The result determination module is configured to perform at least one of the following steps: Based on the comparison results of the first recognized text and the second recognized text corresponding to each test speech, the speech recognition accuracy of the voice interaction system is determined; Based on the comparison results of the first response operation and the second response operation corresponding to each test voice, the operation decision accuracy of the voice interaction system is determined; Based on the comparison results between the first response voice and the second response voice corresponding to each test voice, the voice quality parameters of the voice interaction system are determined.
[0031] In one possible implementation, the multilingual test results include the average response latency of the voice interaction system; the result determination module is further configured to determine the response latency corresponding to each test voice, wherein the response latency is the time interval between the completion time of the test voice playback and the time of acquiring the response voice of the voice interaction system, or the response latency is the difference between the time interval between the completion time of the test voice playback and the time of acquiring the response voice of the voice interaction system and a preset time interval; and the average response latency of the voice interaction system is determined based on the response latency corresponding to each test voice.
[0032] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein at least one piece of program code is stored in the computer-readable storage medium, the at least one piece of program code being loaded and executed by a vehicle's processor to implement the multilingual testing method for the voice interaction system described above.
[0033] According to another aspect of the embodiments of this application, a computer program product is provided, the computer program product including computer program code stored in a computer-readable storage medium, a processor of a multilingual testing system reading the computer program code from the computer-readable storage medium, and the processor executing the computer program code to implement the multilingual testing method of the voice interaction system as described above.
[0034] This application provides a multilingual testing scheme for a voice interaction system. It can play multiple test voices belonging to various target test languages within a vehicle cabin, acquiring first and second response data for each test voice. The first response data is determined by the voice interaction system, and the second response data is determined by an artificial intelligence model. The comparison between the first and second response data for each test voice reflects the correctness of the voice interaction system's response process. Therefore, based on the comparison results, the multilingual test result of the voice interaction system is determined. During the testing process, testers do not need to speak the test voices belonging to various target test languages, nor do they need to manually judge the correctness of the voice interaction system's response data, reducing labor and time costs and improving the efficiency of multilingual testing. Attached Figure Description
[0035] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application.
[0037] Figure 2 This is a flowchart of a multilingual testing method for a voice interaction system provided in an embodiment of this application.
[0038] Figure 3 This is a schematic diagram of an exemplary test process provided in an embodiment of this application.
[0039] Figure 4This is a schematic diagram of another exemplary test process provided in the embodiments of this application.
[0040] Figure 5 This is a schematic diagram of the structure of a multilingual testing system provided in an embodiment of this application.
[0041] Figure 6 This is a schematic diagram of the structure of a multilingual testing device for a voice interaction system provided in an embodiment of this application.
[0042] Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0044] It is understood that the terms “first,” “second,” etc., used in this application may be used to describe various concepts herein, but unless otherwise stated, these concepts are not limited by these terms. These terms are used only to distinguish one concept from another.
[0045] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0046] It should be noted that all information (including but not limited to information used for processing, stored information, and displayed information) and data (including but not limited to data used for processing, stored data, and displayed data) involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the interactive information involved in this application was obtained with full authorization.
[0047] Figure 1 This is a schematic diagram of the structure of an implementation environment provided in an embodiment of this application. See also... Figure 1 The implementation environment includes: vehicle 110 and multilingual testing system 120.
[0048] The following is an introduction to vehicle 110.
[0049] Vehicle 110 includes a vehicle infotainment controller 111 and an interactive device 112, which are connected via a wired or wireless network. The interactive device 112 includes a microphone, a speaker, a communication module, etc. The interactive device 112 can interact with the outside world; for example, the microphone can collect voice, the speaker can play voice, and the communication module can communicate with other devices.
[0050] The vehicle controller 111 operates a voice interaction system. This system is an in-vehicle system that uses voice as the interaction medium. After collecting voice data, it employs speech recognition, semantic understanding, and speech synthesis technologies to respond to the collected voice and determine the necessary operations for the vehicle 110 and the required voice responses. Consequently, the driver or passengers in the vehicle 110 can control the vehicle 110 through the voice interaction system by speaking naturally, eliminating the need for manual operation and improving both convenience and driving safety.
[0051] Furthermore, the voice interaction system supports multiple languages, including but not limited to Chinese, English, German, French, Spanish, and Arabic. Accordingly, the driver or passengers in vehicle 110 can control vehicle 110 by speaking in any of these languages using natural speech.
[0052] The following is an introduction to the multilingual testing system 120.
[0053] The multilingual testing system 120 is connected to the vehicle infotainment controller 111 via a wired or wireless network. The multilingual testing system 120 is used to perform multilingual testing on the voice interaction system. Located inside the vehicle cabin, the multilingual testing system 120 can play test voices belonging to multiple target test languages within the cabin and acquire first response data obtained after the voice interaction system responds to the test voices. It also acquires second response data by using an artificial intelligence model to respond to the test voices. The multilingual test result is determined based on the comparison between the first and second response data. This multilingual test result indicates the voice interaction system's response to test voices belonging to multiple target test languages, reflecting the quality of the voice interaction system's processing of voices in multiple target test languages.
[0054] Optionally, the multilingual testing system 120 includes a testing unit and a control unit.
[0055] The control unit coordinates the work of various units and generates multilingual test results for the voice interaction system. The control unit can be an industrial-grade industrial computer, or include CPU (Central Processing Unit), DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array), etc.
[0056] The test unit and control unit are connected via a wired or wireless network. The test unit is located inside the vehicle cabin, while the control unit can be located inside the vehicle cabin or anywhere outside the vehicle cabin.
[0057] For example, the test unit includes multiple robotic arms and multiple speakers, with each speaker located on one robotic arm and at least one speaker configured on each robotic arm. The robotic arms can move within the vehicle cabin, stop moving after reaching the target test position, and play test voice through the speakers on the robotic arms at the target test position, thereby testing the voice interaction system's ability to collect voice generated at the target test position.
[0058] For example, the test unit is configured with a multi-channel audio interface, each audio interface being connected to a speaker for outputting the test audio to be played to the speaker for playback.
[0059] For example, a four-degree-of-freedom robotic arm is used in the test unit, and the positioning accuracy of the four-degree-of-freedom robotic arm can reach ±1cm.
[0060] For example, the frequency response range of each speaker is: [100Hz (Hertz) - 16kHz (kilohertz)] ±2dB (decibels), which means that each speaker can play speech with frequencies within this frequency response range, covering the frequency range of human speech.
[0061] Optionally, the test unit also includes a response acquisition module, which can acquire response data from the voice interaction system.
[0062] For example, the response acquisition module includes a microphone array comprising multiple microphones through which the response speech output by the voice interaction system can be acquired. For instance, the microphone array includes eight microphones with a sampling rate of 48 kHz and a dynamic range (the range of sound pressure levels that can be picked up) greater than 100 dB.
[0063] For example, there is a one-to-one correspondence between microphones and speakers, with each robotic arm equipped with a pair of microphones and speakers.
[0064] For example, the response acquisition module includes a communication module, which is capable of receiving response operations and screen interfaces sent by the voice interaction system.
[0065] Optionally, the target test location includes, but is not limited to, the driver's seat, passenger seat, and rear seats within the vehicle's cabin. By moving the robotic arm to the target test location such as the driver's seat, passenger seat, or rear seats and then stopping the movement before playing the test voice through a speaker on the robotic arm, a realistic sound field can be generated by simulating the voices spoken by users in the driver's seat, passenger seat, and rear seats in real driving scenarios. This improves the realism and allows for more accurate test results.
[0066] Optionally, the multilingual testing system 120 is configured with a testing software platform, which is developed based on Python 3.9 (a programming language) or other programming languages. The multilingual testing system 120 is also configured with an audio processing library, which can convert test text into playable test speech. This audio processing library includes libraries such as LibROSA (an audio feature extraction library) and PyAudio (an audio streaming library).
[0067] Optionally, the multilingual testing system 120 is equipped with an artificial intelligence model. This model can respond to test voice and obtain response data, including response actions and response voice. The response actions are the actions predicted by the artificial intelligence model that the voice interaction system should perform after receiving the test voice. The response voice is the voice predicted by the artificial intelligence model that the voice interaction system should reply to after receiving the test voice.
[0068] Among them, the artificial intelligence model includes convolutional neural network model, deep learning network model or large language model, or the artificial intelligence model is at least one of the following types: BERT (Bidirectional Encoder Representations from Transformers) multilingual model and Sentence Piece word segmentation library.
[0069] Optionally, the multilingual testing system 120 is equipped with a result analysis engine. The result analysis engine can compare the first response data with the second response data, analyze the comparison results, and determine at least one performance parameter of the voice interaction system, such as speech recognition accuracy, operation decision accuracy, speech quality parameters, average response latency, etc.
[0070] Optionally, the multilingual testing system 120 is equipped with data visualization tools. These tools can visualize the response data corresponding to the test speech, displaying the visualized test results. The data visualization tools include Matplotlib (MATLAB Plot Library) and Seaborn (a high-level wrapper library based on Matplotlib).
[0071] Based on the above embodiments, Figure 2 This is a flowchart illustrating a multilingual testing method for a voice interaction system provided in an embodiment of this application. The execution entity of this embodiment is a multilingual testing system, as described above. Figure 1 The multilingual testing system 120 is shown. See also... Figure 2 The method includes the following steps 201-204.
[0072] 201. The multilingual testing system plays multiple test voices in the vehicle cabin. These multiple test voices belong to various target test languages to be tested.
[0073] The voice interaction system supports multiple languages and can respond to voices in multiple languages. To test the cross-language capability of the voice interaction system, a multilingual testing system can be configured with these multiple languages. During multilingual testing, multiple target test languages are selected from the multiple languages, and test voices in these multiple target test languages are played in the vehicle cabin.
[0074] In this context, "test speech" refers to the test speech being expressed in the target test language. Each test speech may include content belonging to the same target test language, or each test speech may include content belonging to at least two target test languages. That is, each test speech played in the vehicle cabin may belong to one target test language or multiple target test languages, as long as the set of test speech to which all played test speech belongs includes all of these multiple target test languages.
[0075] Optionally, the multilingual testing system connects to a management terminal, which displays multiple test languages for testers to select. Testers then select the chosen target test languages and send them to the multilingual testing system, which receives them. All test languages are supported by the voice interaction system, and may include all languages supported by the voice interaction system, or only some of the languages supported by the voice interaction system.
[0076] Optionally, the multilingual testing system is configured with a test corpus containing multiple test texts and recording the language of each test text. The multilingual testing system can then retrieve test texts belonging to the target test language from the test corpus, convert these test texts into corresponding test speech, and play the test speech within the vehicle cabin.
[0077] The multilingual testing system can randomly select test text belonging to the target test language, or the language testing system can be connected to a management terminal, which displays multiple test texts belonging to the target test language for testers to select. The selected test text is then sent to the multilingual testing system, which receives the selected test text.
[0078] 202. The multilingual testing system acquires the first response data corresponding to each test voice; the first response data is obtained by the voice interaction system responding to the test voice.
[0079] The multilingual testing system can play one or more test voice recordings. For each test voice recording, the voice interaction system collects the recording, responds to it, and obtains the first response data. The voice interaction system sends the first response data to the multilingual testing system, which then receives the corresponding first response data.
[0080] 203. The multilingual testing system uses an artificial intelligence model to respond to each test speech separately, and obtains the second response data corresponding to each test speech.
[0081] The multilingual testing system is equipped with an artificial intelligence model that features voice interaction capabilities. This model can respond to collected voice input and generate response data. The response data includes response actions and response voice. The response actions refer to the actions predicted by the AI model and in accordance with the voice instruction, which should be performed on the vehicle. The response voice refers to the voice predicted by the AI model, which should be used to respond to the given voice.
[0082] When the multilingual testing system plays test audio in the vehicle cabin, it will also respond to the played test audio through an artificial intelligence model to obtain second response data.
[0083] Optionally, the test speech is obtained by converting the test text into test text by the multilingual testing system. In order to avoid additional errors caused by converting the test speech into test text again, when the test speech is played in the vehicle cabin, the multilingual testing system will not convert the test speech into the corresponding test text. Instead, it will directly obtain the test text of the previously converted test speech, and respond to the test text directly through the artificial intelligence model to obtain the response data corresponding to the test text. The response data corresponding to the test text will be used as the second response data corresponding to the test speech.
[0084] Optionally, in order to ensure that the artificial intelligence model can respond accurately to the collected speech, it is necessary to train the artificial intelligence model first.
[0085] The training process of an artificial intelligence model includes: obtaining an initialized artificial intelligence model, obtaining sample data, which includes sample speech and sample response data labeled for the sample speech, the sample response data representing the correct response data of the sample speech, and training the artificial intelligence model based on the sample data so that the trained artificial intelligence model has the ability to accurately respond to the input speech.
[0086] The sample response data includes the sample response operation and sample response speech corresponding to the sample speech, and may also include the sample recognition text corresponding to the sample speech. The sample recognition text refers to the recognized text obtained by recognizing the sample speech, which can represent the content of the sample speech in text format. This sample recognition text can be determined by the annotator based on the content of the sample speech. The sample response operation is the action that the sample speech instructs the vehicle to perform, such as turning on the air conditioning or opening the window. This sample response operation can be determined by the annotator based on the semantics of the sample speech. The sample response speech is the voice used by the vehicle to respond to the sample speech; for example, if the sample speech is "Please turn on the air conditioning," then the sample response speech is "Air conditioning is on." After the annotator determines the content to be responded to based on the sample speech and the sample response operation, the annotator speaks the content while recording the speech to obtain the sample response speech. Alternatively, after the annotator determines the text corresponding to the content to be responded to based on the sample speech and the sample response operation, the annotator uses speech synthesis technology to convert the text into the corresponding speech to obtain the sample response speech.
[0087] Since the sample response data already specifies the response data to be generated when sample speech is collected, training an artificial intelligence model based on the sample speech and sample response data will enable the artificial intelligence model to have accurate speech interaction capabilities. Subsequently, any speech input into the trained artificial intelligence model will allow the model to obtain the corresponding response data.
[0088] For example, training an artificial intelligence model based on sample data includes: inputting sample speech into the artificial intelligence model, having the artificial intelligence model respond to the sample speech to obtain predicted response data, and adjusting the model parameters in the artificial intelligence model based on the sample response data and the predicted response data to reduce the difference between the predicted response data determined by the trained artificial intelligence model and the sample response data. After one or more adjustments, when the training stopping condition is met, the trained artificial intelligence model can be obtained.
[0089] 204. The multilingual testing system determines the multilingual test results of the voice interaction system based on the comparison results of the first response data and the second response data corresponding to each test voice.
[0090] The first response data is the data obtained from the voice interaction system's response and can represent the system's voice interaction capabilities. The second response data is the data obtained by the multilingual testing system itself through its artificial intelligence model and can be used as reference response data. By comparing the first and second response data corresponding to the same test voice, we can determine whether the first response data of the voice interaction system is correct, whether there are any problems, and the reasons for any problems. Therefore, based on the comparison results of the first and second response data for each test voice, the multilingual test results of the voice interaction system can be determined.
[0091] This application provides a multilingual testing method for a voice interaction system. It can play multiple test voices belonging to various target test languages within a vehicle cabin, acquiring first and second response data for each test voice. The first response data is determined by the voice interaction system, and the second response data is determined by an artificial intelligence model. The comparison between the first and second response data for each test voice reflects the correctness of the voice interaction system's response process. Therefore, based on the comparison results, the multilingual test result of the voice interaction system is determined. During the testing process, testers do not need to speak the test voices belonging to various target test languages, nor do they need to manually judge the correctness of the voice interaction system's response data, reducing labor and time costs and improving the efficiency of multilingual testing.
[0092] Moreover, testers do not need to master multiple target testing languages, are not limited by the testers' language abilities, improve the coverage of testing languages, enhance flexibility, and are not affected by the testers' pronunciation, thus improving the accuracy of multilingual testing.
[0093] Based on the above embodiments, step 201 may optionally include the following two possible implementation methods.
[0094] In the first possible implementation, the multilingual testing system plays test audio belonging to each target test language sequentially within the vehicle cabin, according to the order of the multiple target test languages.
[0095] In other words, multiple target test languages are arranged sequentially, with each target test language corresponding to one or more test voices. The multilingual testing system plays one test voice at a time, or multiple test voices in parallel at a time, with multiple test voices belonging to the same target test language. After all test voices belonging to one target test language have been played, the system switches to playing test voices belonging to the next target test language, and so on, until all test voices belonging to multiple target test languages have been played.
[0096] In this embodiment, test voices belonging to different target test languages are played alternately within the vehicle cabin. This eliminates the mixing of different languages, reducing the complexity of speech recognition in the voice interaction system and enabling accurate evaluation of the test results for each target test language. Furthermore, it facilitates precise location of the problematic target test language should a problem occur in the voice interaction system.
[0097] In the second possible implementation, the multilingual testing system plays multiple test voices in parallel within the vehicle's cabin, with each test voice corresponding to a different target test language.
[0098] Each target test language corresponds to one or more test voices. The multilingual testing system selects multiple test voices belonging to different target test languages and plays them in parallel. After the multiple test voices have been played in parallel, the next group of multiple test voices is played in parallel.
[0099] For example, the multilingual testing system selects multiple test voices belonging to different target test languages each time to form a test voice combination. This step is repeated to obtain multiple test voice combinations. In the order of the multiple test voice combinations, multiple test voices in the same test voice combination are played in parallel each time.
[0100] For example, the multilingual testing system is configured with n speakers, where n is an integer greater than 1, meaning that the multilingual testing system can play a maximum of n test voices in parallel at any given time. Each time, the multilingual testing system selects n test voices belonging to different target test languages to form a test voice combination, and so on, obtaining one or more test voice combinations. Then, each time a test voice combination is selected, the n test voices in that combination are played in parallel through the n speakers.
[0101] In this embodiment, test voices belonging to multiple target test languages are played in parallel within the vehicle cabin. Multiple target test languages can be tested in one test, which can improve testing efficiency, significantly shorten the total testing time, and also simulate the scenario where multiple language sound sources exist simultaneously while the vehicle is in motion. This verifies the anti-interference capability of the voice interaction system and its ability to distinguish between different languages, and reveals the problems of the voice interaction system in multilingual mixed scenarios.
[0102] Based on the two possible implementations mentioned above, when the multilingual testing system plays a test voice message in the vehicle cabin each time, the voice interaction system collects the test voice message, responds to it, obtains first response data, and sends it to the multilingual testing system. Furthermore, after confirming the response voice message, the voice interaction system also outputs that response voice message, which the multilingual testing system can collect through its microphone. The multilingual testing system can then determine that the test voice message corresponding to the first response data is the test voice message being played. The time interval between the completion time of the test voice message playback and the time of collecting the response voice message is used as the response delay of the voice interaction system. Alternatively, the difference between the time interval between the completion time of the test voice message playback and the time of collecting the response voice message and a preset time interval can be used as the response delay of the voice interaction system. The preset time interval can be set by default by the multilingual testing system; for example, the preset time interval could be the sum of the average time it takes for voice messages played in the vehicle cabin to reach the microphone of the voice interaction system and the average time it takes for voice messages output by the voice interaction system to reach the microphone of the multilingual testing system.
[0103] In a multilingual testing system, multiple test voices are played in parallel within the vehicle cabin each time. The voice interaction system collects these multiple test voices, responds to them, and obtains combined response data. This combined response data is generated by the voice interaction system based on the multiple test voices played in parallel and is a combination of response data corresponding to the multiple test voices. Accordingly, step 202 includes: the multilingual testing system receiving the combined response data sent by the voice interaction system, and extracting the first response data corresponding to each test voice from the combined response data.
[0104] Optionally, when the voice interaction system collects a combined test voice consisting of multiple test voices, it can split the combined test voice into multiple voices to obtain multiple test voices, and respond to each test voice separately to obtain response data corresponding to each test voice. The response data corresponding to multiple test voices are then combined to obtain combined response data, which is sent to the multilingual testing system. The multilingual testing system splits the combined response data to obtain multiple first response data. Subsequently, the multilingual model matches the multiple test voices with the multiple first response data to determine the first response data corresponding to each test voice. The first response data corresponding to different test voices are different, thus establishing a matching relationship between each test voice and each first response data, distinguishing the first response voices corresponding to different test voices, so that in step 204, the first response data corresponding to the same test voice can be compared with the second response data.
[0105] For example, the multilingual testing system splits the combined response data into multiple first response data, determines the content correlation between each test speech and each first response data, and assigns the multiple first response data to multiple test speech based on the content correlation between each test speech and each first response data, so that the multiple first response data correspond one-to-one with the multiple test speech.
[0106] For example, from multiple first response data, the response data with the highest content relevance to each test speech is determined as the associated response data for each test speech. If the associated response data for different test speech are all different, the associated response data for each test speech is determined as the first response data corresponding to that test speech. If the associated response data for any two test speech are the same, and a first response data point has no associated test speech, that associated response data is assigned to the test speech with the highest content relevance, and the first response data of the unassociated test speech is assigned to the other of those two test speech.
[0107] Alternatively, if at least two test voices have identical associated response data, but at least one test voice has no associated first response data, then for the at least two test voices, their associated response data, and the first response data of the unassociated test voices, the first response data with the highest content relevance is assigned to the test voice with the highest content relevance, according to the degree of content relevance. After removing this set of first response data and test voices, the above steps are repeated for the remaining test voices' first response data until the last remaining first response data is assigned to the last remaining test voice.
[0108] For example, for any test speech and any first response data, the test text corresponding to the test speech is obtained, the first response data is obtained, and the content relevance between the test text and the first response data is determined as the content relevance between the test speech and the first response data. The content relevance can be determined by an artificial intelligence model, a text matching algorithm, or other methods.
[0109] In this embodiment of the application, when test voices belonging to multiple target test languages are played in parallel in the vehicle cabin, not only is the testing efficiency improved, but the first response data corresponding to different test voices can also be accurately extracted from the combined response data of the voice interaction system. This avoids the problem of parallel response data confusion, provides a data foundation for subsequent comparative analysis, and avoids the accuracy of test results being affected by disordered response data.
[0110] Based on the above embodiments, in one possible implementation, the method is executed by a multilingual testing system, which includes multiple robotic arms and multiple speakers, with each speaker located on one robotic arm. Accordingly, step 201 includes: determining multiple target test locations to be tested; each target test location being located within the vehicle cabin; moving each robotic arm to its corresponding target test location; and playing test audio belonging to each target test language through each speaker.
[0111] The multiple target test locations can include the driver's seat, front passenger seat, and rear seats within the vehicle's cabin. These multiple target test locations can be determined by the multilingual testing system by default, or the multilingual testing system can be connected to a management terminal, which displays multiple test locations for the tester to select. The tester then selects the chosen multiple target test locations and sends them to the multilingual testing system, which receives them.
[0112] For example, if the number of the plurality of target test positions is less than or equal to the number of robotic arms, the plurality of robotic arms, which are equal in number to the plurality of target test positions, are controlled to move to the corresponding target test positions respectively, and test voice is played through each speaker on each robotic arm.
[0113] For example, if the number of multiple target test positions is greater than the number of robotic arms, the multiple target test positions are divided into at least two groups, and the number of target test positions in each group is equal to or less than the number of robotic arms. According to the order of the multiple groups of target test positions, the multiple robotic arms are controlled to move to the corresponding target test position in one group each time, and the test voice is played through each speaker on each robotic arm.
[0114] For example, a multilingual testing system determines multiple test voices and multiple target test locations based on multiple target test languages. It then determines the test voice to be played at each target test location, establishing a correspondence between target test locations and test voices. Following this correspondence, a robotic arm is controlled to move to the target test location, and a speaker on the robotic arm plays the test voice corresponding to that location. Furthermore, after playback, the correspondence between target test locations and test voices can be updated, and the test voice can be played again according to the updated correspondence, allowing for testing of various combinations of target test locations and target test languages.
[0115] In this embodiment, the robotic arm is moved to the corresponding target test position, and test audio is played through a speaker on the robotic arm. This simulates the scenario of different users in different positions in a real in-vehicle environment, improving the realism of the test scenario and the accuracy of the test results. Moreover, the target test position is independently controllable, allowing for testing of different scenarios by changing the target test position, thus demonstrating strong reusability.
[0116] Based on the above embodiments, in one possible implementation, the multilingual testing system is configured with test corpora corresponding to multiple languages. The test corpora include test texts belonging to the corresponding languages. Furthermore, the same test corpus can include test texts belonging to multiple interaction types. Here, interaction type refers to the type of interaction scenario to which the test text is applicable. For example, interaction types include, but are not limited to: control commands, information queries, entertainment controls, and chat communication. A control command indicates that the test text is used to control the vehicle to perform an operation, such as "turn on the air conditioner"; an information query indicates that the test text is used to query information, such as "what's the weather like today?"; and entertainment controls indicate that the test text is used to control the vehicle to play multimedia content, such as "play jazz music".
[0117] Step 201 above includes: for each target test language, determining the target test corpus corresponding to the target test language, the target test corpus including test texts belonging to the target test language and belonging to multiple interaction types; obtaining test texts belonging to the target interaction type to be tested from the target test corpus; converting the test texts into corresponding test speech, and then playing the converted test speech.
[0118] The target interaction type can be one or more, and the target interaction type can be determined by the multilingual testing system by default. Alternatively, the multilingual testing system can be connected to a management terminal, which displays multiple interaction types for testers to select. Testers can then select the target interaction types from the multiple interaction types and send them to the multilingual testing system, which will then receive the target interaction types.
[0119] In this embodiment, test texts capable of testing multiple target interaction types cover various interaction scenarios of the voice interaction system, avoiding omissions in interaction scenarios and preventing the omission of test functions due to testing a single interaction type, thus improving the comprehensiveness of the test. Moreover, the target interaction type is independently controllable, allowing for testing of different interaction scenarios by changing the target interaction type, resulting in strong reusability.
[0120] Based on the above embodiments, in one possible implementation, step 203 includes: for each test speech, performing speech recognition on the test speech to obtain a first recognized text corresponding to the test speech; responding to the first recognized text through an artificial intelligence model to obtain a first response operation matching the first recognized text; and obtaining a first response speech matching the first response operation.
[0121] In this embodiment, the process of responding to voice includes the following three stages: the first stage is the voice recognition stage, which recognizes the voice and converts it into corresponding recognized text; the second stage is the operation decision stage, which determines the operation that the vehicle should perform based on the recognized text; and the third stage is the voice response stage, which determines the voice response that should be given based on the response operation. In addition to these three stages, the process of responding to voice may also include other stages, such as an interface display stage.
[0122] Therefore, after acquiring the test speech, the voice interaction system will execute the above three stages to obtain first response data, which includes at least one of the second recognized text, the second response operation, and the second response speech. Correspondingly, the multilingual testing system will also execute the above three stages to obtain second response data, which includes at least one of the first recognized text, the first response operation, and the first response speech.
[0123] Accordingly, step 204 above includes at least one of (1) to (3) below.
[0124] (1) Based on the comparison results of the first and second recognized texts corresponding to each test speech, determine the speech recognition accuracy of the speech interaction system.
[0125] The first recognized text is the text obtained by the multilingual testing system from the test speech, and the second recognized text is the text obtained by the voice interaction system from the test speech. By comparing the first recognized text and the second recognized text corresponding to the same test speech, the voice interaction system's ability to recognize speech can be demonstrated.
[0126] If the first and second recognized texts corresponding to the same test speech are the same, or if they are different but have the same semantic meaning, then the voice interaction system has correctly recognized the test speech. Conversely, if the first and second recognized texts corresponding to the same test speech have different semantic meanings, then the voice interaction system has misrecognized the test speech.
[0127] The multilingual testing system determines the number of correctly recognized test voices and the number of incorrectly recognized test voices based on the comparison results between the first and second recognized texts corresponding to each test voice, thereby determining the voice recognition accuracy of the voice interaction system, which serves as the multilingual test result of the voice interaction system.
[0128] The multilingual testing system can aggregate test voices belonging to multiple target test languages and calculate their speech recognition accuracy together, which serves as the multilingual test result for the voice interaction system. Alternatively, it can separate test voices belonging to different target test languages and determine the speech recognition accuracy separately for each type of test voice, also as the multilingual test result for the voice interaction system.
[0129] Optionally, for the first and second identified texts, a text comparison model can be used to determine whether their semantics are the same. This text comparison model can be a BERT model, a convolutional neural network model, a large language model, or other artificial intelligence models.
[0130] Optionally, the process of training the text comparison model includes: obtaining an initialized text comparison model; obtaining sample data, which includes any two texts and sample labels. The sample labels indicate whether the two texts have the same semantics; if the sample label is the first label, it means the two texts have the same semantics; if the sample label is the second label, it means the two texts have different semantics. The text comparison model is then trained based on this sample data so that the trained text comparison model has the ability to accurately compare whether the semantics of two texts are the same.
[0131] For example, training the text comparison model based on the sample data includes: inputting two texts from the sample data into the text comparison model, comparing them through the text comparison model to obtain predicted labels, and adjusting the model parameters in the text comparison model based on the sample labels and the predicted labels, so that the predicted labels determined by the trained text comparison model are closer to the sample labels. After one or more adjustments, when the training stopping condition is met, the trained text comparison model can be obtained.
[0132] Optionally, after the multilingual testing system determines the comparison results of the first and second recognized texts corresponding to multiple test voices, it can determine whether there are any problems with the voice recognition function of the voice interaction system and the reasons for the problems based on the comparison results.
[0133] For example, if the speech recognition accuracy rate of a voice interaction system is greater than the speech recognition accuracy rate threshold, it indicates that the speech recognition function of the voice interaction system is normal and there is no problem. However, if the speech recognition accuracy rate of a voice interaction system is not greater than the speech recognition accuracy rate threshold, it indicates that there is a problem with the speech recognition function of the voice interaction system.
[0134] Alternatively, based on multiple target test languages, if the speech recognition accuracy of the voice interaction system for a certain target test language is not greater than the speech recognition accuracy threshold, it indicates that the voice interaction system has a problem in recognizing speech belonging to that target test language.
[0135] (2) Based on the comparison results of the first response operation and the second response operation corresponding to each test voice, determine the operation decision accuracy of the voice interaction system.
[0136] The first response operation is the operation that the multilingual testing system determines the vehicle needs to perform based on the test voice, and the second response operation is the operation that the voice interaction system determines the vehicle needs to perform based on the test voice. By comparing the first response operation and the second response operation corresponding to the same test voice, the ability of the voice interaction system to make correct decisions can be demonstrated.
[0137] If the first and second response actions corresponding to the same test voice are the same, it indicates that the response action determined by the voice interaction system based on the test voice is correct. Conversely, if the first and second response actions corresponding to the same test voice are different, it indicates that the response action determined by the voice interaction system based on the test voice is incorrect.
[0138] The multilingual testing system determines the number of test voices that are correctly identified and the number of test voices that are incorrectly identified based on the comparison between the first and second response operations corresponding to each test voice. This determines the operational decision accuracy of the voice interaction system and serves as the multilingual test result of the voice interaction system.
[0139] The multilingual testing system can aggregate test voice recordings belonging to multiple target test languages and statistically analyze their operational decision accuracy as the multilingual test result of the voice interaction system. Alternatively, it can separate test voice recordings belonging to different target test languages and determine the operational decision accuracy separately for each type of test voice recording, also as the multilingual test result of the voice interaction system.
[0140] Optionally, after the multilingual testing system determines the comparison results of the first response operation and the second response operation corresponding to multiple test voices, it can determine whether there is a problem with the operation decision function of the voice interaction system and the reason for the problem based on the comparison results.
[0141] For example, if the operation decision accuracy rate of a voice interaction system is greater than the operation decision accuracy rate threshold, it indicates that the operation decision function of the voice interaction system is normal and there is no problem. However, if the operation decision accuracy rate of the voice interaction system is not greater than the operation decision accuracy rate threshold, it indicates that there is a problem with the operation decision function of the voice interaction system.
[0142] Alternatively, based on multiple target test languages, if the speech recognition accuracy of the voice interaction system for a certain target test language is not greater than the speech recognition accuracy threshold, and the operation decision accuracy is not greater than the operation decision accuracy threshold, it indicates that the voice interaction system has a problem in recognizing speech belonging to that target test language, which affects the accuracy of response operations.
[0143] (3) Based on the comparison results of the first response speech and the second response speech corresponding to each test speech, determine the speech quality parameters of the voice interaction system.
[0144] The first response voice is the voice that the multilingual testing system determines should be replied to based on the test voice, and the second response voice is the voice interaction system's voice that determines should be replied to based on the test voice. By comparing the first response voice and the second response voice corresponding to the same test voice, the ability of the voice interaction system to reply with the correct voice can be demonstrated.
[0145] If the first and second response voices corresponding to the same test voice are the same, or if the first and second response voices corresponding to the same test voice are different but have the same semantic meaning, it means that the response voice determined by the voice interaction system based on the test voice is correct. Conversely, if the first and second response voices corresponding to the same test voice have different semantic meanings, it means that the response voice determined by the voice interaction system based on the test voice is incorrect.
[0146] The multilingual testing system determines the number of test voices that are correctly identified by the response voice and the number of test voices that are incorrectly identified by the response voice based on the comparison results of the first response voice and the second response voice corresponding to each test voice. This determines the voice quality parameters of the voice interaction system and serves as the multilingual test result of the voice interaction system.
[0147] The multilingual testing system can aggregate test voice recordings belonging to multiple target test languages, statistically analyze their voice quality parameters, and use this as the multilingual test result for the voice interaction system. Alternatively, it can separate test voice recordings belonging to different target test languages, determine voice quality parameters separately for each type of target test language, and also use this as the multilingual test result for the voice interaction system.
[0148] Optionally, for the first response speech and the second response speech, speech recognition technology can be used to convert the first response speech into the corresponding first text and the second response speech into the corresponding second text. Then, a text comparison model can be used to determine whether the first text and the second text have the same semantic meaning. If the first text and the second text have the same semantic meaning, it means that the first response speech and the second response speech have the same semantic meaning. If the first text and the second text have different semantic meanings, it means that the first response speech and the second response speech have different semantic meanings. The process of training the text comparison model is the same as the process of training the text comparison model in item (1) above, and will not be repeated here.
[0149] Optionally, for the first and second response speech, a speech comparison model can be used to determine whether their semantics are the same. This speech comparison model can be a BERT model, a convolutional neural network model, a large language model, or other artificial intelligence models.
[0150] Optionally, the process of training the speech contrast model includes: obtaining an initialized speech contrast model; obtaining sample data, which includes any two speech items and sample labels. The sample labels indicate whether the two speech items have the same semantics. If the sample label is a first label, it means the two speech items have the same semantics; if the sample label is a second label, it means the two speech items have different semantics. The speech contrast model is then trained based on this sample data so that the trained speech contrast model has the ability to accurately compare whether the semantics of two speech items are the same.
[0151] For example, training the speech contrast model based on the sample data includes: inputting two speech samples into the speech contrast model, comparing them using the speech contrast model to obtain predicted labels, and adjusting the model parameters in the speech contrast model based on the sample labels and the predicted labels, so that the predicted labels determined by the trained speech contrast model are closer to the sample labels. After one or more adjustments, when the training stopping condition is met, the trained speech contrast model can be obtained.
[0152] Optionally, after the multilingual testing system determines the comparison results of the first response voice and the second response voice corresponding to multiple test voices, it can determine whether there is a problem with the voice response function of the voice interaction system and the reason for the problem based on the comparison results.
[0153] For example, if the voice quality parameters of a voice interaction system are greater than the voice quality parameter threshold, it indicates that the voice response function of the voice interaction system is normal and there is no problem. However, if the voice quality parameters of a voice interaction system are not greater than the voice quality parameter threshold, it indicates that the voice response function of the voice interaction system has a problem.
[0154] Alternatively, it can be differentiated according to multiple target test languages. If the speech recognition accuracy of the voice interaction system for a certain target test language is greater than the speech recognition accuracy threshold, and the operation decision accuracy is greater than the operation decision accuracy threshold, but the speech quality parameter is not greater than the speech quality parameter threshold, it indicates that there is a problem with the speech synthesis function of the voice interaction system, and the quality of the synthesized response speech is poor.
[0155] In this embodiment, there is no need for manual judgment of whether the first response data of the voice interaction system is correct, nor is there a need for manual annotation of the standard response data of each test voice. Instead, the multilingual testing system automatically generates the second response data as the standard response data, which saves annotation costs, reduces the workload of testers, realizes automated testing, eliminates errors caused by manual judgment, and improves the accuracy of test results.
[0156] Based on the above embodiments, in one possible implementation, where the process of responding to speech includes a speech recognition stage, an operation decision stage, and a speech response stage, the training process of the artificial intelligence model includes: obtaining an initialized artificial intelligence model; obtaining sample data, which includes sample speech and sample response data labeled for the sample speech, the sample response data representing the correct response data of the sample speech, the sample response data including the sample recognition text corresponding to the sample speech, the sample response operation, and the sample response speech; and training the artificial intelligence model based on the sample data so that the trained artificial intelligence model has the ability to accurately respond to the input speech.
[0157] The sample recognition text refers to the recognized text obtained by recognizing the sample speech, which represents the content of the sample speech in text format. This sample recognition text can be determined by the annotator based on the content of the sample speech. The sample response operation is the action the sample speech instructs the vehicle to perform, such as turning on the air conditioning or opening a window. This sample response operation can be determined by the annotator based on the semantics of the sample speech. The sample response voice is the voice used by the vehicle to reply to the sample speech; for example, if the sample speech is "Please turn on the air conditioning," then the sample response voice is "Air conditioning is on." After the annotator determines the content that should be replied to based on the sample speech and the sample response operation, the annotator speaks that content while recording the speech to obtain the sample response voice. Alternatively, after the annotator determines the text corresponding to the content that should be replied to based on the sample speech and the sample response operation, speech synthesis technology is used to convert the text into the corresponding speech to obtain the sample response voice.
[0158] Since the sample response data already specifies the correct text to be recognized, the response operation that the vehicle should perform, and the response voice to be given when sample speech is collected, training an artificial intelligence model based on the sample speech and sample response data will enable the artificial intelligence model to have accurate speech recognition, operation decision-making, and speech synthesis capabilities. Subsequently, any speech can be input into the trained artificial intelligence model, and the corresponding response data can be obtained through the trained artificial intelligence model.
[0159] For example, training an artificial intelligence model based on sample data includes: inputting sample speech into the artificial intelligence model; performing speech recognition on the sample speech using the artificial intelligence model to obtain predicted recognized text; determining a predicted response operation based on the predicted recognized text; and determining a predicted response speech based on the predicted response operation. Based on the sample recognized text, the sample response operation, and the sample response speech, as well as the predicted recognized text, the predicted response operation, and the predicted response speech, the model parameters in the artificial intelligence model are adjusted to reduce the differences between the predicted recognized text and the sample recognized text, the differences between the predicted response operation and the sample response operation, and the differences between the predicted response speech and the sample response speech determined by the trained artificial intelligence model. After one or more adjustments, when the training stopping condition is met, the trained artificial intelligence model is obtained.
[0160] Based on the above embodiments, in one possible implementation, for each test voice recording played, the multilingual testing system determines the response delay corresponding to that test voice recording. This response delay is either the time interval between the completion time of the test voice recording and the time when the response voice recording from the voice interaction system is acquired, or the difference between the time interval between the completion time of the test voice recording and the time when the response voice recording from the voice interaction system is acquired, and a preset time interval. The multilingual testing system calculates the average response delay from the response delays corresponding to multiple test voice recordings, and uses this average response delay as the multilingual test result of the voice interaction system.
[0161] The multilingual testing system can aggregate test voice recordings belonging to multiple target test languages and calculate the average response latency together as the multilingual test result of the voice interaction system. Alternatively, it can separate test voice recordings belonging to different target test languages and calculate the average response latency separately for each type of test voice recording, also as the multilingual test result of the voice interaction system.
[0162] In this embodiment, considering that the response latency of the voice interaction system affects the user experience and driving safety, not only the accuracy of the voice interaction system is tested, but also the average response latency of the voice interaction system is tested. This enriches the testing dimensions, enhances the comprehensiveness of the multilingual test results, and facilitates subsequent optimization of the voice interaction system based on the average response latency to improve the user experience.
[0163] Based on the above embodiments, in one possible implementation, the multilingual testing system plays test audio in the vehicle cabin while simultaneously playing ambient noise. This ambient noise simulates background sounds generated during vehicle operation, such as road noise, wind noise, and air conditioning noise. The voice interaction system then collects both the test audio and ambient noise. After removing the ambient noise from the collected audio, the system obtains the test audio and responds to it to obtain first response data. If the process of removing ambient noise encounters problems, the final first response data may be inaccurate. Therefore, the comparison between the first response data and the second response data can also reflect the voice interaction system's ability to remove ambient noise.
[0164] The ambient noise can be manually recorded by the tester and uploaded to the multilingual testing system, or it can be automatically generated by the multilingual testing system.
[0165] The multilingual testing system plays ambient noise through speakers. These speakers can be one or more fixed speakers, or they can be randomly selected during the test. Furthermore, the target test location of the speaker playing the ambient noise can be a fixed test location, or it can be randomly updated during the test.
[0166] The method for determining the target test position of the robotic arm when the multilingual testing system plays ambient noise is the same as the method for determining the target test position when playing test voice in the above embodiment, and will not be repeated here.
[0167] In this embodiment, multilingual testing of the voice interaction system is performed while playing ambient noise, which is closer to the real in-vehicle environment and improves the realism of the test scenario. This allows for testing of the voice interaction system's ability to remove ambient noise interference and easily exposes problems with the voice interaction system when it is affected by ambient noise.
[0168] It should be noted that the solutions in the above embodiments can be combined in any form to form optional solutions for the embodiments of this application, which will not be elaborated here.
[0169] Based on the above embodiments, this application provides an exemplary testing process. Figure 3 This is a schematic diagram of an exemplary test process provided in an embodiment of this application. See also... Figure 3 The testing process includes the following steps 301-309.
[0170] 301. Configure test parameters for the multilingual testing system.
[0171] The test parameters include: target test language, test text belonging to the target test language, target test location, and target interaction type. The process of configuring the target test language, test text, and target test location is detailed in the above embodiment and will not be repeated here.
[0172] Optionally, the multilingual testing system is configured with a test plan, which includes multiple sets of test parameters. Each set of test parameters represents a test procedure, and the parameter items in different test parameters are not entirely the same. For example, in different test parameters, the test text corresponding to the same target test language may be different, or the target test location corresponding to the same test text may be different. For each set of test parameters, perform the following steps 302-307 to conduct the test.
[0173] For example, if Chinese and English are selected as the target test languages, and the test corpus includes the following 10 test scenarios (i.e., interaction types), with 20 test cases (i.e., test text) for each scenario, then the test corpus for Chinese and English will each contain 200 test cases. Examples of test cases for Chinese and English are shown in Table 1.
[0174]
[0175] 302. The multilingual testing system loads the test corpus and converts the test text in the test corpus into test speech.
[0176] 303. The multilingual testing system moves the robotic arm in the movable testing device to the target testing position.
[0177] 304. The multilingual testing system controls the speaker on the robotic arm to play test audio.
[0178] Multiple speakers can play multiple test voices in parallel.
[0179] 305. The multilingual testing system collects the first response data generated by the voice interaction system based on the test voice.
[0180] 306. The multilingual testing system uses an artificial intelligence model to respond to the test speech and obtain second response data.
[0181] 307. The multilingual testing system determines the multilingual test results by comparing the first and second response data corresponding to each test speech, including speech recognition accuracy, operation decision accuracy, speech quality parameters, and average response latency.
[0182] 308. The multilingual testing system determines whether all test parameters have been completed. If yes, proceed to step 309; otherwise, repeat steps 303-307.
[0183] 309. The multilingual testing system generates multilingual test reports based on the established multilingual test results.
[0184] Multilingual test reports can include multilingual test results, as well as a summary of the performance parameters from the multilingual test results.
[0185] Optionally, the multilingual testing system visualizes the various performance parameters in the multilingual test results to obtain performance comparison charts for multiple target test languages. In the performance comparison charts, the performance parameters of each target test language are compared. The horizontal axis represents multiple target test languages, and the vertical axis represents the performance parameters of each target test language, including speech recognition accuracy, operation decision accuracy, speech quality parameters, and average response latency.
[0186] Optionally, the multilingual testing system visualizes the various performance parameters in the multilingual test results to obtain performance comparison charts for multiple target test locations. In the performance comparison charts, the performance parameters of each target test location are compared. The horizontal axis represents multiple target test locations, and the vertical axis represents the performance parameters of each target test location, including speech recognition accuracy, operation decision accuracy, speech quality parameters, and average response latency.
[0187] Optionally, the multilingual testing system analyzes various performance parameters to pinpoint problems in the voice interaction system, identifying problematic functions and their causes, and adding this information to the multilingual test report. For example, if the system finds that the English recognition accuracy in the passenger seat is low, the issue is likely due to a sensitivity problem with the microphone in the passenger seat area.
[0188] For example, the multilingual testing system uses an artificial intelligence model to analyze various performance parameters, locate problems in the voice interaction system, identify problematic functions in the voice interaction system, and determine the causes of the problems.
[0189] Optionally, the multilingual testing system acquires various performance parameters under ambient noise conditions and also acquires various performance parameters without ambient noise conditions. Based on these performance parameters, it conducts an environmental noise impact assessment to determine the voice interaction system's ability to resist environmental noise interference.
[0190] Based on the above embodiments, this application provides another exemplary testing process. Figure 4 This is a schematic diagram of another exemplary test process provided in an embodiment of this application. See also... Figure 4 The testing process includes the following stages.
[0191] I. Test preparation phase, including the following steps 401-404.
[0192] 401. Install the test device in the multilingual test system inside the vehicle cabin, and connect the test device to the power supply and data cable.
[0193] 402. Start the test management platform in the multilingual test system and establish a communication connection with the vehicle's voice interaction system.
[0194] The communication connection can be of the following types: CAN (Controller Area Network), LIN (Local Interconnect Network), Ethernet, etc.
[0195] The test management platform can run on the test device, or the multilingual test system includes the test device and a cloud server connected to the test device, with the test management platform running on the cloud server.
[0196] 403. Calibrate the position of the test device and set four target test positions (driver's seat, passenger seat, left rear seat, right rear seat).
[0197] 404. Select the target test language set and load the test corpus corresponding to each target test language.
[0198] II. Test parameter configuration stage, including the following steps 405-407.
[0199] 405. Set the ambient noise parameters.
[0200] For example, the sound pressure range for urban road noise is set to 55-65 dB, the sound pressure range for highway noise is 65-75 dB, and the sound pressure range for noise generated by stationary vehicles is 45-55 dB.
[0201] 406. Define the target interaction type.
[0202] For example, target interaction types can be set to include: control commands ("turn on the air conditioner"), information queries ("how is the weather today?"), entertainment controls ("play jazz music"), etc.
[0203] 407. Set the threshold for performance parameters.
[0204] For example, setting the speech recognition accuracy threshold to 95% means that the speech recognition accuracy of the voice interaction system must be greater than 95% to pass the test; setting the response latency threshold to 2 seconds means that the average response latency of the voice interaction system must be less than 2 seconds to pass the test; and setting the operation decision accuracy threshold to 90% means that the operation decision accuracy of the voice interaction system must be greater than 90% to pass the test.
[0205] III. Multilingual test execution phase, including the following steps 408.
[0206] 408. The multilingual testing system tests the test text belonging to each target test language in turn, according to the order of the multiple target test languages.
[0207] The testing process for each test text includes: (1) Generate ambient background noise of a specified intensity.
[0208] (2) Play the test text corresponding to the test speech through the speaker (volume is 65-75dB, simulating the intensity of human voice).
[0209] (3) Collect the response voice output by the voice interaction system through a microphone array.
[0210] (4) Record the response delay of the voice interaction system (from the end of playing the test voice to the start of the voice interaction system outputting the response voice).
[0211] (5) Save the response voice and recognized text of the voice interaction system.
[0212] IV. Results Analysis and Report Generation Stage, including the following steps 409-412.
[0213] 409. Rate the response of the voice interaction system to each test voice.
[0214] For example, for each test speech, a score is given based on the comparison between the first response data and the second response data.
[0215] 410. Based on the comparison results of the first response data and the second response data corresponding to multiple test voices, statistically analyze at least one performance parameter.
[0216] For example, performance parameters include: speech recognition accuracy, operation decision accuracy, average response latency, and speech quality parameters.
[0217] 411. Visualize the performance parameters of multiple target testing languages to obtain performance comparison charts for multiple target testing languages. Visualize the performance parameters of multiple target testing locations to obtain performance comparison charts for multiple target testing locations.
[0218] 412. Generate a test report, which includes the score for each test speech, the statistical results of each performance parameter, and a performance comparison chart for each performance.
[0219] Based on the solutions of the above embodiments, the following compares the solutions of related technologies and the embodiments of this application.
[0220] In related technologies, manual testing is usually used, where testers speak test voices in multiple test languages. After each test voice is collected by the tester, the voice interaction system generates response data based on the test voice. The tester then determines the multilingual test result of the voice interaction system by judging whether the response data corresponding to each test voice is correct.
[0221] The related technologies have the following defects 1-4.
[0222] 1. Low testing efficiency: Relying on manual testing, which involves speaking test audio in multiple languages and manually judging the correctness of response data, consumes a lot of manpower and time, has a long testing cycle, and results in low testing efficiency.
[0223] 2. Low test language coverage: Relying on manual testing, which is limited by the language proficiency of testers, it is impossible to cover all language combination scenarios, resulting in a high rate of missed tests.
[0224] 3. Significant environmental interference: Background noise exists in the artificial testing environment, and the test voice spoken by humans may have accent differences, which affects the accuracy of the test and leads to poor reliability of the test results.
[0225] 4. Difficulty in locating problems: When an anomaly is discovered, it is difficult to quickly locate the problematic function, resulting in low efficiency in troubleshooting.
[0226] The embodiments of this application provide an efficient, comprehensive, and automated multilingual testing method, which can achieve the following technical effects 1-5.
[0227] 1. Improved testing efficiency: It can play test audio in multiple languages in parallel, realizing multilingual parallel testing capabilities, which greatly improves testing efficiency and saves testing time.
[0228] 2. Improved test language coverage: No longer relying on manual testing, and not limited by the languages that testers are proficient in, it can cover any preset test language, resulting in a significant improvement in test language coverage.
[0229] 3. Reduced testing costs: Reduced reliance on multilingual testers, thus lowering labor costs.
[0230] 4. Enhanced accuracy of results: By playing ambient noise, a simulated real in-vehicle environment is constructed to evaluate the anti-interference capability of the voice interaction system, improving the reliability of test results and eliminating the influence of human factors, thus improving the accuracy of test results.
[0231] 5. Precise problem localization: A standardized multilingual testing system has been established, covering the entire process of language recognition, operational decision-making, and speech synthesis. This provides a precise problem localization mechanism to quickly identify faulty modules in the voice interaction system.
[0232] Figure 5 This is a schematic diagram of the structure of a multilingual testing system provided in an embodiment of this application. See also... Figure 5 The multilingual testing system is used to perform multilingual testing on the voice interaction system. The multilingual testing system includes a test unit 501 and a control unit 502.
[0233] Test unit 501 is used to play multiple test voices in the vehicle cabin. The multiple test voices belong to multiple target test languages to be tested. The control unit 502 is used to acquire the first response data corresponding to each test voice; the first response data is obtained by the voice interaction system responding to the test voice. The control unit 502 is also used to respond to each test voice through an artificial intelligence model to obtain the second response data corresponding to each test voice. The control unit 502 is also used to determine the multilingual test results of the voice interaction system based on the comparison results of the first response data and the second response data corresponding to each test voice.
[0234] In one possible implementation, the test unit 501 is used to play test voice belonging to each of the multiple target test languages sequentially in the vehicle cabin; or, in the vehicle cabin, multiple test voices are played in parallel each time, with each test voice corresponding to one of the multiple target test languages.
[0235] In one possible implementation, the control unit 502 is configured to receive combined response data sent by the voice interaction system, the combined response data being generated by the voice interaction system based on multiple test voices played in parallel; and to extract first response data corresponding to each test voice from the combined response data.
[0236] In one possible implementation, the control unit 502 is used to split the combined response data to obtain multiple first response data; determine the content correlation degree between each test speech and each first response data; and assign the multiple first response data to multiple test speeches based on the content correlation degree between each test speech and each first response data, so that the multiple first response data correspond one-to-one with the multiple test speeches.
[0237] In one possible implementation, the test unit 501 includes multiple robotic arms and multiple speakers, with each speaker located on one of the robotic arms; The control unit 502 is also used to determine multiple target test locations to be tested; each target test location is located inside the vehicle cabin; Test unit 501 is used to move each robotic arm to the corresponding target test position; The control unit 502 is also used to play test voice belonging to each target test language through each speaker.
[0238] In one possible implementation, the control unit 502 is further configured to, for each target test language, determine the target test corpus corresponding to the target test language, the target test corpus including test text belonging to the target test language and belonging to multiple interaction types; obtain test text belonging to the target interaction type to be tested from the target test corpus; and convert the test text into the corresponding test speech.
[0239] In one possible implementation, the control unit 502 is configured to perform speech recognition on each test speech to obtain a first recognized text corresponding to the test speech; respond to the first recognized text through an artificial intelligence model to obtain a first response operation matching the first recognized text; and obtain a first response speech matching the first response operation.
[0240] In one possible implementation, the second response data includes at least one of the first recognized text, the first response operation, and the first response speech, and the first response data includes at least one of the second recognized text, the second response operation, and the second response speech; Control unit 502 is configured to perform at least one of the following steps: Based on the comparison results of the first and second recognized texts corresponding to each test speech, the speech recognition accuracy of the voice interaction system is determined. Based on the comparison results of the first response operation and the second response operation corresponding to each test voice, the operation decision accuracy of the voice interaction system is determined. Based on the comparison results of the first response voice and the second response voice corresponding to each test voice, the voice quality parameters of the voice interaction system are determined.
[0241] In one possible implementation, the multilingual test results include the average response delay of the voice interaction system; the control unit 502 is further configured to determine the response delay corresponding to each test voice, wherein the response delay is the time interval between the completion time of the test voice playback and the time of acquiring the response voice of the voice interaction system, or the response delay is the difference between the time interval between the completion time of the test voice playback and the time of acquiring the response voice of the voice interaction system and a preset time interval; and the average response delay of the voice interaction system is determined based on the response delay corresponding to each test voice.
[0242] This application provides a multilingual testing system for a voice interaction system. It can play multiple test voices belonging to various target test languages within a vehicle cabin, acquiring first and second response data for each test voice. The first response data is determined by the voice interaction system, and the second response data is determined by an artificial intelligence model. The comparison between the first and second response data for each test voice reflects the correctness of the voice interaction system's response process. Therefore, based on the comparison results, the multilingual test result of the voice interaction system is determined. During the testing process, testers do not need to speak the test voices belonging to various target test languages, nor do they need to manually judge the correctness of the voice interaction system's response data, reducing labor and time costs and improving the efficiency of multilingual testing.
[0243] Figure 6 This is a schematic diagram of the structure of a multilingual testing device for a voice interaction system provided in an embodiment of this application. See also... Figure 6 The multilingual testing device for the voice interaction system includes the following modules.
[0244] The voice playback module 601 is used to play multiple test voices in the vehicle cabin. The multiple test voices belong to various target test languages to be tested. The data acquisition module 602 is used to acquire the first response data corresponding to each test voice; the first response data is obtained by the voice interaction system responding to the test voice. The voice response module 603 is used to respond to each test voice separately through an artificial intelligence model to obtain the second response data corresponding to each test voice. The result determination module 604 is used to determine the multilingual test results of the voice interaction system based on the comparison results of the first response data and the second response data corresponding to each test voice.
[0245] In one possible implementation, the voice playback module 601 is used to play test voices belonging to each target test language sequentially in the vehicle cabin according to the order of multiple target test languages; or, in the vehicle cabin, multiple test voices are played in parallel each time, with each test voice corresponding to one of the multiple target test languages.
[0246] In one possible implementation, the data acquisition module 602 is used to receive combined response data sent by the voice interaction system, the combined response data being generated by the voice interaction system based on multiple test voices played in parallel; and to extract the first response data corresponding to each test voice from the combined response data.
[0247] In one possible implementation, the data acquisition module 602 is used to split the combined response data to obtain multiple first response data; determine the content correlation between each test speech and each first response data; and assign the multiple first response data to multiple test speeches based on the content correlation between each test speech and each first response data, so that the multiple first response data correspond one-to-one with the multiple test speeches.
[0248] In one possible implementation, the multilingual testing apparatus for the voice interaction system includes multiple robotic arms and multiple speakers, with each speaker located on one robotic arm. The voice playback module 601 is used to determine multiple target test locations to be tested; each target test location is located inside the vehicle cabin; each robotic arm is moved to the corresponding target test location; and test voice belonging to each target test language is played through each speaker.
[0249] In one possible implementation, the apparatus further includes: a speech determination module, configured to determine, for each target test language, a target test corpus corresponding to the target test language, the target test corpus including test text belonging to the target test language and belonging to multiple interaction types; obtain test text belonging to the target interaction type to be tested from the target test corpus; and convert the test text into the corresponding test speech.
[0250] In one possible implementation, the voice response module 603 is used to perform voice recognition on each test voice to obtain a first recognized text corresponding to the test voice; respond to the first recognized text through an artificial intelligence model to obtain a first response operation matching the first recognized text; and obtain the first response voice matching the first response operation.
[0251] In one possible implementation, the second response data includes at least one of the first recognized text, the first response operation, and the first response speech, and the first response data includes at least one of the second recognized text, the second response operation, and the second response speech; Result determination module 604 is used to perform at least one of the following steps: Based on the comparison results of the first and second recognized texts corresponding to each test speech, the speech recognition accuracy of the voice interaction system is determined. Based on the comparison results of the first response operation and the second response operation corresponding to each test voice, the operation decision accuracy of the voice interaction system is determined. Based on the comparison results of the first response voice and the second response voice corresponding to each test voice, the voice quality parameters of the voice interaction system are determined.
[0252] In one possible implementation, the multilingual test results include the average response latency of the voice interaction system; the result determination module 604 is further configured to determine the response latency corresponding to each test voice, wherein the response latency is the time interval between the completion time of the test voice playback and the time of acquiring the response voice of the voice interaction system, or the response latency is the difference between the time interval between the completion time of the test voice playback and the time of acquiring the response voice of the voice interaction system and a preset time interval; and the average response latency of the voice interaction system is determined based on the response latency corresponding to each test voice.
[0253] This application provides a multilingual testing device for a voice interaction system. It can play multiple test voices belonging to various target test languages within a vehicle cabin, acquiring first and second response data for each test voice. The first response data is determined by the voice interaction system, and the second response data is determined by an artificial intelligence model. The comparison between the first and second response data for each test voice reflects the correctness of the voice interaction system's response process. Therefore, based on the comparison results, the multilingual test result of the voice interaction system is determined. During the testing process, testers do not need to speak the test voices belonging to various target test languages, nor do they need to manually judge the correctness of the voice interaction system's response data, reducing labor and time costs and improving the efficiency of multilingual testing.
[0254] Figure 7 This is a schematic diagram of a computer device provided in an embodiment of this application. The computer device is implemented as the multilingual testing system described above, to execute the multilingual testing method of the voice interaction system provided in this embodiment. The computer device 700 can vary significantly due to different configurations or performance, and may include one or more processors 701 and one or more memories 702. The memory 702 stores at least one computer program, which is loaded and executed by the processor 701 to implement the multilingual testing method of the voice interaction system provided in the various method embodiments described above. Of course, the computer device may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The computer device may also include other components for implementing device functions, which will not be elaborated here.
[0255] This application also provides a computer-readable storage medium storing at least one line of program code, which is loaded and executed by a vehicle's processor to implement the multilingual testing method for the voice interaction system described in the above embodiments. The computer-readable storage medium can be a memory. For example, it can be a ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, or optical data storage terminal, etc.
[0256] This application also provides a computer program product, which includes computer program code stored in a computer-readable storage medium. The vehicle's processor reads the computer program code from the computer-readable storage medium and executes the computer program code to implement the multilingual testing method for the voice interaction system as described in the above embodiments.
[0257] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0258] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A multilingual testing method for a voice interaction system, characterized in that, The method includes: Multiple test voices are played inside the vehicle cabin, and the multiple test voices belong to various target test languages to be tested; The first response data corresponding to each test voice is obtained; the first response data is obtained by the voice interaction system responding to the test voice. By using an artificial intelligence model, each test speech is responded to separately, and second response data corresponding to each test speech is obtained; Based on the comparison results of the first response data and the second response data corresponding to each test voice, the multilingual test result of the voice interaction system is determined.
2. The method according to claim 1, characterized in that, The process involves playing multiple test voice recordings within the vehicle cabin. These multiple test voice recordings belong to various target test languages to be tested, including: Inside the vehicle cabin, test audio for each of the multiple target test languages is played sequentially according to their order; or, Inside the vehicle cabin, multiple test voices are played in parallel each time, and each of the multiple test voices corresponds one-to-one with the multiple target test languages.
3. The method according to claim 2, characterized in that, The step of acquiring the first response data corresponding to each test voice message includes: Receive combined response data sent by the voice interaction system, wherein the combined response data is generated by the voice interaction system based on multiple test voices played in parallel; From the combined response data, extract the first response data corresponding to each test voice.
4. The method according to claim 3, characterized in that, The step of extracting the first response data corresponding to each test speech from the combined response data includes: The combined response data is split to obtain multiple first response data; Determine the content correlation degree between each test speech and each first response data; Based on the content correlation between each test speech and each first response data, the plurality of first response data are respectively assigned to the plurality of test speech, so that the plurality of first response data corresponds one-to-one with the plurality of test speech.
5. The method according to any one of claims 1 to 4, characterized in that, The method is executed by a multilingual testing system, which includes multiple robotic arms and multiple speakers, with each speaker located on one robotic arm. The process involves playing multiple test voice recordings within the vehicle cabin. These multiple test voice recordings belong to various target test languages to be tested, including: Identify multiple target test locations to be tested; each target test location is located within the vehicle's cabin. Move each robotic arm to its corresponding target test position; Each speaker plays a test audio message belonging to each target test language.
6. The method according to any one of claims 1 to 4, characterized in that, The process involves playing multiple test voice recordings within the vehicle cabin. These multiple test voice recordings belong to various target test languages to be tested, including: For each target test language, a target test corpus corresponding to the target test language is determined. The target test corpus includes test texts belonging to the target test language and belonging to multiple interaction types. Obtain test text belonging to the target interaction type to be tested from the target test corpus; The test text is converted into the corresponding test speech.
7. The method according to any one of claims 1 to 4, characterized in that, The process involves using an artificial intelligence model to respond to each test speech, obtaining second response data corresponding to each test speech, including: For each test speech, speech recognition is performed on the test speech to obtain the first recognized text corresponding to the test speech; The artificial intelligence model responds to the first identified text to obtain a first response operation that matches the first identified text. Obtain the first response voice that matches the first response operation.
8. The method according to claim 7, characterized in that, The second response data includes at least one of the first recognized text, the first response operation, and the first response speech; the first response data includes at least one of the second recognized text, the second response operation, and the second response speech. The determination of the multilingual test result of the voice interaction system based on the comparison result between the first response data and the second response data corresponding to each test voice includes at least one of the following: Based on the comparison results of the first recognized text and the second recognized text corresponding to each test speech, the speech recognition accuracy of the voice interaction system is determined; Based on the comparison results of the first response operation and the second response operation corresponding to each test voice, the operation decision accuracy of the voice interaction system is determined; Based on the comparison results between the first response voice and the second response voice corresponding to each test voice, the voice quality parameters of the voice interaction system are determined.
9. The method according to any one of claims 1 to 4, characterized in that, The multilingual test results include the average response latency of the voice interaction system; the method further includes: Determine the response delay corresponding to each test voice, wherein the response delay is the time interval between the time when the test voice is played and the time when the response voice of the voice interaction system is acquired, or the response delay is the difference between the time interval between the time when the test voice is played and the time when the response voice of the voice interaction system is acquired and a preset time interval. The average response delay of the voice interaction system is determined based on the response delay corresponding to each test voice.
10. A multilingual testing system, characterized in that, The multilingual testing system includes: a testing unit and a control unit; The test unit is used to play multiple test voices in the vehicle cabin, and the multiple test voices belong to multiple target test languages to be tested. The control unit is used to acquire first response data corresponding to each test voice; the first response data is obtained by the voice interaction system responding to the test voice. The control unit is also used to respond to each test voice through an artificial intelligence model to obtain second response data corresponding to each test voice. The control unit is further configured to determine the multilingual test result of the voice interaction system based on the comparison result of the first response data and the second response data corresponding to each test voice.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by the processor of the multilingual testing system to implement the multilingual testing method for the voice interaction system as described in any one of claims 1 to 9.
12. A computer program product, characterized in that, The computer program product includes computer program code stored in a computer-readable storage medium. The processor of the multilingual testing system reads the computer program code from the computer-readable storage medium and executes the computer program code to implement the multilingual testing method for the voice interaction system as described in any one of claims 1 to 9.