Quality evaluation method for vehicle voice interaction function, electronic equipment and medium

By generating diverse voice requests and performing spatiotemporal alignment processing, and combining embedded vectors to evaluate the vehicle's voice interaction function, the problem of incomplete evaluation in existing technologies is solved, efficient and accurate multilingual evaluation is achieved, the cost of manual labeling is reduced, and the robustness of the voice interaction function and user experience are guaranteed.

CN120612923APending Publication Date: 2025-09-09GUANGZHOU XIAOPENG MOTORS TECH CO LTD

Patent Information

Application Number
CN202510875081.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In existing technologies, the quality assessment of vehicle voice interaction functions is too one-sided and difficult to fully reflect its effectiveness and reliability. In addition, multilingual testing relies on manual labeling, which is costly and has limited coverage scenarios.

Method used

By generating multi-language, multi-scenario, multi-timbre, and multi-emotion voice requests, combining system operation logs, voice recognition results, executed operations, and feedback voice for spatiotemporal alignment processing, and using embedded vectors and pre-trained models to evaluate the quality of voice interaction functions, the reliance on manual annotation of native speakers is reduced.

Benefits of technology

It achieves full-dimensional evaluation of vehicle voice interaction functions, improves the accuracy and coverage of the evaluation, reduces the cost of multilingual testing, and ensures the stable operation and user experience of the voice interaction function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612923A_ABST
    Figure CN120612923A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle voice interaction function quality evaluation method, electronic equipment and a computer readable storage medium. The method comprises the steps of obtaining a test case set; according to a test case in the test case set, generating a voice request corresponding to the test case; the quality of the voice interaction function of the vehicle is evaluated according to the test case and the response result of the vehicle for the voice request, and the response result comprises a system operation log, a voice recognition result, execution operation, voice feedback and response time. Therefore, according to the test case in the test case set, the voice request corresponding to the test case is generated, and the quality of the voice interaction function of the vehicle is evaluated according to the response result of the voice request, so that the voice interaction function of the vehicle can be updated, maintained and the like based on the quality evaluation result; robust operation of the vehicle voice interaction function is guaranteed to a certain extent, and then the voice interaction function and the use experience of the vehicle are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of vehicle technology, and in particular to a method for evaluating the quality of a vehicle's voice interaction function, an electronic device, and a computer-readable storage medium. Background Art

[0002] To ensure the effectiveness and reliability of vehicle voice interaction functions, quality assessments are typically conducted, and updates and maintenance of these functions are implemented based on the assessment results. However, the quality of vehicle voice interaction functions is often assessed based on metrics such as voice recognition accuracy and response time, which is rather one-sided and fails to effectively reflect the quality of vehicle voice interaction functions. Summary of the Invention

[0003] The present application provides a method for evaluating the quality of a vehicle's voice interaction function, an electronic device, and a computer-readable storage medium.

[0004] The present application provides a method for evaluating the quality of a vehicle voice interaction function, the method comprising:

[0005] Get the test case set;

[0006] generating, according to a test case in the test case set, a voice request corresponding to the test case;

[0007] The quality of the vehicle's voice interaction function is evaluated based on the test case and the vehicle's response to the voice request, wherein the response result includes the system operation log, voice recognition result, execution operation, feedback voice and response time.

[0008] Thus, in the embodiment of the present application, when a test case set is obtained, a voice request corresponding to the test case can be generated based on the test case in the test case set. Then, based on the test case and the response result of the vehicle to the voice request after obtaining the voice request, the quality of the vehicle's voice interaction function can be evaluated, so that the vehicle's voice interaction function can be updated and maintained based on the quality evaluation results, to a certain extent, ensuring the stable operation of the vehicle's voice interaction function, and thus ensuring the user experience of the voice interaction function and the vehicle. In addition, the quality of the vehicle's voice interaction function can be evaluated based on the system operation log, voice recognition results, execution operations, feedback voice and response time. Compared with the situation where the quality of the vehicle's voice interaction function is evaluated only by indicators such as voice recognition accuracy and response time, the functional quality evaluation of the vehicle's voice interaction can be performed through multi-dimensional data, thereby ensuring the accurate evaluation of the quality of the vehicle's voice interaction function, and thus ensuring the effectiveness and reliability of updates and maintenance based on the quality evaluation results.

[0009] In certain embodiments, evaluating the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request includes:

[0010] Performing spatiotemporal alignment processing on the obtained system operation log, the speech recognition result, the execution operation, and the feedback speech;

[0011] The quality of the vehicle's voice interaction function is evaluated based on the test case, the response time, the system operation log after time-space alignment processing, the voice recognition result, the execution operation, and the feedback voice.

[0012] In this way, the acquired system operation logs, speech recognition results, executed operations, and voice feedback are temporally and spatially aligned. The quality of the vehicle's voice interaction function is evaluated based on test cases, response times, and the temporally and spatially aligned system operation logs, speech recognition results, executed operations, and voice feedback. Synchronizing response result data can reduce errors in multimodal data association and avoid misjudgments caused by time desynchronization. Furthermore, by eliminating the impact of vehicle vibration on video capture, the consistency of multi-view images can be ensured, avoiding misjudgments caused by operations not captured due to perspective shifts.

[0013] In some embodiments, generating a voice request corresponding to a test case in the test case set includes:

[0014] According to the test case, a plurality of voice requests having different voice attributes are generated, wherein the voice attributes include at least one of language category, dialect category, audio timbre, utterance scene, and intonation.

[0015] In this way, multiple voice requests with varying voice attributes are generated based on test cases. These attributes include at least one of language, dialect, audio timbre, context, and intonation. By combining TTS technology with real human voices, voice requests can be generated in multiple languages, contexts, timbres, and emotions, including ambient noise. This simulates voice commands from different users and scenarios, enhancing the comprehensiveness of the quality assessment of in-vehicle voice interaction functions.

[0016] In some embodiments, the voice attribute includes a language category, and evaluating the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request includes:

[0017] Determining a first embedding vector corresponding to the test case according to predetermined mapping data;

[0018] performing encoding processing on the speech recognition result to determine a second embedding vector corresponding to the speech recognition result;

[0019] The quality of the voice interaction function of the vehicle is evaluated based on the test case, the system operation log, the execution operation, the feedback voice, the response time, and the difference between the first embedding vector and the second embedding vector.

[0020] In this way, based on predetermined mapping data, a first embedding vector corresponding to the test case is determined; the speech recognition results are encoded and processed to determine a corresponding second embedding vector; and the quality of the vehicle's voice interaction function is evaluated based on the test case, system operation logs, executed operations, feedback speech, response time, and the difference between the first and second embedding vectors. This method maps voice requests in different languages ​​to the same high-dimensional space, and measures semantic similarity through vector distance, enabling cross-language evaluation based on semantic equivalence rather than lexical correspondence. This improves the accuracy of vehicle voice interaction quality evaluation, and eliminates the need for extensive native speaker manual annotation, reducing the cost of multilingual testing.

[0021] In some embodiments, the voice attributes include multiple dialect categories under the same language category, and evaluating the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request includes:

[0022] Performing conversion processing on the feedback speech to determine a target speech under a target dialect category;

[0023] The quality of the voice interaction function of the vehicle is evaluated based on the test case, the system operation log, the execution operation, the feedback voice, the response time and the target voice.

[0024] In this way, feedback speech is converted to determine the target speech within the target dialect category. The quality of the vehicle's voice interaction function is evaluated based on test cases, system operation logs, executed operations, feedback speech, response time, and the target speech. This conversion of feedback speech into the target speech based on a pre-trained dialect-to-Mandarin conversion model overcomes dialect differences, standardizes dialect semantics, unifies evaluation benchmarks, expands scenario coverage, and improves the comprehensiveness and robustness of the evaluation system.

[0025] In some embodiments, obtaining a test case set includes:

[0026] According to the predetermined vehicle voice interaction function information, a plurality of the test cases are generated to construct the test case set.

[0027] In this way, multiple test cases are generated based on the predetermined vehicle voice interaction function information to form a test case set. This test case set, based on the predetermined vehicle voice interaction function information, can cover the core functional scenarios of vehicle voice interaction, achieving comprehensive evaluation. Furthermore, automatically generating test cases with expected standards based on the function list provides an objective benchmark for evaluation, improves evaluation efficiency, and facilitates systematic data analysis.

[0028] In certain embodiments, evaluating the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request includes:

[0029] When the voice request is played to the vehicle, the response result is obtained by recording devices arranged inside and outside the vehicle cabin space;

[0030] The quality of the voice interaction of the vehicle is evaluated based on the response result and the test case.

[0031] In this way, when a voice request is played to the vehicle, the response is captured by recording devices installed inside and outside the vehicle cabin. Based on the response results and test cases, the quality of the vehicle's voice interaction is evaluated. By recording responses from different angles using multiple recording devices, multiple interaction points within the vehicle are covered, enabling collaborative analysis and evaluation of multimodal data. The system also automatically calibrates the recording devices, reducing manual inspection costs.

[0032] In certain embodiments, evaluating the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request includes:

[0033] The quality of the vehicle's voice interaction is evaluated based on the response result, the test case, and the pre-trained generative model.

[0034] In this way, the quality of a vehicle's voice interaction is evaluated based on response results, test cases, and a pre-trained generative model. This model integrates video, audio, and other data to evaluate the in-vehicle voice system. This allows for determining ASR recognition accuracy, verifying correct operation execution, checking response relevance, measuring response speed, and comparing against thresholds to determine whether the vehicle's response meets expectations, enabling accurate evaluation of the quality of the vehicle's voice interaction functionality. It can also locate hidden defects and generate quantitative metrics such as pass rates, providing data support for system optimization. This enables an upgrade from single-dimensional sampling to full-dimensional verification, addressing issues such as subjectivity and insufficient coverage in manual testing, and improving evaluation efficiency and reliability.

[0035] An embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the steps of the above method are implemented.

[0036] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program. When the computer program is executed by one or more processors, the steps of the above method are implemented.

[0037] The electronic device and computer-readable storage medium provided by the embodiments of the present application can obtain a test case set; generate a voice request corresponding to the test case based on the test case in the test case set; and evaluate the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request, wherein the response result includes the system operation log, voice recognition result, execution operation, feedback voice, and response time. In this way, when a test case set is obtained, a voice request corresponding to the test case can be generated based on the test case in the test case set, and then the quality of the vehicle's voice interaction function can be evaluated based on the test case and the response result to the voice request after the vehicle obtains the voice request, so that the vehicle's voice interaction function can be updated, maintained, etc. based on the quality evaluation results, thereby ensuring the stable operation of the vehicle's voice interaction function to a certain extent, and thus ensuring the user experience of the voice interaction function and the vehicle.

[0038] Additional aspects and advantages of the embodiments of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0040] Figure 1 This is one of the flow charts of a method for evaluating the quality of a vehicle voice interaction function according to certain embodiments of the present application;

[0041] Figure 2 This is a second flow chart of a method for evaluating the quality of a vehicle voice interaction function according to certain embodiments of the present application;

[0042] Figure 3 This is a third flow chart of a method for evaluating the quality of a vehicle voice interaction function according to certain embodiments of the present application;

[0043] Figure 4 This is a fourth flow chart of a method for evaluating the quality of a vehicle voice interaction function according to certain embodiments of the present application;

[0044] Figure 5 This is a fifth flow chart of a method for evaluating the quality of a vehicle voice interaction function according to certain embodiments of the present application;

[0045] Figure 6 This is the sixth flow chart of the quality evaluation method of the vehicle voice interaction function in certain embodiments of the present application;

[0046] Figure 7 Schematic diagram of evaluation tasks for a quality evaluation method for a vehicle voice interaction function according to certain embodiments of the present application;

[0047] Figure 8 This is the seventh flow chart of the quality evaluation method of the vehicle voice interaction function in certain embodiments of the present application;

[0048] Figure 9 This is a schematic diagram of mobile phone arrangement for a method for evaluating the quality of a vehicle voice interaction function according to certain embodiments of the present application;

[0049] Figure 10 is a monitoring diagram of a quality evaluation method for a vehicle voice interaction function according to certain embodiments of the present application;

[0050] Figure 11 This is the eighth flow chart of the quality evaluation method of the vehicle voice interaction function in certain embodiments of the present application;

[0051] Figure 12 is a schematic diagram of a result report of a quality evaluation method for a vehicle voice interaction function according to certain embodiments of the present application;

[0052] Figure 13 is a schematic diagram of human-machine collaboration of a method for evaluating the quality of a vehicle voice interaction function according to certain embodiments of the present application;

[0053] Figure 14 It is a system diagram of a method for evaluating the quality of vehicle voice interaction functions in certain embodiments of the present application. DETAILED DESCRIPTION

[0054] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and should not be understood as limiting the embodiments of the present application.

[0055] In related technologies, in order to ensure the effectiveness and reliability of the vehicle voice interaction function, the vehicle voice interaction function can usually be quality evaluated, and the vehicle voice interaction function can be updated and maintained based on the quality evaluation results.

[0056] The quality of a vehicle's voice interaction function is a core criterion for measuring the practicality, convenience, and reliability of the system in driving scenarios. It encompasses multiple dimensions, including technical performance, driving adaptability, and safety assurance. However, current evaluation systems often base their assessments on metrics such as voice recognition accuracy and response time, which is rather one-sided and fails to effectively reflect the quality of a vehicle's voice interaction function. For example, a user's voice request for "navigate to the airport" may be correctly recognized as text, but a train station map may be loaded onto the vehicle's control panel. This hidden flaw cannot be verified through on-screen operation videos.

[0057] Furthermore, most current vehicle voice interaction testing systems are still manually controlled, with rigid test cases and limited semantic coverage. Furthermore, multilingual testing requires native speakers, increasing labor costs. For example, when faced with complex commands like "Navigate to the highest-rated Italian restaurant" that include dynamic ratings and location attributes, manual testing is unable to exhaust all possibilities.

[0058] Based on the above questions, please refer to Figure 1 The present application provides a method for evaluating the quality of a vehicle voice interaction function, the method comprising:

[0059] 01: Get the test case set;

[0060] 02: Generate voice requests corresponding to the test cases in the test case set;

[0061] 03: Evaluate the quality of the vehicle's voice interaction function based on test cases and the vehicle's response to voice requests. The response results include system operation logs, voice recognition results, executed operations, feedback voice, and response time.

[0062] The embodiment of the present application provides a quality evaluation device for a vehicle voice interaction function. The quality evaluation method for a vehicle voice interaction function of the embodiment of the present application can be implemented by the quality evaluation device for a vehicle voice interaction function of the embodiment of the present application. Specifically, the quality evaluation device for a vehicle voice interaction function includes an acquisition module, a generation module, and an evaluation module. The acquisition module is used to acquire a test case set. The generation module is used to generate voice requests corresponding to the test cases based on the test cases in the test case set. The evaluation module is used to evaluate the quality of the vehicle's voice interaction function based on the test cases and the vehicle's response results to the voice requests, wherein the response results include system operation logs, voice recognition results, execution operations, feedback voice, and response time.

[0063] The embodiment of the present application also provides an electronic device, which includes a memory and a processor. The quality evaluation method of the vehicle voice interaction function of the embodiment of the present application can be implemented by the electronic device of the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is used to obtain a test case set. The processor is also used to generate a voice request corresponding to the test case based on the test case in the test case set. The processor is also used to evaluate the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request, wherein the response result includes the system operation log, voice recognition result, execution operation, feedback voice and response time.

[0064] Specifically, the voice interaction function is a function that enables vehicle control and information exchange based on user voice requests, including a pre-set list of functions such as navigation settings, multimedia control, vehicle control, information query, and open domain voice search. In the embodiments of the present application, a test case set for the quality assessment of the vehicle voice interaction function can be automatically generated based on the vehicle voice interaction function list.

[0065] Among them, the test case set covers the common functions, boundary conditions and abnormal scenarios of vehicle voice interaction to ensure the comprehensiveness of the evaluation, which is used to clarify the expected system behavior and evaluation criteria and provide a judgment basis for subsequent evaluation. In some embodiments, each test case in the test case set consists of at least one text, the vehicle expected result corresponding to the text, and the evaluation criteria corresponding to the text. For example, the text of the test case is "Turn on the air conditioner", and its vehicle expected result is "The vehicle air conditioner is turned on", and the evaluation criteria is "If the vehicle air conditioner is turned on, it is qualified, otherwise it is unqualified".

[0066] Based on the test cases in the test case set, using TTS (Text to Speech) technology, manual data collection, and other techniques, the test case text can be converted into test audio in multiple languages, multiple scenarios, multiple timbres, multiple emotions, and containing environmental noise, namely, voice requests in the embodiments of this application. Voice requests are used to simulate voice commands from different users in different scenarios.

[0067] In response to a voice request, the vehicle's voice interaction function can generate a response result. The response result includes system operation logs, voice recognition results, executed operations, feedback voice, and response time. This response result can be recorded by a recording device, such as a mobile phone or log collector.

[0068] Among them, the system operation log is the underlying operation data of the vehicle system collected by the log collector, including system process status, interface call records, crash logs, performance indicators and other data information, which is used to evaluate the quality of the voice interaction system in combination with other data. The voice recognition result is the text result of the vehicle interaction function recognizing and converting the voice request. In some embodiments, it can be displayed in text form on the vehicle control screen and can be recorded by a mobile phone. The execution operation is a video of the specific action performed by the vehicle hardware or functional module recorded by the recording device. Feedback voice is the voice response audio in response to user instructions. The response time refers to the time interval from playing the voice request to the system completing the response.

[0069] The quality of the vehicle's voice interaction functionality can be assessed by comparing the vehicle's response to voice requests, including system operation logs, voice recognition results, executed actions, feedback, and response time, with the test case text and its clearly defined expected system behavior and evaluation criteria. For example, if the test case text is "Check the weather," the semantic consistency between the system's voice response and the user's command can be analyzed by comparing the feedback voice to see if it contains weather information, thereby assessing the quality of the vehicle's voice interaction functionality.

[0070] Compared with evaluating the quality of vehicle voice interaction functions only through indicators such as voice recognition accuracy and response time, collaborative testing of multimodal data such as test cases and system operation logs, voice recognition results, execution operations, feedback voice and response time can improve the accuracy of the evaluation of the quality of vehicle voice interaction functions and achieve full-dimensional evaluation of the quality of vehicle voice interaction functions.

[0071] In the embodiment of the present application, a test case set can be automatically generated based on the voice interaction function list. Then, based on the test case set and TTS and / or artificial audio acquisition technology, diversified voice requests corresponding to the test cases can be generated to simulate voice requests from different users in different scenarios. In response to the voice request, the system operation log, voice recognition results, execution operation, feedback voice, response time and other response results are collected. By comparing the test case with the response results, it is possible to determine whether the vehicle's voice recognition results are accurate, whether the execution operation is correct, whether the feedback voice has relevant lines, and whether the response result is timely, thereby achieving an accurate evaluation of the quality of the vehicle's voice interaction function.

[0072] In addition, the response results can be summarized to generate evaluation reports such as test case pass rate, error rate, and problem classification statistics, thereby ensuring that updates and maintenance based on quality evaluation results are effective and reliable.

[0073] In summary, in the embodiments of the present application, when a test case set is obtained, a voice request corresponding to the test case can be generated based on the test case in the test case set. Then, based on the test case and the response result of the vehicle to the voice request after the voice request is obtained, the quality of the vehicle's voice interaction function can be evaluated, so that the vehicle's voice interaction function can be updated and maintained based on the quality evaluation results, etc., to a certain extent, the stable operation of the vehicle's voice interaction function can be guaranteed, and thus the user experience of the voice interaction function and the vehicle can be guaranteed. In addition, the quality of the vehicle's voice interaction function can be evaluated based on the system operation log, voice recognition results, execution operations, feedback voice and response time. Compared with the situation where the quality of the vehicle's voice interaction function is evaluated only by indicators such as voice recognition accuracy and response time, the functional quality evaluation of the vehicle's voice interaction can be performed through multi-dimensional data, thereby ensuring the accurate evaluation of the quality of the vehicle's voice interaction function, and thus ensuring the effectiveness and reliability of updates and maintenance based on the quality evaluation results.

[0074] See also Figure 2 In some embodiments, step 03 (evaluating the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request) includes:

[0075] 0301: Perform spatiotemporal alignment processing on the obtained system operation logs, speech recognition results, executed operations, and feedback voice;

[0076] 0302: Evaluate the quality of the vehicle's voice interaction function based on test cases, response time, and system operation logs after spatiotemporal alignment processing, speech recognition results, executed operations, and feedback voice.

[0077] In certain embodiments, the evaluation module is further configured to perform spatiotemporal alignment on the acquired system operation logs, speech recognition results, executed operations, and voice feedback. The evaluation module is further configured to evaluate the quality of the vehicle's voice interaction function based on test cases, response times, and the spatiotemporal alignment of the acquired system operation logs, speech recognition results, executed operations, and voice feedback.

[0078] In certain embodiments, the processor is further configured to perform spatiotemporal alignment on the acquired system operation logs, speech recognition results, executed operations, and voice feedback. The processor is further configured to evaluate the quality of the vehicle's voice interaction function based on test cases, response times, and the spatiotemporal alignment of the acquired system operation logs, speech recognition results, executed operations, and voice feedback.

[0079] Specifically, spatiotemporal alignment processing is used to synchronize system operation logs, voice recognition results, execution operations, and feedback voice in the time dimension, providing more accurate data basis for subsequent evaluation of the quality of the vehicle's voice interaction function to ensure that the accuracy and real-time requirements of the vehicle's voice interaction function are met.

[0080] It is understandable that there may be deviations in the system operation log, voice recognition results, execution operations, and feedback voice due to different recording devices and collection timings. For example, the system operation log is collected by the log collector, and the voice recognition results, execution operations, and feedback voice are recorded by the recording device. Each device is affected by factors such as processing time, algorithm complexity, and delay. The data timing collected by each device may have deviations, which may cause different modal data to be unable to be accurately associated with the same interaction event. For example, in the case of a long delay time, the execution operation of the subsequent response result may be incorrectly associated with the current voice request.

[0081] By adding timestamps to the recording device, various types of response data can be mapped to the same timeline, ensuring that the data is aligned in the time dimension to correct deviations caused by network delays and avoid misjudgments due to time asynchrony.

[0082] Based on the system operation logs, voice recognition results, execution operations, and feedback voice after time-space alignment processing, they can be compared with test cases more accurately to determine whether the vehicle's voice recognition results are accurate, whether the execution operations are correct, whether the feedback voice has relevant lines, and whether the response results are timely based on the response time, so as to evaluate the quality of the vehicle's voice interaction function.

[0083] In this way, the acquired system operation logs, speech recognition results, executed operations, and voice feedback are temporally and spatially aligned. The quality of the vehicle's voice interaction function is evaluated based on test cases, response times, and the temporally and spatially aligned system operation logs, speech recognition results, executed operations, and voice feedback. Synchronizing response result data can reduce errors in multimodal data association and avoid misjudgments caused by time desynchronization. Furthermore, by eliminating the impact of vehicle vibration on video capture, the consistency of multi-view images can be ensured, avoiding misjudgments caused by operations not captured due to perspective shifts.

[0084] See also Figure 3 In some embodiments, step 02 (generating a voice request corresponding to a test case in a test case set) includes:

[0085] 021: Generate multiple voice requests with different voice attributes according to the test case, where the voice attributes include at least one of language category, dialect category, audio timbre, vocalization scene, and intonation.

[0086] In some embodiments, the generation module is further configured to generate a plurality of voice requests with different voice attributes based on the test case, wherein the voice attributes include at least one of language category, dialect category, audio timbre, vocalization scene, and intonation.

[0087] In some embodiments, the processor is further configured to generate a plurality of voice requests with different voice attributes based on the test case, wherein the voice attributes include at least one of language category, dialect category, audio timbre, vocalization scene, and intonation.

[0088] Specifically, the voice attributes include at least one of language category, dialect category, audio timbre, vocalization scene, and intonation. Based on the text of the test case, multiple voice requests of different language categories, dialect categories, audio timbre, vocalization scene, or intonation can be obtained by performing speech conversion processing using TTS technology and / or manual reading.

[0089] Using TTS technology, we can capture synonymous commands in different languages ​​or dialects, thereby generating voice requests in these languages ​​and dialects to simulate the language habits of different users. For example, we can use TTS technology to generate synonymous voice requests in different languages, such as "Play jazz" in English, "play jazz" in Chinese, and "Spiele Jazz" in German, from test case text. We can also generate synonymous voice requests in different dialects, such as "Drive to the nearest convenience store" in Cantonese, "Drive to the nearest convenience store" in Sichuanese, and "Run to the nearest convenience store" in Northeastern Chinese.

[0090] Voice synthesis technology can also be used to adjust timbre parameters or collect real human voice samples to generate multi-timbre, multi-emotional test audio to cover different user characteristics. For example, a high-pitched, drawn-out girl's voice saying, "Please help me navigate to the Bund~" or a relaxed, elderly man's voice saying, "Navigation to the Bund, thank you."

[0091] Pure speech can also be superimposed with dynamic noise, such as wind noise and heavy rain, at different signal-to-noise ratios to generate test audio containing environmental noise, simulating real driving scenarios. In particular, a pre-trained adversarial network can be used to dynamically generate a noisy environment to simulate extreme driving scenarios such as heavy rain or tunnel echoes.

[0092] Algorithms can also be used to adjust intonation to simulate voice commands with varying emotions and assess system adaptability. Specifically, based on a pre-trained spatiotemporal perception model, the system can automatically switch city dialects or scene-specific commands based on the vehicle's GPS location, enabling dynamic scene adaptation. For example, a rising tone at the end of a sentence can indicate a question, such as "Play music?", or "Play music!" with the emphasis on "play."

[0093] In this way, multiple voice requests with varying voice attributes are generated based on test cases. These attributes include at least one of language, dialect, audio timbre, context, and intonation. By combining TTS technology with real human voices, voice requests can be generated in multiple languages, contexts, timbres, and emotions, including ambient noise. This simulates voice commands from different users and scenarios, enhancing the comprehensiveness of the quality assessment of in-vehicle voice interaction functions.

[0094] See also Figure 4 In some embodiments, the voice attribute includes a language category. Step 03 (evaluating the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request) includes:

[0095] 0303: Determine a first embedding vector corresponding to the test case according to predetermined mapping data;

[0096] 0304: Encode the speech recognition result to determine a second embedding vector corresponding to the speech recognition result;

[0097] 0305: Evaluate the quality of the vehicle's voice interaction function based on test cases, system operation logs, executed operations, feedback voice, response time, and the difference between the first embedding vector and the second embedding vector.

[0098] In certain embodiments, the evaluation module is further configured to determine a first embedding vector corresponding to a test case based on predetermined mapping data. The evaluation module is further configured to encode the speech recognition result and determine a second embedding vector corresponding to the speech recognition result. The evaluation module is further configured to evaluate the quality of the vehicle's voice interaction function based on the test case, system operation logs, executed operations, feedback speech, response time, and the difference between the first embedding vector and the second embedding vector.

[0099] In certain embodiments, the processor is further configured to determine a first embedding vector corresponding to a test case based on predetermined mapping data. The processor is further configured to encode the speech recognition result and determine a second embedding vector corresponding to the speech recognition result. The processor is further configured to evaluate the quality of the vehicle's voice interaction function based on the test case, system operation logs, executed operations, feedback speech, response time, and the difference between the first embedding vector and the second embedding vector.

[0100] Specifically, mapping data refers to the corresponding system responses mapped by extracting semantic commonalities between commands in different languages ​​through manual annotation and machine learning. For example, the commands "Play jazz" in English, "play jazz" in Chinese, and "Spiele Jazz" in German all have the same semantics and can be mapped to the response of launching a music player and playing jazz.

[0101] The first embedding vector refers to the semantic embedding vector corresponding to the user instruction of the test case and its response.

[0102] The second embedding vector refers to the semantic embedding vector corresponding to the machine response to the user instructions in multiple different languages ​​of the speech recognition result, which can be obtained by encoding the speech recognition result.

[0103] Understandably, semantic embedding vectors are high-dimensional, dense vectors composed of real numbers, with most vectors being zero and each dimension corresponding to implicit features of speech semantics, such as intent, object, and action. For example, in the semantic embedding vector of the command "Navigate to the Forbidden City," one dimension represents the "location search intent" and another represents the "Forbidden City entity." In the same high-dimensional vector space, the cosine similarity or Euclidean distance between vectors can directly reflect the degree of semantic proximity. That is, the distance between semantically equivalent multilingual command vectors is small, while the distance between semantically unrelated command vectors is large.

[0104] By comparing the differences between the first and second embedding vectors, the embodiments of the present application can calculate semantic equivalence based on vector distance. For example, the equivalence between "Play jazz" and "Come to the first jazz music" can achieve cross-language evaluation of semantic equivalence rather than lexical correspondence. Compared with the related art multilingual vehicle voice interaction evaluation that relies on literal translation for matching, the embodiments of the present application have the ability to parse multi-level semantics and cross-language understanding, which can improve the accuracy of vehicle voice interaction quality evaluation. In addition, the quality evaluation of vehicle voice interaction does not need to rely on a large amount of native language manual annotation, which can also reduce the cost of multilingual testing.

[0105] In this way, based on predetermined mapping data, a first embedding vector corresponding to the test case is determined; the speech recognition results are encoded and processed to determine a corresponding second embedding vector; and the quality of the vehicle's voice interaction function is evaluated based on the test case, system operation logs, executed operations, feedback speech, response time, and the difference between the first and second embedding vectors. This method maps voice requests in different languages ​​to the same high-dimensional space, and measures semantic similarity through vector distance, enabling cross-language evaluation based on semantic equivalence rather than lexical correspondence. This improves the accuracy of vehicle voice interaction quality evaluation, and eliminates the need for extensive native speaker manual annotation, reducing the cost of multilingual testing.

[0106] See also Figure 5 In some embodiments, the voice attributes include multiple dialect categories within the same language category. Step 03 (evaluating the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request) includes:

[0107] 0306: Convert the feedback speech to determine the target speech under the target dialect category;

[0108] 0307: Evaluate the quality of the vehicle's voice interaction function based on test cases, system operation logs, executed operations, feedback voice, response time and target voice.

[0109] In certain embodiments, the evaluation module is further configured to convert the feedback speech to determine a target speech within a target dialect category. The evaluation module is further configured to evaluate the quality of the vehicle's voice interaction function based on test cases, system operation logs, executed operations, feedback speech, response time, and target speech.

[0110] In certain embodiments, the processor is further configured to convert the feedback speech to determine a target speech within a target dialect category. The processor is further configured to evaluate the quality of the vehicle's voice interaction function based on test cases, system operation logs, executed operations, feedback speech, response time, and target speech.

[0111] Specifically, the target dialect category refers to the default language category processed by the system, and the target dialect category is set according to the actual situation. For example, Mandarin can be used as the default processing language for system voice interaction.

[0112] The target speech refers to the response speech converted into the target dialect. For example, if the voice request is in Cantonese, the response speech will also be in Cantonese. The unique tonal characteristics of Cantonese can easily lead to errors in speech recognition results, so it is converted to Mandarin to ensure that the evaluation system is based on a unified semantic benchmark. The quality of the vehicle's voice interaction function is then evaluated based on the target speech and other response data.

[0113] The conversion of feedback speech can be implemented based on a pre-trained dialect-to-Mandarin conversion model. This model ensures that the evaluation system is based on a unified semantic benchmark, addressing semantic understanding and evaluation coverage issues in dialect scenarios. For example, the Cantonese command "turn on the air conditioner" is converted by the module into a semantic embedding vector for the Mandarin "turn on the air conditioner." The system then evaluates semantic understanding by comparing this vector with the vehicle's response. If the vehicle correctly activates the air conditioner, the interaction quality is determined to meet the standard. This avoids rigid test cases, the inability to cover complex semantic scenarios, and recognition errors caused by differences in dialect pronunciation, thereby ensuring the comprehensiveness and robustness of the evaluation.

[0114] It should be noted that, when the voice request generated by the test case is the target voice under the target dialect category, the feedback voice does not need to be converted and the above method steps can be skipped.

[0115] In this way, feedback speech is converted to determine the target speech within the target dialect category. The quality of the vehicle's voice interaction function is evaluated based on test cases, system operation logs, executed operations, feedback speech, response time, and the target speech. This conversion of feedback speech into the target speech based on a pre-trained dialect-to-Mandarin conversion model overcomes dialect differences, standardizes dialect semantics, unifies evaluation benchmarks, expands scenario coverage, and improves the comprehensiveness and robustness of the evaluation system.

[0116] See also Figure 6 In some embodiments, step 01 (obtaining a test case set) includes:

[0117] 011: Based on the predetermined vehicle voice interaction function information, generate multiple test cases to construct a test case set.

[0118] In some embodiments, the acquisition module is further used to generate multiple test cases based on predetermined vehicle voice interaction function information to construct a test case set.

[0119] In some embodiments, the processor is further configured to generate a plurality of test cases based on predetermined vehicle voice interaction function information to construct a test case set.

[0120] Specifically, based on pre-determined vehicle voice interaction function information, such as navigation settings, multimedia control, vehicle control, information query, and open-domain voice search, the generated test case set can be ensured to cover the core functional scenarios of vehicle voice interaction, such as common functions such as "play music" and "navigate to a certain location", boundary conditions such as extreme voice command length and dialect recognition, and abnormal scenarios such as vehicle response results and command ambiguity handling in the absence of a network, thereby ensuring the accuracy and comprehensiveness of the evaluation.

[0121] In some embodiments, based on the predetermined vehicle voice interaction function information, test cases can also be selected according to the evaluation requirements to Figure 7 For example, Figure 7 The following are evaluation tasks in different languages. The corresponding test case sets can be automatically generated based on the functions required for evaluation to accurately evaluate the quality of voice interaction functions.

[0122] Among them, each test case in the test case set consists of at least one text, the vehicle expected result corresponding to the text, and the evaluation standard corresponding to the text. For example, the text of the test case is "Turn on the air conditioner", and its vehicle expected result is "The vehicle air conditioner is turned on", and the evaluation standard is "If the vehicle air conditioner is turned on, it is qualified, otherwise it is unqualified."

[0123] Automatically generating test cases with expected standards based on the functional checklist provides an objective benchmark for evaluation, eliminating the influence of human factors in related technologies. This also ensures standardization of the evaluation process, supports the system's automated evaluation process, improves evaluation efficiency, and facilitates systematic data analysis.

[0124] In this way, multiple test cases are generated based on the predetermined vehicle voice interaction function information to form a test case set. This test case set, based on the predetermined vehicle voice interaction function information, can cover the core functional scenarios of vehicle voice interaction, achieving comprehensive evaluation. Furthermore, automatically generating test cases with expected standards based on the function list provides an objective benchmark for evaluation, improves evaluation efficiency, and facilitates systematic data analysis.

[0125] See also Figure 8 In some embodiments, step 03 (evaluating the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request) includes:

[0126] 0308: When a voice request is played to the vehicle, the response result is obtained through the recording devices installed inside and outside the vehicle cabin space;

[0127] 0309: Evaluate the quality of the vehicle's voice interaction based on response results and test cases.

[0128] In certain embodiments, the evaluation module is further configured to, when a voice request is played to the vehicle, capture responses using recording devices located inside and outside the vehicle cabin. The evaluation module is further configured to evaluate the quality of the vehicle's voice interaction based on the responses and test cases.

[0129] In certain embodiments, the processor is further configured to, when a voice request is played to the vehicle, capture responses from recording devices located inside and outside the vehicle cabin. The processor is further configured to evaluate the quality of the vehicle's voice interaction based on the responses and test cases.

[0130] Specifically, the recording equipment includes a recording device and a log collector. The recording device, or mobile phone, is used to record video and audio information of the vehicle performing corresponding operations in response to voice requests, namely, voice recognition results, executed operations, and voice feedback. The log collector is used to collect low-level operating data of the vehicle system, namely system operation logs and response time.

[0131] The following Figure 9 The following uses the five mobile phones shown as an example to explain the layout of the recording equipment:

[0132] During the environmental preparation process, the vehicle is parked in a noise-controlled environment, and the vehicle power supply is ensured to be normal before starting the vehicle system.

[0133] according to Figure 9 Five mobile phones and brackets are installed in the positions shown, among which mobile phone 1 is installed in the middle position of the vehicle, facing the front of the vehicle, and is used to record the response of the display screen and the front of the vehicle; mobile phone 2 is installed in the middle of the vehicle, facing the rear of the vehicle, and is used to record the response of the rear seats, that is, the rear of the vehicle; mobile phone 3 is installed at the right side passenger window of the vehicle, facing the left main driver's window, and is used to record the response of the left main driver's window and seat; mobile phone 4 is installed near the left main driver's window of the vehicle, facing the right side passenger window, and is used to record the response of the right main driver's window and seat; mobile phone 5 is installed in the front of the outside of the vehicle, facing the inside of the vehicle, and is used to record the response of the headlights in front of the vehicle.

[0134] Five mobile phones can be connected to audio devices to simultaneously record video and audio from different angles as the vehicle responds to voice requests and performs corresponding actions. This generates the voice recognition results, executed actions, voice feedback, and response time. In particular, dynamic noise compensation algorithms can be used to correct the viewing angle in real time if the phone moves while the vehicle is in motion. Multi-device data fusion ensures stable capture in dynamic scenes. A log collector also records the vehicle's response results in a system operation log.

[0135] It should be noted that the use of a mobile phone as the recording device in the embodiments of this application is for illustrative purposes only and should not be construed as limiting the recording device to mobile phones. In other examples, a combination of a vehicle's built-in camera and a distributed microphone array can be used to record video and audio information about the vehicle's response to voice requests and the corresponding operation. This allows direct acquisition of vehicle status data via the vehicle's CAN bus, reducing reliance on external devices. This is not a limitation and should be considered based on actual circumstances.

[0136] After the above environment preparations are completed, the normal data flow is confirmed by calibrating the shooting angles of each camera phone, adjusting the volume and position of the audio playback device, testing the network connection stability and performing the system connectivity test. In this way, when playing voice requests to the vehicle, accurate response results can be obtained through the recording devices set inside and outside the vehicle cabin space to evaluate the quality of the vehicle's voice interaction. In particular, in the process of evaluating the quality of the vehicle's voice interaction, such as Figure 10 For example, the real-time monitoring interface can provide real-time preview of video and data streams, display real-time test progress and results, current test pass rate and problem statistics, and allow manual intervention and task suspension.

[0137] In this way, when a voice request is played to the vehicle, the response is captured by recording devices installed inside and outside the vehicle cabin. Based on the response results and test cases, the quality of the vehicle's voice interaction is evaluated. By recording responses from different angles using multiple recording devices, multiple interaction points within the vehicle are covered, enabling collaborative analysis and evaluation of multimodal data. The system also automatically calibrates the recording devices, reducing manual inspection costs.

[0138] See also Figure 11 In some embodiments, step 03 (evaluating the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request) includes:

[0139] 0310: Evaluate the quality of the vehicle's voice interaction based on response results, test cases, and pre-trained generative models.

[0140] In some embodiments, the evaluation module is further used to evaluate the quality of the vehicle's voice interaction based on the response results, test cases, and pre-trained generative models.

[0141] In some embodiments, the processor is further configured to evaluate the quality of the vehicle's voice interaction based on response results, test cases, and a pre-trained generative model.

[0142] Specifically, the generative model is used to synchronously analyze and process video, audio, and log data, namely response results and test cases, to determine whether the vehicle's response meets expectations. It includes the model itself and the speech generation engine, video analysis engine, and semantic understanding engine. It can determine the accuracy of ASR (Automatic Speech Recognition) recognition, evaluate the correctness of operation execution, check the correctness and relevance of reply content, measure response speed, and compare it with a threshold.

[0143] ASR recognition accuracy is assessed by comparing the voice request with the text output of the ASR speech recognition result. For example, the system can determine whether "Navigate to a gas station" is correctly recognized to avoid errors such as "Traffic Police Station."

[0144] The correctness of operation execution is assessed by combining video footage (i.e., the executed operation) with system operation logs to verify whether the system executed the corresponding operation according to the instruction. For example, after the command "Open music player," check whether the central control screen loads the playback interface and whether there is a record of "music module started" in the system operation log.

[0145] The correctness and relevance of the response content are checked by analyzing the semantic consistency between the feedback voice and the voice request in the test case. For example, when the command "check the weather" is given, the response contains correct weather information and not irrelevant content.

[0146] Response speed is quantitatively measured by comparing the time interval between command playback and system response with a threshold to determine whether real-time performance meets the standard. For example, if the response time for a single analysis is less than or equal to 30 seconds, the response time is considered within the normal range.

[0147] By inputting the response results and test cases into the generative model, the quality of the vehicle's voice interaction can be evaluated and an evaluation report as shown in the figure is generated. Figure 12 As shown, the evaluation report can summarize quantitative indicators such as pass rate, error rate, and test coverage, and present the problem types in the form of charts or data. The following bar chart presents the four main types of problems and their quantities:

[0148] 8 ASR recognition items: Voice recognition errors, such as hearing "gas station" as "traffic police team", are the basis for affecting functional interaction; 4 execution operations: Correct recognition but deviation in execution results, such as "find a restaurant" but displaying all POIs, are functional logic loopholes; 5 reply content and 1 response speed.

[0149] Based on the identification of common failure modes and potential root causes, a detailed list of problems is generated. Multimodal data such as video, audio, and system operation logs can be combined to locate defects, provide preliminary analysis and optimization directions, ensure the effectiveness and reliability of updates and maintenance based on quality assessment results, and improve the comprehensiveness and efficiency of assessments.

[0150] In particular, to ensure the accuracy of the quality assessment of vehicle voice interaction functions, the vehicle voice interaction function quality assessment method of the embodiment of this application can also be implemented with the assistance of human-machine collaboration. Voiceprint recognition and permission classification can also be used to provide differentiated operation permissions to testers with different roles to ensure compliance of the testing process.

[0151] The following Figure 13 Take the following as an example to explain the human-machine collaboration at each stage:

[0152] During the test preparation phase, the recording equipment and other environmental preparations are done manually, and the test cases automatically generated by the system based on the voice interaction function are manually checked to ensure the accuracy of the test cases.

[0153] During the test execution phase, manual sampling of a certain percentage of the generative model judgment results is performed regularly, and the evaluation progress is monitored in real time. The generative model is also regularly updated to continuously improve the judgment accuracy.

[0154] During the result analysis phase, key functions and high-risk scenarios maintain 100% manual review, test items with generative model judgment confidence below the threshold are automatically marked for manual review, and the generative model judgment capability is strengthened by establishing an error case library.

[0155] In this way, the quality of a vehicle's voice interaction can be evaluated based on response results, test cases, and a pre-trained generative model. By integrating video, audio, and test case data, the model can determine ASR recognition accuracy, verify correct operation execution, check response relevance, measure response speed, and compare against thresholds. This allows for accurate evaluation of the vehicle's voice interaction function, determining whether the vehicle's response meets expectations. This can also pinpoint hidden defects and generate quantitative metrics such as pass rates, providing data support for system optimization. This enables an upgrade from single-dimensional sampling to full-dimensional verification, addressing the subjective nature and insufficient coverage of manual testing, and improving evaluation efficiency and reliability.

[0156] The following Figure 14 Take the vehicle voice interaction system architecture under test as an example to explain:

[0157] In the data collection layer, five camera phones and an audio capture device capture video and audio of the vehicle's responses to voice requests, capturing the voice recognition results, executed actions, feedback, and response time. A log collection device, or log collector, simultaneously collects system operation logs. The data collection layer synchronizes the collected response results in time and space, and then feeds them into the AI ​​processing layer to evaluate the quality of vehicle voice interaction.

[0158] In the AI ​​processing layer, the multimodal large model, namely the generative model and the speech generation engine, video analysis engine, and semantic understanding engine can analyze and process the data collected by the data acquisition layer, determine whether the response of the vehicle system meets expectations, and evaluate indicators such as execution correctness, speech recognition accuracy, and response speed.

[0159] The evaluation control layer automatically generates evaluation reports based on data from the AI ​​processing layer, including pass rate, error rate, and problem classification statistics. It also provides preliminary analysis and optimization suggestions for discovered issues. It also automatically generates a multilingual test case library and evaluation task management based on the in-vehicle voice system function list to control the quality evaluation process of vehicle voice interaction.

[0160] The system's evaluation process covers automatic generation of multilingual test cases, diversified audio generation, simultaneous recording and response of multiple mobile phones, multimodal large model analysis, and automatic report generation, and is supplemented by human-computer collaboration mechanisms and quality assurance measures to effectively improve evaluation efficiency and accuracy.

[0161] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned method for evaluating the quality of a vehicle voice interaction function.

[0162] It is understood that a computer program includes computer program code. The computer program code may be in source code form, object code form, executable file, or some intermediate form. Computer-readable storage media may include any entity or device capable of carrying computer program code, recording media, USB flash drives, removable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution media.

[0163] In the description of this specification, the descriptions with reference to the terms "particularly", "further", "particularly", "understandably", etc. are intended to mean that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms are not intended to refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0164] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code that includes one or more executable requests for implementing a specific logical function or step of a process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0165] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A method for evaluating the quality of a vehicle's voice interaction function, characterized in that: include: Get the test case set; generating, according to a test case in the test case set, a voice request corresponding to the test case; The quality of the vehicle's voice interaction function is evaluated based on the test case and the vehicle's response to the voice request, wherein the response result includes the system operation log, voice recognition result, execution operation, feedback voice and response time.

2. The method according to claim 1, characterized in that The evaluating the quality of the voice interaction function of the vehicle according to the test case and the response result of the vehicle to the voice request includes: Performing spatiotemporal alignment processing on the obtained system operation log, the speech recognition result, the execution operation, and the feedback speech; The quality of the vehicle's voice interaction function is evaluated based on the test case, the response time, the system operation log after time-space alignment processing, the voice recognition result, the execution operation, and the feedback voice.

3. The method according to claim 1, characterized in that Generating a voice request corresponding to a test case in the test case set includes: According to the test case, a plurality of voice requests having different voice attributes are generated, wherein the voice attributes include at least one of language category, dialect category, audio timbre, utterance scene, and intonation.

4. The method according to claim 3, characterized in that The voice attribute includes a language category, and the evaluating the quality of the vehicle's voice interaction function based on the test case and the vehicle's response to the voice request includes: Determining a first embedding vector corresponding to the test case according to predetermined mapping data; performing encoding processing on the speech recognition result to determine a second embedding vector corresponding to the speech recognition result; The quality of the voice interaction function of the vehicle is evaluated based on the test case, the system operation log, the execution operation, the feedback voice, the response time, and the difference between the first embedding vector and the second embedding vector.

5. The method according to claim 3, characterized in that The voice attributes include multiple dialect categories under the same language category. The quality of the voice interaction function of the vehicle is evaluated based on the test case and the vehicle's response to the voice request, including: Performing conversion processing on the feedback speech to determine a target speech under a target dialect category; The quality of the voice interaction function of the vehicle is evaluated based on the test case, the system operation log, the execution operation, the feedback voice, the response time and the target voice.

6. The method according to claim 1, characterized in that The obtaining of the test case set includes: According to the predetermined vehicle voice interaction function information, a plurality of the test cases are generated to construct the test case set.

7. The method according to claim 1, characterized in that The evaluating the quality of the voice interaction function of the vehicle according to the test case and the response result of the vehicle to the voice request includes: When the voice request is played to the vehicle, the response result is obtained by recording devices arranged inside and outside the vehicle cabin space; The quality of the voice interaction of the vehicle is evaluated based on the response result and the test case.

8. The method according to claim 1, characterized in that The evaluating the quality of the voice interaction function of the vehicle according to the test case and the response result of the vehicle to the voice request includes: The quality of the vehicle's voice interaction is evaluated based on the response result, the test case, and the pre-trained generative model.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Voice system test method and device, storage medium and electronic equipment

    CN114999457A

  • Intelligent voice interaction test system and method

    CN115188367A

  • Test method and device of voice interaction equipment, electronic equipment and storage medium

    CN116504224A

  • Vehicle voice interaction system test method and device, electronic equipment and storage medium

    CN119274541A

  • Vehicle-mounted voice interaction test method, device and equipment, storage medium and vehicle

    CN119889288A

Cited By

  • Intelligent terminal dialect recognition test method and device

    CN121687011A