A vehicle voice function test method and system

CN122799892APending Publication Date: 2026-09-22ZHEJIANG LINGAI FUTURE TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610984204.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]本申请实施例提供一种车辆语音功能测试方法及系统,旨在解决相关技术中智能车辆的语音功能的测试效率较低的问题

Benefits of technology

[0015]本申请通过语料文本生成音频信号,在目标测试音区输出测试音频,并根据车辆的实际反馈结果确定车辆语音功能的验证结果,缩短回归测试周期,提升测试效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799892A_ABST
    Figure CN122799892A_ABST
Patent Text Reader

Abstract

The application discloses a vehicle voice function test method and system, and belongs to the technical field of vehicle electronic testing. The vehicle voice function test method comprises the following steps: obtaining corpus text and corresponding expected feedback results; converting the corpus text into an audio signal; outputting test audio according to the audio signal in a target test audio area; obtaining actual feedback results obtained after the vehicle receives the test audio; and obtaining verification results of vehicle voice functions according to the actual feedback results and the expected feedback results. The audio signal is generated through the corpus text, the test audio is output in the target test audio area, and the verification results of the vehicle voice functions are determined according to the actual feedback results of the vehicle, so that the regression test cycle is shortened, and the test efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle electronic testing technology, specifically to a method and system for testing vehicle voice functions. Background Technology

[0002] The voice function of intelligent vehicles can use a Large Language Model (LLM) as the core technology engine for semantic understanding and content generation. Users can control functions such as air conditioning temperature adjustment, seat massage start / stop, navigation route planning, and window opening control through natural language. Before the vehicle leaves the factory, the voice function needs to be tested to determine whether it can function properly.

[0003] In related technologies, voice commands are read aloud manually and the vehicle is verified to execute the voice commands correctly. However, this process results in a long regression testing cycle and low testing efficiency. Summary of the Invention

[0004] This application provides a method and system for testing vehicle voice functions, aiming to solve the problem of low testing efficiency of voice functions of intelligent vehicles in related technologies.

[0005] In a first aspect, embodiments of this application provide a method for testing vehicle voice function, the method comprising the following steps: Obtain the corpus text and the corresponding expected feedback results; Convert the text corpus into audio signals; The test audio is output in the target test sound range based on the audio signal; Obtain the actual feedback results after the vehicle receives the test audio; The verification results of the vehicle's voice function are obtained based on the actual feedback results and the expected feedback results.

[0006] In some embodiments, obtaining the corpus text and the corresponding expected feedback result includes: Obtain the voice control requirements form; The corpus text and corresponding expected feedback results are extracted from the voice control requirements table.

[0007] In some embodiments, the voice control requirements table includes a function description, example templates of statements, expected actions to be performed, and expected voice text to be played. The corpus text and corresponding expected feedback results are extracted from the voice control requirements table, including: Based on a large language model, the corpus text is obtained from functional descriptions and example templates of statements, and the expected feedback result is obtained from the expected execution action and the expected broadcast speech text.

[0008] In some embodiments, a simulated mouth is provided inside the vehicle; The test audio is output in the target test range based on the audio signal, including: The audio signal is sent to the simulated mouthpiece, which controls the simulated mouthpiece to output test audio to the target test sound zone.

[0009] In some embodiments, the actual feedback results include the actual voice-broadcast text, the actual action performed, and the log message; the expected feedback results include the expected voice-broadcast text, the expected action performed, and the expected background execution logic. The verification results of the vehicle's voice function are obtained based on the actual feedback results and the expected feedback results, including: The voice broadcast verification result is obtained based on the actual voice broadcast text and the expected voice broadcast text; The execution action verification results are obtained based on the actual and expected execution actions. The background execution verification results are obtained based on the log messages and the expected background execution logic. The verification results of the vehicle's voice function are obtained based on the voice broadcast verification results, the action execution verification results, and the background execution verification results.

[0010] In some embodiments, obtaining a voice broadcast verification result based on the actual voice broadcast text and the expected voice broadcast text includes: Based on a large language model, the actual speech broadcast text is converted into actual semantic data, and the expected speech broadcast text is converted into expected semantic data. The actual semantic data is compared with the expected semantic data. If the deviation between the actual semantic data and the expected semantic data is less than the preset semantic deviation threshold, the voice broadcast verification result is determined to be verified as passed.

[0011] In some embodiments, the actual action performed includes obtaining the actual display interface after the jump by recognizing the image on the central control display screen, wherein the image on the central control display screen is acquired by an image acquisition device; the expected action performed includes the expected display interface. The execution action verification results are obtained based on the actual and expected execution actions, including: The actual displayed interface is compared with the expected displayed interface. If the actual displayed interface is the same as the expected displayed interface, the verification result of the action is determined to be successful.

[0012] In some embodiments, the background execution verification result is obtained based on the log message and the expected background execution logic, including: Obtain the actual backend execution logic from the log messages; The actual backend execution logic is compared with the expected backend execution logic. If the actual backend execution logic is the same as the expected backend execution logic, the backend execution verification result is determined to be successful.

[0013] In some embodiments, the verification result of the vehicle's voice function is obtained based on the voice broadcast verification result, the action execution verification result, and the background execution verification result, including: If the verification results for voice broadcast, execution, and background execution are all passed, the verification result for the vehicle's voice function is determined to be passed.

[0014] Secondly, embodiments of this application provide a vehicle voice function testing system, which includes: The corpus generation module is used to obtain the corpus text and the corresponding expected feedback results; The text-to-speech module is used to convert text corpus into audio signals; An acoustic playback device used to output test audio in the target test sound zone based on an audio signal; The actual feedback result acquisition module is used to acquire the actual feedback result obtained after the vehicle receives the test audio; The test discrimination module is used to obtain the verification results of the vehicle's voice function based on the actual feedback results and the expected feedback results.

[0015] This application generates audio signals from corpus text, outputs test audio in the target test audio region, and determines the verification results of the vehicle's voice function based on the actual feedback from the vehicle, thereby shortening the regression test cycle and improving test efficiency. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating a vehicle voice function testing method provided by an exemplary embodiment of this disclosure; Figure 2 This is a flowchart illustrating step S105 of a vehicle voice function testing method provided in an exemplary embodiment of this disclosure; Figure 3 This is a flowchart of a vehicle voice function testing method provided by an exemplary embodiment of this disclosure; Figure 4This is a schematic diagram of the structure of a vehicle voice function testing system provided in an exemplary embodiment of this disclosure; Figure 5 This is another structural schematic diagram of a vehicle voice function testing system provided in an exemplary embodiment of this disclosure; Figure 6 This is a schematic diagram of the intelligent cockpit layout structure of a vehicle providing a vehicle voice function testing system according to an exemplary embodiment of this disclosure.

[0018] Explanation of icon numbers: 100. Vehicle voice function testing system; 101. Corpus generation module; 102. Text-to-speech module; 103. Acoustic playback device; 104. Actual feedback result acquisition module; 105. Test discrimination module; 106. Test management module; 107. Speech-to-text module; 108. Image recognition module; 109. Log acquisition module; 110. Communication bus message monitoring module; 111. Microphone; 112. Bus interface card; 113. Network board; 114. Camera. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.

[0021] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.

[0022] The use of "applies to" or "configured to" in this application implies open and inclusive language, which does not exclude the applicability to or configuration to devices performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values ​​may in practice be based on additional conditions or values ​​beyond those stated.

[0023] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0024] Firstly, embodiments of this application provide a method for testing vehicle voice functions, such as... Figure 1 As shown, the vehicle voice function testing method includes the following steps: S101. Obtain the corpus text and the corresponding expected feedback results.

[0025] Specifically, the corpus text consists of test statements simulating user voice input, such as voice commands like "set the air conditioning to 24 degrees" or "turn on the seat heating." The expected feedback is the standard response produced by the vehicle's voice function after executing the voice commands in the corpus text, such as a voice announcement saying "The air conditioning is set to 24 degrees," the central control display switching to the air conditioning temperature setting interface, and the background log recording the air conditioning application's startup event. The corpus text and expected feedback provide the foundation for subsequent verification steps.

[0026] S102. Convert the text corpus into an audio signal.

[0027] Specifically, the audio signal is the electrical signal that supplies audio output to the acoustic playback device. Text corpus can be converted into audio signals using a speech synthesis model. For example, the Qwen3-TTS speech synthesis model can convert text corpus into a highly realistic audio signal, generating speech with emotional and prosodic variations based on the input text. This better simulates the tone and manner of a real user issuing commands, facilitating the output of test audio from the audio output device that closely resembles the real timbre or intonation.

[0028] Speech synthesis models can be deployed on internal servers to prevent the leakage of corpus text data and ensure data security.

[0029] The same text corpus can generate completely consistent audio signals, eliminating the differences in volume and speaking speed caused by human reading, ensuring the accuracy of test results, and making the test results reproducible.

[0030] S103. Output test audio in the target test sound zone according to the audio signal.

[0031] Specifically, the target test sound zone is a specific location simulating a user's voice, such as the driver's seat, passenger seat, left rear seat, right rear seat, etc. Audio signals are sent to the acoustic playback device in the target test sound zone, where test audio is output so that the vehicle receives the test audio and provides corresponding actual feedback.

[0032] By outputting test audio in the target test zone to simulate the real sound source location, it is possible to accurately verify whether the voice system can perform differentiated responses based on the sound source location. For example, the test audio output from the driver's position is "open the window" and only the driver's side window is lowered.

[0033] S104. Obtain the actual feedback results after the vehicle receives the test audio.

[0034] Specifically, the actual feedback result is the vehicle's actual response to the voice control commands in the test audio after receiving it. The actual feedback result can take various forms, such as actual voice broadcasts, interface transitions on the central control display, or background execution logs or messages. By collecting actual feedback results from multiple dimensions, a data foundation is provided for the verification results of subsequent steps, improving the accuracy of the verification.

[0035] S105. Based on the actual feedback results and the expected feedback results, obtain the verification results of the vehicle's voice function.

[0036] Specifically, the verification result is the final conclusion obtained by comparing the actual feedback result with the expected feedback result. The verification result includes verification passed and verification failed. The verification result is evaluated through multiple different dimensions, and it is only judged as passed when all dimensions meet the expectations, thereby improving the accuracy of the verification result.

[0037] This application embodiment generates audio signals from corpus text, outputs test audio in the target test audio region, and determines the verification result of the vehicle's voice function based on the actual feedback from the vehicle, thereby shortening the regression test cycle and improving test efficiency.

[0038] In some embodiments, obtaining the corpus text and the corresponding expected feedback results includes: obtaining a voice control demand table; and extracting the corpus text and the corresponding expected feedback results from the voice control demand table.

[0039] Specifically, the voice control requirements table is a structured data table used to uniformly manage test benchmark information related to all voice functions of the vehicle. Each row in the voice control requirements table corresponds to a specific voice control test case.

[0040] The voice control requirements table includes functional hierarchy, functional attributes, functional descriptions, example phrase templates, expected actions, and expected voice prompts. It describes the function, example phrases, and corresponding feedback for each voice control test case. The table contains over 100 voice control test cases, covering multiple categories such as page control, vehicle control, navigation settings, and media playback.

[0041] The voice control requirements table allows for the extraction of corpus text and corresponding expected feedback results, serving as a benchmark for testing. The voice control requirements table can be maintained according to testing needs, allowing for the addition or removal of voice control test cases at any time to adapt to voice testing of different vehicles.

[0042] In some embodiments, extracting the corpus text and the corresponding expected feedback result from the voice control demand table includes: obtaining the corpus text based on the large language model, according to the function description and the example template of the statement, and obtaining the expected feedback result based on the expected execution action and the expected broadcast voice text.

[0043] Specifically, the function description explains the voice control function and describes the target semantic range of the text corpus to be generated. The example template is a simplified, scalable rectangular pattern or regular expression used to guide the large language model in generating diverse voice control commands. The expected action is the standard action that the vehicle should perform when correctly recognizing the voice control command, such as physical action execution or interface transition. The expected voice broadcast is the voice text that the vehicle will respond to after performing the corresponding action, including responses for normal execution and abnormal events.

[0044] In some examples, the functional descriptions, example templates, expected actions, and expected voice text in the voice control requirements table are shown in the following table: Table 1. Examples of Voice Control Requirements

[0045] The large language model can be Qwen3.5-Max, which automatically generates corpus text that conforms to spoken language habits based on the example templates in the voice control requirements table. The large language model can be deployed on an internal server, and communication with the large language model can be achieved through an embedded large language model call interface, avoiding leakage of corpus text data and ensuring data security.

[0046] This embodiment provides an example of a prompt word format: "You are a smart car user. Please generate as many questions as possible that match the function description [{function description}] and example statements [{example statement template}]. Exhaustively list all possible combinations, with results separated by semicolons. Only return the results." The returned text is parsed to form a database of corpus text, and the expected feedback result is obtained based on the expected action and the expected broadcast speech text.

[0047] By combining a large language model with functional descriptions and example templates, dozens to hundreds of high-quality corpora can be generated in seconds. The generated results are constrained by prompt words, which ensures diversity while avoiding deviation from functional boundaries or the generation of dangerous instructions.

[0048] In some embodiments, a simulated mouth is installed inside the vehicle. This simulated mouth is a high-precision artificial sound source device used to simulate the speech produced by a real person inside the vehicle. The simulated mouth has a built-in drive circuit, eliminating the need for an external amplifier and allowing direct output of test audio. The simulated mouth can also accurately simulate the sound field radiation characteristics of a human mouth, ensuring that the emitted sound is close to that of a real person.

[0049] The simulated voice can be placed in the driver's seat headrest and the passenger seat headrest to simulate the control voices of the driver and passenger. It can also be placed in the rear seat headrest to simulate the control voices of the rear passengers.

[0050] The simulated mouthpieces are connected via a multi-channel audio output interface, with two channels connected to the driver's simulated mouthpiece and the passenger's simulated mouthpiece, respectively. When testing the target test range for the driver's seat, the audio stream is routed to the driver's simulated mouthpiece channel via software control; when testing the target test range for the passenger's seat or testing the rear seat sound range using sound field radiation characteristics, the audio is switched to the corresponding channel for playback, thus traversing the four-range differentiated response test cases.

[0051] The test audio is output in the target test sound area based on the audio signal, including: sending the audio signal to the simulated mouth and controlling the simulated mouth to output the test audio to the target test sound area.

[0052] Specifically, the audio signal is sent to the simulated mouth through a communication line, and the simulated mouth outputs the test audio to the target test sound area. The simulated mouth can simulate the voice of a real person and simulate the voices from multiple different positions such as the driver, passenger, and back seat, enabling differentiated testing of different target test sound areas, closely resembling the actual use environment, thereby improving test accuracy.

[0053] In some embodiments, the actual feedback results include the actual voice broadcast text, the actual executed actions, and log messages. The actual voice broadcast text is the text of the actual voice broadcast content responded by the vehicle after receiving the voice control command from the test audio. The actual executed actions are the operations actually performed by the vehicle after receiving the voice control command from the test audio. The log messages are the operation records generated after the vehicle performs the operations.

[0054] The expected feedback results include the expected voice broadcast text, the expected action, and the expected background execution logic. The expected voice broadcast text is the standard voice broadcast content that the vehicle should respond to after receiving the voice control command from the test audio. The expected action is the standard operation that the vehicle should perform after receiving the voice control command from the test audio. The expected background execution logic is the running logic that the vehicle's background processes should execute after receiving the voice control command from the test audio.

[0055] The verification results of the vehicle's voice function are obtained based on the actual feedback results and the expected feedback results, such as... Figure 2 As shown, it includes S1051-S1054.

[0056] S1051. Obtain the voice broadcast verification result based on the actual voice broadcast text and the expected voice broadcast text.

[0057] Specifically, the voice broadcast verification result is determined based on the similarity between the actual voice broadcast text and the expected voice broadcast text. If the similarity exceeds the threshold, the voice broadcast verification result is considered successful; if the similarity is below the threshold, the voice broadcast verification result is considered unsuccessful.

[0058] S1052. Obtain the execution action verification results based on the actual and expected execution actions.

[0059] Specifically, the verification result is determined based on whether the actual action and the expected action are consistent. If they are consistent, the verification result is "verification passed"; if they are inconsistent, the verification result is "verification failed".

[0060] S1053. Obtain the background execution verification result based on the log message and the expected background execution logic.

[0061] Specifically, the system determines whether the actual commands executed in the background and the data transmitted in the log messages are consistent with the expected background execution logic. If they are consistent, the background execution verification result is "verification passed"; if they are inconsistent, the background execution verification result is "verification failed".

[0062] S1054. Based on the voice broadcast verification result, the action execution verification result, and the background execution verification result, obtain the verification result of the vehicle voice function.

[0063] Specifically, the vehicle voice function verification result is successful if all three verification results—voice broadcast verification result, action execution verification result, and background execution verification result—pass. The vehicle voice function verification result is unsuccessful if any one of these verifications fails.

[0064] In some embodiments, obtaining a voice broadcast verification result based on the actual voice broadcast text and the expected voice broadcast text includes: converting the actual voice broadcast text into actual semantic data based on a large language model, and converting the expected voice broadcast text into expected semantic data; comparing the actual semantic data and the expected semantic data, and determining that the voice broadcast verification result is verified as passed if the deviation between the actual semantic data and the expected semantic data is less than a preset semantic deviation threshold.

[0065] Specifically, the actual voice broadcast text is obtained by converting the vehicle's voice response into text format using a speech recognition model. The speech recognition model can be the Canary Qwen 2.5B model deployed on an internal server, optimized for speech recognition in multilingual and high-noise environments, accurately capturing the vehicle's broadcast content even in background noise.

[0066] Actual semantic data consists of structured or vectorized representations extracted from the actual spoken text, used to quantify the semantics of the spoken text. Expected semantic data consists of structured or vectorized representations extracted from the expected spoken text, used to quantify the semantics of the expected spoken text. Actual and expected semantic data can be represented in the form of semantic vectors, and the similarity between these semantic vectors is calculated as the deviation between the actual and expected semantic data.

[0067] The large language model can be the Qwen3.5-Max large language model, deployed on an internal server to prevent the leakage of corpus text data and ensure data security.

[0068] If the deviation between the actual semantic data and the expected semantic data is less than the preset semantic deviation threshold, the voice broadcast verification result is determined to be verification passed; if the deviation between the actual semantic data and the expected semantic data is greater than or equal to the preset semantic deviation threshold, the voice broadcast verification result is determined to be verification failed.

[0069] By using a large language model to convert actual speech text and expected speech text into actual semantic data and expected semantic data, it can accurately handle synonyms or near-synonyms, eliminate the limitations of literal matching, and adjust the strictness of the deviation through the semantic deviation threshold, flexibly adapting to different test scenarios.

[0070] In some embodiments, the actual action performed includes obtaining the actual display interface after the jump by recognizing the image on the central control display screen, wherein the image on the central control display screen is acquired by an image acquisition device; the expected action performed includes the expected display interface.

[0071] The image acquisition device can be a high-definition camera, fixedly installed in front of the central control display screen to capture images of the screen. By recognizing the images on the central control display screen, interface category labels can be obtained for the interfaces displayed in the image, such as the main interface, air conditioning control interface, navigation interface, seat position interface, etc. These interface category labels are pre-defined.

[0072] An image recognition model can identify the interface category label of the screen displayed on the central control screen. This model can be a ResNet64-based classification model. During training, screenshots of all voice-controlled function interfaces are collected and labeled with their corresponding category labels. The model can then classify an image within one second and output the interface category label of the currently displayed interface.

[0073] The actual display interface is the interface displayed on the central control screen after the vehicle receives the voice control command from the test audio. The expected display interface is the standard display interface that the central control screen should display after the vehicle receives the voice control command from the test audio. For ease of comparison, both the actual and expected display interfaces are represented by interface category labels.

[0074] The execution action verification result is obtained based on the actual execution action and the expected execution action, including: comparing the actual display interface with the expected display interface, and determining that the execution action verification result is verified as passed if the actual display interface is the same as the expected display interface.

[0075] Specifically, it determines whether the interface category label of the actual displayed interface is the same as the interface category label of the expected displayed interface. If they are the same, the verification result of the action is determined to be successful; otherwise, the verification result of the action is determined to be unsuccessful.

[0076] In some embodiments, obtaining the background execution verification result based on the log message and the expected background execution logic includes: obtaining the actual background execution logic based on the log message; comparing the actual background execution logic with the expected background execution logic; and determining that the background execution verification result is verified as passed if the actual background execution logic is the same as the expected background execution logic.

[0077] Specifically, log messages include application logs and bus messages. Application logs can be obtained through application background data, for example, by establishing a Secure Shell (SSH) connection with the vehicle's infotainment system via a network board, executing predefined control commands after key login, and capturing the current application package name and interface information in real time to verify whether the corresponding function page has been successfully invoked. Bus messages can be obtained by collecting communication data from the communication bus, for example, by monitoring the vehicle's control network segment through a Controller Area Network (CAN) bus interface card. When voice commands involve physical controls such as headlights and driving modes, the system confirms that the execution layer has completed closed-loop control by comparing the communication bus identification code and changes in data field signal values.

[0078] The actual background execution logic can be obtained from the log messages. This actual background execution logic is a description of the real behavior of the vehicle application in the background, as parsed from the log messages. The expected background execution logic is the standard behavior that the vehicle application should perform.

[0079] Determine whether the actual backend execution logic is the same as the expected backend execution logic. If they are the same, the backend execution verification result is determined to be verification passed; otherwise, the backend execution verification result is determined to be verification failed.

[0080] In some embodiments, the verification result of the vehicle voice function is obtained based on the voice broadcast verification result, the action execution verification result, and the background execution verification result, including: if the voice broadcast verification result, the action execution verification result, and the background execution verification result are all verified as passed, the verification result of the vehicle voice function is determined to be verified as passed.

[0081] Specifically, the vehicle voice function verification result is successful if all three verification results—voice broadcast verification result, action execution verification result, and background execution verification result—pass. The vehicle voice function verification result is unsuccessful if any one of these verifications fails.

[0082] By employing multiple verification methods for cross-validation, a complete closed-loop verification system encompassing auditory, visual, and underlying execution is achieved. This reduces instances where tests appear to pass but are not actually executed, thereby improving the accuracy of the tests.

[0083] The following explanation uses a specific example.

[0084] like Figure 3 As shown, a voice control requirement table is obtained, and the corpus text is recognized through a large language model. The corpus text is converted into an audio signal through a speech synthesis model. The audio signal is sent to the simulated mouths at the driver's and passenger's seats, and the simulated mouths output test audio in the target test sound range.

[0085] The vehicle's smart cockpit is equipped with a microphone, bus interface card, network board, and camera. The microphone records the vehicle's voice responses, which are then converted into actual spoken text using a voice recognition model. The bus interface card reads real-time bus messages. The network board establishes a secure shell protocol connection with the vehicle's infotainment system, executes predefined control commands, and captures the current application package name and interface information in real time to obtain the vehicle's status. The camera captures images of the central control display screen, and an image recognition model identifies the interface category labels displayed on the screen.

[0086] Verification is performed based on the actual voice broadcast text, interface category labels, application logs, and bus messages to obtain voice broadcast verification results, execution action verification results, and background execution verification results. A comprehensive judgment is then made to obtain the final verification result.

[0087] Secondly, embodiments of this application provide a vehicle voice function testing system 100, such as... Figure 4 As shown, the vehicle voice function testing system 100 includes a corpus generation module 101, a text-to-speech module 102, an acoustic playback device 103, an actual feedback result acquisition module 104, and a test discrimination module 105.

[0088] The corpus generation module 101 is used to obtain the corpus text and the corresponding expected feedback results.

[0089] Specifically, the corpus text consists of test statements simulating user voice input, such as voice commands like "set the air conditioning to 24 degrees" or "turn on the seat heating." The expected feedback is the standard response produced by the vehicle's voice function after executing the voice commands in the corpus text, such as a voice announcement saying "The air conditioning is set to 24 degrees," the central control display switching to the air conditioning temperature setting interface, and the background log recording the air conditioning application's startup event. The corpus text and expected feedback provide the foundation for subsequent verification steps.

[0090] The text-to-speech module 102 is used to convert text corpus into audio signals.

[0091] Specifically, the audio signal is the electrical signal that supplies audio output to the acoustic playback device 103. The corpus text can be converted into an audio signal using a speech synthesis model. For example, the Qwen3-TTS speech synthesis model can convert the corpus text into a highly realistic audio signal, generating speech with emotional and prosodic variations based on the input text. This better simulates the tone and manner of a real user issuing commands, facilitating the output of test audio from the audio output device that closely resembles the real timbre or intonation.

[0092] Speech synthesis models can be deployed on internal servers to prevent the leakage of corpus text data and ensure data security.

[0093] The same text corpus can generate completely consistent audio signals, eliminating the differences in volume and speaking speed caused by human reading, ensuring the accuracy of test results, and making the test results reproducible.

[0094] The acoustic playback device 103 is used to output test audio in the target test sound zone according to the audio signal.

[0095] Specifically, the target test sound zone is a specific location simulating a user's voice, such as the driver's seat, passenger seat, left rear seat, right rear seat, etc. By sending audio signals to the audio output device in the target test sound zone, test audio is output in the target test sound zone so that the vehicle can receive the test audio and provide corresponding actual feedback results.

[0096] By outputting test audio in the target test zone to simulate the real sound source location, it is possible to accurately verify whether the voice system can perform differentiated responses based on the sound source location. For example, the test audio output from the driver's position is "open the window" and only the driver's side window is lowered.

[0097] The actual feedback result acquisition module 104 is used to acquire the actual feedback result obtained after the vehicle receives the test audio.

[0098] Specifically, the actual feedback result is the vehicle's actual response to the voice control commands executed in the test audio. The actual feedback result can take various forms, such as actual voice broadcasts, navigation on the central control screen, or background execution logs or messages. By collecting actual feedback results from multiple dimensions, a data foundation is provided for the verification results of subsequent steps, improving the accuracy of the verification.

[0099] The test discrimination module 105 is used to obtain the verification results of the vehicle's voice function based on the actual feedback results and the expected feedback results.

[0100] Specifically, the verification result is the final conclusion obtained by comparing the actual feedback result with the expected feedback result. The verification result includes verification passed and verification failed. The verification result is evaluated through multiple different dimensions, and it is only judged as passed when all dimensions meet the expectations, thereby improving the accuracy of the verification result.

[0101] This application embodiment generates audio signals from corpus text, outputs test audio in the target test audio region, and determines the verification result of the vehicle's voice function based on the actual feedback from the vehicle, thereby shortening the regression test cycle and improving test efficiency.

[0102] In some embodiments, such as Figure 5 As shown, the vehicle voice function testing system 100 also includes a test management module 106, which is used to maintain and manage the voice control requirements table.

[0103] Specifically, the voice control requirements table is a structured data table used to uniformly manage test benchmark information related to all voice functions of the vehicle. Each row in the voice control requirements table corresponds to a specific voice control test case.

[0104] The voice control requirements table includes functional hierarchy, functional attributes, functional description, example templates for speech, expected execution actions, and expected voice text, which are used to explain the function, example speech, and corresponding feedback of each voice control test case.

[0105] The voice control requirements table allows for the extraction of corpus text and corresponding expected feedback results, serving as a benchmark for testing. The voice control requirements table can be maintained according to testing needs, allowing for the addition or removal of voice control test cases at any time to adapt to voice testing of different vehicles.

[0106] In some embodiments, the vehicle voice function testing system 100 further includes a speech-to-text module 107, which converts the vehicle's response speech into text form through a speech recognition model to obtain the actual voice broadcast text.

[0107] Specifically, the speech recognition model can be the Canary Qwen 2.5B model deployed on an internal server, optimized for speech recognition in multilingual and high-noise environments, and can accurately obtain the broadcast content responded by the vehicle in the background noise.

[0108] In some embodiments, the vehicle voice function testing system 100 further includes an image recognition module 108 for recognizing the interface category label of the interface displayed in the image on the central control display screen.

[0109] Specifically, the image recognition model can be a ResNet64-based classification model. During the training phase, screenshots of all voice-controlled functional interfaces are collected and labeled with corresponding interface category tags to train the image recognition model. This model can classify an image within one second and output the interface category tag for the currently displayed interface.

[0110] In some embodiments, the vehicle voice function testing system 100 further includes a log acquisition module 109, used to acquire application logs through application background data.

[0111] Specifically, a Secure Shell (SSH) connection can be established between the network board and the vehicle's infotainment system. After logging in with a key, predefined control commands can be executed, and the current application package name and interface information can be captured in real time to verify whether the corresponding function page has been successfully invoked.

[0112] In some embodiments, the vehicle voice function testing system 100 further includes a communication bus message monitoring module 110, which is used to collect communication data of the communication bus and obtain bus messages.

[0113] Specifically, the system can monitor the vehicle control network segment via the Controller Area Network (CAN) bus interface card. When voice commands involve physical controls such as headlights and driving modes, the system confirms the completion of closed-loop control at the execution layer by comparing the communication bus identification code and changes in data field signal values.

[0114] In some embodiments, such as Figure 6 As shown, the vehicle's smart cockpit is equipped with a microphone 111, a bus interface card 112, a network board 113, and a camera 114. The microphone 111 records the vehicle's voice responses, which are then converted into actual voice-over text using a voice recognition model. The bus interface card 112 reads real-time bus messages. The network board 113 establishes a secure shell protocol connection with the vehicle's infotainment system, executes predefined control commands, and captures the current application package name and interface information in real time to obtain the vehicle's status. The camera 114 captures images of the central control display screen, and an image recognition model identifies the interface category labels displayed on the screen.

[0115] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0116] The above provides a detailed description of a vehicle voice function testing method and system provided by the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for testing vehicle voice function, characterized in that, The vehicle voice function testing method includes the following steps: Obtain the corpus text and the corresponding expected feedback results; Convert the corpus text into an audio signal; The test audio is output in the target test sound region according to the audio signal; Obtain the actual feedback results obtained by the vehicle after receiving the test audio; The verification results of the vehicle voice function are obtained based on the actual feedback results and the expected feedback results.

2. The vehicle voice function testing method according to claim 1, characterized in that, Obtain the corpus text and the corresponding expected feedback results, including: Obtain the voice control requirements form; The corpus text and the corresponding expected feedback results are extracted from the voice control requirements table.

3. The vehicle voice function testing method according to claim 2, characterized in that, The voice control requirements table includes a function description, example templates for speech, expected actions to be performed, and expected voice text to be played. The corpus text and the corresponding expected feedback results are extracted from the voice control requirements table, including: Based on the large language model, the corpus text is obtained according to the functional description and the example template of the statement, and the expected feedback result is obtained according to the expected execution action and the expected broadcast speech text.

4. The vehicle voice function testing method according to claim 1, characterized in that, The vehicle is equipped with a simulated mouth. Output test audio in the target test range based on the audio signal, including: The audio signal is sent to the simulated mouthpiece, and the simulated mouthpiece is controlled to output test audio to the target test sound zone.

5. The vehicle voice function testing method according to claim 1, characterized in that, The actual feedback results include the actual voice broadcast text, the actual executed actions, and log messages; the expected feedback results include the expected voice broadcast text, the expected executed actions, and the expected background execution logic. The verification results of the vehicle voice function are obtained based on the actual feedback results and the expected feedback results, including: The voice broadcast verification result is obtained based on the actual voice broadcast text and the expected voice broadcast text; The execution action verification result is obtained based on the actual execution action and the expected execution action; The background execution verification result is obtained based on the log message and the expected background execution logic. The verification result of the vehicle voice function is obtained based on the voice broadcast verification result, the execution action verification result, and the background execution verification result.

6. The vehicle voice function testing method according to claim 5, characterized in that, The voice broadcast verification result is obtained based on the actual voice broadcast text and the expected voice broadcast text, including: Based on the large language model, the actual speech broadcast text is converted into actual semantic data, and the expected speech broadcast text is converted into expected semantic data. The actual semantic data and the expected semantic data are compared. If the deviation between the actual semantic data and the expected semantic data is less than a preset semantic deviation threshold, the voice broadcast verification result is determined to be verified as passed.

7. The vehicle voice function testing method according to claim 5, characterized in that, The actual execution action includes obtaining the actual display interface after the jump by recognizing the image on the central control display screen, wherein the image on the central control display screen is acquired by an image acquisition device; The expected action to be performed includes the expected display interface; The execution action verification result is obtained based on the actual executed action and the expected executed action, including: The actual display interface is compared with the expected display interface. If the actual display interface is the same as the expected display interface, the verification result of the execution action is determined to be successful.

8. The vehicle voice function testing method according to claim 5, characterized in that, The background execution verification result is obtained based on the log message and the expected background execution logic, including: The actual background execution logic is obtained based on the log messages. The actual background execution logic is compared with the expected background execution logic. If the actual background execution logic is the same as the expected background execution logic, the background execution verification result is determined to be verified as passed.

9. The vehicle voice function testing method according to claim 5, characterized in that, The verification result of the vehicle voice function is obtained based on the voice broadcast verification result, the execution action verification result, and the background execution verification result, including: If the voice broadcast verification result, the action execution verification result, and the background execution verification result are all verified as passed, the verification result of the vehicle voice function is determined to be passed.

10. A vehicle voice function testing system, characterized in that, The vehicle voice function testing system includes: The corpus generation module is used to obtain the corpus text and the corresponding expected feedback results; A text-to-speech module is used to convert the corpus text into audio signals; An acoustic playback device for outputting test audio in the target test sound zone according to the audio signal; The actual feedback result acquisition module is used to acquire the actual feedback result obtained by the vehicle after receiving the test audio; The test discrimination module is used to obtain the verification result of the vehicle voice function based on the actual feedback result and the expected feedback result.