Vehicle voice interaction evaluation system based on AI data big model

The vehicle-mounted voice interaction evaluation system based on AI data models utilizes a simulated robot head and a 3D track network combined with AI models for multi-dimensional evaluation, solving the problems of single testing and strong subjectivity in existing technologies, and realizing a comprehensive, objective evaluation and optimization of vehicle-mounted voice interaction.

CN121789643APending Publication Date: 2026-04-03NANJING DOROTHY INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing testing and evaluation technologies for in-vehicle voice interaction systems suffer from limitations such as single testing dimensions, insufficient environmental simulation, subjective and one-sided evaluation, and a lack of continuous optimization loop based on massive data and intelligent algorithms. These limitations make it difficult to simulate real human-computer interaction scenarios and conduct multi-dimensional automated and objective evaluations.

Method used

The vehicle-mounted voice interaction evaluation system, based on an AI data big model, uses a simulated robot head and a three-dimensional track network to simulate real user scenarios. It combines the AI ​​big model to perform multi-dimensional automatic evaluation, including semantic accuracy, response latency, and emotion matching scores, and forms a closed-loop optimization link through dynamic parameter iterative optimization.

Benefits of technology

It enables comprehensive, objective, efficient, and automated testing and evaluation of in-vehicle voice interaction performance, improves the realism and coverage of the testing environment, and provides accurate performance optimization data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789643A_ABST
    Figure CN121789643A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI data large model analysis, and particularly discloses an AI data large model-based vehicle voice interaction evaluation system, which comprises a simulation robot head, a pickup and a player are arranged on the simulation robot head, a plurality of moving tracks are arranged in a vehicle, corresponding sliding chutes of the moving tracks are formed in the lower side of the simulation robot head, and the moving tracks are arranged in the sliding chutes. The evaluation system further comprises the evaluation method for the vehicle voice interaction. According to the invention, through the simulation robot head capable of freely moving on the three-dimensional track network in the vehicle, an acoustic scene in which a real user interacts with the vehicle at different positions is simulated. And playing the classified basic corpora through a simulated head to test the in-vehicle infotainment, and performing multi-dimensional automatic evaluation on the voice response of the in-vehicle infotainment by using an AI large model. The dynamic iteration optimization closed loop of test corpus generation and AI model parameters is realized, the multi-modal verification and whole-process monitoring functions are expanded, and the comprehensive, objective and efficient automatic test and evaluation of the vehicle voice interaction performance are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of AI data large model analysis technology, and specifically discloses a vehicle-mounted voice interaction evaluation system based on AI data large model. Background Technology

[0002] As a core function of smart cockpits, the performance of in-vehicle voice interaction systems is directly related to user experience. Currently, the testing and evaluation of in-vehicle voice systems mainly rely on manual testing or simple automated scripts. Manual testing methods typically involve testers issuing voice commands according to a pre-set test case list in a real vehicle or bench environment, and subjectively judging the accuracy, speed, and fluency of the system's response. This method suffers from low efficiency, high cost, limited test coverage, strong subjectivity, and difficulty in quantifying and reproducing results. Existing automated testing tools mostly focus on functional verification, such as triggering tests by recording and playing back audio, but they often cannot simulate the real in-vehicle acoustic environment, differences in user pronunciation, and complex multi-turn dialogue scenarios, and are even less able to objectively evaluate deeper indicators such as semantic accuracy and emotional consistency of the response.

[0003] Regarding hardware testing equipment, existing technologies typically employ a single microphone or simple microphone array in a fixed location for recording, or use static "dummy" testing equipment. These devices cannot simulate the dynamic scenarios of real users interacting with the vehicle's infotainment system from different positions within the vehicle (e.g., driver's seat, passenger seat, rear seat), leading to discrepancies between the picked-up audio signals and what the human ear hears. Furthermore, they cannot accurately assess the sound field performance and speech clarity of the vehicle's speakers at different locations. In addition, current testing systems lack deep integration with high-performance AI models. Test corpora are often static, limited in size, and have singular evaluation criteria, primarily focused on recognition accuracy or simple response time, failing to meet the increasingly complex evaluation requirements of natural language understanding, contextual understanding, personalized services, and emotional interaction.

[0004] Existing testing and evaluation technologies for in-vehicle voice interaction systems suffer from significant drawbacks, including limited testing dimensions, insufficient environmental simulation, subjective and biased evaluation, and a lack of continuous optimization loops based on massive data and intelligent algorithms. Therefore, there is an urgent need for a comprehensive evaluation system and methodology that can highly simulate real human-computer interaction scenarios, achieve multi-dimensional automated and objective evaluation, and utilize large AI models for in-depth analysis and self-iterative optimization. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a vehicle-mounted voice interaction evaluation system based on a large AI data model to solve the above-mentioned technical problems.

[0006] To achieve the above objectives, the present invention provides the following technical solution: The vehicle-mounted voice interaction evaluation system based on an AI data big data model includes a simulated robot head equipped with a microphone and a player. Multiple moving tracks are located inside the vehicle, and corresponding grooves for these tracks are located on the underside of the simulated robot head. The evaluation system also includes an evaluation method for vehicle-mounted voice interaction, with the following steps: S1, pre-value input: A portion of basic corpus is compiled and input into the evaluation system; S2, the evaluation system is connected to the vehicle-mounted voice system; S3, the basic corpus is classified; S4, the vehicle-mounted voice system is tested; S5, the vehicle-mounted system and the evaluation system are connected to the AI ​​big data model; S6, the response results of the vehicle-mounted voice interaction are evaluated using the AI ​​big data model. It also includes step S7, dynamic parameter iterative optimization: adjusting the weight parameters of the AI ​​big model in reverse according to the evaluation results, and generating new test data to inject into the vehicle voice system to form a closed-loop optimization link.

[0007] Preferably, the evaluation criteria for the AI ​​big model in step S6 include: semantic accuracy score: calculating the BERT semantic similarity between the vehicle system response content and the expected answer, with a threshold of ≥0.85 for qualification; response delay score: classifying the delay as excellent (≤1s), good (1-2s), or poor (≥2s) based on the interval between the end of the voice command and the first frame of audio response from the vehicle system; and emotion matching score: verifying whether the tone of the response matches the emotional needs of the command scenario by comparing acoustic features (pitch, speech rate, pauses) with the emotion classification model.

[0008] Preferably, the corpus is classified into nine categories: natural common sense, contextual understanding, scenario-based, personalized, knowledge base information, intent service, emotional state, legal test, and emergency call, forming nine different partitions.

[0009] Preferably, the microphone of the simulated robot head includes a ring array microphone group installed in the ear area of ​​the head, and the player is a multi-band speaker embedded in the mouth cavity of the head, and the sound wave emission direction of the player is at an angle of 30°-45° to the receiving direction of the microphone.

[0010] Preferably, the moving track includes a transverse track on the roof, a longitudinal track on the floor, and a vertical track on the seat back, and a universal steering joint is provided at the intersection of each track to enable the simulated robot head to move freely in the three-dimensional space inside the vehicle.

[0011] Preferably, the slide is equipped with an electric drive device, which uses a pressure sensor to detect the positional offset of the simulated robot head in real time and dynamically adjusts the moving speed using a PID algorithm; the end of the slide is also equipped with an electromagnetic locking mechanism for fixing the test position of the simulated robot head.

[0012] Preferably, the AI ​​big model is a pre-trained speech-semantic joint model, which includes: a speech recognition module: extracting response audio features of the vehicle's voice system based on the Conformer architecture; a semantic understanding module: parsing the contextual relevance of voice commands through a multi-head attention mechanism; and a scoring decision module: outputting a comprehensive evaluation result including semantic accuracy score, response latency score, and sentiment matching score.

[0013] Preferably, the eyes of the simulated robot's head are embedded with infrared cameras to collect the content displayed on the vehicle's infotainment screen and perform multimodal consistency verification with the voice response results.

[0014] Preferably, the system further includes a process monitoring module, which is used to record and associate stored data in real time during the testing process.

[0015] The working principle and beneficial effects of this solution are as follows: This invention utilizes a simulated robotic head that can move freely on a three-dimensional track network within the vehicle to simulate the acoustic scenarios of real users interacting with the vehicle's infotainment system from different locations. The system plays categorized basic language data through the simulated head to test the infotainment system and uses a large AI model to automatically evaluate the system's voice response from multiple dimensions, including semantic accuracy, response latency, and emotional matching. Furthermore, the solution achieves a closed loop of dynamic iterative optimization between test data generation and AI model parameters, and expands multimodal verification and full-process monitoring functions, thereby enabling comprehensive, objective, and efficient automated testing and evaluation of the infotainment system's voice interaction performance. Attached Figure Description

[0016] Figure 1 This is a side view of the in-vehicle sliding rail installation position of the vehicle-mounted voice interaction evaluation system based on AI data large model according to the present invention; Figure 2 This is a top view of the in-vehicle sliding rail installation position of the vehicle-mounted voice interaction evaluation system based on AI data large model according to the present invention; Figure 3 This is a schematic diagram of the universal steering joint in an embodiment of the vehicle-mounted voice interaction evaluation system based on AI data large model of the present invention; Figure 4 This is a schematic diagram of the slide rail and slider in an embodiment of the vehicle-mounted voice interaction evaluation system based on AI data large model of the present invention; Figure 5 This is a schematic diagram showing the positional relationship between the slide rail, slide plate, and electric drive device in an embodiment of the vehicle-mounted voice interaction evaluation system based on AI data large model of the present invention. Detailed Implementation

[0017] In the description of this invention, it should be understood that the terms "front", "rear", "left", "right", "up", "down", "vertical", "horizontal", "high", "low", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting the scope of protection of this invention.

[0018] The following detailed description illustrates the specific implementation method: Example 1 The in-vehicle voice interaction evaluation system based on AI data large model described in this embodiment consists of three main parts in its core architecture: a hardware simulation platform, a central processing and control unit, and a cloud-based AI large model service.

[0019] The core of the hardware simulation platform is the simulated robot head. This head is made of lightweight, high-strength composite materials, its shape and size mimicking an adult human head, with an internal hollow cavity structure. In the "ear" region of the head, a precisely installed ring array microphone group consists of six high-fidelity MEMS microphones arranged in a ring at equal intervals. This allows for 360-degree, unobstructed audio acquisition from the vehicle's speakers in different directions, and beamforming technology can simulate the directional pickup characteristics of the human ear. In the "mouth" cavity of the head, a multi-band speaker is embedded, its sound outlet acoustically designed to simulate the sound source location and spectral characteristics of a real human voice. Specifically, a fixed 35° angle is designed between the speaker's main emission axis and the ring array microphone's main receiving axis. This design effectively reduces the risk of self-excitation feedback during testing and more realistically simulates the relative positional relationship between the mouth speaking and the ear hearing.

[0020] To achieve full cabin coverage for testing, a three-dimensional network of moving tracks was pre-installed inside the cabin. This includes: transverse tracks running along the front-to-back direction of the roof; longitudinal tracks laid in the central aisle and on both sides of the vehicle floor; and vertical tracks embedded in the backrests of the driver and passenger seats. These tracks are connected at key spatial nodes via omnidirectional joints, allowing the simulated robot head to smoothly change direction. A corresponding slide integrated under the head houses a silent electric drive unit. This drive unit moves the head along the tracks according to commands from the central control unit. Pressure sensors integrated within the slides detect minute positional shifts or vibrations of the head in real time and dynamically adjust the motor speed using a PID control algorithm, ensuring smooth movement and precise positioning. When the head reaches the preset test position, an electromagnetic locking mechanism at the end of the slide is energized to generate a strong magnetic force, firmly attaching the head to the track to eliminate mechanical vibration interference during the test.

[0021] The central processing and control unit (CPU) is typically an industrial control computer located inside the vehicle, responsible for scheduling the entire testing process. It first executes steps S1 and S3: retrieving test cases from a built-in basic corpus. This corpus has been pre-classified into nine logical partitions, such as: the "Natural Knowledge" partition containing "What's the weather like today?"; the "Contextual Understanding" partition containing the continuous instructions "Turn on the air conditioner—it's too cold—raise it two degrees"; and the "Emergency Call" partition containing "Call emergency services for me" simulating an accident. The unit then converts the classified corpus into an audio stream using software.

[0022] Next, steps S2 and S4 are executed: the central unit establishes a connection with the vehicle's voice system under test via wired or wireless means. At the start of the test, it first controls the simulated robot's head to move to a designated coordinate, and then plays the test voice command through a speaker inside the head. The vehicle's voice system picks up and processes the command, providing a voice response. Simultaneously, the circular array microphones on the head begin recording the vehicle's response audio, and the central unit precisely records the time interval from the end of the command playback to the detection of the first frame of the response audio, i.e., the response delay.

[0023] Subsequently, steps S5 and S6 are executed: the central unit packages the recorded response audio, recorded delay time, original test command text, and expected answer text, and uploads them to the cloud-based AI large model service via a high-speed network. This large model is a pre-trained speech-semantic joint model. Its speech recognition module, based on the Conformer architecture, first converts the response audio into text. The semantic understanding module uses a multi-head attention mechanism to deeply analyze the semantic relationship between the vehicle's response text and the expected answer text. The scoring decision module comprehensively calculates the results: for example, it uses the BERT model to calculate the semantic vector similarity between the two to obtain a semantic accuracy score; it gives a response delay score of "good" based on a delay time of 1.5 seconds; and it analyzes the acoustic features of the response audio through a sentiment classification model to determine whether its tone matches the command scenario (e.g., the tone should be apologetic rather than cheerful when reporting a navigation error), and gives a sentiment matching score. Finally, a report containing scores for each item and a comprehensive evaluation is generated and sent back to the central unit.

[0024] The advantages of this embodiment are as follows: This embodiment constructs a highly biomimetic, fully automated, and comprehensive vehicle-mounted voice interaction testing platform. Firstly, this is reflected in the physical simulation realism and spatial coverage integrity of the testing environment. Through the precise acoustic design of a human-like head and the synergy of a three-dimensional track network, the system can accurately simulate the sound field environment of real users in different seats and postures within the vehicle, solving the pain point that traditional fixed testing equipment cannot reflect spatial position differences. Secondly, the level of automation and intelligence in the testing process has been significantly improved. From corpus classification, instruction issuance, location movement and data collection, the entire process is automatically scheduled by the central unit, eliminating the subjectivity and inefficiency of manual testing. Finally, the objectivity and depth of the evaluation system have been fundamentally transformed. With the help of cloud-based AI models, the system has surpassed the traditional superficial evaluation that only focuses on recognition rate and latency, and has achieved quantitative scoring of complex cognitive aspects such as semantic accuracy and emotional matching. This provides unprecedentedly accurate data support and insights for the performance optimization of the vehicle voice system.

[0025] Example 2 This embodiment further expands upon the first embodiment by incorporating dynamic acoustic environment simulation and anti-interference testing. The hardware simulation platform integrates an in-vehicle environmental noise simulation system. This system includes auxiliary speakers placed throughout the vehicle, capable of dynamically injecting different types and intensities of background noise during testing, based on the needs of each test case. For example, when testing the "navigate to airport" command, a pre-set mixture of high-speed wind noise, tire road noise, and air conditioning vent noise can be played simultaneously, with the sound pressure level adjustable within the range of 60-75 dB(A). The circular array microphones on the simulated robot's head simultaneously collect these interfering noises when picking up the vehicle's response. Before sending the audio to the AI ​​large model, the central processing unit can choose to enable or bypass the noise reduction preprocessing module. The AI ​​large model's scoring decision module can therefore output an additional anti-interference robustness score to assess whether the vehicle's voice system meets the recognition and response performance standards in complex noise environments. Simultaneously, the infrared camera on the robot's head plays a crucial role in this scenario, confirming whether the vehicle's screen can still correctly display the navigation route under strong noise interference, thus achieving audiovisual consistency verification.

[0026] Furthermore, this system can be configured with multiple simulated robot heads, each assigned a different role (e.g., "driver - owner," "passenger - family member," "rear passenger - passenger"). These heads can move independently on their respective tracks under the coordinated control of the central unit. The system supports writing complex multi-turn dialogue scripts. For example, the driver says, "Play Jay Chou's song," the vehicle system responds, the passenger then says, "Turn it up," and then the rear passenger asks, "What's the name of this song?" When evaluating the system, the AI ​​model not only assesses the accuracy of individual responses but also analyzes, through its semantic understanding module's multi-head attention mechanism, whether the vehicle system accurately understands the roles corresponding to different sound source locations and whether it correctly maintains the context of the entire dialogue (e.g., knowing that "this song" refers to the previously played Jay Chou song), thus providing a specific score for the vehicle system's dialogue management capabilities.

[0027] The advantages of this embodiment are: the simulation of environmental noise enhances the realism and comprehensiveness of the test; the multi-role interaction test uncovers the deep capabilities of the vehicle-machine dialogue system; and it greatly enhances the depth and value of the scheme protected by the claims.

[0028] The above-mentioned method for using the vehicle-mounted voice interaction evaluation system based on AI data models includes the following steps: S1, Predicted Input: Organize a portion of the basic corpus and input the basic predictions into the evaluation system; S2 connects the assessment system to the vehicle's voice system; S3 categorizes the basic corpus; S4, testing the vehicle's voice system; S5 connects the vehicle infotainment system and the evaluation system to the AI ​​big model; The S6 uses a large AI model to evaluate the response of the vehicle's voice commands.

[0029] The above descriptions are merely embodiments of the present invention, and common knowledge regarding specific structures and characteristics in the solutions is not described in detail here. It should be noted that those skilled in the art can make various modifications and improvements without departing from the structure of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications and improvements will not affect the effectiveness of the implementation of the present invention or its practicality.

Claims

1. A vehicle-mounted voice interaction evaluation system based on an AI data large model, comprising a simulated robot head, wherein the simulated robot head is equipped with a microphone and a player, the vehicle interior is provided with multiple moving tracks, and the lower side of the simulated robot head is provided with corresponding grooves for the moving tracks, characterized in that, The evaluation system also includes an evaluation method for in-vehicle voice interaction, the steps of which are as follows: S1, Predicted Input: Organize a portion of the basic corpus and input the basic predictions into the evaluation system; S2 connects the assessment system to the vehicle's voice system; S3 categorizes the basic corpus; S4, testing the vehicle's voice system; S5 connects the vehicle infotainment system and the evaluation system to the large AI model; The S6 uses a large AI model to evaluate the response of the vehicle's voice commands. The vehicle-mounted voice interaction evaluation system based on AI data big model as described in claim 1 is characterized by further including step S7, dynamic parameter iterative optimization: adjusting the weight parameters of the AI ​​big model in reverse according to the evaluation results, and generating new test data to inject into the vehicle-mounted voice system to form a closed-loop optimization link.

2. The vehicle-mounted voice interaction evaluation system based on AI data large model according to claim 1, characterized in that: The evaluation criteria for the large AI model in step S6 include: Semantic accuracy score: Calculate the BERT semantic similarity between the vehicle system response content and the expected answer, and a threshold of ≥0.85 is considered acceptable; Response delay rating: Based on the interval between the end of the voice command and the first frame of audio response from the vehicle system, it is graded as excellent (≤1s), good (1-2s), and poor (≥2s). Emotional matching score: By comparing acoustic features (pitch, speech rate, pauses) with the emotional classification model, it is verified whether the tone of response matches the emotional needs of the instruction scenario.

3. The vehicle-mounted voice interaction evaluation system based on AI data large model according to claim 1, characterized in that: The corpus is categorized into nine types: natural common sense, contextual understanding, scenario-based, personalized, knowledge base information, intent service, emotional state, legal testing, and emergency call. These nine corpus categories form nine different partitions.

4. The vehicle-mounted voice interaction evaluation system based on AI data large model according to claim 1, characterized in that: The microphone of the simulated robot head includes a ring array microphone array, which is installed in the ear area of ​​the head.

5. The player is a multi-band speaker, embedded in the oral cavity of the head, and the sound wave emission direction of the player forms an angle of 30°-45° with the receiving direction of the microphone.

6. The vehicle-mounted voice interaction evaluation system based on AI data large model according to claim 1, characterized in that: The moving track includes a transverse track on the roof, a longitudinal track on the floor, and a vertical track on the seat back. A universal joint is provided at the intersection of the tracks to enable the simulated robot head to move freely in the three-dimensional space inside the vehicle.

7. The vehicle-mounted voice interaction evaluation system based on AI data large model according to claim 1, characterized in that: The slide is equipped with an electric drive device, which uses a pressure sensor to detect the positional offset of the simulated robot head in real time and dynamically adjusts the movement speed using a PID algorithm; the end of the slide is also equipped with an electromagnetic locking mechanism to fix the test position of the simulated robot head.

8. The vehicle-mounted voice interaction evaluation system based on AI data large model according to claim 1, characterized in that: The AI ​​large model is a pre-trained speech-semantic joint model, which includes: Speech recognition module: Extracts response audio features of the vehicle's voice system based on the Conformer architecture; Semantic understanding module: parses the contextual relevance of voice commands through a multi-head attention mechanism; Scoring Decision Module: Outputs a comprehensive evaluation result including semantic accuracy score, response delay score, and sentiment matching score.

9. The vehicle-mounted voice interaction evaluation system based on AI data large model according to claim 1, characterized in that: The simulated robot's head has infrared cameras embedded in its eyes, which are used to collect the content displayed on the vehicle's infotainment screen and perform multimodal consistency verification with the voice response results.

10. The vehicle-mounted voice interaction evaluation system based on AI data large model according to claim 1, characterized in that: The system also includes a process monitoring module, which is used to record and associate stored data in real time during the testing process.