A method and device for automatic voice recognition in an intelligent cockpit
By playing test voices in the smart cockpit and using a large language model algorithm to evaluate the voice interaction system, identifying feedback values and comparing them with standard values, the problems of low recognition efficiency and insufficient accuracy in existing technologies are solved, and efficient voice recognition and system optimization are achieved.
Patent Information
- Application Number
- CN202510724553.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The existing intelligent cockpit voice interaction system has low recognition efficiency and insufficient accuracy in real environments. Affected by factors such as vehicle noise and reverberation, the recognition task is highly repetitive and has low recognition efficiency.
By playing test voices, the response information of the smart cockpit is obtained, and the performance of the voice interaction system is evaluated using a large language model algorithm. The feedback values are identified and compared with the standard values, and evaluation feedback results are generated to guide system upgrades and optimizations, including adjusting the response generation logic and improving the user interaction interface.
It improves the accuracy and efficiency of smart cockpit voice recognition, achieves accurate recognition of voice interaction performance, and improves the system's response quality and user satisfaction.
Smart Images

Figure CN120236581B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of intelligent cockpit voice interaction and voice recognition, and in particular to an intelligent cockpit voice automatic recognition method and device. Background Art
[0002] With the continuous advancement of automotive technology, the Internet of Vehicles (IoV) and artificial intelligence (AI) technologies are gaining momentum, enabling an increasing number of features to be integrated into vehicles. The intelligent cockpit (ICC) aims to integrate multiple IT and AI technologies to create a new, integrated in-vehicle digital platform, providing drivers with an intelligent experience and promoting driving safety. Currently, voice interaction, a hallmark of the intelligent cockpit, is integrated with various in-vehicle applications and has become a core function in the cockpit ecosystem. The entire voice interaction process encompasses voice reception, voice recognition, and semantic understanding. Errors in any of these steps can lead to overall interaction failure. Furthermore, acoustic factors such as driving noise, wind noise, and interior reverberation can also affect voice recognition performance. Overcoming these constraints in real-world environments is crucial to the quality of the in-vehicle voice interaction system experience.
[0003] To determine the quality of a car's voice interaction function, it's necessary to test and analyze the various performance indicators of the in-vehicle voice interaction system to generate recognition results. However, existing human-computer interaction recognition systems often involve numerous tasks, are repetitive, and have low recognition efficiency and accuracy. Therefore, providing a more accurate and efficient automated voice recognition method for intelligent cockpits, thereby improving both the accuracy and efficiency of intelligent cockpit voice recognition, has become a pressing technical challenge in this field. Summary of the Invention
[0004] The purpose of this application is to provide a method and device for automatic speech recognition in an intelligent cockpit, which can improve the accuracy and efficiency of intelligent cockpit speech recognition.
[0005] To achieve the above objectives, this application provides the following solutions.
[0006] In a first aspect, the present application provides a method for automatic speech recognition in an intelligent cockpit, which specifically includes the following steps.
[0007] Play the test audio.
[0008] Acquire response information of the smart cockpit based on the test voice; wherein the response information includes a vehicle screen image of the smart cockpit.
[0009] The response information is identified to obtain a feedback value of the smart cockpit; wherein the feedback value represents an identification index of the interactive performance of the smart cockpit in responding to the test voice.
[0010] The feedback value is compared with a standard value corresponding to the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit.
[0011] Optionally, the intelligent cockpit automatic voice recognition method further includes the following steps.
[0012] An evaluation method based on a large language model algorithm is adopted to evaluate the response performance of the voice interaction system of the smart cockpit and generate evaluation feedback results.
[0013] The evaluation feedback results are fed back to the development or maintenance process of the voice interaction system of the smart cockpit to guide the upgrade and optimization of the voice interaction system of the smart cockpit; the upgrade and optimization include adjusting the response generation logic, improving the user interaction interface prompts and / or fine-tuning the large language model that supports response generation.
[0014] Optionally, an evaluation method based on a large language model algorithm is used to evaluate the response performance of the voice interaction system of the smart cockpit and generate evaluation feedback results, which specifically includes the following steps.
[0015] Receive instructions from the user to the voice interaction system of the smart cockpit .
[0016] Receive the voice interaction system of the smart cockpit for the instruction Generated Response .
[0017] Process the instructions using a pre-configured large language model evaluation chain and the response , and obtain the evaluation feedback results.
[0018] Optionally, identifying the response information and obtaining a feedback value of the smart cockpit specifically includes the following steps.
[0019] Calculate the matching degree between the vehicle screen image and the current database in the model database; wherein the model database includes a database corresponding to the response interface of the voice interaction system of the smart cockpit and the original text of the test voice; the matching degree label of the vehicle screen image and the current database in the model database The smaller the size, the higher the matching degree between the vehicle screen image and the current database in the model database. The calculation formula is as follows.
[0020] .
[0021] .
[0022] in, Y a and Y b Respectively represent the vehicle screen image data and the data in the current database, and Respectively f The vehicle screen image data of the feature dimension and the data in the current database, F is the characteristic dimension of the current database data, X for Y a and Y b The distance between It is a preset threshold for measuring the matching degree.
[0023] Optionally, the large language model evaluation chain includes an LLM model and a prompt template, wherein the prompt template is used to guide the evaluation, and the LLM model is used to evaluate the response according to the evaluation dimensions defined in the prompt template. Performing an evaluation and generating the evaluation feedback result.
[0024] The evaluation dimensions include semantic correctness, state change confirmation, and unambiguous expression.
[0025] The evaluation feedback results include the scores and opinions of each evaluation dimension, the comprehensive score and validity flags , wherein the comprehensive score Used to indicate the comprehensive score of all the evaluation dimensions, the validity mark Used to indicate whether this interactive assessment test has passed;
[0026] The comprehensive score The calculation formula is as follows.
[0027] .
[0028] .
[0029] .
[0030] in, Represents the evaluation results output by the large language model evaluation chain, including the scores and opinions of each evaluation dimension. Indicates the comprehensive score, Represents the process of performing evaluation through the large language model evaluation chain, Indicates the instructions given by the user to the voice interaction system of the smart cockpit. R Indicates that the voice interaction system of the intelligent cockpit is for instructions The generated response, Prompt templates and criteria for assessment, 、 、 They represent the score of semantic correctness, the score of state change confirmation, and the score of unambiguous expression respectively; A label indicating the degree of match between the vehicle screen image and the current database in the model database; Indicates combining image matching labels The comprehensive rating of , Represents the weight parameter for semantic evaluation and image matching.
[0031] The validity flag The calculation formula is as follows.
[0032] .
[0033] in, Indicates the preset threshold for the comprehensive score.
[0034] Alternatively, the validity flag The calculation formula is as follows.
[0035] .
[0036] in, 、 、 Represents the combined image matching label The preset threshold for the comprehensive score, the preset threshold for semantic correctness, the preset threshold for state change confirmation, and the preset threshold for unambiguous expression.
[0037] If the comprehensive score and validity flags If there is a matching deviation with the actual feedback result on the vehicle screen or the matching deviation is greater than the threshold, the large language model parameters are modified through the loss function, where the calculation formula of the loss function is as follows.
[0038] .
[0039] in, is the loss function value, To match the sample size, Determine the category type for the large language model, Give the real comprehensive score to the large language model. It is a symbol of the true effectiveness of the large language model.
[0040] If the loss function value is less than a preset value, it is determined that the large language model parameters are correct.
[0041] If the loss function value is greater than or equal to the preset value, the large language model parameters are modified.
[0042] Optionally, the vehicle screen image includes a response interface of the smart cockpit based on the test voice.
[0043] Identifying the response information and obtaining the feedback value of the smart cockpit specifically includes the following steps.
[0044] Identify the vehicle screen image and obtain the awakening state of the voice interaction system of the smart cockpit; wherein the awakening state includes successful awakening and failed awakening.
[0045] Comparing the feedback value with a standard value corresponding to the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit specifically includes the following steps.
[0046] The wake-up status of the voice interaction system of the smart cockpit is counted to obtain the wake-up rate of the voice interaction system of the smart cockpit.
[0047] The wake-up rate of the voice interaction system of the smart cockpit is compared with a preset wake-up rate threshold to obtain an identification result of the voice interaction performance of the smart cockpit.
[0048] Optionally, the vehicle screen image includes text information obtained by the smart cockpit recognizing the test voice.
[0049] Identifying the response information and obtaining the feedback value of the smart cockpit specifically includes the following steps.
[0050] Identify the vehicle screen image and obtain text information obtained by the smart cockpit recognizing the test voice.
[0051] Comparing the feedback value with a standard value corresponding to the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit specifically includes the following steps.
[0052] The text information obtained by the smart cockpit recognizing the test voice is compared with the original text of the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit.
[0053] Optionally, obtaining response information of the smart cockpit based on the test voice specifically includes the following steps.
[0054] Obtain interaction information of the smart cockpit based on the test voice; wherein the interaction information includes voice information and / or text information displayed on the vehicle screen of the smart cockpit.
[0055] Identifying the response information and obtaining the feedback value of the smart cockpit specifically includes the following steps.
[0056] The interaction information is identified to obtain feedback information of the smart cockpit.
[0057] Comparing the feedback value with a standard value corresponding to the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit specifically includes the following steps.
[0058] The feedback information is compared with standard interaction information corresponding to the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit.
[0059] Optionally, obtaining response information of the smart cockpit based on the test voice specifically includes the following steps.
[0060] Obtain a response time of the smart cockpit based on the test voice.
[0061] Identifying the response information and obtaining the feedback value of the smart cockpit specifically includes the following steps.
[0062] Calculate the response duration of the smart cockpit based on the test voice; wherein the response duration represents the duration from the test voice being played to the smart cockpit responding.
[0063] Comparing the feedback value with a standard value corresponding to the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit specifically includes the following steps.
[0064] The response time is compared with a preset time threshold to obtain a recognition result of the voice interaction performance of the smart cockpit.
[0065] In a second aspect, the present application provides an intelligent cockpit automatic speech recognition device, which includes the following modules.
[0066] The test voice playback module is used to play the test voice.
[0067] A response information acquisition module is used to obtain response information of the smart cockpit based on the test voice; wherein the response information includes the vehicle screen image of the smart cockpit.
[0068] A feedback result recognition module is used to recognize the response information and obtain a feedback value of the smart cockpit; wherein the feedback value represents an identification index of the interactive performance of the smart cockpit in response to the test voice.
[0069] The recognition result calculation module is used to compare the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the smart cockpit.
[0070] According to the specific embodiments provided in this application, this application has the following technical effects.
[0071] The present application provides a method and device for automatic recognition of voice in an intelligent cockpit, which comprises the following steps: playing a test voice; obtaining response information of the intelligent cockpit based on the test voice; identifying the response information, and obtaining a feedback value of the intelligent cockpit; wherein the response information includes an image of the vehicle screen of the intelligent cockpit, and the feedback value represents an identification index of the interactive performance of the intelligent cockpit in response to the test voice; comparing the feedback value with a standard value corresponding to the test voice, and obtaining a recognition result of the voice interaction performance of the intelligent cockpit; that is, playing a test voice and obtaining response information of the intelligent cockpit based on the test voice, identifying the response information and obtaining a feedback value of the intelligent cockpit, and comparing the feedback value with a standard value corresponding to the test voice, and obtaining a recognition result of the voice interaction performance of the intelligent cockpit, thereby achieving accurate recognition of the voice interaction performance of the intelligent cockpit, and utilizing automatic testing of the test voice to improve the accuracy and efficiency of voice recognition in the intelligent cockpit. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0073] Figure 1 A flowchart of an automated voice recognition method for an intelligent cockpit is provided in accordance with one embodiment of the present application.
[0074] Figure 2 A schematic diagram of the structure of an intelligent cockpit automatic voice recognition device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0075] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0076] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0077] like Figure 1 As shown, this embodiment proposes an intelligent cockpit voice automation recognition method, which specifically includes the following steps.
[0078] Step S1: Play test speech.
[0079] Specifically, this application operates the voice verification tool software on the host computer, creates a task (including the recognition of multiple corpora, for example, 100), reads the recognized test voice and plays it using an artificial mouth, for example: "Hello, turn on the air conditioner."
[0080] An artificial mouth is a special artificial sound source, also known as a simulated mouth or artificial mouth. It consists of a small speaker mounted on a specially shaped baffle. The baffle is designed to mimic the average directivity and radiation pattern of a human mouth. The artificial mouth also requires a frequency compensation network or acoustic compression to achieve a specific frequency response. The artificial mouth is used to simulate the electroacoustic characteristics of microphones used in actual working conditions for testing or calibrating transmitters, thereby creating a sound source that closely resembles a real human mouth.
[0081] Step S2: Obtain response information of the smart cockpit based on the test voice.
[0082] The response information includes an image of the smart cockpit's on-board screen. After playing each test voice using an artificial mouth, the application uses a high-definition camera to capture the smart cockpit's voice interaction system's response based on the test voice. For example, the test voice, "Hello, turn on the air conditioner," normally displays the corresponding text and an interface for turning on the air conditioner on the smart cockpit's on-board screen after receiving the test voice. The application uses a high-definition camera to capture the on-board screen image to determine whether the smart cockpit's voice interaction system has responded and whether the response is correct.
[0083] Step S3: Identify the response information and obtain the feedback value of the smart cockpit.
[0084] The feedback value represents an indicator of the smart cockpit's interactive performance in responding to the test voice. After obtaining the smart cockpit's response information, this application identifies the response information to obtain the smart cockpit's feedback value. For example, by identifying the vehicle screen image to obtain text information on the vehicle screen or interface information for turning on the air conditioner, the application determines whether the smart cockpit's voice interaction system has received the test voice and responded correctly.
[0085] Step S4: Compare the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the smart cockpit.
[0086] After obtaining feedback from the smart cockpit, this embodiment compares this feedback value with the standard value corresponding to the test speech to obtain a recognition result of the smart cockpit's voice interaction performance. For example, after obtaining text information or interface information on the smart cockpit's vehicle screen, the standard text and standard interface information for turning on the air conditioner are compared to determine whether the smart cockpit correctly turns on the air conditioner, that is, whether the smart cockpit's voice interaction system successfully interacts. This allows the interaction success rate of the voice interaction system (one of the recognition indicators) to be obtained by counting multiple test speech sounds. Similarly, this application can obtain multiple recognition indicators for the smart cockpit's voice interaction system and, based on these multiple recognition indicators, comprehensively obtain the recognition result of the smart cockpit's voice interaction system.
[0087] This embodiment plays a test voice; obtains response information of the smart cockpit based on the test voice; identifies the response information to obtain a feedback value of the smart cockpit; wherein the response information includes an image of the vehicle screen of the smart cockpit, and the feedback value represents an identification index of the interactive performance of the smart cockpit in response to the test voice; compares the feedback value with the standard value corresponding to the test voice to obtain an identification result of the voice interactive performance of the smart cockpit; that is, by playing the test voice and obtaining the response information of the smart cockpit based on the test voice, identifying the response information to obtain the feedback value of the smart cockpit, and by comparing the feedback value with the standard value corresponding to the test voice, an identification result of the voice interactive performance of the smart cockpit is obtained, thereby achieving accurate identification of the voice interactive performance of the smart cockpit, and utilizing automatic testing of the test voice to improve test efficiency, thereby quickly identifying the voice interactive performance and improving the accuracy and efficiency of automatic voice recognition in the smart cockpit.
[0088] In an exemplary embodiment, the vehicle screen image includes a response interface of the smart cockpit based on the test voice; wherein, the specific implementation method of the above-mentioned step S3 may be: identifying the vehicle screen image to obtain the wake-up state of the voice interaction system of the smart cockpit; wherein, the wake-up state includes successful wake-up and failure to wake up; correspondingly, the specific implementation method of step S4 may be: counting the wake-up state of the voice interaction system of the smart cockpit to obtain the wake-up rate of the voice interaction system of the smart cockpit; comparing the wake-up rate of the voice interaction system of the smart cockpit with the preset wake-up rate threshold to obtain the recognition result of the voice interaction performance of the smart cockpit.
[0089] Specifically, this application identifies interface information in the vehicle screen image to obtain the smart cockpit's response interface based on a test voice, and determines whether the smart cockpit's voice interaction system has been awakened based on this response interface. Furthermore, by counting the awakening status of the smart cockpit's voice interaction system based on multiple test voices, the application calculates the smart cockpit's voice interaction system's wakeup rate, which is the ratio of the number of test voices that successfully awakened the voice interaction system to the total number of test voices played. After obtaining the smart cockpit's voice interaction system's wakeup rate, the application compares the wakeup rate with a preset wakeup rate threshold to determine whether the smart cockpit's voice interaction system's wakeup rate is qualified. For example, if the wakeup rate is lower than 95%, the wakeup rate performance indicator is considered low.
[0090] In an exemplary embodiment, the vehicle screen image includes text information obtained by the smart cockpit recognizing the test voice; wherein, the specific implementation method of the above step S3 may be: recognizing the vehicle screen image, and obtaining the text information obtained by the smart cockpit recognizing the test voice; correspondingly, the specific implementation method of step S4 may be: comparing the text information obtained by the smart cockpit recognizing the test voice with the original text of the test voice, and obtaining the recognition result of the voice interaction performance of the smart cockpit.
[0091] Specifically, this application identifies text information in the vehicle screen image to obtain text instructions obtained by the smart cockpit recognizing test speech, and compares the text instructions with the original text of the test speech to obtain the voice recognition accuracy of the smart cockpit voice interaction system. The application can also calculate the voice recognition accuracy based on the voice recognition accuracy of multiple test speech lines, that is, the ratio of the number of correctly recognized test speech lines to the total number of test speech lines played. In addition, a correctness threshold can be set. After calculating the recognition accuracy of the smart cockpit voice recognition system, the recognition accuracy can be compared with the correctness threshold to determine whether the voice recognition accuracy of the smart cockpit voice interaction system is qualified.
[0092] In an exemplary embodiment, the specific implementation method of the above-mentioned step S3 can be: calculating the matching degree between the vehicle screen image and the current database in the model database; wherein the model database includes a database of the response interface of the voice interaction system corresponding to the smart cockpit and the original text of the test voice; based on the matching degree, searching for the feedback value of the smart cockpit in a database in the model database.
[0093] Since the car screen image of the smart cockpit may contain text information or interface information, after obtaining the car screen image, this application can determine the search database of the feedback value based on the matching degree between the car screen image and the corresponding original text database and the response interface database to improve the search efficiency.
[0094] Specifically, after obtaining the vehicle screen image, this application calculates the matching degree between the vehicle screen image and the current database in the model database (database of original text or database of response interface, etc.), and the specific method is as follows.
[0095] .
[0096] .
[0097] in, Y a and Y b Respectively represent the vehicle screen image data and the data in the current database, Respectively f The vehicle screen image data of the feature dimension and the data in the current database, F is the characteristic dimension of the current database data, for and The distance between is a pre-set threshold for measuring matching degree. is the label of the matching degree between the vehicle screen image and the current database in the model database, that is, Y a and Y b The label of the matching degree.
[0098] When 0≤ <1, it is determined that the vehicle screen image data matches the current database. The smaller the value, the higher the matching degree.
[0099] In an exemplary embodiment, a feedback mechanism for evaluating the responsiveness of a smart cockpit's voice interaction system is provided. This mechanism primarily employs a large language model algorithm-based evaluation method to assess the responsiveness of the smart cockpit's voice interaction system, thereby helping to improve the system's responsiveness. Specifically, the mechanism includes the following steps.
[0100] (1) An evaluation method based on a large language model algorithm is used to evaluate the response performance of the voice interaction system of the smart cockpit and generate evaluation feedback results.
[0101] (2) Feedback the evaluation feedback results to the development or maintenance process of the smart cockpit voice interaction system to guide the upgrade and optimization of the smart cockpit voice interaction system; the upgrade and optimization includes adjusting the response generation logic, improving the user interaction interface prompts and / or fine-tuning the large language model that supports response generation.
[0102] Among them, an evaluation method based on a large language model algorithm is adopted to evaluate the response performance of the voice interaction system of the smart cockpit and generate evaluation feedback results, which specifically includes the following steps.
[0103] (11) Receive instructions from the user to the voice interaction system of the smart cockpit (corresponding to the code ).
[0104] (12) Receive the voice interaction system of the smart cockpit in response to the instruction Generated Response (corresponding to the code ).
[0105] (13) Using the pre-configured large language model evaluation chain (corresponding to and function) to process the instruction and the response , and obtain the evaluation feedback results.
[0106] In this embodiment, the large language model evaluation chain includes an LLM model (such as a large language model supported by aliyun_bailian or openrouter) and a prompt template (Prompt Template), wherein the prompt template is used to guide the evaluation, and the LLM model is used to evaluate the response according to the evaluation dimensions defined in the prompt template. Perform evaluation and generate structured evaluation feedback results (corresponding to the Evaluation Result object in the code).
[0107] In this embodiment, the evaluation dimensions include semantic correctness (SC), state change confirmation (SCC), and unambiguous expression (UE).
[0108] In this embodiment, the evaluation feedback results include scores and opinions for each evaluation dimension, a comprehensive score and a validity flag (corresponding to the code and ), wherein the comprehensive score Used to indicate the comprehensive score of all the evaluation dimensions, the validity mark Used to indicate whether this interactive evaluation test has passed. Scores for each evaluation dimension 、 、 Generated by the large language model based on its internal understanding and evaluation criteria, with a value range of [0, 1].
[0109] In this embodiment, the comprehensive score The calculation formula is as follows.
[0110] .
[0111] .
[0112] .
[0113] in, Represents the evaluation results output by the large language model evaluation chain, including the scores and opinions of each evaluation dimension. Indicates the comprehensive score, Represents the process of performing evaluation through the large language model evaluation chain, Indicates the instructions given by the user to the voice interaction system of the smart cockpit. R Indicates that the voice interaction system of the intelligent cockpit is for instructions The generated response, Prompt templates and criteria for assessment, 、 、 They represent the score of semantic correctness, the score of state change confirmation, and the score of unambiguous expression respectively; A label indicating the degree of match between the vehicle screen image and the current database in the model database; Indicates combining image matching labels The comprehensive rating of , Represents the weight parameter for semantic evaluation and image matching.
[0114] In this embodiment, the validity flag It is expressed as the following formula.
[0115] .
[0116] in, Indicates the preset threshold for the comprehensive score.
[0117] In another embodiment, the validity flag It can also be expressed as the following formula.
[0118] .
[0119] in, 、 、 Represents the combined image matching label The preset threshold for the comprehensive score, the preset threshold for semantic correctness, the preset threshold for state change confirmation, and the preset threshold for unambiguous expression.
[0120] In this embodiment, the scores of each evaluation dimension will be included 、 、 , comprehensive score , validity mark and suggestions for improvement (corresponding to ) feedback is fed back into the development or maintenance process of the smart cockpit's voice interaction system. This feedback is used to guide upgrades and optimizations to the smart cockpit's voice interaction system, such as adjusting response generation logic, improving user interface prompts, and / or fine-tuning the large language model that supports response generation. By utilizing this evaluation and feedback mechanism for iterative optimization, the quality of interactive voice responses, user satisfaction, and task completion efficiency are continuously improved.
[0121] In an exemplary embodiment, the specific implementation method of the above-mentioned step S3 may be: if the matching degree is less than a preset matching degree threshold, then calculating the loss function value of the vehicle screen image and the current database; if the loss function value is less than the preset value, then searching for the feedback value of the smart cockpit in the current database; if the loss function value is greater than or equal to the preset value, then searching for the feedback value of the smart cockpit in other databases; wherein the other databases and the current database are different databases.
[0122] If the comprehensive score and validity flags If there is a matching deviation with the actual feedback result on the vehicle screen or the matching deviation is greater than a certain threshold, the large language model parameters are modified through the loss function, where the calculation formula of the loss function is as follows.
[0123] .
[0124] in, is the loss function value, To match the sample size, Determine the category type for the large language model, Give the real comprehensive score to the large language model. It is a symbol of the true effectiveness of the large language model.
[0125] If the loss function value is less than a preset value, it is determined that the large language model parameters are correct.
[0126] If the loss function value is greater than or equal to the preset value, the large language model parameters are modified.
[0127] In this embodiment, the large language model technology is applied to the intelligent cockpit voice automatic recognition scenario. By feeding back the test results to the voice interaction system of the intelligent cockpit, the voice interaction system automatically uses an evaluation method based on the large language model algorithm to evaluate the response performance and generate an evaluation feedback result. The evaluation feedback result can be fed back to the development or maintenance process of the voice interaction system of the intelligent cockpit to guide the upgrade and optimization of the voice interaction system of the intelligent cockpit, thereby helping to improve the recognition rate of the interactive voice, and improve the response quality of the interactive voice, user satisfaction and task completion efficiency.
[0128] Below, the implementation process of this application is described in detail: This application selects the required test database, selects the recognition area of the high-definition camera, normalizes the volume of the test corpus in the test database, and transmits it to the artificial mouth for playback after processing.
[0129] Volume normalization is performed using average volume normalization, as shown below.
[0130] .
[0131] in, is the sample size, For the The amplitude value of the samples, Indicates the average volume level of the sample volumes.
[0132] The calculation formula for normalized audio is as follows.
[0133] .
[0134] in, Indicates the desired average volume level.
[0135] This embodiment uses a normalization method to normalize the volume of the test speech, which can effectively improve the test accuracy and speech recognition effect.
[0136] In this embodiment, after receiving the voice command from the artificial mouth, the voice interaction system of the smart cockpit is awakened and recognizes the voice command, executes the command content, and the screen displays the recognized command content.
[0137] In this embodiment, a high-definition camera is used for real-time monitoring to obtain text or image information within the selected area. The voice interaction system of the smart cockpit is awakened, for example, to obtain image information, or recognize voice commands, for example, to obtain text information.
[0138] Specifically, the captured image is grayscaled to obtain the grayscale value Gray , expressed as the following formula.
[0139] .
[0140] in, 、 、 They are the values of the red, green, and blue color channels respectively.
[0141] Then each pixel in the grayscale image is binarized according to the grayscale, which is expressed as the following formula.
[0142] .
[0143] in, represents the binarized image, It is a preset threshold value, 255 represents white and 0 represents black.
[0144] To denoise the binarized image, a Gaussian filter is used, as shown below.
[0145] .
[0146] in, x and y Represent the horizontal and vertical coordinates of the distance from the center, is the standard deviation of the Gaussian distribution, which controls the width of the kernel.
[0147] The denoised images are resized so that all input images have the same size, which is represented as follows.
[0148] .
[0149] in, For the output image, is the input image, is the mean value of the image pixels.
[0150] Feature extraction is performed on the resized image to extract information that is helpful for character recognition from the processed image. For example, edge detection is performed using the Sobel operator, as shown below.
[0151] .
[0152] .
[0153] .
[0154] .
[0155] in, represents the image matrix, and Respectively represent the gradient matrices in the horizontal and vertical directions at the point where the image is calculated, Indicates the gradient magnitude at the point where the image is calculated, Indicates the gradient direction at the point where the image is calculated.
[0156] The image after edge detection is recognized. The specific recognition process is implemented using a convolutional neural network to automatically learn features and introduce a language model to handle grammar or dictionary correction, as well as possible semantic analysis, to improve recognition accuracy.
[0157] The text or image captured by the high-definition camera is recognized and the corresponding recognition result is obtained. The recognition result is compared with the local text true value or image true value to determine whether the recognition result and the true value are consistent. If they are consistent, it means that the recognition is successful, and if they are inconsistent, it means that the recognition has failed.
[0158] Finally, based on multiple performance indicators such as the determined voice command recognition rate, a test report is generated.
[0159] In an exemplary embodiment, the specific implementation method of the above-mentioned step S2 may be: obtaining the interaction information of the smart cockpit based on the test voice; wherein the interaction information includes voice information and / or text information displayed on the vehicle screen of the smart cockpit; correspondingly, the specific implementation method of step S3 may be: identifying the interaction information and obtaining feedback information of the smart cockpit; correspondingly, the specific implementation method of step S4 may be: comparing the feedback information with the standard interaction information corresponding to the test voice to obtain the recognition result of the voice interaction performance of the smart cockpit.
[0160] Specifically, after playing the test voice, the present application obtains the interactive information of the smart cockpit, specifically including the interactive voice of the voice interaction system or the interactive text displayed on the vehicle screen. The present application obtains the interactive voice of the voice interaction system through a microphone and the interactive text of the voice interaction system through a high-definition camera, and obtains the interactive information of the voice interaction system of the smart cockpit by identifying the interactive voice or interactive text. The interactive performance of the voice interaction system of the smart cockpit, that is, the interactive accuracy, is judged by comparing the interactive information with the standard interactive information corresponding to the test voice, and the interactive accuracy corresponding to multiple test voices is calculated by counting the interactive accuracies of multiple test voices to calculate the interactive accuracy rate of the voice interaction system of the smart cockpit, that is, the number of correctly interacted test voices and the total number of played test voices.
[0161] In an exemplary embodiment, the specific implementation method of the above-mentioned step S2 may be: obtaining the response time of the smart cockpit based on the test voice; correspondingly, the specific implementation method of step S3 may be: calculating the response duration of the smart cockpit based on the test voice; wherein the response duration represents the duration from the test voice being played to the smart cockpit responding; correspondingly, the specific implementation method of step S4 may be: comparing the response duration with a preset time threshold to obtain the recognition result of the voice interaction performance of the smart cockpit.
[0162] Specifically, this application records the playback time of a test voice and obtains the smart cockpit's response information (including wake-up, recognition, and interaction times) after the test voice is played to determine its response time. Based on the playback and response times, the smart cockpit's response duration is calculated. By comparing the response duration with a preset time threshold, a response speed indicator for the smart cockpit's voice interaction system is obtained. This application can obtain multiple response times for multiple test voices and then take a weighted average to determine the response time of the smart cockpit's voice interaction system. Alternatively, different response times can be obtained for different commands.
[0163] Based on the same inventive concept, embodiments of the present application also provide an intelligent cockpit voice automation recognition device for implementing the aforementioned intelligent cockpit voice automation recognition method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the intelligent cockpit voice automation recognition device provided below can be found in the aforementioned limitations of the intelligent cockpit voice automation recognition method and will not be further elaborated here.
[0164] Figure 2 The figure shows a schematic diagram of the structure of an intelligent cockpit voice automatic recognition device. Figure 2 As shown, the intelligent cockpit voice automatic recognition device includes the following modules.
[0165] The test voice playing module N1 is used to play the test voice.
[0166] The response information acquisition module N2 is used to obtain the response information of the smart cockpit based on the test voice; wherein the response information includes the vehicle screen image of the smart cockpit.
[0167] The feedback result recognition module N3 is used to recognize the response information and obtain the feedback value of the smart cockpit; wherein the feedback value represents the recognition index of the interactive performance of the smart cockpit in responding to the test voice.
[0168] The recognition result calculation module N4 is used to compare the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the smart cockpit.
[0169] In this embodiment, the test voice is played by the test voice playing module N1; the response information acquisition module N2 obtains the response information of the smart cockpit based on the test voice; the feedback result recognition module N3 recognizes the response information and obtains the feedback value of the smart cockpit; wherein, the response information includes the vehicle screen image of the smart cockpit, and the feedback value represents the recognition index of the interactive performance of the smart cockpit in response to the test voice; the recognition result calculation module N4 compares the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the smart cockpit; that is, by playing the test voice and obtaining the response information of the smart cockpit based on the test voice, the response information is recognized to obtain the feedback value of the smart cockpit, and by comparing the feedback value with the standard value corresponding to the test voice, the recognition result of the voice interaction performance of the smart cockpit is obtained, thereby realizing accurate recognition of the voice interaction performance of the smart cockpit, and utilizing automatic testing of the test voice to improve test efficiency, so that the voice interaction performance can be quickly recognized.
[0170] In an exemplary embodiment, the vehicle screen image includes a response interface of the smart cockpit based on the test voice; wherein, the above-mentioned feedback result recognition module N3 can be further configured to: recognize the vehicle screen image to obtain the wake-up status of the voice interaction system of the smart cockpit; wherein, the wake-up status includes successful wake-up and failure to wake up; correspondingly, the recognition result calculation module N4 can be further configured to: count the wake-up status of the voice interaction system of the smart cockpit to obtain the wake-up rate of the voice interaction system of the smart cockpit; compare the wake-up rate of the voice interaction system of the smart cockpit with the preset wake-up rate threshold to obtain the recognition result of the voice interaction performance of the smart cockpit.
[0171] In an exemplary embodiment, the vehicle screen image includes text information obtained by the smart cockpit recognizing the test voice; wherein, the above-mentioned feedback result recognition module N3 can be further configured to: recognize the vehicle screen image and obtain the text information obtained by the smart cockpit recognizing the test voice; correspondingly, the recognition result calculation module N4 can be further configured to: compare the text information obtained by the smart cockpit recognizing the test voice with the original text of the test voice to obtain the recognition result of the voice interaction performance of the smart cockpit.
[0172] In an exemplary embodiment, the feedback result recognition module N3 can be further configured to: calculate the matching degree between the vehicle screen image and the current database in the model database; wherein the model database includes a database of the response interface of the voice interaction system corresponding to the smart cockpit and the original text of the test voice; based on the matching degree, search for the feedback value of the smart cockpit in a database in the model database.
[0173] In an exemplary embodiment, the feedback result identification module N3 can be further configured as follows: if the matching degree is less than a preset matching degree threshold, the loss function value of the vehicle screen image and the current database is calculated; if the loss function value is less than the preset value, the feedback value of the smart cockpit is searched in the current database; if the loss function value is greater than or equal to the preset value, the feedback value of the smart cockpit is searched in other databases; wherein the other databases and the current database are different databases.
[0174] In an exemplary embodiment, the response information acquisition module N2 can be further configured to obtain the interaction information of the smart cockpit based on the test voice; wherein the interaction information includes voice information and / or text information displayed on the vehicle screen of the smart cockpit; correspondingly, the feedback result recognition module N3 can be further configured to recognize the interaction information and obtain feedback information of the smart cockpit; correspondingly, the recognition result calculation module N4 can be further configured to compare the feedback information with the standard interaction information corresponding to the test voice to obtain the recognition result of the voice interaction performance of the smart cockpit.
[0175] In an exemplary embodiment, the response information acquisition module N2 can be further configured to obtain the response time of the smart cockpit based on the test voice; correspondingly, the feedback result recognition module N3 can be further configured to calculate the response duration of the smart cockpit based on the test voice; wherein, the response duration represents the duration from the test voice being played to the smart cockpit responding; correspondingly, the recognition result calculation module N4 can be further configured to compare the response duration with a preset time threshold to obtain the recognition result of the voice interaction performance of the smart cockpit.
[0176] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0177] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for automatic voice recognition in an intelligent cockpit, characterized in that: The intelligent cockpit voice automatic recognition method includes: Play test audio; Obtaining response information of the smart cockpit based on the test voice; wherein the response information includes a vehicle screen image of the smart cockpit; Identify the response information and obtain the feedback value of the smart cockpit; wherein the feedback value represents the recognition index of the interactive performance of the smart cockpit in response to the test voice; identify the response information and obtain the feedback value of the smart cockpit, specifically including: calculating the matching degree between the vehicle screen image and the current database in the model database; based on the matching degree, search for the feedback value of the smart cockpit in a database in the model database; wherein the model database includes a database corresponding to the response interface of the voice interaction system of the smart cockpit and the original text of the test voice respectively; a label of the matching degree between the vehicle screen image and the current database in the model database The smaller the value, the higher the matching degree between the vehicle screen image and the current database in the model database. <1, it is determined that the vehicle screen image data matches the current database, and the label The calculation formula is: ; ; in, Y a and Y b Respectively represent the vehicle screen image data and the data in the current database, and Respectively f The vehicle screen image data of the feature dimension and the data in the current database, F is the characteristic dimension of the current database data, X for Y a and Y b The distance between A preset threshold for measuring matching degree; Comparing the feedback value with a standard value corresponding to the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit; The intelligent cockpit voice automatic recognition method further includes: An evaluation method based on a large language model algorithm is used to evaluate the response performance of the voice interaction system of the smart cockpit and generate evaluation feedback results; specifically, the method includes: receiving instructions from the user to the voice interaction system of the smart cockpit; ; Receive the voice interaction system of the smart cockpit for the instruction Generated Response ; Process the instructions using a pre-configured large language model evaluation chain and the response , get the evaluation feedback results; Feedback of the evaluation results to the development or maintenance process of the smart cockpit's voice interaction system to guide upgrades and optimizations to the smart cockpit's voice interaction system; such upgrades and optimizations may include adjusting response generation logic, improving user interface prompts, and / or fine-tuning the large language model that supports response generation. The large language model evaluation chain includes an LLM model and a prompt template, wherein the prompt template is used to guide the evaluation, and the LLM model is used to evaluate the response according to the evaluation dimension defined in the prompt template. Performing an evaluation and generating the evaluation feedback result; The evaluation dimensions include semantic correctness, confirmation of state changes, and unambiguous expression; The evaluation feedback results include the scores and opinions of each evaluation dimension, the comprehensive score and validity flags , wherein the comprehensive score Used to indicate the comprehensive score of all the evaluation dimensions, the validity mark Used to indicate whether this interactive assessment test has passed; The comprehensive score The calculation formula is: ; ; ; in, Represents the evaluation results output by the large language model evaluation chain, including the scores and opinions of each evaluation dimension. Indicates the comprehensive score, Represents the process of performing evaluation through the large language model evaluation chain, Indicates the instructions given by the user to the voice interaction system of the smart cockpit. R Indicates that the voice interaction system of the intelligent cockpit is for instructions The generated response, Prompt templates and criteria for assessment, 、 、 They represent the score of semantic correctness, the score of state change confirmation, and the score of unambiguous expression respectively; A label indicating the degree of match between the vehicle screen image and the current database in the model database; Indicates combining image matching labels The comprehensive rating of , Represents the weight parameter for semantic evaluation and image matching; The validity flag The calculation formula is: ; in, Indicates the preset threshold value of the comprehensive score; Alternatively, the validity flag The calculation formula is: ; in, 、 、 Represents the combined image matching label The preset threshold for the comprehensive score, the preset threshold for semantic correctness, the preset threshold for state change confirmation, and the preset threshold for unambiguous expression.
2. The intelligent cockpit voice automatic recognition method according to claim 1, characterized in that: The vehicle screen image includes a response interface of the smart cockpit based on the test voice; Identifying the response information and obtaining a feedback value of the smart cockpit specifically includes: Identify the vehicle screen image to obtain the wake-up state of the voice interaction system of the smart cockpit; wherein the wake-up state includes successful wake-up and failure to wake up; Comparing the feedback value with a standard value corresponding to the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit, specifically including: Counting the wake-up status of the voice interaction system of the smart cockpit to obtain the wake-up rate of the voice interaction system of the smart cockpit; The wake-up rate of the voice interaction system of the smart cockpit is compared with a preset wake-up rate threshold to obtain an identification result of the voice interaction performance of the smart cockpit.
3. The intelligent cockpit voice automatic recognition method according to claim 1, characterized in that: The vehicle screen image includes text information obtained by the smart cockpit recognizing the test voice; Identifying the response information and obtaining a feedback value of the smart cockpit specifically includes: Recognize the vehicle screen image and obtain text information obtained by the smart cockpit recognizing the test voice; Comparing the feedback value with a standard value corresponding to the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit, specifically including: The text information obtained by the smart cockpit recognizing the test voice is compared with the original text of the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit.
4. The method for automatic voice recognition in an intelligent cockpit according to claim 1, characterized in that: Obtaining the response information of the smart cockpit based on the test voice, specifically including: Acquiring interaction information of the smart cockpit based on the test voice; wherein the interaction information includes voice information and / or text information displayed on a screen of the smart cockpit; Identifying the response information and obtaining a feedback value of the smart cockpit specifically includes: Identifying the interaction information and obtaining feedback information from the smart cockpit; Comparing the feedback value with a standard value corresponding to the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit, specifically including: The feedback information is compared with standard interaction information corresponding to the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit.
5. The intelligent cockpit voice automatic recognition method according to claim 1, characterized in that: Obtaining the response information of the smart cockpit based on the test voice, specifically including: Obtaining a response time of the smart cockpit based on the test voice; Identifying the response information and obtaining a feedback value of the smart cockpit specifically includes: Calculating the response duration of the smart cockpit based on the test voice; wherein the response duration represents the duration from the test voice being played to the smart cockpit responding; Comparing the feedback value with a standard value corresponding to the test voice to obtain a recognition result of the voice interaction performance of the smart cockpit, specifically including: The response time is compared with a preset time threshold to obtain a recognition result of the voice interaction performance of the smart cockpit.
6. An intelligent cockpit voice automatic recognition device, characterized in that: The intelligent cockpit voice automation recognition device uses the intelligent cockpit voice automation recognition method according to claim 1, and the intelligent cockpit voice automation recognition device includes: Test voice playback module, used to play test voice; A response information acquisition module, configured to acquire response information of the smart cockpit based on the test voice; wherein the response information includes an image of the vehicle screen of the smart cockpit; a feedback result recognition module, configured to recognize the response information and obtain a feedback value of the smart cockpit; wherein the feedback value represents an identification index of the interactive performance of the smart cockpit in response to the test speech; The recognition result calculation module is used to compare the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the smart cockpit.
Citation Information
Patent Citations
Vehicle-mounted control screen voice recognition process testing method, electronic equipment and system
CN109616106A
Test method, system and device of intelligent voice interaction system and medium
CN118135998A
Intelligent automobile virtual simulation test method based on large language model
CN118586281A