Intelligent cockpit voice automatic recognition method and device
By playing test voice in the smart cockpit and obtaining response information, identifying and comparing feedback values to evaluate voice interaction performance, and using large language model algorithms for evaluation, the problem of low recognition efficiency and poor accuracy of the smart cockpit voice interaction system in the real environment is solved, and efficient and accurate speech recognition and system optimization are achieved.
Patent Information
- Application Number
- CN202510724553.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The existing smart cockpit voice interaction system has low recognition efficiency and low accuracy in real environments, making it difficult to overcome the influence of acoustic factors such as driving noise, wind noise and in-vehicle reverb.
By playing test voice, the response information of the smart cockpit is obtained, including the car screen image, the feedback value is recognized, the feedback value is compared with the standard value to identify the voice interaction performance, and the response performance can be evaluated using a large language model algorithm to generate evaluation feedback results to guide the system optimization.
Accurate recognition of the voice interaction performance of the smart cockpit is achieved, the accuracy and efficiency of voice recognition is improved, and the system's response quality and user satisfaction are improved through automatic testing.
Smart Images

Figure CN120236581A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of intelligent cockpit voice interaction and voice recognition, and in particular to an intelligent cockpit voice automatic recognition method and device. Background Art
[0002] With the continuous development of automobile technology, Internet of Vehicles and artificial intelligence technology are gradually emerging, and more and more functions are being installed in cars. The intelligent cockpit aims to integrate a variety of IT and artificial intelligence technologies to create a new integrated digital platform in the car, provide drivers with an intelligent experience, and promote driving safety. At present, the voice interaction function, as a representative function of the car's intelligent cockpit, is combined with a variety of applications in the car and has become a core function of the cockpit ecosystem. The entire process of voice interaction includes voice reception, voice recognition, semantic understanding, etc. Errors in any link will lead to overall interaction failure. In addition, acoustic factors such as vehicle driving noise, wind noise, and in-car reverberation will also affect the effect of voice recognition. Whether these constraints can be overcome in a real environment is the key to the quality of the in-vehicle voice interaction system experience.
[0003] In order to determine the quality of the car's voice interaction function, it is necessary to test and count the various performance indicators of the vehicle-mounted voice interaction system to form a recognition result. However, the existing human-computer interaction recognition tasks are numerous, with a lot of repetitive work, low recognition efficiency, and low recognition accuracy. Therefore, how to provide a more accurate and efficient intelligent cockpit voice automatic recognition method to improve the accuracy and efficiency of intelligent cockpit voice recognition has become a technical problem that needs to be solved in this field. Summary of the invention
[0004] The purpose of this application is to provide a method and device for automatic recognition of intelligent cockpit speech, which can improve the accuracy and efficiency of intelligent cockpit speech recognition.
[0005] To achieve the above objectives, this application provides the following solutions.
[0006] In a first aspect, the present application provides a smart cockpit voice automation recognition method, which specifically includes the following steps.
[0007] Play the test audio.
[0008] Acquire response information of the smart cockpit based on the test voice; wherein the response information includes a vehicle screen image of the smart cockpit.
[0009] The response information is identified to obtain a feedback value of the smart cockpit; wherein the feedback value represents an identification index of the interactive performance of the smart cockpit in response to the test voice.
[0010] Compare the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0011] Optionally, the intelligent cockpit voice automatic recognition method further includes the following steps.
[0012] Adopt an evaluation method based on the large language model algorithm to evaluate the response performance of the voice interaction system of the intelligent cockpit, and generate an evaluation feedback result.
[0013] Feed back the evaluation feedback result to the development or maintenance process of the voice interaction system of the intelligent cockpit to guide the upgrade and optimization of the voice interaction system of the intelligent cockpit; the upgrade and optimization include adjusting the response generation logic, improving the user interface prompt, and / or fine-tuning the large language model supporting the response generation.
[0014] Optionally, adopting an evaluation method based on the large language model algorithm to evaluate the response performance of the voice interaction system of the intelligent cockpit and generate an evaluation feedback result specifically includes the following steps.
[0015] Receive the instruction given by the user to the voice interaction system of the intelligent cockpit .
[0016] Receive the response generated by the voice interaction system of the intelligent cockpit for the instruction generated .
[0017] Process the instruction and the response using a pre-configured large language model evaluation chain to obtain an evaluation feedback result.
[0018] Optionally, identifying the response information to obtain the feedback value of the intelligent cockpit specifically includes the following steps.
[0019] Calculate the matching degree between the in-vehicle computer screen image and the current database in the model database; wherein, the model database includes databases corresponding to the response interface of the voice interaction system of the intelligent cockpit and the original text of the test voice respectively; the smaller the label of the matching degree between the in-vehicle computer screen image and the current database in the model database, the higher the matching degree between the in-vehicle computer screen image and the current database in the model database, and the formula for the label is as follows.
[0020] .
[0021] .
[0022] Among them, Y a and Y b respectively represent the in-vehicle infotainment (IVI) screen image data and the data in the current database, and respectively represent f the IVI screen image data of the feature dimension and the data in the current database, F is the feature dimension of the current database data, X is Y a and Y b the distance between, is a threshold for measuring the matching degree set in advance.
[0023] Optionally, the large language model evaluation chain includes an LLM model and a prompt template, where the prompt template is used to guide the evaluation, and the LLM model is used to evaluate the response according to the evaluation dimensions defined in the prompt template, and generate the evaluation feedback result.
[0024] The evaluation dimensions include semantic correctness, status change confirmation, and unambiguous expression.
[0025] The evaluation feedback result includes the scores and opinions of each evaluation dimension, the comprehensive score and the validity flag , where the comprehensive score is used to represent the comprehensive score situation of all evaluation dimensions, and the validity flag is used to indicate whether the interactive evaluation test passes this time; The comprehensive score has the following calculation formula.
[0026] .
[0027] .
[0028] .
[0029] Among them, represents the evaluation result output by the large language model evaluation chain, including the scores and opinions of each evaluation dimension, represents the comprehensive score, represents the process of performing the evaluation through the large language model evaluation chain, represents the instruction given by the user to the voice interaction system of the intelligent cockpit, R represents the response generated by the voice interaction system of the intelligent cockpit for the instruction The prompt templates and criteria for evaluation 、 、 respectively represent the scores for semantic correctness, the scores for status change confirmation, and the scores for unambiguous expressions; is the label for the matching degree between the in-vehicle infotainment (IVI) screen image and the current database in the model database; represents the combined score considering the image matching label ; , represents the weight parameter for semantic evaluation and image matching.
[0030] The validity flag has the following calculation formula.
[0031] 。
[0032] Among them, represents the preset threshold for the combined score.
[0033] Alternatively, the validity flag has the following calculation formula.
[0034] 。
[0035] Among them, 、 、 respectively represent the preset thresholds for the combined score considering the image matching label , the preset threshold for semantic correctness, the preset threshold for status change confirmation, and the preset threshold for unambiguous expressions.
[0036] If there is a matching deviation between the combined score and the validity flag and the actual feedback result of the IVI screen, or the matching deviation is greater than the threshold, then the parameters of the large language model are re-modified through the loss function. Among them, the calculation formula of the loss function is as follows.
[0037] 。
[0038] Among them, is the value of the loss function, is the number of matching samples, is the type of judgment category of the large language model, is the true combined score of the large language model, is the true validity flag of the large language model.
[0039] If the value of the loss function is less than the preset value, it is determined that the parameters of the large language model are correct.
[0040] If the value of the loss function is greater than or equal to the preset value, the parameters of the large language model are modified again.
[0041] Optionally, the in-vehicle screen image includes the response interface of the intelligent cockpit based on the test voice.
[0042] Identifying the response information to obtain the feedback value of the intelligent cockpit specifically includes the following steps.
[0043] Identifying the in-vehicle screen image to obtain the wake-up state of the voice interaction system of the intelligent cockpit; wherein, the wake-up state includes successful wake-up and non-wake-up.
[0044] Comparing the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit specifically includes the following steps.
[0045] Counting the wake-up state of the voice interaction system of the intelligent cockpit to obtain the wake-up rate of the voice interaction system of the intelligent cockpit.
[0046] Comparing the wake-up rate of the voice interaction system of the intelligent cockpit with a preset wake-up rate threshold to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0047] Optionally, the in-vehicle screen image includes the text information obtained by the intelligent cockpit recognizing the test voice.
[0048] Identifying the response information to obtain the feedback value of the intelligent cockpit specifically includes the following steps.
[0049] Identifying the in-vehicle screen image to obtain the text information obtained by the intelligent cockpit recognizing the test voice.
[0050] Comparing the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit specifically includes the following steps.
[0051] Comparing the text information obtained by the intelligent cockpit recognizing the test voice with the original text of the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0052] Optionally, obtaining the response information of the intelligent cockpit based on the test voice specifically includes the following steps.
[0053] Obtaining the interaction information of the intelligent cockpit based on the test voice; wherein, the interaction information includes voice information and / or text information displayed on the in-vehicle screen of the intelligent cockpit.
[0054] Identifying the response information to obtain the feedback value of the intelligent cockpit specifically includes the following steps.
[0055] Identify the interaction information to obtain the feedback information of the intelligent cockpit.
[0056] Compare the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit, which specifically includes the following steps.
[0057] Compare the feedback information with the standard interaction information corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0058] Optionally, obtain the response information of the intelligent cockpit based on the test voice, which specifically includes the following steps.
[0059] Obtain the response moment of the intelligent cockpit based on the test voice.
[0060] Identify the response information to obtain the feedback value of the intelligent cockpit, which specifically includes the following steps.
[0061] Calculate the response duration of the intelligent cockpit based on the test voice; wherein, the response duration represents the duration from the playback of the test voice to the response of the intelligent cockpit.
[0062] Compare the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit, which specifically includes the following steps.
[0063] Compare the response duration with a preset time threshold to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0064] In a second aspect, the present application provides an intelligent cockpit voice automatic recognition device, and the intelligent cockpit voice automatic recognition device includes the following modules.
[0065] A test voice playback module for playing test voices.
[0066] A response information acquisition module for acquiring the response information of the intelligent cockpit based on the test voice; wherein, the response information includes the in-vehicle screen image of the intelligent cockpit.
[0067] A feedback result recognition module for identifying the response information to obtain the feedback value of the intelligent cockpit; wherein, the feedback value represents the recognition index of the interaction performance of the intelligent cockpit in response to the test voice.
[0068] A recognition result calculation module for comparing the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0069] According to the specific embodiments provided by the present application, the present application has the following technical effects.
[0070] The present application provides a method and device for automatic voice recognition in an intelligent cockpit. By playing a test voice; obtaining response information of the intelligent cockpit based on the test voice; identifying the response information to obtain a feedback value of the intelligent cockpit; wherein, the response information includes an image of the in-vehicle screen of the intelligent cockpit, and the feedback value represents an identification index of the interaction performance of the intelligent cockpit in response to the test voice; comparing the feedback value with a standard value corresponding to the test voice to obtain an identification result of the voice interaction performance of the intelligent cockpit; that is, by playing the test voice and obtaining the response information of the intelligent cockpit based on the test voice, identifying the response information to obtain the feedback value of the intelligent cockpit, and comparing the feedback value with the standard value corresponding to the test voice, an identification result of the voice interaction performance of the intelligent cockpit can be obtained, so as to accurately identify the voice interaction performance of the intelligent cockpit, and improve the accuracy and efficiency of voice recognition in the intelligent cockpit by using automatic testing of the test voice. Description of the Drawings
[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0072] Figure 1 It is a schematic flowchart of a method for automatic voice recognition in an intelligent cockpit provided by an embodiment of the present application.
[0073] Figure 2 It is a schematic structural diagram of a device for automatic voice recognition in an intelligent cockpit provided by an embodiment of the present application. Detailed Embodiments
[0074] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0075] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0076] As Figure 1 shown, this embodiment proposes a method for automatic voice recognition in an intelligent cockpit. The method for automatic voice recognition in the intelligent cockpit specifically includes the following steps.
[0077] Step S1: Play the test voice.
[0078] Specifically, in this application, by operating the voice verification tool software on the host computer, a task is created (including recognition of multiple pieces of corpus, such as 100 pieces), the recognized test voice is read and played using an artificial mouth. For example: "Hello, turn on the air conditioner."
[0079] Among them, the artificial mouth is a special artificial sound source, also known as a simulation mouth or artificial mouth, which is composed of a small speaker installed on a baffle with a special shape. The design of the baffle shape is to simulate the average directivity and radiation pattern of the human mouth. The artificial mouth also needs to be equipped with a frequency compensation network or sound compression to achieve a certain sound frequency response. The artificial mouth is used to simulate the electroacoustic characteristics of the microphone under actual working conditions for testing or calibration, so as to obtain a sound source close to the real human mouth.
[0080] Step S2: Obtain the response information of the intelligent cockpit based on the test voice.
[0081] Among them, the response information includes the image of the in-vehicle screen of the intelligent cockpit. After this application uses the artificial mouth to play each test voice, it uses a high-definition camera to obtain the response information of the voice interaction system of the intelligent cockpit based on this test voice. For example, for the above test voice "Hello, turn on the air conditioner", under normal circumstances, after the voice interaction system receives this test voice, the corresponding text will be displayed on the in-vehicle screen of the intelligent cockpit and the interface for turning on the air conditioner will be shown. This application obtains the in-vehicle screen image through a high-definition camera to determine whether the voice interaction system of the intelligent cockpit has a response and whether the response is correct.
[0082] Step S3: Identify the response information to obtain the feedback value of the intelligent cockpit.
[0083] Among them, the feedback value represents the recognition index of the interaction performance of the intelligent cockpit in response to the test voice. After this application obtains the response information of the intelligent cockpit, it obtains the feedback value of the intelligent cockpit by identifying this response information. For example, by identifying the text information or the interface information for turning on the air conditioner on the in-vehicle screen through the in-vehicle screen image, to determine whether the voice interaction system of the intelligent cockpit has received the test voice and whether it has responded correctly.
[0084] Step S4: Compare the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0085] After obtaining the feedback value of the intelligent cockpit in this embodiment, the recognition result of the voice interaction performance of the intelligent cockpit is obtained by comparing the feedback value with the standard value corresponding to the test voice. For example, after obtaining the text information or interface information on the in-vehicle screen of the intelligent cockpit, by comparing the standard text and standard interface information for turning on the air conditioner, it is determined whether the intelligent cockpit correctly turns on the air conditioner, that is, it is determined whether the voice interaction system of the intelligent cockpit interacts successfully, so that the interaction success rate of the voice interaction system (one of the recognition metrics) can be obtained by counting multiple test voices. By analogy, this application can obtain multiple recognition metrics of the voice interaction system of the intelligent cockpit, and based on the multiple recognition metrics, the recognition result of the voice interaction system of the intelligent cockpit is comprehensively obtained.
[0086] In this embodiment, a test voice is played; response information of the intelligent cockpit based on the test voice is obtained; the response information is recognized to obtain the feedback value of the intelligent cockpit; wherein, the response information includes the in-vehicle screen image of the intelligent cockpit, and the feedback value represents the recognition metric of the interaction performance of the intelligent cockpit in response to the test voice; the feedback value is compared with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit; that is, by playing the test voice and obtaining the response information of the intelligent cockpit based on the test voice, the response information is recognized to obtain the feedback value of the intelligent cockpit, and by comparing the feedback value with the standard value corresponding to the test voice, the recognition result of the voice interaction performance of the intelligent cockpit is obtained, so that the voice interaction performance of the intelligent cockpit can be accurately recognized, and the automatic test of the test voice is used to improve the test efficiency, so that the voice interaction performance can be quickly recognized, and the accuracy and efficiency of the voice automation recognition of the intelligent cockpit are improved.
[0087] In an exemplary embodiment, the in-vehicle screen image includes the response interface of the intelligent cockpit based on the test voice; wherein, the specific implementation manner of the above step S3 may be: recognizing the in-vehicle screen image to obtain the wake-up state of the voice interaction system of the intelligent cockpit; wherein, the wake-up state includes successful wake-up and non-wake-up; correspondingly, the specific implementation manner of step S4 may be: counting the wake-up state of the voice interaction system of the intelligent cockpit to obtain the wake-up rate of the voice interaction system of the intelligent cockpit; comparing the wake-up rate of the voice interaction system of the intelligent cockpit with a preset wake-up rate threshold to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0088] Specifically, the present application identifies the interface information in the in-vehicle infotainment (IVI) screen image to obtain the response interface of the intelligent cockpit based on the test voice, determines whether the voice interaction system of the intelligent cockpit is awakened based on this response interface, and calculates the wake-up rate of the voice interaction system of the intelligent cockpit by counting the wake-up states of the voice interaction system of the intelligent cockpit based on multiple test voices, that is, the ratio of the number of test voices that successfully awaken the voice interaction system to the total number of played test voices. After obtaining the wake-up rate of the voice interaction system of the intelligent cockpit, the wake-up rate is compared with a preset wake-up rate threshold to determine whether the wake-up rate of the voice interaction system of the intelligent cockpit is qualified. For example, if the wake-up rate is lower than 95%, it is considered that the performance index of the wake-up rate is low.
[0089] In an exemplary embodiment, the IVI screen image includes the text information obtained by the intelligent cockpit recognizing the test voice; wherein, the specific implementation manner of the above step S3 can be: recognizing the IVI screen image to obtain the text information obtained by the intelligent cockpit recognizing the test voice; correspondingly, the specific implementation manner of step S4 can be: comparing the text information obtained by the intelligent cockpit recognizing the test voice with the original text of the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0090] Specifically, the present application identifies the text information in the IVI screen image to obtain the text command obtained by the intelligent cockpit recognizing the test voice, and compares the text command with the original text of the test voice to obtain the speech recognition accuracy of the voice interaction system of the intelligent cockpit, and the speech recognition accuracy rate can be calculated based on the speech recognition accuracies of multiple test voices, that is, the ratio of the number of test voices correctly recognized to the total number of played test voices. And a correct rate threshold can be set. After calculating the recognition correct rate of the speech recognition system of the intelligent cockpit, the recognition correct rate is compared with the correct rate threshold to determine whether the speech recognition correct rate of the voice interaction system of the intelligent cockpit is qualified.
[0091] In an exemplary embodiment, the specific implementation manner of the above step S3 can be: calculating the matching degree between the IVI screen image and the current database in the model database; wherein, the model database includes databases corresponding to the response interface of the voice interaction system of the intelligent cockpit and the original text of the test voice respectively; and searching for the feedback value of the intelligent cockpit in one of the databases in the model database based on the matching degree.
[0092] Since the IVI screen image of the intelligent cockpit may be text information or interface information, after obtaining the IVI screen image, the present application can determine the search database of the feedback value according to the matching degree between the IVI screen image and the databases of the corresponding original text and the response interface, so as to improve the search efficiency.
[0093] Specifically, after obtaining the in-vehicle infotainment (IVI) screen image, this application calculates the matching degree between the IVI screen image and the current database in the model database (such as the database of the original text or the database of the response interface, etc.). The specific method is as follows.
[0094] 。
[0095] 。
[0096] Among them, Y a and Y b respectively represent the data of the IVI screen image and the data in the current database, respectively represent f the data of the IVI screen image in the feature dimension and the data in the current database, F is the feature dimension of the data in the current database, is and the distance between, is a pre-set threshold for measuring the matching degree, is the label of the matching degree between the IVI screen image and the current database in the model database, that is Y a and Y b the label of the matching degree.
[0097] When 0 ≤ < 1, it is determined that the data of the IVI screen image and the current database are matched, the smaller the
[0098] value, the higher the matching degree.
[0098] In an exemplary embodiment, an evaluation feedback mechanism designed for the response performance of the voice interaction system of the intelligent cockpit is also provided. The evaluation feedback mechanism mainly uses an evaluation method based on the large language model algorithm to evaluate the response performance of the voice interaction system of the intelligent cockpit, so as to help improve the response performance of the voice interaction system of the intelligent cockpit. Specifically, it includes the following steps.
[0099] (1) Use an evaluation method based on the large language model algorithm to evaluate the response performance of the voice interaction system of the intelligent cockpit, and generate an evaluation feedback result.
[0100] (2) Feed the evaluation feedback result back to the development or maintenance process of the voice interaction system of the intelligent cockpit to guide the upgrade and optimization of the voice interaction system of the intelligent cockpit; the upgrade and optimization include adjusting the response generation logic, improving the user interface prompt, and / or fine-tuning the large language model that supports the response generation (Fine-Tuning).
[0101] Among them, an evaluation method based on the large language model algorithm is adopted to evaluate the response performance of the voice interaction system of the intelligent cockpit, and an evaluation feedback result is generated, which specifically includes the following steps.
[0102] (11) Receive the instruction issued by the user to the voice interaction system of the intelligent cockpit (corresponding to )
[0103] (12) Receive the response generated by the voice interaction system of the intelligent cockpit for the instruction (corresponding to in the code )
[0104] (13) Use the pre-configured large language model evaluation chain (corresponding to and in the code, function) to process the instruction and the response , and obtain the evaluation feedback result.
[0105] In this embodiment, the large language model evaluation chain includes an LLM model (such as a large language model supported by aliyun_bailian or openrouter) and a prompt template (Prompt Template). Among them, the prompt template is used to guide the evaluation, and the LLM model is used to evaluate the response according to the evaluation dimensions defined in the prompt template, and generate a structured evaluation feedback result (Evaluation Result object corresponding to the code).
[0106] In this embodiment, the evaluation dimensions include semantic correctness (Semantic Correctness, SC), state change confirmation (State Change Confirmation, SCC), and unambiguous expression (Unambiguous Expression, UE).
[0107] In this embodiment, the evaluation feedback result includes the scores and opinions of each evaluation dimension, a comprehensive score and a validity flag (corresponding to and in the code). Among them, the comprehensive score is used to represent the comprehensive score situation of all the evaluation dimensions, and the validity flag is used to indicate whether the current interaction evaluation test passes. The scores of each evaluation dimension , , Generated by the large language model according to its internal understanding and evaluation criteria, with a value range of [0, 1].
[0108] In this embodiment, the comprehensive score is calculated as follows.
[0109] .
[0110] .
[0111] .
[0112] Among them, represents the evaluation result output by the large language model evaluation chain, including the scores and opinions of each evaluation dimension, represents the comprehensive score, represents the process of performing evaluation through the large language model evaluation chain, represents the instruction given by the user to the voice interaction system of the intelligent cockpit, R represents the response generated by the voice interaction system of the intelligent cockpit for the instruction , is the prompt template and standard for evaluation, , , respectively represent the scores of semantic correctness, status change confirmation, and unambiguous expression; is the label for the matching degree between the in-vehicle screen image and the current database in the model database; represents the comprehensive score combined with the image matching label , , represents the weight parameter of semantic evaluation and image matching.
[0113] In this embodiment, the validity flag is expressed as the following formula.
[0114] .
[0115] Among them, represents the preset threshold of the comprehensive score.
[0116] In another embodiment, the validity flag can also be expressed as the following formula.
[0117] .
[0118] Among them, , , respectively represent the combination with the image matching label Preset thresholds for comprehensive scores, semantic correctness, status change confirmation, and unambiguous expressions.
[0119] In this embodiment, the scores for each evaluation dimension , , , comprehensive scores , validity flags and improvement suggestions (corresponding to in the code) of the evaluation feedback results are fed back to the development or maintenance process of the voice interaction system of the intelligent cockpit. The evaluation feedback results are used to guide the upgrade and optimization of the voice interaction system of the intelligent cockpit, such as adjusting the response generation logic, improving the user interface prompts, and / or fine-tuning the large language model that supports the response generation. By using this evaluation feedback mechanism for iterative optimization, the response quality, user satisfaction, and task completion efficiency of the interactive voice are continuously improved.
[0120] In an exemplary embodiment, the specific implementation manner of the above step S3 may be: if the matching degree is less than the preset matching degree threshold, calculate the loss function value of the in-vehicle screen image and the current database; if the loss function value is less than the preset value, search for the feedback value of the intelligent cockpit in the current database; if the loss function value is greater than or equal to the preset value, search for the feedback value of the intelligent cockpit in other databases; where the other database and the current database are different databases.
[0121] If the comprehensive score and the validity flag show a matching deviation or the matching deviation is greater than a certain threshold from the actual feedback result of the in-vehicle screen, then the parameters of the large language model are re-modified through the loss function, where the calculation formula of the loss function is as follows.
[0122] .
[0123] Where is the loss function value, is the number of matching samples, is the type of judgment category of the large language model, is the true comprehensive score of the large language model, is the true validity flag of the large language model.
[0124] If the loss function value is less than the preset value, it is determined that the parameters of the large language model are correct.
[0125] If the loss function value is greater than or equal to the preset value, the parameters of the large language model are re-modified.
[0126] In this embodiment, the large language model technology is applied to the intelligent cockpit voice automatic recognition scenario. By feeding the test results back to the voice interaction system of the intelligent cockpit, the voice interaction system automatically uses an evaluation method based on the large language model algorithm to evaluate the response performance and generate an evaluation feedback result. This evaluation feedback result can be fed back to the development or maintenance process of the voice interaction system of the intelligent cockpit, so as to guide the upgrade and optimization of the voice interaction system of the intelligent cockpit, thereby helping to improve the recognition rate of interactive speech, the response quality of interactive speech, user satisfaction, and task completion efficiency.
[0127] Next, the implementation process of this application will be specifically described: This application selects the required test database, outlines the recognition area of the high-definition camera, performs volume normalization processing on the test corpus in the test database, and transmits the processed data to the artificial mouth for playback.
[0128] The volume normalization is processed by the average volume normalization method, which is expressed as follows.
[0129] 。
[0130] Among them, is the number of samples, is the amplitude value of the th sample, and
[0131] represents the average volume level of the sample volume.
[0132] 。
[0133] Among them, represents the desired average volume level.
[0134] In this embodiment, by using the normalization method to normalize the volume of the test speech, the test accuracy and the speech recognition effect can be effectively improved.
[0135] In this embodiment, after receiving the voice command transmitted from the artificial mouth, the voice interaction system of the intelligent cockpit is awakened, recognizes the voice command, executes the command content, and the recognized command content is displayed on the screen.
[0136] In this embodiment, through real-time monitoring by the high-definition camera, the text or image information within the outlined area is obtained. The voice interaction system of the intelligent cockpit is awakened, for example, to obtain image information, or to recognize voice commands, for example, to obtain text information.
[0137] Specifically, the captured image is grayscale processed to obtain the grayscale value Gray , which is expressed as the following formula.
[0138] .
[0139] Among them, , , are the values of the red, green, and blue color channels respectively.
[0140] Then, each pixel in the grayscale image is binarized according to the grayscale, which is expressed by the following formula.
[0141] .
[0142] Among them, represents the binarized image, is a preset threshold, 255 represents white, and 0 represents black.
[0143] Denoise the binarized image, specifically using a Gaussian filter, which is expressed as follows.
[0144] .
[0145] Among them, x and y represent the abscissa and ordinate distances from the center respectively, is the standard deviation of the Gaussian distribution, which controls the width of the kernel.
[0146] Adjust the size of the denoised image so that all input images have the same size, which is expressed as follows.
[0147] .
[0148] Among them, is the output image, is the input image, is the mean value of the image pixel values.
[0149] Extract features from the resized image to extract information helpful for character recognition from the processed image. For example, use the Sobel operator for edge detection, which is expressed as follows.
[0150] .
[0151] .
[0152] .
[0153] .
[0154] Among them, represents the image matrix, and respectively represent the gradient matrices in the horizontal and vertical directions at the point where the image is calculated, represents the gradient magnitude at the point where the image is calculated, represents the gradient direction at the point where the image is calculated.
[0155] Identify the image after edge detection. The specific identification process is implemented using a convolutional neural network, automatically learning features, and introducing a language model to process grammar or dictionary correction, as well as possible semantic analysis to improve the accuracy of identification.
[0156] Capture the text or image captured by the high-definition camera, and obtain the corresponding recognition result. Compare the recognition result with the text truth or image truth stored locally to determine whether the recognition result is consistent with the truth. If it is consistent, it represents successful recognition; if not, it represents failed recognition.
[0157] Finally, based on multiple performance indicators such as the determined speech command recognition rate, generate a test report.
[0158] In an exemplary embodiment, the specific implementation manner of the above step S2 may be: obtaining the interaction information of the intelligent cockpit based on the test voice; wherein, the interaction information includes voice information and / or text information displayed on the in-vehicle screen of the intelligent cockpit; correspondingly, the specific implementation manner of step S3 may be: identifying the interaction information to obtain the feedback information of the intelligent cockpit; correspondingly, the specific implementation manner of step S4 may be: comparing the feedback information with the standard interaction information corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0159] Specifically, after playing the test voice, this application obtains the interaction information of the intelligent cockpit, specifically including the interaction voice of the voice interaction system or the interaction text displayed on the in-vehicle screen. This application obtains the interaction voice of the voice interaction system through a microphone, and obtains the interaction text of the voice interaction system through a high-definition camera, and identifies the interaction voice or interaction text to obtain the interaction information of the voice interaction system of the intelligent cockpit. By comparing the interaction information with the standard interaction information corresponding to the test voice, the interaction performance of the voice interaction system of the intelligent cockpit, that is, the interaction accuracy, is judged, and by counting the interaction accuracies corresponding to multiple test voices, the interaction accuracy rate of the voice interaction system of the intelligent cockpit is calculated, that is, the number of test voices with correct interactions and the total number of test voices played.
[0160] In an exemplary embodiment, the specific implementation of step S2 may be: obtaining the response moment of the intelligent cockpit based on the test voice; correspondingly, the specific implementation of step S3 may be: calculating the response duration of the intelligent cockpit based on the test voice; where the response duration represents the duration from the start of playing the test voice to the response of the intelligent cockpit; correspondingly, the specific implementation of step S4 may be: comparing the response duration with a preset time threshold to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0161] Specifically, in this application, the playback moment is recorded when playing the test voice, and after playing the test voice, the response information of the intelligent cockpit (including moments such as wake-up, recognition, and interaction) is obtained to determine its response moment, and the response duration of the intelligent cockpit is calculated based on the playback moment and the response moment. By comparing the response duration with a preset time threshold, the response speed index of the voice interaction system of the intelligent cockpit is obtained. This application can obtain multiple response durations for multiple test voices and then calculate the weighted average to obtain the response duration of the voice interaction system of the intelligent cockpit, or different response durations can be obtained for different instructions.
[0162] Based on the same inventive concept, an embodiment of this application also provides an intelligent cockpit voice automatic recognition device for implementing the intelligent cockpit voice automatic recognition method involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the intelligent cockpit voice automatic recognition device provided below can refer to the limitations on the intelligent cockpit voice automatic recognition method in the above text and will not be elaborated here.
[0163] Figure 2 Fig. shows a schematic structural diagram of an intelligent cockpit voice automatic recognition device. As Figure 2 shown, the intelligent cockpit voice automatic recognition device includes the following modules.
[0164] A test voice playback module N1 for playing test voices.
[0165] A response information acquisition module N2 for acquiring the response information of the intelligent cockpit based on the test voice; where the response information includes the image of the in-vehicle screen of the intelligent cockpit.
[0166] A feedback result recognition module N3 for recognizing the response information to obtain the feedback value of the intelligent cockpit; where the feedback value represents the recognition index of the interaction performance of the intelligent cockpit in response to the test voice.
[0167] A recognition result calculation module N4 for comparing the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0168] In this embodiment, the test voice is played through the test voice playback module N1; the response information acquisition module N2 acquires the response information of the intelligent cockpit based on the test voice; the feedback result recognition module N3 recognizes the response information to obtain the feedback value of the intelligent cockpit; wherein, the response information includes the in-vehicle screen image of the intelligent cockpit, and the feedback value represents the recognition index of the interaction performance of the intelligent cockpit in response to the test voice; the recognition result calculation module N4 compares the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit; that is, by playing the test voice and acquiring the response information of the intelligent cockpit based on the test voice, recognizing the response information to obtain the feedback value of the intelligent cockpit, and comparing the feedback value with the standard value corresponding to the test voice, the recognition result of the voice interaction performance of the intelligent cockpit can be obtained, so as to accurately recognize the voice interaction performance of the intelligent cockpit and utilize the automatic test of the test voice to improve the test efficiency, thereby quickly recognizing the voice interaction performance.
[0169] In an exemplary embodiment, the in-vehicle screen image includes the response interface of the intelligent cockpit based on the test voice; wherein, the above-mentioned feedback result recognition module N3 can be further configured to: recognize the in-vehicle screen image to obtain the wake-up state of the voice interaction system of the intelligent cockpit; wherein, the wake-up state includes successful wake-up and non-wake-up; correspondingly, the recognition result calculation module N4 can be further configured to: count the wake-up state of the voice interaction system of the intelligent cockpit to obtain the wake-up rate of the voice interaction system of the intelligent cockpit; compare the wake-up rate of the voice interaction system of the intelligent cockpit with a preset wake-up rate threshold to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0170] In an exemplary embodiment, the in-vehicle screen image includes the text information obtained by the intelligent cockpit recognizing the test voice; wherein, the above-mentioned feedback result recognition module N3 can be further configured to: recognize the in-vehicle screen image to obtain the text information obtained by the intelligent cockpit recognizing the test voice; correspondingly, the recognition result calculation module N4 can be further configured to: compare the text information obtained by the intelligent cockpit recognizing the test voice with the original text of the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0171] In an exemplary embodiment, the above-mentioned feedback result recognition module N3 can be further configured to: calculate the matching degree between the in-vehicle screen image and the current database in the model database; wherein, the model database includes databases corresponding to the response interface of the voice interaction system of the intelligent cockpit and the original text of the test voice respectively; search for the feedback value of the intelligent cockpit in one of the databases in the model database based on the matching degree.
[0172] In an exemplary embodiment, the above-mentioned feedback result recognition module N3 may be further configured as follows: if the matching degree is less than a preset matching degree threshold, calculate the loss function value of the in-vehicle computer screen image and the current database; if the loss function value is less than a preset value, search for the feedback value of the intelligent cockpit in the current database; if the loss function value is greater than or equal to the preset value, search for the feedback value of the intelligent cockpit in other databases; wherein, the other database and the current database are different databases.
[0173] In an exemplary embodiment, the above-mentioned response information acquisition module N2 may be further configured as follows: acquire the interaction information of the intelligent cockpit based on the test voice; wherein, the interaction information includes voice information and / or text information displayed on the in-vehicle computer screen of the intelligent cockpit; correspondingly, the feedback result recognition module N3 may be further configured as follows: recognize the interaction information to obtain the feedback information of the intelligent cockpit; correspondingly, the recognition result calculation module N4 may be further configured as follows: compare the feedback information with the standard interaction information corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0174] In an exemplary embodiment, the above-mentioned response information acquisition module N2 may be further configured as follows: acquire the response moment of the intelligent cockpit based on the test voice; correspondingly, the feedback result recognition module N3 may be further configured as follows: calculate the response duration of the intelligent cockpit based on the test voice; wherein, the response duration represents the duration from the playback of the test voice to the response of the intelligent cockpit; correspondingly, the recognition result calculation module N4 may be further configured as follows: compare the response duration with a preset time threshold to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
[0175] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0176] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An intelligent cockpit voice automatic recognition method, characterized in that, The intelligent cockpit voice automatic recognition method includes: Playing a test voice; Obtaining response information of the intelligent cockpit based on the test voice; wherein, the response information includes the in-vehicle screen image of the intelligent cockpit; Identifying the response information to obtain a feedback value of the intelligent cockpit; wherein, the feedback value represents an identification index of the interaction performance of the intelligent cockpit in response to the test voice; Comparing the feedback value with a standard value corresponding to the test voice to obtain an identification result of the voice interaction performance of the intelligent cockpit.
2. The intelligent cockpit voice automatic recognition method according to claim 1, wherein The intelligent cockpit voice automatic recognition method further includes: Adopting an evaluation method based on a large language model algorithm to evaluate the response performance of the voice interaction system of the intelligent cockpit and generating an evaluation feedback result; Feeding back the evaluation feedback result to the development or maintenance process of the voice interaction system of the intelligent cockpit to guide the upgrade and optimization of the voice interaction system of the intelligent cockpit; the upgrade and optimization include adjusting the response generation logic, improving the user interface prompt, and / or fine-tuning the large language model supporting the response generation.
3. The intelligent cockpit voice automatic recognition method according to claim 2, wherein Adopting an evaluation method based on a large language model algorithm to evaluate the response performance of the voice interaction system of the intelligent cockpit and generating an evaluation feedback result, specifically including: Receive the instruction issued by the user to the voice interaction system of the intelligent cockpit ; Receive the response generated by the voice interaction system of the intelligent cockpit for the instruction generated ; Process the instruction using a pre-configured large language model evaluation chain and the response to obtain an evaluation feedback result.
4. The intelligent cockpit voice automated recognition method according to claim 3, wherein Identifying the response information to obtain a feedback value of the intelligent cockpit, specifically including: Calculate the matching degree between the in-vehicle infotainment (IVI) screen image and the current database in the model database; wherein, the model database includes databases corresponding to the response interface of the voice interaction system of the intelligent cockpit and the original text of the test voice respectively; the label of the matching degree between the IVI screen image and the current database in the model database The smaller the value is, the higher the matching degree between the IVI screen image and the current database in the model database is, and the label The calculation formula is as follows: ; ; Among them, Y a and Y b respectively represent the in-vehicle infotainment (IVI) screen image data and the data in the current database, and respectively represent f the IVI screen image data of the feature dimension and the data in the current database, F is the feature dimension of the current database data, X is Y a and Y b the distance between, is a preset threshold for measuring the matching degree.
5. The intelligent cockpit voice automatic recognition method according to claim 3 or 4, characterized in that The large language model evaluation chain includes an LLM model and a prompt template. Among them, the prompt template is used to guide the evaluation, and the LLM model is used to evaluate the response according to the evaluation dimensions defined in the prompt template to generate the evaluation feedback result; The evaluation dimensions include semantic correctness, status change confirmation, and unambiguous expression; The evaluation feedback results include the scores and opinions of each of the evaluation dimensions, and the comprehensive score and the validity flag , wherein the comprehensive score is used to represent the comprehensive score situation of all the evaluation dimensions, and the validity flag is used to indicate whether the current interactive evaluation test passes; The comprehensive score is calculated by the following formula: ; ; ; Among them, represents the evaluation results output by the large language model evaluation chain, including the scores and opinions of each evaluation dimension, represents the comprehensive score, represents the process of performing evaluation through the large language model evaluation chain, represents the instruction issued by the user to the voice interaction system of the intelligent cockpit, R represents the voice interaction system of the intelligent cockpit for the instruction generated response, is the prompt template and standard for evaluation, 、 、 respectively represent the scores of semantic correctness, status change confirmation, and unambiguous expression; is the label of the matching degree between the in-vehicle screen image and the current database in the model database; represents the combined image matching label comprehensive score, , represents the weight parameter of semantic evaluation and image matching; The validity flag has the following calculation formula: ; Among them, represents a preset threshold for the comprehensive score; Alternatively, the validity flag is calculated as follows: ; Among them, , , respectively represent the preset thresholds of the comprehensive score combining the image matching label , the preset threshold of semantic correctness, the preset threshold of status change confirmation, and the preset threshold of unambiguous expression; If the comprehensive score and the validity flag show a matching deviation or a matching deviation greater than the threshold from the actual feedback result of the in-vehicle screen, the parameters of the large language model are re-modified through the loss function. The calculation formula of the loss function is as follows: ; Among them, is the loss function value, is the number of matching samples, is the number of categories judged by the large language model, is the true comprehensive score of the large language model, is the true validity flag of the large language model; If the loss function value is less than a preset value, it is determined that the large language model parameters are correct; If the loss function value is greater than or equal to the preset value, the large language model parameters are modified again.
6. The intelligent cockpit voice automatic recognition method according to claim 1, characterized in that, The in-vehicle screen image includes the response interface of the intelligent cockpit based on the test voice; Identifying the response information to obtain a feedback value of the intelligent cockpit, specifically including: Identifying the in-vehicle screen image to obtain the wake-up state of the voice interaction system of the intelligent cockpit; wherein, the wake-up state includes successful wake-up and non-wake-up; Comparing the feedback value with a standard value corresponding to the test voice to obtain an identification result of the voice interaction performance of the intelligent cockpit, specifically including: Counting the wake-up state of the voice interaction system of the intelligent cockpit to obtain the wake-up rate of the voice interaction system of the intelligent cockpit; Comparing the wake-up rate of the voice interaction system of the intelligent cockpit with a preset wake-up rate threshold to obtain an identification result of the voice interaction performance of the intelligent cockpit.
7. The intelligent cockpit voice automatic recognition method according to claim 1, wherein, The in-vehicle screen image includes the text information obtained by the intelligent cockpit recognizing the test voice; Identifying the response information to obtain a feedback value of the intelligent cockpit, specifically including: Identifying the in-vehicle screen image to obtain the text information obtained by the intelligent cockpit recognizing the test voice; Comparing the feedback value with a standard value corresponding to the test voice to obtain an identification result of the voice interaction performance of the intelligent cockpit, specifically including: Comparing the text information obtained by the intelligent cockpit recognizing the test voice with the original text of the test voice to obtain an identification result of the voice interaction performance of the intelligent cockpit.
8. The intelligent cockpit voice automatic recognition method according to claim 1, wherein Obtaining response information of the intelligent cockpit based on the test voice, specifically including: Obtain the interaction information of the intelligent cockpit based on the test voice; wherein, the interaction information includes voice information and / or text information displayed on the in-vehicle screen of the intelligent cockpit; Identify the response information to obtain the feedback value of the intelligent cockpit, specifically including: Identify the interaction information to obtain the feedback information of the intelligent cockpit; Compare the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit, specifically including: Compare the feedback information with the standard interaction information corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
9. The intelligent cockpit voice automatic recognition method according to claim 1, wherein Obtain the response information of the intelligent cockpit based on the test voice, specifically including: Obtain the response moment of the intelligent cockpit based on the test voice; Identify the response information to obtain the feedback value of the intelligent cockpit, specifically including: Calculate the response duration of the intelligent cockpit based on the test voice; wherein, the response duration represents the duration from the playback of the test voice to the response of the intelligent cockpit; Compare the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit, specifically including: Compare the response duration with a preset time threshold to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
10. An intelligent cockpit voice automatic recognition device, characterized in that, The intelligent cockpit voice automatic recognition device includes: A test voice playback module for playing test voices; A response information acquisition module for acquiring the response information of the intelligent cockpit based on the test voice; wherein, the response information includes the in-vehicle screen image of the intelligent cockpit; A feedback result recognition module for identifying the response information to obtain the feedback value of the intelligent cockpit; wherein, the feedback value represents the recognition index of the interaction performance of the intelligent cockpit in response to the test voice; A recognition result calculation module for comparing the feedback value with the standard value corresponding to the test voice to obtain the recognition result of the voice interaction performance of the intelligent cockpit.
Citation Information
Patent Citations
Vehicle-mounted control screen voice recognition process testing method, electronic equipment and system
CN109616106A
Vehicle voice function test method and device, electronic equipment and storage medium
CN115579025A
Automatic driving target detection system safety test method based on semantic perception
CN117830769A
Large model test method and device, electronic equipment and storage medium
CN117971661A
Test method, system and device of intelligent voice interaction system and medium
CN118135998A