Optimization method and device for vehicle intelligent cabin voice wake-up model and computer program product
By collecting and analyzing multimode data, forming a false wake-up training data set and optimizing a voice wake-up model, the problems of high false wake-up rate and insufficient data representation in the existing technology are solved, and higher speech recognition accuracy and user experience are achieved.
Patent Information
- Application Number
- CN202510021383.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The existing smart cockpit voice wake-up technology has problems in practical applications such as high false wake-up rate, insufficient representation of training data, and lack of in-depth data analysis, resulting in poor accuracy and user experience of the speech recognition system.
By collecting multi-mode data on user behavior, audio data in the cockpit and environmental data, analyzing and extracting structured data related to speech error wake-up, forming a false wake-up training data set, and importing it into the speech wake-up model for iterative training to optimize the speech wake-up model.
It significantly reduces the voice error wake-up rate, improves the accuracy and user experience of the speech recognition system, makes the voice wake-up model better adapt to actual usage scenarios, and enhances the model's self-learning and continuous optimization capabilities.
Smart Images

Figure CN119964566A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent connected vehicles, and in particular to an optimization method, device and computer program product for a vehicle intelligent cockpit voice wake-up model. Background Art
[0002] In the current field of voice wake-up technology for smart cockpits, the mainstream method focuses on the matching analysis of audio frequencies, that is, activating the wake-up function by comparing the audio signal emitted by the user with the audio frequency characteristics of the preset wake-up word. However, this technical path has exposed significant limitations in practical applications, and it ignores more complex and critical auxiliary judgment conditions, such as the understanding of the context and the verification of the authenticity of the sound source.
[0003] Specifically, the current technical solutions usually follow the following process: First, effective audio detection is performed through VW-VAD (Voice Wake-up Voice Activity Detection) technology to eliminate silent segments and non-wake-up related audio to optimize the allocation of computing resources; then, the feature extraction stage is entered, and the FBank (Filter Bank) algorithm is used to extract the frequency band or frequency sampling features of the audio to provide input data for the acoustic model; then, the acoustic model is responsible for classifying the pronunciation status of every 25 millisecond audio frame; after that, through the decoding process, a decoding path is constructed based on the wake-up word to evaluate the match between the audio and the wake-up word; finally, a confidence model is introduced for secondary verification of the wake-up result.
[0004] Although the above solution realizes the voice wake-up function to a certain extent, it still has the following defects:
[0005] Limitations of training data: The training of the wake-up model is highly dependent on a large amount of test data in a laboratory environment. These data deviate greatly from the complex and changing conditions in real car usage scenarios, resulting in insufficient representativeness and generalization of the training data.
[0006] Single perspective on problem solving: Faced with the complex problem of false voice wake-up, existing technologies often only seek solutions from a single factor, ignoring the combined impact of multiple possible factors, limiting the improvement of false wake-up suppression effects.
[0007] Lack of in-depth data analysis: Few technical solutions conduct in-depth data analysis on the specific reasons that lead to false voice wake-up, and fail to accurately identify the root cause of the problem, making it difficult to formulate targeted optimization strategies to fundamentally solve the problem of false wake-up. Summary of the invention
[0008] The technical problem to be solved by the embodiments of the present invention is to provide a method, device and computer program product for optimizing the vehicle intelligent cockpit voice wake-up model, so as to effectively reduce the voice false wake-up rate and improve the accuracy of the voice recognition system and the user experience.
[0009] In order to solve the above technical problems, the present invention provides a method for optimizing a vehicle intelligent cockpit voice wake-up model, comprising the following steps:
[0010] According to the trigger rules for voice false wake-up data collection, collect multi-mode data of user behavior, in-cabin audio data, and environmental data;
[0011] Parse the collected data to obtain structured data related to voice false wake-up;
[0012] Extracting data associated with voice false wakeup from the structured data;
[0013] Classifying data associated with voice false wake-up to form a false wake-up training data set;
[0014] The false wake-up training data set is imported into the voice wake-up model for iterative training to optimize the voice wake-up model.
[0015] Preferably, the collecting of multi-mode user behavior data, in-cabin audio data and environmental data according to the voice false wake-up data collection triggering rule specifically includes:
[0016] Detect and respond to voice wake-up signals, and identify the wake-up voice to confirm whether it is a valid wake-up command;
[0017] Compare the recognized wake-up voice with the homophone database to determine whether the comparison is successful;
[0018] If the comparison is successful, the next round of voice interaction will be monitored; if the comparison is unsuccessful, it will be determined as a suspected false wake-up, and the collection of multi-modal user behavior data, in-cabin audio data, and environmental data will be triggered;
[0019] If no valid wake-up command is recognized during the next round of voice interaction, it will be judged as a suspected false wake-up, and trigger the collection of multi-modal user behavior data, in-cabin audio data, and environmental data.
[0020] Preferably, the structured data suitable for machine learning specifically includes user feature data and wake-up environment data.
[0021] Preferably, extracting data associated with voice false wakeup from the structured data specifically includes:
[0022] Selecting and extracting features related to voice false wakeup from the structured data;
[0023] Use association rule mining algorithm to analyze the correlation between the extracted features and voice false wake-up;
[0024] Based on the association relationship, data directly associated with the voice false wake-up is mined.
[0025] Preferably, the classifying the data associated with the voice false wakeup to form a false wakeup training data set specifically includes:
[0026] Determine the cause of the false wake-up by analyzing the features related to the false wake-up of voice in the structured data;
[0027] According to the cause of false wake-up, a false wake-up classification algorithm is used to classify the data associated with voice false wake-up mined from structured data into non-false wake-up, false touch, echo self-excitation, human voice wake-up, external sound source, and non-effective sound wake-up.
[0028] Preferably, the method further comprises:
[0029] A secondary confidence algorithm module is set in the voice wake-up model, which is used to calculate the secondary confidence of the current voice wake-up signal using the secondary confidence algorithm when the wake-up score of the current voice wake-up signal exceeds the preset wake-up threshold value;
[0030] The secondary confidence is used to compare with a preset secondary confidence threshold value. If the secondary confidence exceeds the preset secondary confidence threshold value, it is determined that the wake-up is successful.
[0031] Preferably, the step of calculating the secondary confidence of the current voice wake-up signal using the secondary confidence algorithm specifically includes:
[0032] A confidence score that the current voice wake-up signal is a valid wake-up command is calculated using a trained computing model that is trained using features extracted from the false wake-up training data set that reflect audio content, environmental conditions, and user characteristics.
[0033] The present invention also provides an optimization device for a vehicle intelligent cockpit voice wake-up model, comprising:
[0034] The collection module is used to collect multi-mode data of user behavior, in-cabin audio data and environmental data according to the trigger rules of voice false wake-up data collection;
[0035] The parsing module is used to parse the collected data and obtain structured data related to voice false wake-up;
[0036] An extraction module, configured to extract data associated with voice false wakeup from the structured data;
[0037] A classification module, used to classify data associated with voice false wake-up to form a false wake-up training data set;
[0038] The optimization module is used to import the false wake-up training data set into the voice wake-up model for iterative training to optimize the voice wake-up model.
[0039] The present invention also provides an optimization device for a vehicle intelligent cockpit voice wake-up model, comprising:
[0040] one or more processors;
[0041] Memory;
[0042] One or more applications, wherein the one or more applications are stored in the memory and are configured to be executed by the one or more processors, and the one or more applications are configured to execute the optimization method of the vehicle intelligent cockpit voice wake-up model.
[0043] The present invention also provides a computer program product, comprising computer instructions, wherein the computer instructions instruct a computer device to execute operations corresponding to the method.
[0044] The implementation of the present invention has the following beneficial effects: the present invention provides a rich data basis for model optimization by comprehensively collecting multi-modal data such as user behavior, audio and environment in the cockpit; the false awakening data is imported into the model as a counter-example for iterative training, which not only optimizes the model algorithm, but also significantly reduces the false awakening rate and improves the accuracy of awakening. The present invention enables the voice awakening model to better adapt to actual usage scenarios, reduce unnecessary interference, and bring users a more targeted voice interaction experience. At the same time, it enhances the model's self-learning and continuous optimization capabilities, providing strong support for the intelligent development of vehicle smart cockpits. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0046] Figure 1 It is a flow chart of an optimization method of a vehicle intelligent cockpit voice wake-up model according to an embodiment of the present invention.
[0047] Figure 2 It is a schematic diagram of the process of data collection in an embodiment of the present invention.
[0048] Figure 3It is a schematic diagram of the secondary confidence judgment process in an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The following descriptions of the embodiments refer to the accompanying drawings to illustrate specific embodiments in which the present invention may be implemented.
[0050] Please refer to Figure 1 As shown, the first embodiment of the present invention provides a method for optimizing a vehicle intelligent cockpit voice wake-up model, comprising the following steps:
[0051] According to the trigger rules for voice false wake-up data collection, collect multi-mode data of user behavior, in-cabin audio data, and environmental data;
[0052] Parse the collected data to obtain structured data suitable for machine learning;
[0053] Extracting data associated with voice false wakeup from the structured data;
[0054] Classifying data associated with voice false wake-up to form a false wake-up training data set;
[0055] The false wake-up training data set is imported into the voice wake-up model for iterative training to optimize the voice wake-up model.
[0056] Through the above steps, it can be seen that the present invention provides a rich data foundation for model optimization by comprehensively collecting multi-modal data such as user behavior, audio and environment in the cockpit; the false awakening data is imported into the model as a counterexample for iterative training, which not only optimizes the model algorithm, but also significantly reduces the false awakening rate and improves the accuracy of awakening. The present invention enables the voice awakening model to better adapt to actual usage scenarios, reduce unnecessary interference, and bring users a more targeted voice interaction experience. At the same time, it enhances the model's self-learning and continuous optimization capabilities, providing strong support for the intelligent development of vehicle smart cockpits.
[0057] Specifically, please combine Figure 2 As shown, in the embodiment of the present invention, according to the voice false wake-up data collection triggering rule, the specific process of collecting user behavior multi-mode data, in-cabin audio data and environmental data is as follows:
[0058] First, the voice wake-up signal is detected and responded to, and then the wake-up voice is recognized to confirm whether it is a valid wake-up command.
[0059] The recognition result is then compared with the homophone database to determine whether the comparison is successful. If the comparison is successful, the next round of voice interaction will continue to be monitored; if the comparison is unsuccessful, it will be determined as a suspected false wake-up and the data collection mechanism will be triggered.
[0060] It is understandable that in an embodiment of the present invention, the purpose of comparing the recognized voice with the homophone library is mainly to verify whether the recognized wake-up voice matches the preset wake-up word or instruction word. If the comparison is successful, it means that the voice issued by the user matches a valid wake-up word or instruction word in the system. This usually means that the user has successfully interacted with the system through voice commands, and the system should perform corresponding operations according to the recognized instructions. If the comparison is unsuccessful, that is, the wake-up text similarity is lower than the preset threshold (the similarity between the recognized wake-up text and the valid wake-up word is too low), it may mean that the voice issued by the user does not match any valid wake-up word or instruction word in the system. This may be because the user did not issue the correct wake-up word, or the speech recognition system failed to accurately recognize the user's voice. Therefore, it will be judged as a false wake-up, and then data collection will be triggered to collect relevant data for subsequent analysis and model optimization.
[0061] If no valid wake-up command is recognized during the next round of voice interaction, it is also determined as a suspected false wake-up and the data collection mechanism is triggered. Of course, if the next round of voice interaction is still valid, the normal working state is maintained and the user command is continued to be processed. In other words, the voice false wake-up data collection triggering rules of the embodiment of the present invention specifically include: the wake-up text similarity is lower than the preset threshold or the user has no valid wake-up command in the next round of voice interaction.
[0062] As shown in Table 1 below, it shows how to collect key data through data collection strategies for the problem of false wake-up in smart cockpit voice wake-up. Table 1 lists several key influencing factors and data collection strategies for these influencing factors.
[0063] Table 1 Key factors affecting voice wake-up and their corresponding data embedding strategies
[0064]
[0065] Background noise is affected by both the interior of the cabin (such as audio playback, air conditioning, windows) and the external environment, so data needs to be collected from multiple angles, including the equipment inside the cabin (such as audio playback, air conditioning, windows) and the external environment. This data helps to identify and quantify the impact of background noise on speech recognition accuracy.
[0066] User characteristics include the user's voice characteristics (such as accent, speaking speed, tone) and behavioral characteristics. By collecting and analyzing on-site audio data and combining it with the anonymized data obtained by DMS technology, the user's voice characteristics can be analyzed to assist in determining the false triggering of the wake-up system and improve the user experience.
[0067] Scenario-based recognition refers to optimizing the accuracy of speech recognition based on the specific scenario the user is in. By collecting data for the next round of interactions, the system can better understand the user's intentions and needs, thereby improving the accuracy of recognition and the appropriateness of responses.
[0068] The wake-up behavior characteristics involve the user's behavior pattern of waking up the smart cockpit system. By collecting data on the wake-up scene, we can analyze the user's wake-up habits, the frequency of use of wake-up words, etc., so as to optimize the wake-up algorithm and improve the system's response speed.
[0069] As can be seen from the above, the data collection mechanism of the embodiment of the present invention specifically includes the following aspects:
[0070] User behavior multi-modal data collection: Capture and record various user behavior data, such as key operations, gestures, etc., to analyze the correlation between user behavior patterns and false wake-ups.
[0071] In-cabin audio data collection: includes collecting current audio data to record voice features during false wake-up.
[0072] In-cabin environmental data collection: including collecting in-cabin environmental data, such as noise, temperature, etc., to evaluate the impact of environmental factors on speech recognition accuracy.
[0073] Through the above process, the present invention can distinguish between valid user commands and invalid voice inputs, respond quickly when suspected false wake-up is detected, comprehensively collect relevant data, and provide data support for subsequent reporting to the cloud platform for in-depth analysis, thereby continuously optimizing the accuracy of voice wake-up and user experience.
[0074] The collected data can be reported to the cloud platform, where it is parsed and processed to obtain structured data suitable for machine learning. The specific process is as follows:
[0075] (1) Audio data sampling
[0076] Since the collected audio data is a continuous analog signal, the computer cannot process it directly. Therefore, a suitable sampling rate (such as 16kHz, 44.1kHz, etc.) is used to sample the audio signal to convert the continuous audio waveform into a series of discrete numerical points.
[0077] (2) Audio data format conversion
[0078] Convert audio data into a format suitable for processing. The original audio file may be stored in different formats such as WAV, MP3, AAC, etc. In order to process uniformly, these data need to be converted into a common format, such as PCM (Pulse Code Modulation). During the conversion process, the sampling rate, bit depth and other parameters of the audio data must be kept consistent to ensure the accuracy of subsequent processing.
[0079] (3) Audio signal preprocessing
[0080] Audio signals may contain interference factors such as noise and distortion, and preprocessing is required to improve the quality of the audio signals and improve the accuracy of subsequent feature extraction.
[0081] Preprocessing operations include:
[0082] Noise reduction: Use noise reduction algorithms (such as spectral subtraction, Kalman filtering, etc.) to reduce background noise and other interference and improve the clarity of speech signals; Frequency response adjustment: Adjust the frequency characteristics of audio signals to suit specific processing requirements.
[0083] Frequency response adjustment: Adjust the frequency response of the audio signal through filters (such as high-pass filter, low-pass filter, band-pass filter, etc.) to adapt to specific processing requirements.
[0084] Other preprocessing operations: such as audio enhancement, volume normalization, etc., to further improve the quality of the audio signal.
[0085] (4) Feature extraction
[0086] Extract features from the preprocessed audio signal that are helpful for subsequent processing or analysis, such as speech rate, pitch, voiceprint, etc. Specifically, it includes the following aspects:
[0087] Speech rate extraction: measures the rate of speech, that is, the number of syllables spoken per unit time.
[0088] Pitch extraction: Use pitch detection algorithms (such as autocorrelation function method, Fourier transform method, etc.) to extract pitch information from audio signals.
[0089] Voiceprint extraction: Voiceprint features in audio signals are extracted through voiceprint recognition technology (such as Mel-frequency cepstral coefficients MFCC, linear predictive coding LPC, etc.).
[0090] Other feature extraction: such as noise type, background music, etc., to more comprehensively describe the characteristics of the audio signal.
[0091] (5) Feature data processing
[0092] The extracted features are further processed to obtain structured data suitable for machine learning. Specific operations include but are not limited to:
[0093] Data cleaning: Remove outliers and noise from feature data to improve data accuracy.
[0094] Normalization: Scale feature data to a uniform range to reduce scale differences between different features.
[0095] Feature selection: Select the features that are most helpful in distinguishing different speech wakeup events.
[0096] Feature fusion: Multiple features are combined to improve the accuracy of voice wake-up.
[0097] Data encoding: Convert feature data into a format suitable for processing by machine learning models.
[0098] Finally, the processed feature data is stored in a structured form in a database or data file for subsequent analysis and application. These structured data include:
[0099] User characteristic data: such as accent, speaking speed, voiceprint, etc. These data are used for personalized voice recognition and user verification.
[0100] Wake-up environment data: Such as noise levels, this data is used to understand the impact of the wake-up environment on speech recognition accuracy.
[0101] The structured data obtained through the above steps are not all directly related to voice false wake-up. The embodiment of the present invention also needs to mine these structured data to obtain data directly related to voice false wake-up. That is, the data associated with voice false wake-up is mainly extracted by applying algorithms and statistical methods. Specifically, features that may be related to voice false wake-up are selected from the structured data, such as vehicle speed, acceleration, whether the owner is speaking, spectral characteristics of the audio signal in the cabin, ambient noise level, etc. Then, these features are extracted from the original data by applying specific algorithms and statistical methods. For example, signal processing technology can be used to extract the spectral characteristics of the audio signal, or machine learning algorithms can be used to identify the owner's voice commands. Then, algorithms such as association rule mining are used to analyze the correlation between the extracted features and voice false wake-up to help identify which factors are most likely to cause false wake-up events. Finally, data associated with voice false wake-up is mined.
[0102] After obtaining the data associated with voice false wakeup, the embodiment of the present invention adopts a false wakeup classification algorithm to determine the specific cause of false wakeup by analyzing the features related to voice false wakeup, such as audio content, sound source, user behavior, etc., so as to classify the data associated with voice false wakeup mined from the structured data into six types: non-false wakeup, false touch, echo self-excitation, human voice wakeup, external sound source, and non-effective sound wakeup:
[0103] Non-false wakeup: refers to the situation where it is actually a valid wakeup but is mistakenly judged as a false wakeup;
[0104] False touch: refers to the false wake-up caused by the user's unintentional touch or operation;
[0105] Echo self-excitation: refers to false wake-up caused by sound reflection or echo in the cabin;
[0106] Voice wake-up: refers to false wake-up caused by the voice of a non-specific user or other voice interference;
[0107] External sound source: refers to false wake-up caused by music, broadcast, traffic noise and other sound sources outside the cabin;
[0108] Invalid audio wake-up: refers to false wake-up caused by audio that does not match the wake-up word or is of low quality.
[0109] The data of human voice awakening type is further classified into similar words, similar instruction words, and irrelevant words through the false awakening word classification algorithm, and finally the false awakening training data is generated.
[0110] The embodiment of the present invention inputs the generated false wake-up data as counter-example data into the voice wake-up model for iterative training to improve the accuracy and robustness of the voice wake-up model. It should be noted that in the embodiment of the present invention, counter-example data refers to voice data that is mistakenly identified as a wake-up instruction by the voice wake-up model but should not actually trigger wake-up. The counter-example data is contrary to the expected behavior of the voice wake-up model (i.e., responding only to valid wake-up instructions), indicating that the voice wake-up model has defects or deficiencies in recognizing and processing voice wake-up requests.
[0111] Specifically, the counterexample data includes various situations in which the model is falsely awakened, such as awakening due to environmental noise, non-specific human voices, self-excitation of echoes, false touches, or other non-valid sound sources. These false awakening events should not have occurred, but due to the limitations of the model algorithm or the lack of training data, they are mistakenly identified as valid wake-up commands. By collecting these false awakening events as counterexample data and importing them into the voice wake-up model for iterative training, the model can be prompted to learn under what circumstances the wake-up should not be triggered, thereby optimizing its algorithm, improving the accuracy of the wake-up, and reducing the false awakening rate. Counterexample data has therefore become an important basis for model optimization, helping the model identify and correct its previous misjudgments, so that it can respond more accurately to valid wake-up commands in future use, while ignoring interference factors that may cause false awakenings.
[0112] As a further improvement of the embodiment of the present invention, the embodiment of the present invention further adds a secondary confidence confirmation process, that is, by adding an additional confidence judgment step to reduce false wake-up events and improve the accuracy of voice wake-up.
[0113] The secondary confidence algorithm is an algorithm based on machine learning and multi-dimensional data analysis. The algorithm uses the aforementioned counterexample data for learning and combines a variety of related data to calculate the threshold value of the secondary confidence, thereby achieving a more accurate judgment of the wake-up request. The following is the specific process of calculating the threshold value of the secondary confidence:
[0114] After preprocessing the collected counterexample data (including audio signal noise reduction, feature extraction, text transcription, etc.), extract features that can reflect the audio content, environmental conditions, user characteristics, and are associated with false wakeup events;
[0115] Use machine learning algorithms (such as support vector machines, neural networks, etc.) to train the extracted features to build a model that can distinguish between valid awakenings and false awakenings;
[0116] Based on the trained model, predict the data in the validation set or test set and calculate the confidence score of each sample, that is, the probability that the model considers the sample to be a valid awakening;
[0117] According to the relationship between the confidence score and the false awakening rate, the secondary confidence threshold is calculated. The secondary confidence threshold can identify as many valid awakenings as possible while ensuring a low false awakening rate.
[0118] When the obtained secondary confidence threshold is applied to the wake-up judgment, the language wake-up signal is calculated through the secondary confidence algorithm to obtain the corresponding secondary confidence, and then compared with the secondary confidence threshold. If the score exceeds the threshold, it is judged as a valid wake-up and triggers the corresponding response; if the score is lower than the threshold, it is judged as a false wake-up or invalid request and no response is triggered. Please refer to Figure 3 As shown, the wake-up process is described in detail below:
[0119] First, the current voice wake-up signal is collected, and then it is preliminarily processed and recognized locally to determine whether it is a wake-up word.
[0120] The wake-up score is calculated based on the results of the preliminary recognition. The wake-up score is an assessment of the degree of match between the voice input and the wake-up word based on the local algorithm, taking into account multiple factors such as voice clarity, background noise, and the speaker's voiceprint characteristics.
[0121] Determine whether the wake-up score exceeds the preset wake-up threshold: Compare the wake-up score with the preset wake-up threshold. If the wake-up score exceeds the wake-up threshold, proceed to the next step; if the score does not exceed the threshold, it is determined as non-wake-up, the process ends, and the system does not respond.
[0122] Secondary confidence judgment: The current voice wake-up signal is calculated through the secondary confidence algorithm to obtain the corresponding secondary confidence; then it is determined whether the secondary confidence score exceeds the secondary confidence threshold. If the secondary confidence score exceeds the secondary confidence threshold, it is determined to be a successful wake-up; if the secondary confidence score does not exceed the secondary confidence threshold, it is determined to be a false wake-up, the process ends, and the system does not respond.
[0123] The present invention can preliminarily screen out possible wake-ups locally through two-level judgment, and then make more accurate confirmations in the cloud, which can reduce false wake-ups caused by a single judgment standard and improve the accuracy of judgments on wake-up events. In other words, while maintaining a low false wake-up rate, the ability to respond to the user's true intentions is improved, thereby improving the user experience.
[0124] It should also be noted that the optimization method of the vehicle intelligent cockpit voice wake-up model of the embodiment of the present invention can be continuously iterated through vehicle-cloud communication after the vehicle is sold.
[0125] Corresponding to the optimization method of the vehicle intelligent cockpit voice wake-up model described in the first embodiment of the present invention, the second embodiment of the present invention further provides an optimization device for the vehicle intelligent cockpit voice wake-up model, including:
[0126] The collection module is used to collect multi-mode data of user behavior, in-cabin audio data and environmental data according to the trigger rules of voice false wake-up data collection;
[0127] The parsing module is used to parse the collected data and obtain structured data related to voice false wake-up;
[0128] An extraction module, configured to extract data associated with voice false wakeup from the structured data;
[0129] A classification module, used to classify data associated with voice false wake-up to form a false wake-up training data set;
[0130] The optimization module is used to import the false wake-up training data set into the voice wake-up model for iterative training to optimize the voice wake-up model.
[0131] Corresponding to the optimization method of the vehicle intelligent cockpit voice wake-up model described in the first embodiment of the present invention, the third embodiment of the present invention further provides an optimization device for the vehicle intelligent cockpit voice wake-up model, including:
[0132] one or more processors;
[0133] Memory;
[0134] One or more applications, wherein the one or more applications are stored in the memory and are configured to be executed by the one or more processors, and the one or more applications are configured to execute the voice wake-up method for the automobile smart cockpit.
[0135] Corresponding to the optimization method of the vehicle intelligent cockpit voice wake-up model described in the aforementioned embodiment 1 of the present invention, embodiment 4 of the present invention also provides a computer program product, including computer instructions, and the computer instructions instruct a computer device to perform operations corresponding to the method.
[0136] Preferably, the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor. The processor is the control center of the device, and various parts of the device are connected using various interfaces and lines.
[0137] The memory mainly includes a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function, etc., and the data storage area can store related data, etc. In addition, the memory can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, and a flash card (Flash Card), etc., or the memory can also be other volatile solid-state storage devices.
[0138] It should be noted that the above-mentioned device may include but is not limited to a processor and a memory, which can be understood by those skilled in the art.
[0139] For the working principle and process of the above embodiment, please refer to the description of the above embodiment of the present invention, which will not be repeated here.
[0140] It can be seen from the above description that compared with the prior art, the beneficial effects of the present invention are: the present invention provides a rich data basis for model optimization by comprehensively collecting multi-modal data such as user behavior, audio and environment in the cockpit; the false awakening data is imported into the model as a counter-example for iterative training, which not only optimizes the model algorithm, but also significantly reduces the false awakening rate and improves the accuracy of awakening. The present invention enables the voice wake-up model to better adapt to actual usage scenarios, reduce unnecessary interference, and bring users a more targeted voice interaction experience. At the same time, it enhances the model's self-learning and continuous optimization capabilities, providing strong support for the intelligent development of vehicle smart cockpits.
[0141] The above disclosure is only the preferred embodiment of the present invention, which certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.
Claims
1. A method for optimizing a vehicle intelligent cockpit voice wake-up model, characterized in that: The following steps are involved: According to the trigger rules for voice false wake-up data collection, collect multi-mode data of user behavior, in-cabin audio data, and environmental data; Parse the collected data to obtain structured data related to voice false wake-up; Extracting data associated with voice false wakeup from the structured data; Classifying data associated with voice false wake-up to form a false wake-up training data set; The false wake-up training data set is imported into the voice wake-up model for iterative training to optimize the voice wake-up model.
2. The method according to claim 1, characterized in that The method of collecting multi-mode user behavior data, in-cabin audio data, and environmental data according to the voice false wake-up data collection triggering rule specifically includes: Detect and respond to voice wake-up signals, and identify the wake-up voice to confirm whether it is a valid wake-up command; Compare the recognized wake-up voice with the homophone database to determine whether the comparison is successful; If the comparison is successful, the next round of voice interaction will be monitored; if the comparison is unsuccessful, it will be determined as a suspected false wake-up, and the collection of multi-modal user behavior data, in-cabin audio data, and environmental data will be triggered; If no valid wake-up command is recognized during the next round of voice interaction, it will be judged as a suspected false wake-up, and trigger the collection of multi-modal user behavior data, in-cabin audio data, and environmental data.
3. The method according to claim 1, characterized in that The structured data suitable for machine learning specifically includes user feature data and wake-up environment data.
4. The method according to claim 1, characterized in that The extracting data associated with voice false wakeup from the structured data specifically includes: Selecting and extracting features related to voice false wakeup from the structured data; Use association rule mining algorithm to analyze the correlation between the extracted features and voice false wake-up; Based on the association relationship, data directly associated with the voice false wake-up is mined.
5. The method according to claim 4, characterized in that The classifying the data associated with the voice false wake-up to form the false wake-up training data set specifically includes: Determine the cause of the false wake-up by analyzing the features related to the false wake-up of voice in the structured data; According to the cause of false wake-up, a false wake-up classification algorithm is used to classify the data associated with voice false wake-up mined from structured data into non-false wake-up, false touch, echo self-excitation, human voice wake-up, external sound source, and non-effective sound wake-up.
6. The method according to claim 1, characterized in that Also includes: A secondary confidence algorithm module is set in the voice wake-up model, which is used to calculate the secondary confidence of the current voice wake-up signal using the secondary confidence algorithm when the wake-up score of the current voice wake-up signal exceeds the preset wake-up threshold value; The secondary confidence is used to compare with a preset secondary confidence threshold value. If the secondary confidence exceeds the preset secondary confidence threshold value, it is determined that the wake-up is successful.
7. The method according to claim 6, characterized in that The method of using the secondary confidence algorithm to calculate the secondary confidence of the current voice wake-up signal specifically includes: A confidence score that the current voice wake-up signal is a valid wake-up command is calculated using a trained computing model that is trained using features extracted from the false wake-up training data set that reflect audio content, environmental conditions, and user characteristics.
8. An optimization device for a vehicle intelligent cockpit voice wake-up model, characterized in that: include: The collection module is used to collect multi-mode data of user behavior, in-cabin audio data and environmental data according to the trigger rules of voice false wake-up data collection; The parsing module is used to parse the collected data and obtain structured data related to voice false wake-up; An extraction module, configured to extract data associated with voice false wakeup from the structured data; A classification module, used to classify data associated with voice false wake-up to form a false wake-up training data set; The optimization module is used to import the false wake-up training data set into the voice wake-up model for iterative training to optimize the voice wake-up model.
9. An optimization device for a vehicle intelligent cockpit voice wake-up model, characterized in that: include: one or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to execute the optimization method of the vehicle intelligent cockpit voice wake-up model as described in any one of claims 1 to 7.
10. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions instruct a computer device to execute operations corresponding to the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Wake-up word preset confidence threshold adjustment method and system
CN108847219A
Voice wake-up method, storage medium and terminal
CN109256134A
Speech wake-up method and system, electronic equipment and computer readable storage medium
CN110265036A
Method and device for outputting information
CN111640426A
False wake-up corpus determination method and device, electronic equipment and storage medium
CN112233681A