Method and System for Custom Wake-up Word Configuration of In-vehicle Screen Based on Voice Control
By constructing a noise level database and voiceprint feature library, combining real-time noise level and voiceprint features, a wake-up word detection sub-model is constructed, which solves the awakening accuracy and stability of the vehicle-mounted voice wake-up system in complex acoustic environments, and achieves efficient and personalized voice interaction.
Patent Information
- Application Number
- CN202510108179.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The existing vehicle voice wake-up system has unstable wake-up performance in complex acoustic environments, and the custom wake-up words match the user's voiceprints inadequately, which fails to effectively deal with the time-varying and diversity of noise in the car, resulting in a decrease in wake-up accuracy.
Build a noise database and divide the noise levels, collect user-defined wake word voice samples, match vocal print features, build a wake word detection sub-model, combine real-time noise level and vocal print features for wake word recognition, locate the wake word speaker through a microphone array, and set wake-up angle and distance threshold for verification.
It improves the recognition accuracy of wake-up words and system response speed, reduces the false wake-up rate, adapts to various driving environments, improves user experience and optimizes resource usage.
Smart Images

Figure CN119541501B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and more specifically, to a method and system for customizing wake-up words for a vehicle-mounted screen based on voice control. Background Art
[0002] With the rapid development of artificial intelligence technology, voice interaction has become one of the important ways of human-computer interaction. Especially in the field of intelligent driving, controlling in-vehicle devices through voice commands can maximize driving safety. An in-vehicle voice assistant can complete tasks such as navigation, music playback, and air-conditioning control through voice interaction when the driver focuses on the road conditions. Among them, voice wake-up is a key link in realizing in-vehicle voice interaction.
[0003] In existing in-vehicle voice wake-up systems, the wake-up is mainly triggered by using a preset wake-up word or a user-defined wake-up word. For example, the patent application with the publication number CN113611294A discloses a voice wake-up method that executes wake-up when the user's voice matches a combined wake-up word by configuring multiple combined wake-up words. This method supports preset wake-up words, user-defined wake-up words, and multi-wake-up word wake-up, but does not consider the influence of complex acoustic environments such as in-vehicle noise on wake-up performance. Another example is the patent with the authorization announcement number CN109360552B, which proposes a method for automatically filtering wake-up words. By comparing the user's voice with the wake-up word audio, meaningless wake-up words in the user's voice are obtained and blocked, thereby improving the accuracy of semantic parsing. This method focuses on the semantic understanding stage after wake-up, but is prone to false wake-up in a complex acoustic environment.
[0004] In summary, existing voice wake-up technologies do not fully consider the influence of in-vehicle noise on wake-up performance, and the matching degree between user-defined wake-up words and the user's voiceprint is insufficient; in-vehicle noise is time-varying and diverse, and the wake-up system needs to maintain stable wake-up performance in various complex noise environments, while existing methods generally lack an adaptation mechanism for in-vehicle noise; factors such as the user's voice quality, speech rate, and emotion will all affect the acoustic characteristics of user-defined wake-up words. If there is a large difference between the user-defined wake-up word model and the user's actual pronunciation, the wake-up accuracy will decrease; existing user-defined wake-up word methods usually only train a unified acoustic model and lack the ability to model the user's personalized voiceprint characteristics. Summary of the Invention
[0005] To overcome the above-mentioned defects of the prior art, the present invention provides a method and system for configuring a custom wake-up word for a vehicle-mounted screen based on voice control. The method first collects in-vehicle noise data to construct a noise database, and divides the noise records into multiple noise levels; then obtains a custom wake-up word voice sample, matches and obtains the user's voiceprint feature; and then constructs a wake-up word detection sub-model for different noise levels. In the wake-up recognition stage, according to the noise level to which the in-vehicle audio data belongs, the corresponding wake-up word detection sub-model is selected, and at the same time, the user's voiceprint feature is combined for matching verification, so as to accurately recognize the custom wake-up word in a complex in-vehicle noise environment, and significantly improve the wake-up accuracy rate and the system response speed.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control, comprising:
[0008] Collect the in-vehicle noise data of the target vehicle to construct an in-vehicle noise database, where the in-vehicle noise database includes n1 noise records; divide the n1 noise records in the in-vehicle noise database into N1 noise levels, and record the noise level labels of each noise record; obtain the custom wake-up word voice sample of the user, extract the voiceprint feature vector of the custom wake-up word voice sample, and label it as the custom wake-up word voiceprint feature vector; measure the similarity between the custom wake-up word voiceprint feature vector and each voiceprint feature vector in the pre-constructed voiceprint feature library, and match to obtain the first voiceprint feature with the highest similarity to the custom wake-up word voiceprint feature.
[0009] According to the in-vehicle noise database and the first voiceprint feature, construct wake-up word detection sub-models corresponding to N1 noise levels.
[0010] Real-time collect the in-vehicle audio data, and perform wake-up word recognition based on the in-vehicle audio data, the wake-up word detection sub-model and the first voiceprint feature to trigger the wake-up of the vehicle-mounted screen.
[0011] Further, for each of the n1 noise records, each noise record includes attribute data of a noise segment; the attribute data of the noise segment includes a noise segment number, noise segment data, a noise type label, a noise energy value, and a noise spectrum feature vector.
[0012] The step of dividing the n1 noise records in the in-vehicle noise database into N1 noise levels includes: clustering the n1 noise records according to the noise energy value and the noise spectrum feature vector of the n1 noise records in the in-vehicle noise database.
[0013] Further, the clustering of the n1 noise records includes:
[0014] Step S1210: Calculate the energy mean value E based on the noise energy value and the noise spectrum feature vector of each noise record in the in-vehicle noise database. i , the spectral centroid F i , and the spectral dispersion D i , to form the noise acoustic feature vector [E i , F i , D i for each noise record; where E i is the energy mean value of the i-th noise record, F i is the spectral centroid of the i-th noise record, D i is the spectral dispersion of the i-th noise record, and [E i , F i , D i represents the noise acoustic feature vector of the i-th noise record;
[0015] Step S1220: Use the noise acoustic feature vector [E i , F i , D i as the feature description to cluster n1 noise records and obtain N1 noise cluster centers;
[0016] Step S1230: Calculate the average silhouette coefficient SC. If SC ≤ SC', adjust the number of noise cluster centers N1 and return to Step S1220 for re-clustering until SC > SC', and then output the clustering result; SC' is a preset silhouette coefficient threshold.
[0017] Furthermore, the construction method of the voiceprint feature library is as follows: Collect voice samples of multiple users, extract the voiceprint feature vectors of the voice samples, and construct a voiceprint feature library; the voiceprint feature vectors include fundamental frequency, formant, and speech rate.
[0018] The first voiceprint feature obtained with the highest similarity to the custom wake-up word voiceprint feature vector includes:
[0019] Measure the similarity between the custom wake-up word voiceprint feature vector and each voiceprint feature vector in the voiceprint feature library to obtain a similarity score; sort according to the similarity score and select the voiceprint feature vector in the voiceprint feature library with the highest score, which is marked as the first voiceprint feature.
[0020] Furthermore, the construction of N1 wake-up word detection sub-models corresponding to N1 noise levels includes:
[0021] Based on the N1 noise levels in the in-vehicle noise database, construct N1 noise data subsets, and each noise data subset corresponds to one noise level;
[0022] Traverse N1 noise levels. For each noise level, construct an independent wake word detection sub-model respectively; obtain all the wake word detection sub-models corresponding to the N1 noise levels. The output of the wake word detection sub-model is a custom wake word detection result represented by a boolean value. If it is 1, it means the custom wake word is detected; if it is 0, it means the custom wake word is not detected.
[0023] The construction of the N1 noise data subsets includes:
[0024] Traverse n1 noise records in the in-vehicle noise database, and divide them into the corresponding N1 noise data subsets according to the noise level labels of each noise record.
[0025] Further, the construction of an independent wake word detection sub-model for each noise level respectively includes:
[0026] According to the preset data set division ratio, divide the j-th noise data subset into a training set, a validation set, and a test set; 1 ≤ j ≤ N1;
[0027] Fuse the training set in the j-th noise data subset with the first voiceprint feature to construct a wake word detection training set for noise level j;
[0028] Use the wake word detection training set for noise level j as the input to train the initial wake word detection sub-model at noise level j.
[0029] Further, the wake word recognition based on the in-vehicle audio data, the wake word detection sub-model, and the first voiceprint feature to trigger the wake-up of the car machine screen includes:
[0030] Step S3100, determine the noise level to which the in-vehicle audio data belongs, and select the wake word detection sub-model corresponding to the noise level of the in-vehicle audio data to perform custom wake word detection, and determine whether the custom wake word is detected;
[0031] Step S3200, if the custom wake word is not detected, jump to step S3100 to continue the next round of in-vehicle audio data collection and detection; if the custom wake word is detected, collect wake word audio data through the microphone array, locate the wake word speaker, and obtain the horizontal azimuth angle of the wake word speaker relative to the vehicle-mounted microphone array and the distance ;
[0032] Step S3300, set the wake-up angle threshold range and the distance threshold range ; If and , it is determined that the detected custom wake-up word comes from a reasonable position inside the vehicle and is recognized as a wake-up word to be confirmed; otherwise, it is regarded as a false wake-up, and the process jumps to step S3100 to continue the next round of in-vehicle audio data collection and custom wake-up word detection;
[0033] Step S3400: Extract acoustic features from the wake-up word audio data determined to be a wake-up word to be confirmed to construct a second voiceprint feature; calculate the similarity score SIM between the second voiceprint feature and the first voiceprint feature. If the similarity score SIM is greater than the preset voiceprint verification threshold, the wake-up word to be confirmed is recognized as a valid wake-up word, triggering the wake-up of the in-vehicle screen; otherwise, it is recognized as an invalid wake-up word, rejecting the wake-up of the in-vehicle screen and jumping to step S3100 to continue the next round of in-vehicle audio data collection and custom wake-up word detection.
[0034] Furthermore, the determination of the noise level to which the in-vehicle audio data belongs includes:
[0035] Extract features from the real-time collected in-vehicle audio data to obtain the real-time energy value Es, the real-time spectral centroid Fs, and the real-time spectral dispersion Ds, and construct a real-time acoustic feature vector [Es, Fs, Ds];
[0036] Compare the real-time acoustic feature vector [Es, Fs, Ds] with N1 noise clustering centers, calculate the Euclidean distances between the real-time acoustic feature vector and each noise clustering center, and select the noise level corresponding to the noise clustering center with the smallest Euclidean distance as the noise level to which the real-time collected in-vehicle audio data belongs.
[0037] Furthermore, the acquisition of the horizontal azimuth angle and distance of the wake-up word speaker relative to the in-vehicle microphone array includes:
[0038] Collect wake-up word audio data through the microphone array. There are M voice signals in the wake-up word audio data, and the wake-up word audio data of the m-th voice signal is denoted as , where t is time, ;
[0039] Perform voice activity detection on the wake-up word audio data to extract the wake-up word voice segments received by each microphone, find the start and end times of the wake-up word voice segments, and denote the wake-up word voice segments as ; Select one microphone from the microphone array as the reference microphone, and the other M - 1 microphones as non-reference microphones, and estimate the time delay between the wake-up word voice segments of the non-reference microphones and the reference microphone, where , and , is the The time delay of the wake-up word speech segment of the non-reference microphone relative to the reference microphone; based on the geometric layout of the microphone array, construct M-1 equations:
[0040]
[0041] in It is The position vectors of the non-reference microphones relative to the reference microphone, is the speed of sound, is the horizontal azimuth of the sound source, is the pitch angle of the sound source;
[0042] Solve the M-1 equations simultaneously to obtain the horizontal azimuth angle of the wake-up word speaker relative to the microphone array: and distance ,in .
[0043] A vehicle screen custom wake-up word configuration system based on voice control, which is used to implement the above-mentioned vehicle screen custom wake-up word configuration method based on voice control, the system includes:
[0044] Noise level classification module: used to collect the in-vehicle noise data of the target vehicle and build an in-vehicle noise database, wherein the in-vehicle noise database includes n1 noise records; and classify the n1 noise records in the in-vehicle noise database into N1 noise levels;
[0045] The first voiceprint feature acquisition module is used to obtain the user's custom wake-up word voice sample, extract the voiceprint feature vector of the custom wake-up word voice sample, and mark it as the custom wake-up word voiceprint feature vector; measure the similarity between the custom wake-up word voiceprint feature vector and each voiceprint feature vector in the pre-built voiceprint feature library, and match to obtain the first voiceprint feature with the highest similarity to the custom wake-up word voiceprint feature vector;
[0046] Model building module: used to build a wake-up word detection sub-model corresponding to N1 noise levels based on the in-vehicle noise database and the first voiceprint feature;
[0047] Wake-up word recognition module: used to collect in-car audio data in real time, perform wake-up word recognition based on in-car audio data, wake-up word detection sub-model and first voiceprint features, and trigger the car screen to wake up.
[0048] An electronic device includes a memory, a central processing unit, and a computer program stored in the memory and executable on the central processing unit. When the central processing unit executes the computer program, the method for configuring a custom wake-up word for a vehicle screen based on voice control is implemented.
[0049] A computer-readable storage medium has a computer program stored thereon, and when the computer program is executed, it implements the above-mentioned method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control.
[0050] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0051] Improve recognition accuracy. By constructing a noise level database and a custom wake-up word voiceprint feature library, the wake-up word can be accurately recognized in different noise environments, and an acoustic feature vector and a similarity matching algorithm are adopted to effectively reduce the false wake-up rate. Enhance the system response speed. Real-time collect and process in-vehicle audio data, quickly judge the wake-up word, achieve instant response, and at the same time use an independent wake-up word detection sub-model to optimize for different noise levels and improve the detection efficiency. Adapt to diverse environments. Through the clustering analysis of noise data, the system can adapt to the noise changes in various driving environments, and design a flexible noise level division mechanism to keep the system stable in a complex acoustic environment. Enhance the user experience. Allow users to customize the wake-up word, improve personalization and convenience, and avoid false wake-up and invalid operations through accurate sound source localization, thus enhancing the interaction experience. Optimize resource usage. Through clustering and model optimization, reduce the consumption of computing resources, improve the overall efficiency of the system, and achieve efficient noise data management and retrieval for subsequent analysis and optimization. These advantages make the present invention have significant technological breakthroughs and application values in the field of in-vehicle voice interaction. Description of the Drawings
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0053] Figure 1 It is a principle flow chart of the method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control in the present invention;
[0054] Figure 2 It is a method flow chart for clustering n1 noise records in the method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control in the present invention;
[0055] Figure 3 It is a method flow chart for constructing a wake-up word detection sub-model corresponding to N1 noise levels in the method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control in the present invention;
[0056] Figure 4The flowchart of the method for constructing independent wake-up word detection sub-models for each noise level in the vehicle-mounted screen custom wake-up word configuration method based on voice control of the present invention;
[0057] Figure 5 The flowchart of the method for setting the wake-up angle threshold range and distance threshold range in the vehicle-mounted screen custom wake-up word configuration method based on voice control of the present invention;
[0058] Figure 6 The functional module diagram of the vehicle-mounted screen custom wake-up word configuration system based on voice control in the present invention. Detailed implementation manners
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0060] Embodiment 1
[0061] Please refer to Figure 1 As shown, this embodiment provides a vehicle-mounted screen custom wake-up word configuration method based on voice control, including:
[0062] Step S1000, collect the in-vehicle noise data of the target vehicle, construct an in-vehicle noise database, and the in-vehicle noise database includes n1 noise records; divide the n1 noise records in the in-vehicle noise database into N1 noise levels; obtain the user's custom wake-up word voice sample, measure the similarity between the voiceprint features of the custom wake-up word voice sample and each voiceprint feature in the preset voiceprint feature library, and match the first voiceprint feature with the highest similarity to the voiceprint feature of the custom wake-up word;
[0063] Further, step S1000 includes:
[0064] Step S1100, collect the in-vehicle noise data of the target vehicle, construct an in-vehicle noise database; the in-vehicle noise database includes n1 noise records, and each noise record includes the attribute data of a noise segment; the attribute data of the noise segment includes the noise segment number, noise segment data, noise type label, noise energy value, and noise spectrum feature vector;
[0065] Specifically, in step S1100, the in-vehicle noise in the real environment is acquired through in-vehicle acoustic sensors, comprehensively covering various noise interferences in the driving scenario. First, the in-vehicle noise data of the target vehicle is collected by using an in-vehicle acoustic sensor array, including but not limited to engine noise, wind noise, external traffic noise, in-vehicle human voice noise, etc.; among them, the acoustic sensor array includes a microphone array, MEMS sensors, etc., and is arranged at positions such as the in-vehicle driving area and the passenger area; among them, the engine noise reflects the working condition of the vehicle and is one of the main noise sources affecting voice interaction; the wind noise is the airflow noise caused by the vehicle speed and will increase with the increase of the vehicle speed; the external traffic noise includes the noise of other vehicles, horn sounds, etc., which is the background noise introduced by the external environment; the in-vehicle human voice noise comes from the conversation sounds, breathing sounds, etc. of the passengers and is the non-speech noise generated by in-vehicle activities.
[0066] Preprocess the collected in-vehicle noise data, including operations such as noise segmentation, energy normalization, and noise reduction filtering; noise segmentation is to use speech endpoint detection algorithms, such as the double-threshold method, energy-entropy ratio method, etc., to divide the noise data into different noise segments; energy normalization is to normalize the energy of each noise segment to make its amplitude range consistent; noise reduction filtering is to use an adaptive filter, such as LMS, RLS, etc., to enhance the noise reduction of the noise; through noise segmentation, the noise can be extracted from the continuous audio stream for subsequent directional processing. Energy normalization eliminates the amplitude differences caused by factors such as noise recording equipment and distance, making the noise data have a unified energy benchmark. Noise reduction filtering removes the random noise components while retaining the essential characteristics of the noise, improving the signal-to-noise ratio. Based on the preprocessed in-vehicle noise data, construct a structured noise library, which on the one hand provides a unified data source for noise grading, and on the other hand makes the noise data easier to retrieve, manage and apply, providing a data basis for the research and development, testing and iteration of the wake-up system.
[0067] Based on the preprocessed in-vehicle noise data, a structured in-vehicle noise database is constructed; the noise segment numbers in the in-vehicle noise database uniquely identify each noise segment, facilitating data retrieval and management; such as 001, 002, which are used to distinguish different noise segments. The noise segment data is used to store the audio data of each noise segment for analysis and processing, such as an audio file containing 5 seconds of engine noise. The noise type label identifies the source or category of the noise, facilitating classification and processing, such as "engine noise", "wind noise", "traffic noise", "human voice noise". The noise energy value represents the loudness or intensity of the noise, used for comparison and normalization processing, such as 85 dB, indicating the sound pressure level of the noise. The noise spectrum feature vector describes the frequency characteristics of the noise, used for signal processing and pattern recognition, a spectrum feature vector that shows the peak at a specific frequency to help identify the noise type. Therefore, step S1100 collects and optimizes the in-vehicle noise data in the real environment, laying a solid data foundation for subsequent noise grading and model adaptation, which is the key to improving the environmental adaptability and robustness of the wake-up system.
[0068] Step S1200, according to the noise energy values and noise spectrum feature vectors of n1 noise records in the in-vehicle noise database, cluster the n1 noise records, divide the n1 noise records in the in-vehicle noise database into N1 noise levels, and record the noise level labels of each noise record;
[0069] Further, as Figure 2 shown, step S1200 includes:
[0070] Step S1210, according to the noise energy value and noise spectrum feature vector of each noise record in the in-vehicle noise database, calculate the energy mean E i , the spectral centroid F i , and the spectral dispersion D i , to form the noise acoustic feature vector [E i , F i , D i of each noise record; where, E i is the energy mean of the i-th noise record, F i is the spectral centroid of the i-th noise record, D i is the spectral dispersion of the i-th noise record, and [E i , F i , D i represents the noise acoustic feature vector of the i-th noise record;
[0071] Step S1220, with the noise acoustic feature vector [E i , F i , D iFor feature description, cluster n1 noise records to obtain N1 noise cluster centers;
[0072] In step S1230, calculate the average silhouette coefficient SC. If SC ≤ SC', adjust the number of noise cluster centers N1, return to step S1220 for reclustering until SC > SC', and output the clustering result; SC' is a preset silhouette coefficient threshold.
[0073] The calculation of the average silhouette coefficient SC includes:
[0074]
[0075] Where:
[0076] : The th noise cluster center, with a value range from 1 to N1.
[0077] : The cluster center number to which the th noise record belongs.
[0078] : The th noise record and the th cluster center, the Euclidean distance between them, measuring the distance between the noise record and the cluster center , and the calculation formula is:
[0079] ;
[0080] Where are respectively the energy mean, spectral centroid, and spectral dispersion of the cluster center .
[0081] The numerator represents the distance between the th noise record and the nearest distance of other categories minus the distance between it and the center of its own category. The larger the numerator, the farther the boundary distance of the noise record from other categories, the higher the cohesion with its own category, and the better the clustering effect.
[0082] The denominator is used to normalize the silhouette coefficients of different records. The denominator takes the larger value between the distance between the th noise record and the center of its own category and its nearest distance from other categories. This can balance the influence of the intra-cluster distance and the inter-cluster distance on the silhouette coefficient.
[0083] Average the silhouette coefficients of all noise records to obtain the average silhouette coefficient SC, with a value range of When the SC is closer to 1, it indicates better cohesion and separation of the clustering, and higher rationality in the classification of the noise level. When SC is greater than the given threshold SC', the current noise clustering result can be considered excellent.
[0084] The function of this formula is to quantitatively evaluate the effect of noise clustering. The closer the number of noise levels N1 is to the true noise distribution, the more accurately the noise clustering center C k can reflect the acoustic characteristics of each noise level, and the higher the value of the average silhouette coefficient SC will be. By iteratively adjusting the value of N1 to make SC reach a larger value, the noise clustering result can be made more reasonable, thus providing better data support for subsequent applications such as hierarchical noise reduction and adaptive wake-up.
[0085] It is worth mentioning that since the silhouette coefficient calculates the relative distances of each noise record to its own category and other categories, it is robust to the order of noise records and the encoding of category labels. Even if the order of noise records changes, or the clustering categories are re-numbered, as long as the clustering result remains unchanged, the SC value remains the same. This is beneficial to improving the reliability of clustering evaluation.
[0086] In summary, this average silhouette coefficient formula comprehensively considers the two aspects of the tightness and dispersion of noise clustering, and gives a quantitative evaluation of the rationality of noise classification. The average silhouette coefficient changes with the number of clustering centers N1. When N1 gradually approaches the true number of noise levels, SC will reach a larger value. Further, the value of N1 can be automatically adjusted by a parameter optimization algorithm to maximize SC, so as to achieve the adaptive optimization of noise level classification. This formula is of great significance for both the mining of in-vehicle noise data and the improvement of the environmental adaptability of in-vehicle voice interaction technology.
[0087] Specifically, the energy mean (E i ) reflects the average energy level of the i-th noise record. Its calculation method is to square and sum the noise energy values at all moments of this noise record, and then divide by the length of the record. The larger the energy mean, the higher the loudness of the noise. The spectral centroid (F i ) represents the central position of the spectral distribution of the i-th noise record. It is calculated by summing the products of the frequencies of each frequency point and the squares of their corresponding spectral amplitudes, and then dividing by the total sum of the squares of the spectral amplitudes. The larger the spectral centroid, the higher the frequency component of the noise. The spectral dispersion (D iIt is used to measure the degree of dispersion of the spectral distribution of the i-th noise record. It is obtained by calculating the sum of the products of the square of the difference between the frequency of each frequency point and the spectral centroid and the square of the spectral amplitude, and then dividing it by the sum of the squares of the spectral amplitudes. The greater the spectral dispersion, the more dispersed the spectrum of the noise and the richer the frequency components it contains. Exemplarily, the energy mean of a certain engine noise record is 80 dB, the spectral centroid is 500 Hz, and the spectral dispersion is 200 Hz squared. Therefore, its acoustic feature vector can be expressed as [80, 500, 200]. And for a certain wind noise record, the energy mean is 60 dB, the spectral centroid is 2000 Hz, and the spectral dispersion is 1000 Hz squared, and the acoustic feature vector is [60, 2000, 1000]. It can be seen from this that the engine noise has higher energy, lower frequency and is more concentrated, while the wind noise has lower energy but higher frequency and is more dispersed.
[0088] These three features are selected as the acoustic feature vector of the noise because they comprehensively describe the acoustic properties of the noise from three aspects: energy, frequency, and dispersion. This feature description enables noise clustering to group noise records with similar energy levels and spectral distributions into one category, and the resulting noise levels have acoustic homogeneity.
[0089] Using algorithms such as K-means clustering, with the noise acoustic feature vector as the sample feature, clustering analysis is performed on multiple noise records to obtain multiple cluster centers. Each cluster center represents a noise level and contains the average energy, spectral centroid, and spectral dispersion of the noise at that level.
[0090] For example, the clustering result may obtain three noise levels, and their cluster centers are respectively:
[0091] The first level center is [70, 800, 300], representing low-frequency noise, which may include engine noise, etc.;
[0092] The second level center is [65, 1500, 500], representing medium-frequency noise, which may include external traffic noise and human voice noise, etc.;
[0093] The third level center is [50, 4000, 1500], representing high-frequency noise, which may include wind noise, etc.
[0094] The clustering effect is evaluated by calculating the average silhouette coefficient. The silhouette coefficient is an index to measure the compactness and separation of clustering, and its value range is from -1 to 1. The larger the value, the better the clustering quality. For each noise record, the silhouette coefficient is calculated by comparing the average distance of the record to the noise of the same class and the average distance to the nearest noise of a different class. If the silhouette coefficient is close to 1, it means the record is correctly classified; close to -1 may indicate misclassification; close to 0 means the record is between two classes.
[0095] The average silhouette coefficient is the arithmetic mean of all silhouette coefficients, reflecting the overall quality of the clustering result. When the average silhouette coefficient is greater than the preset silhouette coefficient threshold, it can be considered that the clustering effect is ideal and the noise level division is reasonable. Otherwise, the number of clusters needs to be adjusted and clustering is performed again until the ideal effect is achieved.
[0096] For example, if the obtained average silhouette coefficient is 0.8, it indicates that the current three noise levels have high distinguishability in acoustic features and the clustering result is reliable; while if the value is 0.3, it may be necessary to try adjusting the number of noise levels (for example, changing to 2 or 4) and perform clustering again to find the optimal grading scheme.
[0097] Through such iterative optimization, a noise grading result with separable acoustic features is finally achieved. This result reveals the internal structure and grading of the in-vehicle noise data, providing an important basis for subsequent adaptive noise suppression and graded wake-up. Compared with the simple energy threshold grading method, this method makes full use of the frequency domain features of the noise and realizes a more refined and targeted noise grading, laying a foundation for improving the environmental adaptability of in-vehicle voice interaction systems.
[0098] In summary, step S1200 successfully realizes the automatic grading of in-vehicle noise data through the extraction and clustering of noise acoustic features. This method combines signal processing and machine learning technologies, and can extract valuable acoustic patterns and rules from complex in-vehicle noises, providing data support and algorithm guarantee for intelligent and personalized in-vehicle voice interaction.
[0099] In step S1300, a user-defined wake-up word voice sample is obtained, the acoustic feature vector of the user-defined wake-up word voice sample is extracted and marked as the user-defined wake-up word acoustic feature vector; the similarity between the user-defined wake-up word acoustic feature vector and each acoustic feature in the pre-constructed acoustic feature library is measured, and the first acoustic feature with the highest similarity to the user-defined wake-up word acoustic feature vector is obtained by matching.
[0100] Furthermore, step S1300 includes:
[0101] In step S1310, voice samples of multiple users are collected, the acoustic feature vectors of the voice samples are extracted, and an acoustic feature library is constructed; the acoustic feature vectors include fundamental frequency, formant, and speech rate;
[0102] In step S1320, the similarity between the user-defined wake-up word acoustic feature vector and each acoustic feature vector in the acoustic feature library is measured to obtain a similarity score; according to the similarity score ranking, the acoustic feature vector in the acoustic feature library with the highest score is selected and marked as the first acoustic feature.
[0103] Specifically, in step S1300, personalized voiceprint registration of the custom wake-up word is achieved by obtaining the voice sample of the user-defined wake-up word, extracting its voiceprint features, and matching them with the preset voiceprint feature library. First, the system prompts the user to record the voice sample of the custom wake-up word, such as "classmate XX", "hello, little X", etc. The user can use the in-vehicle microphone to record the wake-up word voice multiple times in a quiet environment to obtain as pure voice data as possible.
[0104] In step S1310, rich voiceprint data is obtained by collecting the voice samples of multiple users. The voice samples cover users of different genders, ages, and dialects, including various types such as sentences and words, ensuring the diversity and representativeness of the voiceprint feature library. For each voice sample, the voiceprint feature vector is extracted using voiceprint analysis technology. The fundamental frequency is a measure of the vocal cord vibration frequency and reflects the physiological characteristics of the speaker such as age and gender. It can be extracted by the average magnitude difference function (AMDF) based on the time domain or the cepstrum method based on the frequency domain. The formant is the frequency domain peak generated by the vocal tract resonance and reflects the characteristics of the speaker's articulatory organs. Methods such as linear predictive coding (LPC) can be used to estimate it. Generally, the center frequencies and bandwidths of the first 3-5 formants are extracted as features. The speech rate is the number of pronunciation units (such as syllables, words, etc.) per unit time and reflects the rhythm habit of the speaker. It is obtained by statistical analysis of the endpoint detection and pronunciation unit division of the speech.
[0105] For example, the voiceprint features of a certain user's voice sample can be expressed as [103.5Hz, 650Hz, 1400Hz, 2600Hz, 3.5 syllables / second], corresponding to the 5 feature dimensions of the fundamental frequency, the first 3 formants, and the speech rate respectively. By extracting the features of the voice samples of multiple users, a voiceprint feature library covering group differences can be constructed. This is the basis for voiceprint recognition and matching.
[0106] Through the extraction of the voiceprint features of the voice samples, the voiceprint of each user is transformed into a feature vector with a fixed dimension, which is convenient for comparison and matching. The feature set composed of the voiceprint feature vectors of each user is the user voiceprint feature library. This feature library covers the voiceprint templates of different users and is the basis for voiceprint recognition and verification. Pre-constructing the voiceprint feature library can, on the one hand, speed up the voiceprint matching speed without repeated feature extraction; on the other hand, it is beneficial to collect more comprehensive user voice samples and improve the richness and reliability of the voiceprint features.
[0107] Using a voiceprint similarity measurement method, calculate the similarity between the voiceprint features of the wake-up word and the voiceprint features of each user in the voiceprint library. Commonly used similarity measurements include Euclidean distance, cosine similarity, probability distance, etc. Euclidean distance measures the straight-line distance between two voiceprint feature vectors in the feature space. The smaller the distance, the closer the voiceprints. Cosine similarity calculates the cosine value of the angle between two voiceprint feature vectors. The larger the value, the more consistent the directions, that is, the more similar the voiceprints. Probability distances such as KL divergence measure the difference in the probability distributions of two voiceprint features. The smaller the difference, the better the voiceprint matching. Regardless of the measurement method used, a similarity score can be obtained, which quantifies the similarity between the wake-up word voiceprint and each voiceprint in the library. Calculate the similarity between the wake-up word voiceprint and each voiceprint in the library in turn, and a similarity score vector can be obtained. Sort the similarity scores from high to low. The voiceprint with the highest score is the voiceprint that best matches the current wake-up word voiceprint, and the corresponding user identity is also identified as the current waking user. Through similarity sorting, the result of voiceprint matching is more reliable and the misrecognition rate is low. Based on the matching method of the existing voiceprint library, the cumbersome steps of repeated input of user voices are avoided. The user only needs to input the wake-up word once to conveniently complete voiceprint registration and recognition.
[0108] Exemplarily, the voiceprint features of the user-defined wake-up word "Hello, Little X" are [105Hz, 720Hz, 1300Hz, 2400Hz, 4 syllables / second]. The similarity calculation results with 3 features in the voiceprint library are [0.8, 0.6, 0.3]. Then the first voiceprint feature [110Hz, 700Hz, 1350Hz, 2500Hz, 3.8 syllables / second] corresponding to the highest score of 0.8 is selected as the best match and can be used as a reference and optimization for subsequent wake-up word acoustic modeling.
[0109] Extracting the voiceprint features of the user-defined wake-up word implants the user's personality from the source of voice input, which helps to improve the pertinence of wake-up word recognition. Through similarity matching with the voiceprint feature library, the most suitable acoustic reference can be automatically selected for different users, which has self-adaptability and reduces the workload of manual optimization. Combining multiple acoustic features such as fundamental frequency and formant can more comprehensively reflect the voiceprint attributes than a single feature, improve the robustness of voiceprint representation and matching, and reduce the recognition error rate. The voiceprint feature library can be continuously expanded and optimized by incorporating more users' voice samples, making the matching reference benchmark more abundant and reliable, and having scalability.
[0110] In summary, through voiceprint feature extraction and similarity matching, step S1300 selects the optimal acoustic reference for the user-defined wake-up word, which not only preserves the user's personality but also takes into account the reliability of recognition. It is an intelligent and personalized solution in the voice interaction system. By using voiceprint recognition technology, a voiceprint adaptation mechanism is implanted in the wake-up word setting link, laying a foundation for subsequent wake-up word modeling and recognition, and is of great significance for improving system performance and user experience. Compared with the traditional "one-size-fits-all" wake-up word setting, this method is more flexible and intelligent, demonstrating the development trend and potential of voice interaction technology.
[0111] Step S2000: Construct N1 wake-up word detection sub-models corresponding to N1 noise levels according to the in-vehicle noise database and the first voiceprint feature;
[0112] Further, as Figure 3 shown, step S2000 includes:
[0113] Step S2100: Based on the N1 noise levels in the in-vehicle noise database, construct N1 noise data subsets, with each noise data subset corresponding to one noise level;
[0114] The construction of the N1 noise data subsets includes:
[0115] Traverse the n1 noise records in the in-vehicle noise database, and divide them into the corresponding N1 noise data subsets according to the noise level labels of each noise record;
[0116] Step S2200: For each noise level, construct an independent wake-up word detection sub-model respectively;
[0117] Further, as Figure 4 shown, step S2200 includes:
[0118] Step S2210: Divide the j-th noise data subset into a training set, a validation set, and a test set according to the preset data set division ratio; 1 ≤ j ≤ N1;
[0119] Step S2220: Fuse the training set in the j-th noise data subset with the first voiceprint feature to construct a wake-up word detection training set for noise level j;
[0120] Step S2230: Use the wake-up word detection training set for noise level j as the input to train the initial wake-up word detection sub-model at noise level j;
[0121] Step S2240: Use the validation set in the j-th noise data subset to optimize the model parameters and obtain the final wake-up word detection sub-model at noise level j;
[0122] Step S2250: Use the test set in the j-th noise data subset to evaluate the performance metrics of the wake word detection sub-model at the final noise level j.
[0123] Step S2300: Traverse N1 noise levels, execute Step S2200, and obtain wake word detection sub-models corresponding to all N1 noise levels. The output of the wake word detection sub-model is a custom wake word detection result represented by a boolean value. If it is 1, it means the custom wake word is detected; if it is 0, it means the custom wake word is not detected.
[0124] Specifically, Step S2000 aims to build multiple dedicated wake word detection sub-models for the wake word recognition problem in different noise environments, thereby improving the environmental adaptability and recognition performance of the system.
[0125] In Step S2100, based on the previously obtained N1 noise levels, the original in-vehicle noise database is divided into N1 noise data subsets. Each subset corresponds to a specific type of noise or noise intensity, such as "high-speed wind noise", "medium engine noise", etc. By grouping n1 noise records according to the noise level labels, N1 data subsets with consistent internal noise characteristics and high external discrimination are finally obtained. This way of dataset division enables subsequent model training to be more targeted, which is beneficial to improving the fine-grainedness and accuracy of wake word detection.
[0126] The core of Step S2200 is to train a wake word detection sub-model for each noise level. In Step S2210, the hold-out method is first used to divide each noise data subset into a training set, a validation set, and a test set according to a certain ratio (such as 7:2:1). Among them, the training set is used to train the model parameters, the validation set is used for parameter tuning and selecting the optimal model, and the test set is used to evaluate the actual performance of the model. This dataset division helps to improve the generalization ability and robustness of the model.
[0127] Step S2220 aims to fuse the custom wake word voiceprint features registered by the user with the noise data to construct a personalized wake word detection training set. The specific approach is to splice the custom wake word voiceprint feature vector of the user with each noise record in the training set to form a new training sample. For example, a training sample originally containing 5 seconds of noise, after being spliced with the extracted 40-dimensional MFCC voiceprint feature, becomes a 45-dimensional new feature vector. This feature fusion enables the training data to contain both the acoustic information of the environmental noise and the target wake word, so that the trained model can accurately recognize the user's wake word voice in complex noise. Compared with traditional wake word detection models, this method makes full use of the personalized features of the user's voiceprint and significantly improves the sensitivity and accuracy of wake-up.
[0128] In step S2230, a combined model of CNN and LSTM is selected to train the wake word detection sub-model. CNN (Convolutional Neural Network) can automatically extract local features in the speech signal, and LSTM (Long Short-Term Memory Network) can model the long-term dependencies of speech. The combination of the two can comprehensively characterize the acoustic and temporal features of the wake word speech. The input of the model is a noisy speech segment fused with voiceprint features, and the output is a custom wake word detection result represented by a boolean value. During training, the manually annotated wake word segments are used as positive samples, and other speech segments and pure noise segments are used as negative samples to minimize the cross-entropy loss function between the sample prediction result and the true label, and iteratively optimize the model parameters. After multiple rounds of training, the model can accurately locate and identify the target wake word from the noisy audio stream.
[0129] For example, a positive sample in the training set may be a 3-second speech of the user saying "Hello, Xiao X" under wind noise of about 80 decibels, which is spliced with the user's voiceprint features and input into the model for training. The model outputs "1", that is, the custom wake word is detected, which is consistent with the manual annotation, indicating that the model has learned to recognize the wake word sound of this user under medium wind noise.
[0130] In step S2240, the model is tuned using the validation set. By methods such as grid search, different hyperparameters (such as the convolutional kernel size of CNN and the number of hidden layer units of LSTM) are traversed, the parameter combination with the best performance on the validation set is selected, and the early stopping method is used to prevent overfitting, and finally a wake word detection sub-model with stable performance and good generalization is obtained. This process ensures that the optimization direction of the model is consistent with the performance in the real environment, and reduces the difference between the training set and the actual application.
[0131] In step S2250, the performance metrics of the model, such as accuracy, recall rate, false wake-up rate, etc., are evaluated using the test set. By analyzing these metrics, the actual performance of the wake word detection sub-model at this noise level can be comprehensively judged, providing a basis for subsequent model selection and improvement. An ideal sub-model should have a high accuracy in recognizing the target wake word, a low false wake-up rate, and also be able to adapt to the changes in the acoustic environment within this noise level.
[0132] After the iteration of step S2300, finally N1 wake word detection sub-models for different noise levels are obtained. Each sub-model is trained with specific noise data and can robustly detect the user's custom wake word in a specific in-vehicle noise environment. Compared with a single general wake model, this multi-model mechanism of dividing and conquering and teaching students in accordance with their aptitude has obvious advantages in dealing with complex noise and personalized wake word recognition.
[0133] In summary, through key technologies such as noise dataset division, voiceprint feature fusion, and hierarchical modeling, step S2000 constructs an adaptive wake-word detection scheme in a vehicle environment. This scheme can dynamically select the optimal wake-word detection sub-model according to the real-time changes in in-vehicle noise, thus achieving a good balance in aspects such as noise suppression, wake-up sensitivity, and false wake-up rate, laying a solid foundation for the practical application of in-vehicle voice interaction systems.
[0134] In step S3000, in-vehicle audio data is collected in real time, and wake-word recognition is performed based on the in-vehicle audio data, the wake-word detection sub-model, and the first voiceprint feature to trigger the wake-up of the vehicle console screen.
[0135] Furthermore, step S3000 includes:
[0136] In step S3100, in-vehicle audio data is collected in real time, the noise level to which the in-vehicle audio data belongs is judged, and the wake-word detection sub-model corresponding to the noise level to which the in-vehicle audio data belongs is selected for custom wake-word detection to judge whether a custom wake-word is detected;
[0137] Furthermore, step S3100 includes:
[0138] In step S3110, feature extraction is performed on the in-vehicle audio data collected in real time to obtain the real-time energy value Es, the real-time spectral centroid Fs, and the real-time spectral dispersion Ds, and a real-time acoustic feature vector [Es, Fs, Ds] is constructed;
[0139] In step S3120, the real-time acoustic feature vector [Es, Fs, Ds] is compared with N1 noise clustering centers, the Euclidean distances between the real-time acoustic feature vector and each noise clustering center are calculated, and the noise level corresponding to the noise clustering center with the smallest Euclidean distance is selected as the noise level to which the in-vehicle audio data collected in real time belongs;
[0140] In step S3130, according to the noise level to which the in-vehicle audio data belongs, the corresponding wake-word detection sub-model is selected for custom wake-word detection.
[0141] Specifically, in step S3100, an in-vehicle microphone array is used to continuously collect in-vehicle audio signals to achieve real-time monitoring of in-vehicle audio. The in-vehicle microphone array consists of multiple microphone units distributed at different positions in the vehicle to pick up in-vehicle sounds from all directions. The audio data collected in real time is processed in frames, each frame contains L sampling points, and a certain proportion of overlap is allowed between adjacent frames to smooth the temporal changes of audio features. For example, when collecting audio signals at a sampling rate of 16 kHz, the frame length is taken as L = 400 sampling points (corresponding to 25 ms), and the frame shift is taken as 160 sampling points (corresponding to 10 ms), then the overlap between adjacent frames is 50%.
[0142] The audio frames are pre-processed by pre-emphasis, framing, and windowing to extract the features of the audio frames and obtain the real-time energy value Es, the real-time spectrum center of gravity Fs, and the real-time spectrum dispersion Ds. Pre-emphasis is to filter the audio signal with a first-order high-pass filter to enhance the high-frequency components and compensate for the high-frequency attenuation of the channel and the microphone. Framing is to divide the pre-emphasized audio signal frame by frame according to the frame length L to obtain a series of audio frames. Windowing is to apply smoothing window functions such as Hamming window and Hanning window to each frame of audio to reduce signal mutations at the edge of the frame. After pre-processing, the noise components in the audio frame are suppressed and the speech components are enhanced.
[0143] Step S3100 realizes real-time noise estimation and adaptive wake-up of in-car audio through feature extraction and noise classification, which is the key to improving the environmental adaptability of the in-vehicle voice wake-up system. This step combines the noise modeling of step S1000 and the hierarchical wake-up of step S2000 to dynamically optimize the wake-up word detection process in a complex noise environment, thereby maximizing the accuracy and real-time performance of the wake-up system. Compared with the traditional fixed model wake-up method, the noise hierarchical adaptive wake-up mechanism introduced in this step can significantly improve the wake-up quality in harsh noise environments and achieve more natural and smooth human-computer voice interaction.
[0144] Step S3200: If no custom wake-up word is detected, jump to step S3100 to continue the next round of in-vehicle audio data collection and detection; if a custom wake-up word is detected, the wake-up word audio data is collected through the microphone array, the wake-up word speaker is located, and the horizontal azimuth angle of the wake-up word speaker relative to the vehicle microphone array is obtained. and distance ;
[0145] Further, step S3200 includes:
[0146] Step S3210: If a custom wake-up word is detected, the wake-up word audio data is collected through the microphone array. The wake-up word audio data has a total of M voice signals, and the wake-up word audio data of the mth voice signal is recorded as , t is time, ;
[0147] Step S3220: wake-up word audio data Perform voice endpoint detection and extract the voice received by each microphone. Find the start and end time of the wake-up word voice segment in , the wake-up word speech segment is recorded as ; Select one microphone from the microphone array as the reference microphone, and the other M-1 microphones as non-reference microphones, and estimate the time delay of the wake-up word speech segment between the non-reference microphone and the reference microphone ,in , is the time delay of the wake-up word voice segment of the th non-reference microphone relative to the reference microphone; based on the geometric layout of the microphone array, M-1 equations are constructed:
[0148] ;
[0149] where is the position vector of the th non-reference microphone relative to the reference microphone, is the speed of sound, is the horizontal azimuth angle of the sound source, is the elevation angle of the sound source;
[0150] Step S3230, solve the M-1 equations simultaneously to obtain the horizontal azimuth angle of the wake-up word speaker relative to the vehicle-mounted microphone array and the distance , where .
[0151] Specifically, when a custom wake-up word is detected, the microphone array starts to synchronously collect wake-up word audio data, including M voice signals , corresponding to the received signals of the wake-up word on M microphones. Through voice endpoint detection and cross-correlation analysis, the time boundaries of the wake-up word voice segment can be estimated, as well as the time delay between the non-reference microphone and the reference microphone for the wake-up word voice segment.
[0152] The time delay reflects the time difference between the wake-up word voice arriving at the th non-reference microphone and the reference microphone, and contains information about the sound source azimuth angle and distance . By solving the system of equations simultaneously, the estimated values of the azimuth angle and elevation angle of the sound source can be obtained. Further, substituting the estimated azimuth angle into the far-field model of the microphone array, the estimated value of the sound source distance can be obtained. The far-field model of the microphone array is common knowledge in the fields of acoustics and signal processing and is widely used in sound source localization and audio signal processing.
[0153] For example: Suppose the vehicle-mounted microphone array consists of 4 microphone units , respectively located at the front, rear, left, and right positions inside the vehicle, forming a rectangular array. When the wake word "Hello, Xiao X" is detected, 4 microphones simultaneously collect audio data with a length of 100 ms. The signal received by the reference microphone 1 is denoted as , and the signal received by the non-reference microphone 2 is denoted as , and so on. The speed of sound is generally taken as 340 m / s. Cross-correlation operations are performed on these 4 voice signals to estimate the following time delays:
[0154] ;
[0155] Substitute into the system of equations:
[0156] ;
[0157] ;
[0158] Among them, the known microphone position vectors are:
[0159] ;
[0160] Solve simultaneously to obtain:
[0161] ;
[0162] Substitute the estimated azimuth angle into the far-field model to obtain the distance estimate value:
[0163] ;
[0164] Therefore, the speaker of the wake word is located at the azimuth of 60° in the front right of the vehicle, about 0.4 m away from the microphone array. This indicates that the wake word comes from the passenger area inside the vehicle, and the next step of voiceprint verification can be continued. If the positioning result shows that the wake word comes from a relatively far position outside the vehicle, it can be determined as a false wake-up caused by environmental noise, and the wake-up request should be rejected.
[0165] Step S3300, set the wake-up angle threshold range and the distance threshold range ; If and , then it is determined that the detected custom wake word comes from a reasonable position inside the vehicle and is recognized as a wake word to be confirmed. Otherwise, it is regarded as a false wake-up caused by external noise, the wake-up is rejected, and the process jumps to step S3100 to continue the next round of in-vehicle audio data collection and custom wake word detection;
[0166] Furthermore, as shown in Figure 5 , step S3300 includes:
[0167] Step S3310: Establish a three-dimensional in-vehicle space model based on the spatial structure parameters and seat layout of the vehicle model; in the three-dimensional in-vehicle space model, calibrate the installation positions of the on-vehicle microphone arrays.
[0168] Step S3320: In the three-dimensional in-vehicle space model, delimit wake-up regions, each wake-up region is represented by a spatial polygon, and record the corner coordinates of the polygon, where is a positive integer.
[0169] Step S3330: Project the wake-up regions onto the horizontal plane with the microphone array as the origin, and extract the azimuth span and distance span of each projected polygon, where is the total number of wake-up regions, is the minimum azimuth of the projected polygon, representing the leftmost boundary of the wake-up region outward from the microphone array on the horizontal plane; is the maximum azimuth of the projected polygon, representing the rightmost boundary of the wake-up region outward from the microphone array on the horizontal plane; is the shortest distance of the projected polygon, representing the nearest boundary of the wake-up region starting from the microphone array; is the longest distance of the projected polygon, representing the farthest boundary of the wake-up region starting from the microphone array;
[0170] Step S3340: Take the union of the azimuth spans of each to determine the wake-up angle threshold range , ; take the union of the distance spans of each to determine the wake-up distance threshold range .
[0171] Specifically, based on sound source localization, step S3300 introduces prior information on the vehicle interior layout and usage scenarios, sets the rationality requirements for the position of the wake word speaker, and further improves the reliability of wake-up confirmation. By modeling the three-dimensional space inside the vehicle and depicting the position distribution characteristics of the driver and passengers, a three-dimensional passenger position template can be obtained. Mapping the position of the speaker obtained by microphone array localization into the vehicle interior space and comparing it with the passenger position template can determine whether the position is within a reasonable wake-up range. On this basis, distance and angle threshold conditions are set to form a three-dimensional wake-up permission area. Only when the localization result falls within the permission area is it considered that the wake word comes from a legitimate user inside the vehicle; otherwise, it is regarded as a false wake-up caused by external noise interference. This method makes full use of the prior information of the vehicle interior scenario, sets wake-up limit conditions from the perspective of the spatial position of the sound source, can effectively distinguish between external noise wake-up and real wake-up inside the vehicle, and reduce the occurrence of false wake-ups. Compared with traditional wake-up methods, this step can more comprehensively describe the spatial characteristics of in-vehicle voice interaction, utilize acoustic space information to dynamically adjust the wake-up strategy, and make in-vehicle voice wake-up more flexible and accurate.
[0172] Set the wake-up angle threshold range and the distance threshold range to determine whether it comes from a reasonable position inside the vehicle according to the azimuth and distance of the sound source detected by the wake word. The reasonable position here refers to the effective reception range of the vehicle-mounted microphone, usually the seat areas of the driver and the co-driver. represents the horizontal angle between the sound source and the microphone array, and are respectively the lower and upper limits of the in-vehicle wake-up angle, such as 60° to 120°, indicating that only sounds within a certain range directly in front are responded to; d represents the distance from the sound source to the microphone array, and are the lower and upper limits of the in-vehicle wake-up distance, such as 0.5 meters to 1.5 meters, indicating that only sounds within a certain distance are responded to.
[0173] When the detected wake word satisfies ∈ and d ∈ When it is detected, it can be preliminarily determined that it is a valid wake-up initiated by a user inside the vehicle, and its voiceprint needs to be further verified; otherwise, it may be the audio of a user in other positions, other devices, or a false wake-up caused by external vehicle noise, and it should be directly rejected to avoid unnecessary voiceprint comparison. Exemplarily, if the sound source of the detected wake-up word "Hello, Xiao X" is located in the co-pilot seat, the horizontal angle with the microphone is 100°, and the distance is 1 meter, which belongs to the preset wake-up angle range [60°, 120°] and distance range [0.5m, 1.5m], it is temporarily determined as a wake-up word to be confirmed and needs to enter the voiceprint verification process; if the sound source of the wake-up word "Hello, Xiao X" is located in the rear seat, the horizontal angle with the microphone is 150° or the distance is 2 meters, then it exceeds the preset angle or distance range, and it can be determined as a false wake-up, and the wake-up should be rejected and a new round of wake-up word detection should be started.
[0174] This judgment of the wake-up position based on angle and distance makes full use of the beamforming ability of the microphone array, can effectively reduce false wake-ups caused by sounds in non-target areas such as the rear seats and outside the vehicle, and improve the anti-noise interference ability of the system. It is a pre-screening for voiceprint verification, which can reduce unnecessary voiceprint comparison, reduce the energy consumption and response delay of the system. At the same time, the angle threshold and distance threshold can be flexibly set according to factors such as vehicle models and seat layouts, and have a certain degree of self-adaptability.
[0175] Step S3400: Extract acoustic features from the wake-up word audio data determined to be a wake-up word to be confirmed, and construct a second voiceprint feature; calculate the similarity score SIM between the second voiceprint feature and the first voiceprint feature. If the similarity score SIM is greater than the preset voiceprint verification threshold, the wake-up word to be confirmed is identified as a valid wake-up word, triggering the wake-up of the car machine screen; otherwise, it is identified as an invalid wake-up word, rejecting the wake-up of the car machine screen and jumping to step S3100 to continue the next round of in-vehicle audio data collection and custom wake-up word detection.
[0176] Specifically, the second voiceprint feature uses the same feature extraction method as that in step S1000 for constructing the voiceprint feature library, that is, fundamental frequency, formant, speech rate, etc., ensuring the consistency of voiceprint verification. Then, calculate the similarity score SIM between the second voiceprint feature and the first voiceprint feature (i.e., the voiceprint feature of the custom wake-up word registered by the user) obtained in step S1300. The similarity score can be calculated using common feature matching measurement methods such as Euclidean distance and cosine similarity. If the similarity score SIM is greater than the preset voiceprint verification threshold, such as 0.8, it is determined that the speaker of the wake-up word to be confirmed matches the voiceprint of the registered user, and it is identified as a valid wake-up word, triggering the wake-up of the car machine screen; otherwise, if SIM is less than or equal to the verification threshold, it is determined that the speaker does not match the voiceprint of the registered user, and it is identified as an invalid wake-up word of others, rejecting the wake-up of the car machine screen.
[0177] It is worth mentioning that the setting of the voiceprint verification threshold needs to balance security and convenience. The higher the threshold, the stricter the voiceprint verification, the lower the false wake-up rate, but the rejection rate of the user's own wake-up may increase; the lower the threshold, the easier it is to pass the voiceprint verification, and it is more convenient for the user to wake up by themselves, but the false wake-up rate may increase. The threshold can be flexibly adjusted according to the actual application scenario and user preferences. At the same time, as the number of times the user uses the custom wake-up word increases, the first voiceprint feature can be dynamically updated to improve the accuracy of voiceprint matching.
[0178] In summary, steps S3300 and S3400 form two lines of defense through wake-up position judgment and voiceprint verification, which can effectively reduce the false wake-up rate of the custom wake-up word and improve the security and reliability of in-vehicle voice interaction. Compared with the traditional fixed wake-up word scheme, this scheme supports users to independently register personalized wake-up words and provides a better user experience; compared with the simple voiceprint recognition scheme, this method integrates the sound source localization technology, introduces wake-up position judgment before voiceprint verification, reduces unnecessary voiceprint comparisons, and reduces system overhead. Therefore, while improving user convenience, this scheme takes into account the real-time performance, accuracy, and security of the system, providing a new idea for in-vehicle voice interaction.
[0179] Embodiment 2
[0180] Based on Embodiment 1, this embodiment provides a system for configuring custom wake-up words for the in-vehicle screen based on voice control, as Figure 6 shown, including:
[0181] Noise level classification module: used to collect the in-vehicle noise data of the target vehicle, construct an in-vehicle noise database, and the in-vehicle noise database includes n1 noise records; divide the n1 noise records in the in-vehicle noise database into N1 noise levels;
[0182] First voiceprint feature acquisition module: used to obtain the voice sample of the user's custom wake-up word, extract the voiceprint feature vector of the custom wake-up word voice sample, and mark it as the custom wake-up word voiceprint feature vector; measure the similarity between the custom wake-up word voiceprint feature vector and each voiceprint feature vector in the pre-constructed voiceprint feature library, and match the first voiceprint feature with the highest similarity to the custom wake-up word voiceprint feature vector.
[0183] Model construction module: used to construct wake-up word detection sub-models corresponding to N1 noise levels according to the in-vehicle noise database and the first voiceprint feature.
[0184] Wake-up word recognition module: used to collect in-vehicle audio data in real time, and perform wake-up word recognition based on the in-vehicle audio data, wake-up word detection sub-model, and first voiceprint feature to trigger the wake-up of the in-vehicle screen.
[0185] In the noise level classification module, the n1 noise records, each noise record includes the attribute data of a noise segment; the attribute data of the noise segment includes a noise segment number, noise segment data, a noise type label, a noise energy value, and a noise spectrum feature vector;
[0186] In the noise level classification module, the division of the n1 noise records in the in-vehicle noise database into N1 noise levels includes:
[0187] Step S1210, according to the noise energy value and the noise spectrum feature vector of each noise record in the in-vehicle noise database, calculate the energy mean value E i , the spectral centroid F i , and the spectral dispersion D i , to form the noise acoustic feature vector [E i ,F i ,D i of each noise record; where E i is the energy mean value of the i-th noise record, F i is the spectral centroid of the i-th noise record, D i is the spectral dispersion of the i-th noise record, and [E i ,F i ,D i represents the noise acoustic feature vector of the i-th noise record;
[0188] Step S1220, using the noise acoustic feature vector [E i ,F i ,D i as the feature description, cluster the n1 noise records to obtain N1 noise cluster centers;
[0189] Step S1230, calculate the average silhouette coefficient SC. If SC ≤ SC', adjust the number of noise cluster centers N1, and return to Step S1220 to re-cluster until SC > SC', and output the clustering result; SC' is a preset silhouette coefficient threshold.
[0190] In the first voiceprint feature acquisition module, the obtaining of the first voiceprint feature with the highest similarity to the voiceprint feature vector of the custom wake-up word includes:
[0191] Step S1310, collect voice samples of multiple users, extract the voiceprint feature vectors of the voice samples, and construct a voiceprint feature library; the voiceprint feature vector includes the fundamental frequency, formants, and speech rate;
[0192] Step S1320: Measure the similarity between the custom wake-word voiceprint feature vector and each voiceprint feature vector in the voiceprint feature library to obtain a similarity score; sort according to the similarity score, and select the voiceprint feature vector in the voiceprint feature library with the highest score, which is marked as the first voiceprint feature.
[0193] In the model construction module, the construction of the wake-word detection sub-models corresponding to N1 noise levels includes:
[0194] Step S2100: Based on the N1 noise levels in the in-vehicle noise database, construct N1 noise data subsets, and each noise data subset corresponds to one noise level;
[0195] The construction of the N1 noise data subsets includes:
[0196] Traverse the n1 noise records in the in-vehicle noise database, and divide them into the corresponding N1 noise data subsets according to the noise level labels of each noise record;
[0197] Step S2200: For each noise level, construct an independent wake-word detection sub-model respectively;
[0198] Step S2300: Traverse the N1 noise levels, execute Step S2200, and obtain all the wake-word detection sub-models corresponding to the N1 noise levels. The output of the wake-word detection sub-model is the custom wake-word detection result represented by a boolean value. If it is 1, it means the custom wake-word is detected; if it is 0, it means the custom wake-word is not detected.
[0199] The said Step S2200 includes:
[0200] Step S2210: According to the preset data set division ratio, divide the j-th noise data subset into a training set, a validation set, and a test set; 1 ≤ j ≤ N1;
[0201] Step S2220: Fuse the training set in the j-th noise data subset with the first voiceprint feature to construct a wake-word detection training set for noise level j;
[0202] Step S2230: Use the wake-word detection training set for noise level j as the input to train the initial wake-word detection sub-model at noise level j;
[0203] Step S2240: Use the validation set in the j-th noise data subset to optimize the model parameters and obtain the final wake-word detection sub-model at noise level j;
[0204] Step S2250: Use the test set in the j-th noise data subset to evaluate the performance metrics of the final wake-word detection sub-model at noise level j.
[0205] In the wake word recognition module, the wake word recognition based on in-vehicle audio data, the wake word detection sub-model, and the first voiceprint feature to trigger the wake-up of the in-vehicle screen includes:
[0206] Step S3100: Collect in-vehicle audio data in real time, judge the noise level to which the in-vehicle audio data belongs, and select the wake word detection sub-model corresponding to the noise level to which the in-vehicle audio data belongs to perform custom wake word detection, and judge whether a custom wake word is detected;
[0207] Step S3200: If no custom wake word is detected, jump to step S3100 to continue the next round of in-vehicle audio data collection and detection; if a custom wake word is detected, collect wake word audio data through the microphone array, locate the wake word speaker, and obtain the horizontal azimuth angle of the wake word speaker relative to the vehicle-mounted microphone array and the distance ;
[0208] Step S3300: Set the wake-up angle threshold range and the distance threshold range ; If and , it is determined that the detected custom wake word comes from a reasonable position inside the vehicle and is recognized as a wake word to be confirmed, otherwise it is regarded as a false wake-up caused by external noise, the wake-up is rejected and step S3100 is jumped to continue the next round of in-vehicle audio data collection and custom wake word detection;
[0209] Step S3400: Extract acoustic features from the wake word audio data determined to be a wake word to be confirmed to construct a second voiceprint feature; calculate the similarity score SIM between the second voiceprint feature and the first voiceprint feature. If the similarity score SIM is greater than the preset voiceprint verification threshold, the wake word to be confirmed is recognized as a valid wake word to trigger the wake-up of the in-vehicle screen; otherwise, it is recognized as an invalid wake word, the wake-up of the in-vehicle screen is rejected and step S3100 is jumped to continue the next round of in-vehicle audio data collection and custom wake word detection.
[0210] The said step S3100 includes:
[0211] Step S3110: Extract features from the in-vehicle audio data collected in real time to obtain the real-time energy value Es, the real-time spectral centroid Fs, and the real-time spectral dispersion Ds, and construct a real-time acoustic feature vector [Es, Fs, Ds];
[0212] Step S3120: Compare the real-time acoustic feature vector [Es, Fs, Ds] with N1 noise clustering centers, calculate the Euclidean distances between the real-time acoustic feature vector and each noise clustering center, and select the noise level corresponding to the noise clustering center with the minimum Euclidean distance as the noise level to which the in-vehicle audio data collected in real time belongs.
[0213] Step S3130: According to the noise level to which the in-vehicle audio data belongs, select the corresponding wake word detection sub-model to perform custom wake word detection.
[0214] The said step S3200 includes:
[0215] Step S3210: If a custom wake word is detected, collect wake word audio data through the microphone array. The wake word audio data has M voice signals, and the wake word audio data of the m-th voice signal is denoted as , where t is time, ;
[0216] Step S3220: Perform voice activity detection on the wake word audio data , extract the wake word voice segments received by each microphone, find the start and end times of the wake word voice segments, and denote the wake word voice segment as ; Select one microphone from the microphone array as the reference microphone, and the other M - 1 microphones as non-reference microphones, and estimate the time delay between the wake word voice segments of the non-reference microphones and the reference microphone, where , , is the time delay of the wake word voice segment of the -th non-reference microphone relative to the reference microphone; Based on the geometric layout of the microphone array, construct M - 1 equations:
[0217] ;
[0218] where is the position vector of the -th non-reference microphone relative to the reference microphone, is the speed of sound, is the horizontal azimuth angle of the sound source, is the elevation angle of the sound source;
[0219] Step S3230: Solve the M - 1 equations simultaneously to obtain the horizontal azimuth angle and distance of the wake word speaker relative to the vehicle-mounted microphone array, where .
[0220] The step S3300 includes:
[0221] Step S3310, establish a three-dimensional space model of the vehicle interior according to the spatial structure parameters and seat layout of the vehicle model; in the three-dimensional space model of the vehicle interior, calibrate the installation positions of the in-vehicle microphone arrays.
[0222] Step S3320, in the three-dimensional space model of the vehicle interior, delimit wake-up regions, each wake-up region is represented by a spatial polygon, and record the corner point coordinates of the polygon, where is a positive integer.
[0223] Step S3330, project the wake-up regions onto the horizontal plane with the microphone array as the origin, and extract the azimuth span and distance span of each projected polygon, where is the total number of wake-up regions, is the minimum azimuth angle of the projected polygon, representing the leftmost boundary of the wake-up region outward from the microphone array on the horizontal plane; is the maximum azimuth angle of the projected polygon, representing the rightmost boundary of the wake-up region outward from the microphone array on the horizontal plane; is the shortest distance of the projected polygon, representing the nearest boundary of the wake-up region starting from the microphone array; is the longest distance of the projected polygon, representing the farthest boundary of the wake-up region starting from the microphone array;
[0224] Step S3340, take the union of the azimuth spans of each to determine the wake-up angle threshold range , ; take the union of the distance spans of each to determine the wake-up distance threshold range .
[0225] Embodiment 3
[0226] This embodiment discloses an electronic device, which may include one or more processors and one or more memories. Among them, computer-readable code is stored in the memory, and when the computer-readable code is run by one or more processors, it can execute the above-mentioned method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control.
[0227] The method or system according to the embodiments of the present application can also be implemented by means of the architecture of an electronic device. The electronic device may include a bus, one or more CPUs, a read-only memory (ROM), a random access memory (RAM), a communication port connected to a network, input / output components, a hard disk, etc. The storage device in the electronic device, such as the ROM or the hard disk, can store the method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control provided by the present application. The method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control may, for example, include: collecting in-vehicle noise data of a target vehicle, constructing an in-vehicle noise database, where the in-vehicle noise database includes n1 noise records; dividing the n1 noise records in the in-vehicle noise database into N1 noise levels, and recording the noise level tags of each noise record; obtaining a custom wake-up word voice sample of a user, extracting the voiceprint feature vector of the custom wake-up word voice sample, and marking it as the custom wake-up word voiceprint feature vector; measuring the similarity between the custom wake-up word voiceprint feature vector and each voiceprint feature vector in a pre-constructed voiceprint feature library, and matching to obtain the first voiceprint feature with the highest similarity to the custom wake-up word voiceprint feature vector; constructing wake-up word detection sub-models corresponding to N1 noise levels according to the in-vehicle noise database and the first voiceprint feature; collecting in-vehicle audio data in real time, and performing wake-up word recognition based on the in-vehicle audio data, the wake-up word detection sub-models, and the first voiceprint feature to trigger the wake-up of the vehicle-mounted screen.
[0228] Further, the electronic device may further include a user interface. Of course, the architecture disclosed in the present invention is only exemplary. When implementing different devices, one or more components of the electronic device disclosed in the present invention may be omitted according to actual needs.
[0229] Embodiment 4
[0230] This embodiment discloses a computer-readable storage medium with computer-readable instructions stored thereon. When the computer-readable instructions are run by a processor, the method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control according to the embodiments of the present application can be executed. The storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may, for example, include random access memory (RAM) and cache memory, etc. Non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.
[0231] In addition, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the present application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be run by a processor to execute instructions corresponding to the method steps provided by the present application, such as: collecting in-vehicle noise data of a target vehicle, constructing an in-vehicle noise database, where the in-vehicle noise database includes n1 noise records; dividing the n1 noise records in the in-vehicle noise database into N1 noise levels and recording the noise level labels of each noise record; obtaining a user's custom wake-up word voice sample, extracting the voiceprint feature vector of the custom wake-up word voice sample, and marking it as the custom wake-up word voiceprint feature vector; measuring the similarity between the custom wake-up word voiceprint feature vector and each voiceprint feature vector in the pre-constructed voiceprint feature library, and matching to obtain the first voiceprint feature with the highest similarity to the custom wake-up word voiceprint feature vector; constructing wake-up word detection sub-models corresponding to N1 noise levels according to the in-vehicle noise database and the first voiceprint feature; collecting in-vehicle audio data in real time, and performing wake-up word recognition based on the in-vehicle audio data, the wake-up word detection sub-models, and the first voiceprint feature to trigger the wake-up of the in-vehicle screen. When this computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed.
[0232] The methods, systems, and devices of the present application can be implemented in many ways. For example, the methods, systems, and devices of the present application can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is only for illustration, and the steps of the method of the present application are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present application can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to the present application. Therefore, the present application also covers a recording medium storing a program for executing the method according to the present application.
[0233] In addition, parts of the above technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive elaboration.
[0234] As described above in the specific embodiments, the objectives, technical solutions, and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for configuring a custom wake-up word for a car infotainment screen based on voice control, characterized in that, The method includes: Collecting in-vehicle noise data of a target vehicle to construct an in-vehicle noise database, where the in-vehicle noise database includes n1 noise records; dividing the n1 noise records in the in-vehicle noise database into N1 noise levels and recording the noise level labels of each noise record; obtaining a user-defined wake-up word voice sample, extracting the voiceprint feature vector of the user-defined wake-up word voice sample, and marking it as the user-defined wake-up word voiceprint feature vector; measuring the similarity between the user-defined wake-up word voiceprint feature vector and each voiceprint feature vector in a pre-constructed voiceprint feature library, and matching to obtain the first voiceprint feature with the highest similarity to the user-defined wake-up word voiceprint feature. Constructing wake-up word detection sub-models corresponding to N1 noise levels according to the in-vehicle noise database and the first voiceprint feature. Collecting in-vehicle audio data in real time, performing wake-up word recognition based on the in-vehicle audio data and the wake-up word detection sub-models, and determining whether the user-defined wake-up word is detected; if the user-defined wake-up word is not detected, then performing the next round of in-vehicle audio data collection and wake-up word recognition; if the user-defined wake-up word is detected, then constructing a second voiceprint feature, and determining whether to trigger the wake-up of the in-vehicle infotainment (IVI) screen according to the first voiceprint feature and the second voiceprint feature.
2. The method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control according to claim 1, wherein For each of the n1 noise records, each noise record includes attribute data of a noise segment; the attribute data of the noise segment includes a noise segment number, noise segment data, a noise type label, a noise energy value, and a noise spectrum feature vector. The dividing the n1 noise records in the in-vehicle noise database into N1 noise levels includes: clustering the n1 noise records according to the noise energy values and noise spectrum feature vectors of the n1 noise records in the in-vehicle noise database.
3. The method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control according to claim 2, wherein The clustering of the n1 noise records includes: Step S1210: Calculate the energy mean value E based on the noise energy value and the noise spectrum feature vector of each noise record in the in-vehicle noise database i , the spectrum centroid F i , and the spectrum dispersion D i to form the noise acoustic feature vector [E i , F i , D i for each noise record; where E i is the energy mean value of the i-th noise record, F i is the spectrum centroid of the i-th noise record, D i is the spectrum dispersion of the i-th noise record, and [E i , F i , D i represents the noise acoustic feature vector of the i-th noise record; Step S1220, using the noise acoustic feature vector [E i , F i , D i as the feature description, clustering n1 noise records to obtain N1 noise clustering centers; Step S1230, calculating the average silhouette coefficient SC. If SC ≤ SC', then adjusting the number of noise clustering centers N1, and returning to step S1220 for re-clustering until SC > SC', and outputting the clustering result; SC' is a preset silhouette coefficient threshold.
4. The method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control according to claim 1, wherein The method for constructing the voiceprint feature library is: collecting voice samples of multiple users, extracting the voiceprint feature vectors of the voice samples, and constructing a voiceprint feature library. The voiceprint feature vector includes fundamental frequency, formant, and speech rate. The obtaining the first voiceprint feature with the highest similarity to the user-defined wake-up word voiceprint feature includes: Measuring the similarity between the user-defined wake-up word voiceprint feature vector and each voiceprint feature vector in the voiceprint feature library to obtain similarity scores. Sorting according to the similarity scores, and selecting the voiceprint feature vector in the voiceprint feature library with the highest score, and marking it as the first voiceprint feature.
5. The method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control according to claim 1, wherein The constructing wake-up word detection sub-models corresponding to N1 noise levels includes: Based on the N1 noise levels in the in-vehicle noise database, constructing N1 noise data subsets, where each noise data subset corresponds to one noise level. Traverse N1 noise levels. For each noise level, construct an independent wake word detection sub-model respectively; obtain all N1 wake word detection sub-models corresponding to the noise levels. The output of the wake word detection sub-model is a custom wake word detection result represented by a boolean value. If it is 1, it means the custom wake word is detected; if it is 0, it means the custom wake word is not detected. The construction of the N1 noise data subsets includes: Traverse n1 noise records in the in-vehicle noise database and divide them into the corresponding N1 noise data subsets according to the noise level labels of each noise record.
6. The method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control according to claim 5, wherein For each noise level, constructing an independent wake word detection sub-model respectively includes: According to the preset data set division ratio, divide the j-th noise data subset into a training set, a validation set, and a test set; 1 ≤ j ≤ N1. Fuse the training set in the j-th noise data subset with the first voiceprint feature to construct a wake word detection training set for noise level j. Taking the wake word detection training set for noise level j as the input, train the initial wake word detection sub-model at noise level j.
7. The method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control according to claim 1, wherein The method for wake word recognition based on in-vehicle audio data and the wake word detection sub-model is: determine the noise level to which the in-vehicle audio data belongs, and select the wake word detection sub-model corresponding to the noise level to which the in-vehicle audio data belongs to perform custom wake word detection, and determine whether the custom wake word is detected. The method for constructing the second voiceprint feature and determining whether to trigger the in-vehicle screen wake-up according to the first voiceprint feature and the second voiceprint feature includes: If a custom wake word is detected, the microphone array is used to collect the audio data of the wake word, and the speaker of the wake word is located to obtain the horizontal azimuth angle of the wake word speaker relative to the in-vehicle microphone array and the distance ; Set the wake-up angle threshold range and the distance threshold range ; If and , it is determined that the detected custom wake-up word comes from a reasonable position inside the vehicle and is identified as a wake-up word to be confirmed; otherwise, it is regarded as a false wake-up, and the next round of in-vehicle audio data collection and custom wake-up word detection continues; Extract acoustic features from the wake word audio data determined to be a wake word to be confirmed to construct the second voiceprint feature; calculate the similarity score SIM between the second voiceprint feature and the first voiceprint feature. If the similarity score SIM is greater than the preset voiceprint verification threshold, recognize the wake word to be confirmed as a valid wake word and trigger the in-vehicle screen wake-up; otherwise, recognize it as an invalid wake word, reject the in-vehicle screen wake-up, and continue the next round of in-vehicle audio data collection and custom wake word detection.
8. The method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control according to claim 7, wherein The determination of the noise level to which the in-vehicle audio data belongs includes: Extract features from the real-time collected in-vehicle audio data to obtain the real-time energy value Es, the real-time spectral centroid Fs, and the real-time spectral dispersion Ds, and construct a real-time acoustic feature vector [Es, Fs, Ds]. Compare the real-time acoustic feature vector [Es, Fs, Ds] with N1 noise cluster centers, calculate the Euclidean distance between the real-time acoustic feature vector and each noise cluster center, and select the noise level corresponding to the noise cluster center with the smallest Euclidean distance as the noise level to which the real-time collected in-vehicle audio data belongs.
9. The method for configuring a custom wake-up word for a vehicle-mounted screen based on voice control according to claim 7, wherein The obtained horizontal azimuth angle of the wake word speaker relative to the vehicle-mounted microphone array and distance include: Collect wake-up word audio data through a microphone array. The wake-up word audio data has M voice signals, and the wake-up word audio data of the m-th voice signal is denoted as , where t is time, ; Perform voice activity detection on the wake word audio data to extract the wake word voice segments received by each microphone find the start and end times of the wake word voice segments and record the wake word voice segments as ; select one microphone from the microphone array as the reference microphone, and the other M - 1 microphones as non - reference microphones, and estimate the time delay of the wake word voice segments of the non - reference microphones relative to the reference microphone where is the time delay of the wake word voice segment of the th non - reference microphone relative to the reference microphone; Based on the geometric layout of the microphone array, construct M - 1 equations: ; wherein is the position vector of the th non-reference microphone relative to the reference microphone, is the speed of sound, is the horizontal azimuth angle of the sound source, is the elevation angle of the sound source; Solve the M-1 equations simultaneously to obtain the horizontal azimuth angle of the wake-up word speaker relative to the microphone array: and distance ,in .
10. A vehicle-mounted screen custom wake-up word configuration system based on voice control, which is used to implement the vehicle-mounted screen custom wake-up word configuration method according to any one of claims 1-9, characterized in that, The system includes: Noise level division module: used to collect the in-vehicle noise data of the target vehicle, construct an in-vehicle noise database, and the in-vehicle noise database includes n1 noise records; divide the n1 noise records in the in-vehicle noise database into N1 noise levels. The first voiceprint feature acquisition module: It is used to acquire the custom wake-up word voice samples of the user, extract the voiceprint feature vectors of the custom wake-up word voice samples, and mark them as the custom wake-up word voiceprint feature vectors; measure the similarity between the custom wake-up word voiceprint feature vectors and each voiceprint feature vector in the pre-constructed voiceprint feature library, and match to obtain the first voiceprint feature with the highest similarity to the custom wake-up word voiceprint feature vectors. The model construction module: It is used to construct wake-up word detection sub-models corresponding to N1 noise levels according to the in-vehicle noise database and the first voiceprint feature. The wake-up word recognition module: It is used to collect in-vehicle audio data in real time, and perform wake-up word recognition based on the in-vehicle audio data, the wake-up word detection sub-model and the first voiceprint feature, and trigger the wake-up of the in-vehicle computer screen.
Citation Information
Patent Citations
A method and system for automatically filtering wake words
CN109360552B
Voice wake-up method, device and equipment and medium
CN113611294A
Vehicle-mounted voiceprint awakening method and device, electronic equipment and storage medium
CN117935841A
Vehicle awakening method and device, electronic equipment and storage medium
CN118366460A