Anti-interference voiceprint recognition system and method in vehicle scene and electronic equipment
By constructing a voiceprint model using a microphone array and the GMM-UBM framework in a vehicle scenario, and combining hardware optimization, the interference problem of voiceprint recognition in a vehicle scenario was solved, the recognition accuracy and response speed were improved, and the false wake-up rate was reduced.
Patent Information
- Application Number
- CN202511695448.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-10
AI Technical Summary
Voiceprint recognition in vehicle scenarios faces interference from various sources, changes in the driver's own vocal cords, and interference from special groups on the driver's voice identity verification, resulting in problems such as low recognition accuracy, prolonged response time, and high false wake-up rate.
A microphone array is used to achieve directional sound pickup, and a voiceprint model is built in combination with the GMM-UBM framework. Through the speech acquisition, voiceprint registration, real-time verification and dynamic update modules, combined with hardware optimization such as chip upgrade, computing acceleration and storage compression, the voiceprint recognition in vehicle scenarios is optimized.
It improves the accuracy of voiceprint recognition, shortens the response time, reduces the false wake-up rate, and enhances the voiceprint recognition performance in vehicle scenarios.
Smart Images

Figure CN121506180A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition, and in particular to an anti-interference voiceprint recognition system, an anti-interference voiceprint recognition method for vehicle scenarios, electronic devices, storage media, and vehicles. Background Technology
[0002] Given the unique challenges of voiceprint recognition in in-vehicle infotainment systems—variable in-vehicle noise (tire noise / air conditioning / children crying), different voice sources in different seats, vocal cord changes due to colds, interference from children's high-frequency voices, voiceprint drift caused by the driver's cold, voiceprint confusion between twins, and external noise interference—this paper proposes an anti-interference voiceprint recognition solution for in-vehicle scenarios. This solution optimizes both software and hardware to improve voiceprint recognition accuracy, shorten response time, and reduce false wake-up rates; it also enhances user experience, convenience, and accuracy; saves time; and further reflects the intelligence of the in-vehicle infotainment system. Summary of the Invention
[0003] The purpose of this invention is to provide a deep learning-based vehicle turning and obstacle avoidance method, a deep learning-based vehicle turning and obstacle avoidance system, an electronic device, a storage medium, and a vehicle, thereby solving at least one of a number of technical problems.
[0004] 1. The problem of multiple interference sources interfering with the owner's voice authentication in vehicle scenarios.
[0005] 2. Problems with the owner's own vocal cords leading to issues with the owner's voice verification.
[0006] 3. The problem of interference with voice identity verification for car owners by special groups.
[0007] This invention provides the following solution:
[0008] According to a first aspect of the present invention, an anti-interference voiceprint recognition system for vehicle scenarios is provided, comprising:
[0009] The voice acquisition module uses a microphone array to achieve directional sound pickup;
[0010] The voiceprint registration module is used to collect and store voice samples of the driver's wake-up words in multiple scenarios;
[0011] The voiceprint modeling module is based on the GMM-UBM framework to build and train voiceprint models.
[0012] The real-time verification module integrates a wake-word detection unit, a voice activity detection unit, a voiceprint feature extraction unit, and a matching decision unit to achieve rapid identity verification after wake-up triggering;
[0013] The dynamic update module adjusts model parameters based on changes in the scene and the user's voice status.
[0014] The hardware optimization module improves voiceprint recognition performance through chip upgrades, computing acceleration, storage compression, and power consumption control.
[0015] To combat interference in voiceprint recognition caused by varying in-vehicle noise, different voice sources, and changes in the state of the user's vocal cords, the voice acquisition module, voiceprint registration module, voiceprint modeling module, real-time verification module, dynamic update module, and hardware optimization module work together.
[0016] According to a second aspect of the present invention, an anti-interference voiceprint recognition method in a vehicle scenario is provided, comprising: voice acquisition, voiceprint registration, voiceprint modeling, real-time verification, dynamic updating, and hardware optimization;
[0017] Voice acquisition includes recording multiple wake-up word voice samples at time intervals greater than or equal to a preset minimum, covering various vehicle scenarios;
[0018] Among them, various vehicle scenarios are constructed based on the noise background of driving speed, driving environment, and the working status of on-board equipment;
[0019] Based on preprocessing and feature extraction, voiceprint registration is performed for vehicle owners who are specific speakers.
[0020] Voiceprint model training includes voiceprint modeling;
[0021] Among them, voiceprint modeling is based on the GMM-UBM framework, and a model corresponding to a specific speaker is adaptively generated by combining a general background model with the maximum a posteriori probability.
[0022] Real-time verification includes collecting voice signals through a microphone array and triggering the verification process by detecting a wake word;
[0023] The verification process includes sequentially performing voice activity detection, voiceprint feature extraction, voiceprint matching, and threshold determination.
[0024] During this period, a real-time response of less than 1 second is achieved, which serves as the data collection window for verifying the correctness of the results;
[0025] Dynamic updates include integrating passive triggering, active periodic, and event-driven strategies to dynamically optimize the model based on specific speaker voice changes and scene interference;
[0026] The optimization model includes optimization from both software algorithm and hardware resource aspects.
[0027] Among them, optimizations are made with the aim of reducing in-vehicle noise, changing the position of the voice, and reducing interference caused by changes in the vocal cords of specific speakers;
[0028] The evaluation is conducted using the criteria of improving recognition accuracy, shortening response time, and reducing false wake-up rate as optimization standards.
[0029] Furthermore, it also includes:
[0030] Voiceprint feature extraction includes:
[0031] A front-end processing workflow of frame segmentation, windowing, and pre-emphasis is adopted to extract MFCC, PLP, or x-vector as voiceprint features;
[0032] Frame-level incremental computation is achieved through streaming processing, and hardware acceleration is performed using DSP / NPU.
[0033] Furthermore, it also includes:
[0034] Real-time verification, including:
[0035] Based on the detection of wake words, with the goal of achieving a latency response of less than 200ms, a lightweight CNN / RNN end-to-end model or an HMM-DNN hybrid model is used for model quantization and pruning optimization.
[0036] Furthermore, vehicle-related scenarios also include: scenarios where children's voices interfere;
[0037] For scenarios involving interference with children's voices, based on the fact that children's fundamental frequency is generally >300Hz and their formants are concentrated in 4-6kHz, a frequency band suppression processing above 6kHz is adopted. By comparing the distribution of the first and second formants corresponding to adults and specific speakers, a decision is made on whether to reject the voice source of that specific speaker.
[0038] If a second peak is detected at >2800Hz, the speaker is identified as a child, and the speaker's voice source is denied authentication.
[0039] Furthermore, vehicle scenarios also include: voiceprint drift caused by a cold;
[0040] For scenarios where a speaker experiences voiceprint drift due to a cold, a multi-state model is adopted. During training, aerobic / hoarse voice synthesis data generated based on VAE acoustic transformation is incorporated, and the MFCC weighting coefficients are dynamically adjusted by detecting the fundamental frequency offset in real time.
[0041] Furthermore, vehicle-related scenarios also include: scenarios where children imitate adult voices;
[0042] For scenarios where children imitate adult voices, dynamic spectrogram detection is added to calculate frame-level spectral contrast. If the contrast is lower than a preset threshold, the sound source is determined to be a child imitating, and the identity verification of the sound source is refused.
[0043] Furthermore, vehicle scenarios also include scenarios where the voiceprints of twins are confused;
[0044] For scenarios where the voiceprints of twins are confused, behavioral features are integrated into the voiceprint matching process, including judging by the duration of the wake word pronunciation and judging the correlation between the control behavior of the control terminal and the voice wake-up behavior, to decide whether to perform identity verification.
[0045] Further hardware optimizations include:
[0046] Using the ARMCMSIS-DSP library to achieve a 4x speedup in MFCC extraction;
[0047] The GMM scoring process was optimized using Q15 format fixed-point quantization.
[0048] PCA dimensionality reduction combined with Huffman coding was used to compress the voiceprint model parameters from 300KB to 80KB / user.
[0049] Voiceprint comparison is initiated only after the wake word is triggered. During the comparison, the large CPU core is awakened, while the small core remains in standby mode at other times.
[0050] Furthermore, the threshold decision adopts a dynamic threshold strategy, adjusting the threshold according to the security level. The FAR is set to 0.01 for door disabling scenarios and 0.1 for vehicle-machine interaction scenarios.
[0051] According to a third aspect of the present invention, an electronic device is provided, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0052] The memory stores a computer program, which, when executed by the processor, causes the processor to perform steps such as an anti-interference voiceprint recognition method in a vehicle scenario.
[0053] According to a fourth aspect of the present invention, a computer-readable storage medium is provided storing a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform steps such as an anti-interference voiceprint recognition method in a vehicle scenario.
[0054] According to a fifth aspect of the present invention, a vehicle is provided, comprising:
[0055] Electronic devices used to implement steps of anti-interference voiceprint recognition methods, such as those used in vehicle scenarios;
[0056] The processor runs programs, and when the programs are running, they execute steps such as anti-interference voiceprint recognition methods in vehicle scenarios based on data output from electronic devices.
[0057] Storage medium used to store programs that, when running, execute steps such as an anti-interference voiceprint recognition method in a vehicle scenario based on data output from an electronic device.
[0058] The above solution achieves the following beneficial technical effects:
[0059] This application achieves sound source shielding for children in scenarios where children's voices interfere with their speech by using frequency band suppression processing above 6kHz, based on the fact that children's fundamental frequency is generally >300Hz and their formant peaks are concentrated in 4-6kHz. By comparing the distribution of the first and second formants corresponding to adults and specific speakers, it decides whether to reject the sound source of that specific speaker.
[0060] This application employs a multi-state model, incorporates aerophone / hoarse voice synthesis data generated based on VAE acoustic transformation during training, and dynamically adjusts the MFCC weighting coefficients by detecting the fundamental frequency offset in real time, thereby achieving enhanced vehicle owner recognition in scenarios where the voiceprint drift is caused by a cold in a specific speaker (e.g., a car owner).
[0061] This application enhances vehicle owner recognition in scenarios where twin voiceprints are confused by fusing behavioral features during the voiceprint matching process.
[0062] This application enhances the voiceprint comparison through hardware, only initiating it after the wake word is triggered. During the comparison, the large CPU core is activated, while the small core remains in standby mode at other times, thereby improving the vehicle owner recognition capability. Attached Figure Description
[0063] Figure 1 This is a flowchart of an anti-interference voiceprint recognition method in a vehicle scenario provided by one or more embodiments of the present invention.
[0064] Figure 2 This is a structural diagram of an anti-interference voiceprint recognition system in a vehicle scenario provided by one or more embodiments of the present invention.
[0065] Figure 3 This is a schematic diagram of the basic framework of a voiceprint recognition solution provided in a specific embodiment of the present invention.
[0066] Figure 4 This is a schematic diagram of a real-time wake-up verification process provided in a specific embodiment of the present invention.
[0067] Figure 5 This is a block diagram of an electronic device structure for an anti-interference voiceprint recognition method in a vehicle scenario provided by one or more embodiments of the present invention. Detailed Implementation
[0068] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0069] Figure 2 This is a structural diagram of an anti-interference voiceprint recognition system in a vehicle scenario provided by one or more embodiments of the present invention.
[0070] like Figure 2 The anti-interference voiceprint recognition system for vehicle scenarios shown includes:
[0071] The voice acquisition module uses a microphone array to achieve directional sound pickup;
[0072] The voiceprint registration module is used to collect and store voice samples of the driver's wake-up words in multiple scenarios;
[0073] The voiceprint modeling module is based on the GMM-UBM framework to build and train voiceprint models.
[0074] The real-time verification module integrates a wake-word detection unit, a voice activity detection unit, a voiceprint feature extraction unit, and a matching decision unit to achieve rapid identity verification after wake-up triggering;
[0075] The dynamic update module adjusts model parameters based on changes in the scene and the user's voice status.
[0076] The hardware optimization module improves voiceprint recognition performance through chip upgrades, computing acceleration, storage compression, and power consumption control.
[0077] To combat interference in voiceprint recognition caused by varying in-vehicle noise, different voice sources, and changes in the state of the user's vocal cords, the voice acquisition module, voiceprint registration module, voiceprint modeling module, real-time verification module, dynamic update module, and hardware optimization module work together.
[0078] Specifically, the speech activity detection unit uses either the traditional approach of energy threshold + zero-crossing rate or the RNN-based deep learning approach to ensure that a valid speech segment containing sufficient identity information is extracted in 1-3 seconds.
[0079] The voiceprint matching decision unit supports log-likelihood ratio calculation (GMM-UBM model) or cosine similarity calculation (x-vector model), and can also be adapted to the ECAPA-TDNN end-to-end model to directly output matching scores.
[0080] Figure 1 This is a flowchart of an anti-interference voiceprint recognition method in a vehicle scenario provided by one or more embodiments of the present invention.
[0081] like Figure 1 The anti-interference voiceprint recognition method in the vehicle scenario shown includes: voice acquisition, voiceprint registration, voiceprint modeling, real-time verification, dynamic update and hardware optimization;
[0082] Step S1, voice acquisition includes recording multiple wake-up word voice samples at intervals greater than or equal to a preset minimum time interval, covering various vehicle scenarios;
[0083] Among them, various vehicle scenarios are constructed based on the noise background of driving speed, driving environment, and the working status of on-board equipment;
[0084] Step S2: Based on preprocessing and feature extraction, register the voiceprint of the corresponding vehicle owner as a specific speaker;
[0085] Step S3, voiceprint model training includes voiceprint modeling;
[0086] Among them, voiceprint modeling is based on the GMM-UBM framework, and a model corresponding to a specific speaker is adaptively generated by combining a general background model with the maximum a posteriori probability.
[0087] Step S4, real-time verification includes collecting voice signals through a microphone array and triggering the verification process by detecting a wake word;
[0088] The verification process includes sequentially performing voice activity detection, voiceprint feature extraction, voiceprint matching, and threshold determination.
[0089] During this period, a real-time response of less than 1 second is achieved, which serves as the data collection window for verifying the correctness of the results;
[0090] Step S5, dynamic update includes integrating passive triggering, active periodic and event-driven strategies, and dynamically optimizing the model based on specific speaker voice changes and scene interference;
[0091] The optimization model includes optimization from both software algorithm and hardware resource aspects.
[0092] Among them, optimizations are made with the aim of reducing in-vehicle noise, changing the position of the voice, and reducing interference caused by changes in the vocal cords of specific speakers;
[0093] The evaluation is conducted using the criteria of improving recognition accuracy, shortening response time, and reducing false wake-up rate as optimization standards.
[0094] In this embodiment, it also includes:
[0095] Voiceprint feature extraction includes:
[0096] A front-end processing workflow of frame segmentation, windowing, and pre-emphasis is adopted to extract MFCC, PLP, or x-vector as voiceprint features;
[0097] Frame-level incremental computation is achieved through streaming processing, and hardware acceleration is performed using DSP / NPU.
[0098] In this embodiment, it also includes:
[0099] Real-time verification, including:
[0100] Based on the detection of wake words, with the goal of achieving a latency response of less than 200ms, a lightweight CNN / RNN end-to-end model or an HMM-DNN hybrid model is used for model quantization and pruning optimization.
[0101] In this embodiment, the vehicle scenario also includes: a child's voice interference scenario;
[0102] For scenarios involving interference with children's voices, based on the fact that children's fundamental frequency is generally >300Hz and their formants are concentrated in 4-6kHz, a frequency band suppression processing above 6kHz is adopted. By comparing the distribution of the first and second formants corresponding to adults and specific speakers, a decision is made on whether to reject the voice source of that specific speaker.
[0103] If a second peak is detected at >2800Hz, the speaker is identified as a child, and the speaker's voice source is denied authentication.
[0104] In this embodiment, the vehicle scenario also includes: a voiceprint drift scenario caused by a cold;
[0105] For scenarios where a speaker experiences voiceprint drift due to a cold, a multi-state model is adopted. During training, aerobic / hoarse voice synthesis data generated based on VAE acoustic transformation is incorporated, and the MFCC weighting coefficients are dynamically adjusted by detecting the fundamental frequency offset in real time.
[0106] In this embodiment, the vehicle scenario also includes a scenario where a child imitates an adult's voice;
[0107] For scenarios where children imitate adult voices, dynamic spectrogram detection is added to calculate frame-level spectral contrast. If the contrast is lower than a preset threshold, the sound source is determined to be a child imitating, and the identity verification of the sound source is refused.
[0108] In this embodiment, the vehicle scenario also includes: a scenario where the voiceprints of twins are confused;
[0109] For scenarios where the voiceprints of twins are confused, behavioral features are integrated into the voiceprint matching process, including judging by the duration of the wake word pronunciation and judging the correlation between the control behavior of the control terminal and the voice wake-up behavior, to decide whether to perform identity verification.
[0110] In this embodiment, hardware optimization includes:
[0111] Using the ARMCMSIS-DSP library to achieve a 4x speedup in MFCC extraction;
[0112] The GMM scoring process was optimized using Q15 format fixed-point quantization.
[0113] PCA dimensionality reduction combined with Huffman coding was used to compress the voiceprint model parameters from 300KB to 80KB / user.
[0114] Voiceprint comparison is initiated only after the wake word is triggered. During the comparison, the large CPU core is awakened, while the small core remains in standby mode at other times.
[0115] In this embodiment, the threshold decision adopts a dynamic threshold strategy, which adjusts the threshold according to the security level. The FAR is set to 0.01 for the door disabling scenario and 0.1 for the vehicle-machine interaction scenario.
[0116] It is worth noting that although this system / device only discloses the above-mentioned modules and units, it does not mean that this system / device is limited to the above-mentioned basic functional modules and units. On the contrary, what this invention intends to express is that, based on the above-mentioned basic functional modules, those skilled in the art can add one or more functional modules in combination with the prior art to form an infinite number of embodiments or technical solutions. That is to say, this system is open rather than closed. It cannot be assumed that the scope of protection of the claims of this invention is limited to the above-disclosed basic functional modules just because this embodiment only discloses a few basic functional modules.
[0117] In one specific embodiment, a voiceprint recognition solution that reduces and minimizes interference from voice wake-up is disclosed. This solution employs, for example... Figure 3 The basic framework of the voiceprint recognition solution is shown.
[0118] Specifically, it includes:
[0119] 1. Vehicle owner voiceprint registration stage
[0120] Recording requirements: More than 20 wake words (e.g., "Xiao Mou Xiao Mou"), with a 2-second interval between each word, covering as many scenarios as possible: quiet car interior (e.g., air conditioning off), driving (e.g., tire noise at 60km / h), music playback (e.g., volume 30%), etc., to achieve sufficiently rich sampling;
[0121] 2. Voiceprint Modeling (GMM-UBM Framework)
[0122] The core idea of GMM-UBM is to use Gaussian mixture models to model the statistical distribution of speech features (usually MFCCs and their derived features); it adopts an adaptive approach rather than training an independent model from scratch for each speaker.
[0123] Thanks to UBM's large amount of general data and likelihood ratio normalization, it exhibits good robustness to insufficient training data, channel variations, and limited noise.
[0124] Registering new speakers is completed quickly and adaptively, avoiding costly training from scratch;
[0125] Based on statistical modeling and Bayesian adaptation, the theoretical framework is clear;
[0126] For a long time, it was the mainstream and most advanced method for speaker recognition, with excellent performance;
[0127] GMM-UBM is a milestone in the history of voiceprint recognition. It cleverly utilizes the Universal Background Model (UBM) as powerful prior knowledge, efficiently generates a TargetSpeakerModel through Maximum A posteriori probability adaptation (MAP), and uses log-likelihood ratio (LLR) for normalized scoring, effectively overcoming data sparsity and improving the robustness and efficiency of the system.
[0128] 3. Real-time wake-up verification process
[0129] 1) Real-time wake-up verification is a process used in voiceprint recognition systems, such as... Figure 4 As shown, this technology enables immediate authentication after a user triggers the system with a specific wake word. This process is widely used in smart speakers, in-vehicle systems, and security access control systems, requiring low latency and high robustness. The following are its core processes and key technical points:
[0130] 2) Wake word detection (Keyword Spotting, KWS)
[0131] Function: Real-time monitoring of audio streams and recognition of preset wake words (e.g., "Hey [name]").
[0132] technology:
[0133] Lightweight models: end-to-end models based on CNN / RNN (e.g., TC-ResNet) and hybrid HMM-DNN models.
[0134] Deployment optimization: Model quantization and pruning to meet the real-time requirements of embedded devices (e.g., latency <200ms).
[0135] Output: The subsequent voiceprint verification process is triggered after the wake word is detected.
[0136] 3) Voice Activity Detection (VAD)
[0137] Function: Locates valid speech segments after wake-up and excludes silence / noise.
[0138] method:
[0139] Energy threshold + zero crossing rate: traditional lightweight solution.
[0140] Deep learning models, such as RNN-based VAD (which can achieve higher robustness).
[0141] Key requirement: Ensure that the captured audio contains sufficient identity information (usually 1-3 seconds).
[0142] 4) Voiceprint feature extraction
[0143] process:
[0144] Front-end processing: frame splitting (e.g., 20-30ms / frame), windowing, and pre-emphasis.
[0145] Feature generation: Extract MFCC (current mainstream method), PLP or DeepFeatures (such as x-vector).
[0146] Real-time optimization:
[0147] Streaming processing: frame-level incremental computation avoids waiting for the entire audio segment.
[0148] Hardware acceleration: Utilize DSP / NPU for parallel computation of MFCC.
[0149] 5) Voiceprint matching and decision-making
[0150] Matching method:
[0151] Template comparison (traditional method):
[0152] Target model: GMM-UBM or x-vector pre-stored during registration.
[0153] Similarity calculation: Log-likelihood ratio (GMM-UBM) or cosine similarity (x-vector).
[0154] End-to-end model (modern approach)
[0155] Direct input of features → output of scores (e.g., ECAPA-TDNN).
[0156] Decision-making logic:
[0157] Python
[0158] if similarity_score > threshold:
[0159] return "Verified" else:
[0160] return "Rejected"
[0161] Dynamic threshold: Adjusted according to security level (e.g., 0.01 FAR for doors, 0.1 FAR for speakers).
[0162] 6) Feedback and Implementation
[0163] Real-time response (e.g., <1s):
[0164] Success: The instruction was executed (e.g., open the door, play music).
[0165] Failure: Rejection prompt or transfer to human verification.
[0166] 4. Dynamic Model Update Strategy
[0167] Dynamic model updates are the core strategy for voiceprint recognition systems to maintain continuous performance, especially in complex and ever-changing scenarios where user voices change over time (e.g., with age, illness, or device replacement). The following is a hierarchical technical solution design and key considerations.
[0168] Update strategy design principles:
[0169] Stability-adaptability balance: avoiding model oscillations caused by frequent updates, while preventing data from becoming outdated;
[0170] Computational efficiency: Edge devices require low overhead, while the cloud can perform periodic batch updates;
[0171] Security defenses: mechanisms to prevent injection attacks and model contamination;
[0172] Personalization: Different update frequencies for different users (e.g., children vs. adults);
[0173] 5. Dynamic update trigger mechanism
[0174] The sampling strategies include passive triggering, active periodic sampling, and event-driven sampling.
[0175] 6. In this embodiment, the main scenario problems addressed include:
[0176] 1) Scene identification and solution: There is interference from children's voices.
[0177] Band suppression: Attenuate energy above 6kHz during the pre-emphasis phase (children's fundamental frequency is generally >300Hz, and the resonant peak is concentrated at 4-6kHz).
[0178] Formant verification: The distribution of speaker F1 / F2 (first / second formants) is shown in Table 1:
[0179] Table 1
[0180] speaker F1 range (Hz) F2 range (Hz) adult male 200-800 800-2300 6-year-old children 350-1100 1100-3500
[0181] Judgment logic: If F2 > 2800Hz, reject directly (excluding 99% of children's voiceprints);
[0182] 2) Scenario Assessment and Solution: The driver's cold caused voiceprint drift.
[0183] Solution:
[0184] Multi-state model: Breathy / hoarse sound synthesis data (based on VAE acoustic transformation) are added during training;
[0185] Short-time feature compensation: Real-time detection of fundamental frequency offset and dynamic adjustment of MFCC weighting coefficients.
[0186] 3) Scene identification and solution: "Daughter imitates voice to wake up car system"
[0187] Root cause analysis: Children faking adult low-frequency resonant peaks (deliberately lowering their voice) - Solution:
[0188] Add dynamic spectrogram detection (generally, the spectrum of adult pronunciation is continuous, but children's imitation has discontinuities);
[0189] / / Pseudocode: Spectral Continuity Decision
[0190] if (calc_spectral_contrast(frame) < threshold) {
[0191] reject_as_child();
[0192] }
[0193] Scene identification + solution: Voiceprint confusion between twins;
[0194] Behavioral feature fusion:
[0195] Wake word pronunciation duration (adults > 0.6 seconds, children < 0.4 seconds);
[0196] If the driver wakes up within 0.5 seconds after the brake is applied, the driver is considered the owner (by default, children cannot access the pedals).
[0197] 7. Hardware resource optimization techniques
[0198] 1) Upgrade the chip:
[0199] 2) Accelerated computation:
[0200] MFCC extraction: ARM CMSIS-DSP library is used (Cortex-M7 utilizes SIMD instructions for 4x speedup);
[0201] GMM scoring: Fixed-point quantization (Q15 format) to avoid floating-point operations;
[0202] 3) Storage compression:
[0203] Voiceprint model parameters: PCA dimensionality reduction (60-dimensional → 32-dimensional) + Huffman coding;
[0204] Final space: Compressed from the original 300KB to 80KB / user;
[0205] 4) Power consumption control:
[0206] Voiceprint comparison is only initiated when the wake word is triggered (to save continuous computing power).
[0207] The CPU's large cores are woken up during the comparison process (the small cores are idle at other times).
[0208] Figure 5 This is a block diagram of an electronic device structure for an anti-interference voiceprint recognition method in a vehicle scenario provided by one or more embodiments of the present invention.
[0209] like Figure 5 As shown, this application provides an electronic device, including: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0210] The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of an anti-interference voiceprint recognition method in a vehicle scenario.
[0211] This application also provides a computer-readable storage medium storing a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of an anti-interference voiceprint recognition method in a vehicle scenario.
[0212] This application also provides a vehicle, including:
[0213] Electronic equipment, used to implement an anti-interference voiceprint recognition method in vehicle scenarios;
[0214] The processor runs a program, and when the program runs, it executes the steps of the anti-interference voiceprint recognition method in the vehicle scenario from the data output by the electronic device.
[0215] The storage medium is used to store the program, which, when running, executes the steps of the anti-interference voiceprint recognition method in a vehicle scenario based on data output from the electronic device.
[0216] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.
[0217] The electronic device comprises a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory. The operating system can be any one or more computer operating systems that control the electronic device through processes, such as Linux, Unix, Android, iOS, or Windows. Furthermore, in this embodiment of the invention, the electronic device can be a smartphone, tablet computer, or other handheld device, or a desktop computer, portable computer, or other electronic device; there is no particular limitation in this embodiment.
[0218] In this embodiment of the invention, the executing entity for electronic device control can be an electronic device itself, or a functional module within an electronic device capable of calling and executing a program. The electronic device can obtain the firmware corresponding to the storage medium. This firmware is provided by the supplier, and different storage media may have the same or different firmware; no limitation is made here. After obtaining the firmware corresponding to the storage medium, the electronic device can write this firmware into the storage medium; specifically, it burns the firmware corresponding to the storage medium into the storage medium. The process of burning the firmware into the storage medium can be implemented using existing technology, and will not be elaborated upon in this embodiment of the invention.
[0219] Electronic devices can also obtain reset commands corresponding to the storage media. The reset commands corresponding to the storage media are provided by the supplier. The reset commands corresponding to different storage media can be the same or different, and no restrictions are imposed here.
[0220] At this time, the storage medium of the electronic device is a storage medium on which the corresponding firmware has been written. The electronic device can respond to the reset command corresponding to the storage medium on which the corresponding firmware has been written, thereby resetting the storage medium on which the corresponding firmware has been written according to the reset command. The process of resetting the storage medium according to the reset command can be implemented by existing technology and will not be described in detail in this embodiment of the invention.
[0221] For ease of description, the above devices are described separately by function as various units and modules. Of course, in implementing this application, the functions of each unit and module can be implemented in one or more software and / or hardware.
[0222] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined.
[0223] For the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0224] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0225] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An anti-interference voiceprint recognition system for vehicle scenarios, characterized in that, The anti-interference voiceprint recognition system in the vehicle scenario includes: The voice acquisition module uses a microphone array to achieve directional sound pickup; The voiceprint registration module is used to collect and store voice samples of the driver's wake-up words in multiple scenarios; The voiceprint modeling module is based on the GMM-UBM framework to build and train voiceprint models. The real-time verification module integrates a wake-word detection unit, a voice activity detection unit, a voiceprint feature extraction unit, and a matching decision unit to achieve rapid identity verification after wake-up triggering; The dynamic update module adjusts model parameters based on changes in the scene and the user's voice status. The hardware optimization module improves voiceprint recognition performance through chip upgrades, computing acceleration, storage compression, and power consumption control. To combat interference in voiceprint recognition caused by varying in-vehicle noise, different voice sources, and changes in the state of the user's vocal cords, the voice acquisition module, voiceprint registration module, voiceprint modeling module, real-time verification module, dynamic update module, and hardware optimization module work together.
2. A method for interference-resistant voiceprint recognition in vehicle scenarios, characterized in that, The anti-interference voiceprint recognition method in the vehicle scenario includes: voice acquisition, voiceprint registration, voiceprint modeling, real-time verification, dynamic updating, and hardware optimization; Voice acquisition includes recording multiple wake-up word voice samples at time intervals greater than or equal to a preset minimum, covering various vehicle scenarios; Among them, various vehicle scenarios are constructed based on the noise background of driving speed, driving environment, and the working status of on-board equipment; Based on preprocessing and feature extraction, voiceprint registration is performed for vehicle owners who are specific speakers. Voiceprint model training includes voiceprint modeling; Among them, voiceprint modeling is based on the GMM-UBM framework, and a model corresponding to a specific speaker is adaptively generated by combining a general background model with the maximum a posteriori probability. Real-time verification includes collecting voice signals through a microphone array and triggering the verification process by detecting a wake word; The verification process includes sequentially performing voice activity detection, voiceprint feature extraction, voiceprint matching, and threshold determination. During this period, a real-time response of less than 1 second is achieved, which serves as the data collection window for verifying the correctness of the results; Dynamic updates include integrating passive triggering, active periodic, and event-driven strategies to dynamically optimize the model based on specific speaker voice changes and scene interference; The optimization model includes optimization from both software algorithm and hardware resource aspects. Among them, optimizations are made with the aim of reducing in-vehicle noise, changing the position of the voice, and reducing interference caused by changes in the vocal cords of specific speakers; The evaluation is conducted using the criteria of improving recognition accuracy, shortening response time, and reducing false wake-up rate as optimization standards.
3. The anti-interference voiceprint recognition method in a vehicle scenario according to claim 2, characterized in that, Also includes: Voiceprint feature extraction includes: A front-end processing workflow of frame segmentation, windowing, and pre-emphasis is adopted to extract MFCC, PLP, or x-vector as voiceprint features; Frame-level incremental computation is achieved through streaming processing, and hardware acceleration is performed using DSP / NPU.
4. The anti-interference voiceprint recognition method in a vehicle scenario according to claim 2, characterized in that, Also includes: Real-time verification, including: Based on the detection of wake words, with the goal of achieving a latency response of less than 200ms, a lightweight CNN / RNN end-to-end model or an HMM-DNN hybrid model is used for model quantization and pruning optimization.
5. The anti-interference voiceprint recognition method in a vehicle scenario according to claim 2, characterized in that, The vehicle scenario also includes: a scenario where children's voices interfere with the sound; For scenarios involving interference with children's voices, based on the fact that children's fundamental frequency is generally >300Hz and their formants are concentrated in 4-6kHz, a frequency band suppression processing above 6kHz is adopted. By comparing the distribution of the first and second formants corresponding to adults and specific speakers, a decision is made on whether to reject the voice source of that specific speaker. If a second peak is detected at >2800Hz, the speaker is identified as a child, and the speaker's voice source is denied authentication.
6. The anti-interference voiceprint recognition method in a vehicle scenario according to claim 2, characterized in that, The vehicle scenario also includes: a voiceprint drift scenario caused by a cold; For scenarios where a speaker experiences voiceprint drift due to a cold, a multi-state model is adopted. During training, aerobic / hoarse voice synthesis data generated based on VAE acoustic transformation is incorporated, and the MFCC weighting coefficients are dynamically adjusted by detecting the fundamental frequency offset in real time.
7. The anti-interference voiceprint recognition method in a vehicle scenario according to claim 2, characterized in that, The vehicle scenario also includes a scenario where children imitate adult voices; For scenarios where children imitate adult voices, dynamic spectrogram detection is added to calculate frame-level spectral contrast. If the contrast is lower than a preset threshold, the sound source is determined to be a child imitating, and the identity verification of the sound source is refused.
8. The anti-interference voiceprint recognition method in a vehicle scenario according to claim 2, characterized in that, The vehicle scenario also includes: a scenario where the voiceprints of twins are confused; For scenarios where the voiceprints of twins are confused, behavioral features are integrated into the voiceprint matching process, including judging by the duration of the wake word pronunciation and judging the correlation between the control behavior of the control terminal and the voice wake-up behavior, to decide whether to perform identity verification.
9. The anti-interference voiceprint recognition method in a vehicle scenario according to claim 2, characterized in that, The hardware optimizations include: Using the ARMCMSIS-DSP library to achieve a 4x speedup in MFCC extraction; The GMM scoring process was optimized using Q15 format fixed-point quantization. PCA dimensionality reduction combined with Huffman coding was used to compress the voiceprint model parameters from 300KB to 80KB / user. Voiceprint comparison is initiated only after the wake word is triggered. During the comparison, the large CPU core is awakened, while the small core remains in standby mode at other times.
10. The anti-interference voiceprint recognition method in a vehicle scenario according to claim 2, characterized in that, The threshold decision adopts a dynamic threshold strategy, which adjusts the threshold according to the security level. The FAR is set to 0.01 for door disabling scenarios and 0.1 for vehicle-machine interaction scenarios.