Dynamic voiceprint recognition updating method for multiple scenes
Through multimodal sensor fusion and dynamic feature adaptation mechanism, the adaptability problem of traditional voiceprint recognition in complex scenarios is solved, and the accuracy and robustness of efficient voiceprint recognition in variable environments is achieved.
Patent Information
- Application Number
- CN202510670950.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional voiceprint recognition methods are difficult to adapt to complex and changeable application scenarios, resulting in a decrease in recognition accuracy and the inability to update the voiceprint model in time.
The 16-channel MEMS microphone array is used to synchronize voice signals with the three-axis acceleration sensor at the throat to construct a three-dimensional scene data set, combined with the scene adaptive dual-mode processing architecture, features are extracted through RLS-lattice filtering, Gammatone subband energy weighting and Teager energy operator, feature weights are dynamically adjusted, three-level pipeline architecture is deployed for real-time processing, and biological joint verification is called when the environment changes.
It significantly improves the robustness and accuracy of voiceprint recognition in complex scenarios, can quickly respond to environmental changes, and improves the system's anti-interference ability and recognition accuracy.
Smart Images

Figure CN120496538A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of biometrics, and in particular to a dynamic voiceprint recognition and updating method for multiple scenarios. Background Art
[0002] In today's digital age, voiceprint recognition technology is widely used in many fields, including security verification and intelligent voice interaction. Traditional voiceprint recognition methods focus on single voice signal analysis and rely solely on conventional acoustic features to build voiceprint models. However, in actual applications, the usage scenarios are complex and diverse, covering different acoustic scenarios such as quiet indoor environments, noisy outdoor environments, and in-vehicle environments, as well as different semantic and emotional scenarios such as daily conversations, professional terminology exchanges, and urgent emotional expressions.
[0003] In different acoustic scenarios, factors such as background noise, echo, and reverberation can severely interfere with speech signals, making voiceprint feature extraction difficult and significantly reducing recognition accuracy. Traditional voiceprint recognition methods struggle to adapt to the dynamic changes in these complex scenarios and are unable to update voiceprint models in a timely manner, causing system performance to plummet when switching between different scenarios.
[0004] From the above, we can see that how to improve the adaptability of voiceprint recognition in complex scenarios still needs to be solved. Summary of the Invention
[0005] In order to improve the adaptability of voiceprint recognition in complex scenarios, the present application provides a dynamic voiceprint recognition update method for multiple scenarios.
[0006] In the first aspect, the present application provides a dynamic voiceprint recognition and update method for multiple scenarios, which adopts the following technical solutions:
[0007] A dynamic voiceprint recognition update method for multiple scenarios, comprising:
[0008] The system uses a circularly distributed 16-channel MEMS microphone array and a three-axis throat accelerometer to synchronously collect speech signals, constructing a three-dimensional scene dataset containing acoustic scene fingerprints, semantic context vectors, and emotional biometrics.
[0009] A scene-adaptive dual-mode processing architecture is used for feature extraction: In quiet scenes, RLS-lattice filtering combined with dual-density complex wavelet transform is implemented to extract anti-aliasing acoustic features. In noisy scenes, a robust feature mode fused with Gammatone subband energy weighting and Teager energy operator is activated. Semantic features are simultaneously extracted through an attention-enhanced keyword matching tree. Emotional state encoding vectors are generated by combining fundamental frequency jitter quantification. A corresponding multidimensional feature set is constructed based on acoustic features, semantic features, and emotion encoding vectors.
[0010] The dynamic topological structure of the voiceprint feature library is constructed based on a multi-dimensional feature set. The scene parameter matching degree is calculated based on the Gaussian mixture model. When the scene mutation index exceeds the set threshold, the online clustering reconstruction mechanism is triggered. The Takagi-Sugeno inference mechanism is used to dynamically adjust the fusion weights of acoustic features, semantic features, and emotion coding vectors, and output the optimized weighted feature combination.
[0011] A three-stage pipeline architecture based on weighted feature combination is deployed for real-time processing: a Bark-scale subband noise suppression filter is configured on the FPGA for feature enhancement, a GPU cluster parallelizes the MFCC-PLP hybrid feature matrix for matching, and zero-copy transmission of feature data between the CPU and GPU is achieved via the RDMA network.
[0012] Based on the MFCC-PLP hybrid feature matrix and real-time environmental parameters, a scene adaptation factor is generated. The weighted feature combination is combined with the scene adaptation factor for multi-feature joint matching. When the confidence of multiple consecutive matches is lower than the set threshold, the speech signal is called to activate the biological joint verification of the vocal cord vibration fundamental frequency and lip optical flow features.
[0013] Optionally, the method also includes: the MEMS microphone array adopts a double-layer spiral topology structure, the inner ring microphone unit is equipped with an acoustic waveguide, and the outer ring microphone unit is covered with a nanoporous membrane; the laryngeal three-axis acceleration sensor retains the biological vibration characteristic frequency band through a bandpass filter, and realizes microsecond-level synchronous acquisition with the voice signal.
[0014] Optionally, the method also includes: the MEMS microphone array adopts a double-layer spiral topology structure, the inner ring microphone unit is equipped with an acoustic waveguide, and the outer ring microphone unit is covered with a nanoporous membrane; the laryngeal three-axis acceleration sensor retains the biological vibration characteristic frequency band through a bandpass filter, and realizes microsecond-level synchronous acquisition with the voice signal.
[0015] Optionally, the method also includes: the MEMS microphone array adopts a double-layer spiral topology structure, the inner ring microphone unit is equipped with an acoustic waveguide, and the outer ring microphone unit is covered with a nanoporous membrane; the laryngeal three-axis acceleration sensor retains the biological vibration characteristic frequency band through a bandpass filter, and realizes microsecond-level synchronous acquisition with the voice signal.
[0016] Optionally, the method also includes: the Bark scale subband filter group adopts a non-uniform division method, and filter structures of different orders are set in different frequency bands; when calculating the MFCC-PLP hybrid features, overlapping frame processing is adopted, and the PLP coefficient order is dynamically adjusted according to the current signal-to-noise ratio.
[0017] Optionally, the method also includes: the biological joint verification includes dual verification conditions, respectively verifying the correlation between the fundamental frequency of vocal cord vibration and the fundamental frequency of speech, and the time domain alignment of lip movement and speech energy; when any verification condition is not met, the three-dimensional lip shape reconstruction verification based on optical detection is triggered.
[0018] Optionally, the method also includes: the scene mutation index is calculated using composite judgment conditions, including: divergence changes of acoustic environment fingerprints, changes in semantic keyword distribution entropy values, and changes in emotional impulse response energy variance; when the combined judgment criteria of multiple conditions are met, it is determined to be a scene mutation event.
[0019] In the second aspect, the present application provides a dynamic voiceprint recognition and updating device for multiple scenarios, which adopts the following technical solutions:
[0020] A dynamic voiceprint recognition and updating device for multiple scenarios, comprising:
[0021] The 3D scene dataset construction module uses a circularly distributed 16-channel MEMS microphone array and a three-axis throat accelerometer to synchronously collect voice signals, and is used to construct a 3D scene dataset containing acoustic scene fingerprints, semantic context vectors, and emotional biometrics.
[0022] The multidimensional feature set construction module uses a scene-adaptive dual-mode processing architecture for feature extraction: In quiet scenes, RLS-lattice filtering combined with dual-density complex wavelet transform is implemented to extract anti-aliasing acoustic features. In noisy scenes, a robust feature mode fused with Gammatone subband energy weighting and Teager energy operator is activated. Semantic features are simultaneously extracted through an attention-enhanced keyword matching tree. The emotional state coding vector is generated by combining fundamental frequency jitter quantification. The corresponding multidimensional feature set is constructed based on the acoustic features, semantic features, and emotional coding vector.
[0023] The weighted feature combination output module builds a dynamic topological structure of the voiceprint feature library based on a multidimensional feature set, calculates the matching degree of scene parameters based on a Gaussian mixture model, triggers an online clustering reconstruction mechanism when the scene mutation index exceeds a set threshold, and uses the Takagi-Sugeno inference mechanism to dynamically adjust the fusion weights of acoustic features, semantic features, and emotion encoding vectors to output the optimized weighted feature combination;
[0024] A three-stage pipeline architecture deployment module is used to deploy a three-stage pipeline architecture for real-time processing based on weighted feature combination. The FPGA side is configured with a Bark-scale subband noise suppression filter for feature enhancement. The GPU cluster parallelizes the calculation of the MFCC-PLP hybrid feature matrix for matching operations. Zero-copy transmission of feature data between the CPU and GPU is achieved through the RDMA network.
[0025] The multi-feature joint matching module generates a scene adaptation factor based on the MFCC-PLP hybrid feature matrix and real-time environmental parameters, and combines the weighted feature combination with the scene adaptation factor for multi-feature joint matching. When the confidence of multiple consecutive matches is lower than the set threshold, the speech signal is called to activate the biological joint verification of the vocal cord vibration fundamental frequency and lip optical flow features.
[0026] In a third aspect, the present application provides a dynamic voiceprint recognition and updating device for multiple scenarios, which adopts the following technical solutions:
[0027] A dynamic voiceprint recognition and updating device for multiple scenarios includes a processor running a program of any one of the above-mentioned dynamic voiceprint recognition and updating methods for multiple scenarios.
[0028] In a fourth aspect, the present application provides a storage medium, which adopts the following technical solution:
[0029] A storage medium stores a program of any one of the above-mentioned dynamic voiceprint recognition and updating methods for multiple scenarios.
[0030] In summary, this application includes at least one of the following beneficial technical effects:
[0031] Through multimodal sensor fusion and dynamic feature adaptation mechanism, the robustness of voiceprint recognition in complex scenarios has been significantly improved. By constructing a three-dimensional data system including acoustic fingerprints, semantic context and biometric features, and through a scene-adaptive dual-mode processing architecture, dynamic weight adjustment is achieved in noise suppression, semantic extraction and emotion encoding, effectively responding to problems such as sudden environmental changes, background interference and semantic scene switching, and greatly enhancing the system's feature extraction accuracy and anti-interference ability under non-ideal conditions.
[0032] Through the collaborative optimization of real-time topology reconstruction and three-level heterogeneous computing architecture, full-process scenario adaptation is achieved from hardware-level noise suppression to algorithm-level feature fusion; combined with the biological joint verification mechanism, when the environment suddenly changes or the matching confidence decreases, the biological characteristics of vocal cord vibration and lip movement can be quickly called for multi-dimensional cross-verification, breaking through the performance bottleneck of traditional voiceprint recognition in complex scenarios, and significantly improving the recognition accuracy and system reliability in dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 The present invention is a flowchart of a dynamic voiceprint recognition and updating method for multiple scenarios according to an exemplary embodiment.
[0034] Figure 2 The figure is a structural block diagram of a dynamic voiceprint recognition and updating device for multiple scenarios according to an exemplary embodiment. DETAILED DESCRIPTION
[0035] Embodiments of the present application are described in detail below, examples of which are illustrated in the accompanying drawings.
[0036] Throughout this specification, reference to the terms "certain embodiments," "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0037] The present application embodiment discloses a dynamic voiceprint recognition update method for multiple scenarios, referring to Figure 1 ,include:
[0038] S100 uses a circularly distributed 16-channel MEMS microphone array and a three-axis accelerometer in the throat to synchronously collect voice signals, constructing a three-dimensional scene dataset containing acoustic scene fingerprints, semantic context vectors, and emotional biometrics.
[0039] In an embodiment of the present application, a 16-channel MEMS microphone array with a double-layer spiral topology is adopted, in which the inner ring is equipped with an acoustic waveguide to enhance the far-field sound pickup capability, and the outer ring surface is covered with a nanoporous membrane; it can not only effectively capture sound signals from different directions, but also reduce environmental noise interference through physical means.
[0040] Among them, the inner ring microphone units are equipped with acoustic waveguides, which help focus and guide the sound into the microphone, thereby improving the quality of sound capture at long distances or in low signal-to-noise ratio environments; the outer ring microphone units, microphones covered with nanoporous membranes can effectively absorb and block background noise, especially high-frequency noise, ensuring the clarity of voice signals; synchronous acquisition: all 16 microphone units work together through a precise time synchronization mechanism to ensure that audio signals are received at the same time, so that all-round, high-fidelity sound information can be obtained.
[0041] The 16-channel MEMS microphone array utilizes a double-layer spiral topology. The inner ring is equipped with an acoustic waveguide to enhance far-field sound pickup, while the outer ring is covered with a nanoporous membrane. This design not only effectively captures sound signals from different directions but also physically reduces ambient noise interference.
[0042] Inner-Ring Microphone Units: These units are equipped with an acoustic waveguide that helps focus and direct sound into the microphone, improving the quality of sound capture at long distances or in low signal-to-noise ratio environments.
[0043] Outer ring microphone unit: The microphone covered with a nanoporous membrane can effectively absorb and block background noise, especially high-frequency noise, ensuring the clarity of the voice signal.
[0044] Synchronous acquisition: All 16 microphone units work together through a precise time synchronization mechanism to ensure simultaneous reception of audio signals, thus obtaining full-range, high-fidelity sound information.
[0045] In addition, a three-axis accelerometer is placed in the throat area, as this is the most direct manifestation of vocal cord vibration. The sensor retains the biometric frequency bands (such as the fundamental frequency and harmonic components) through a bandpass filter. The signals within these frequency ranges can accurately reflect the speaker's vocal cord vibration. The sensor and microphone collect data synchronously to ensure microsecond-level synchronization between the two, which is crucial for subsequent analysis of the speaker's emotional state and semantic content.
[0046] The data obtained from the above devices is used to generate three main components: acoustic scene fingerprint, which is based on the sound signal collected by the microphone array and uses a specific algorithm to extract a "fingerprint" that can represent the characteristics of the current acoustic environment; semantic context vector, which uses natural language processing technology to analyze the vocabulary in the speech and extract semantic features that can describe the meaning of the speech content; emotional state coding vector, which combines vocal cord vibration information (provided by a three-axis accelerometer) and speech signals to quantify and generate a coding vector representing the speaker's emotional state.
[0047] The main purpose of this stage is to provide a comprehensive and high-quality basic data set for the subsequent feature extraction and recognition process. By integrating information on acoustic characteristics, semantic content, and emotional state, the system can not only identify who is speaking (i.e., voiceprint recognition), but also understand what they are saying (semantic analysis) and their emotional state (emotional analysis); multi-level data integration enables the system to maintain efficient and accurate performance in more complex and diverse application scenarios. For example, it can still accurately identify the user's identity in noisy environments and understand the user's intentions and emotional changes, providing a solid data support foundation for subsequent feature extraction, model training, and real-time applications.
[0048] The S200 uses a scene-adaptive dual-mode processing architecture for feature extraction: in quiet scenes, RLS-lattice filtering combined with dual-density complex wavelet transform is implemented to extract anti-aliasing acoustic features. In noisy scenes, a robust feature mode fused with Gammatone subband energy weighting and Teager energy operator is activated. Semantic features are simultaneously extracted through an attention-enhanced keyword matching tree, and emotional state coding vectors are generated by combining fundamental frequency jitter quantification. A corresponding multidimensional feature set is constructed based on acoustic features, semantic features, and emotional coding vectors.
[0049] In an embodiment of the present application, the RLS-lattice filter sets a multi-order adaptive structure and adopts a variable step size algorithm to update the coefficients; when the Teager energy operator fusion mode is activated, the auditory masking effect model of the critical frequency band is synchronously constructed, and the weight coefficient of each sub-band is dynamically adjusted according to the real-time signal-to-noise ratio.
[0050] The specific implementation process is as follows:
[0051] For feature extraction in quiet scenes, the RLS-lattice filter adopts a multi-order adaptive structure, which means that it can automatically adjust its internal parameters according to the characteristics of the input signal and use a variable step-size algorithm to update the coefficients, allowing the filter to track signal changes more quickly and accurately, reducing aliasing effects; the double-density complex wavelet transform is an advanced wavelet transform method that can improve frequency resolution without increasing computational complexity. Through this transform, high-quality acoustic features can be extracted from speech signals, making it particularly suitable for processing high-fidelity audio signals.
[0052] Feature extraction in noisy scenarios: Gammatone subband energy weighting can mimic the frequency selectivity of the human auditory system, decomposing the audio signal into multiple subbands and weighting the energy of each subband. This method helps to highlight the main components of the speech signal in noisy environments. The Teager Energy Operator (TEO) is used to calculate the instantaneous energy changes of the speech signal, emphasizing the nonlinear dynamic characteristics of speech. When this mode is activated, an auditory masking effect model for the critical frequency band is also constructed simultaneously. The weight coefficient of each frequency band is dynamically adjusted according to the real-time signal-to-noise ratio to adapt to the current environmental conditions.
[0053] Using natural language processing technology, a keyword matching tree is constructed to identify and extract semantic information. The keyword matching tree structure assigns different attention weights according to the importance of the context, thereby improving the accuracy of keyword recognition. Then, the vocal cord vibration information obtained from the three-axis accelerometer and the speech signal itself are combined to quantify the speaker's emotional state by analyzing indicators such as fundamental frequency jitter to generate an emotional state encoding vector.
[0054] Based on the acoustic features, semantic features, and emotion encoding vectors extracted in the above steps, a multi-dimensional feature set is constructed. This feature set not only includes traditional acoustic features, but also incorporates the results of semantic understanding and sentiment analysis, providing a rich information foundation for subsequent recognition and classification tasks.
[0055] By using RLS-lattice filters and dual-density complex wavelet transforms, aliasing effects can be effectively removed in quiet environments, ensuring high-quality acoustic feature extraction. In noisy environments, the combination of Gammatone band energy weighting and the Teager energy operator enhances the system's robustness to background noise, ensuring accurate speech feature extraction even under adverse conditions. The attention-enhanced keyword matching tree allows the system to more accurately capture key information in speech, improving its understanding of semantic content. By quantifying the speaker's emotional state through fundamental frequency jitter, the system can not only identify who is speaking but also perceive the speaker's emotional state, which is particularly important for applications requiring highly personalized services. The entire architecture design takes into account the needs of various application scenarios. Whether in quiet or noisy environments, the adaptive mechanism can adjust the strategy to ensure optimal performance.
[0056] S300 builds a dynamic topological structure of the voiceprint feature library based on a multi-dimensional feature set, calculates the scene parameter matching degree based on the Gaussian mixture model, triggers the online clustering reconstruction mechanism when the scene mutation index exceeds the set threshold, and uses the Takagi-Sugeno reasoning mechanism to dynamically adjust the fusion weights of acoustic features, semantic features and emotion coding vectors, and outputs the optimized weighted feature combination.
[0057] In an embodiment of the present application, an online clustering reconstruction mechanism sets a dynamic radius adjustment strategy, using different initialization radii in quiet scenes and noisy scenes; the Takagi-Sugeno reasoning mechanism sets a multi-layer fuzzy rule base, and the input parameters include the scene parameter change rate and the feature vector similarity.
[0058] The specific implementation process is as follows: First, using the acoustic features, semantic features, and emotion encoding vectors extracted in the previous step (S200), a multidimensional feature set is constructed. This feature set not only includes traditional acoustic features but also covers information about the understanding of speech content and the speaker's emotional state. Then, these multidimensional features are modeled using GMM to calculate the scene parameter matching degree. GMM is a probabilistic model that can well capture data distribution characteristics and is suitable for describing complex voiceprint feature patterns.
[0059] In addition, a scene mutation index needs to be defined. The scene mutation index comprehensively considers factors such as the change in the divergence of the acoustic environment fingerprint, the change in the entropy value of the semantic keyword distribution, and the change in the energy variance of the emotional impulse response. When this index exceeds the set threshold, it indicates that the current scene has changed significantly and the corresponding adjustment mechanism needs to be triggered.
[0060] To adapt to different scenarios (quiet or noisy), the online clustering reconstruction mechanism uses a dynamic radius adjustment strategy. In quiet scenes, the initial cluster radius is smaller to more precisely capture similar features; in noisy scenes, the initial radius is larger to cover a wider range of feature variation. This differentiated initialization strategy helps improve clustering efficiency and accuracy. Once the scene mutation index exceeds a threshold, the system initiates the online clustering reconstruction mechanism, which reevaluates and adjusts existing feature clusters to ensure they accurately reflect the feature distribution in the current scene.
[0061] The Takagi-Sugeno inference mechanism is used to dynamically adjust the fusion weights of acoustic features, semantic features, and emotion coding vectors. The Takagi-Sugeno inference mechanism sets up a multi-layer fuzzy rule base, in which the input parameters include the rate of change of scene parameters and the similarity of feature vectors. Through these inputs, the system can intelligently decide how to best combine different types of features to optimize the final recognition effect.
[0062] Based on the above analysis, the Takagi-Sugeno inference mechanism outputs a set of optimized weighted feature combinations. The weighted feature combination fully considers the characteristics of the current scenario, making subsequent processing (such as matching and verification) more accurate and effective.
[0063] Through its dynamic topology and online clustering reconstruction mechanism, the system can rapidly respond to environmental changes, maintaining high recognition accuracy regardless of whether the environment changes from quiet to noisy or vice versa. By using GMM to calculate scene parameter matching and adjusting feature weights based on actual conditions, the system maintains efficient voiceprint recognition capabilities even in complex and changing environments. The dynamic radius adjustment strategy and Takagi-Sugeno inference mechanism work together to ensure that the system can flexibly adjust its internal parameters based on actual scenario requirements, thereby providing more personalized and accurate services. This not only improves the robustness and adaptability of the voiceprint recognition system, but also provides an effective solution to address challenges in real-world applications, particularly in the face of constantly changing environmental conditions.
[0064] The S400 deploys a three-stage pipeline architecture based on weighted feature combination for real-time processing: the FPGA side is configured with a Bark-scale subband noise suppression filter for feature enhancement, the GPU cluster parallelly calculates the MFCC-PLP hybrid feature matrix for matching operations, and zero-copy transmission of feature data between the CPU and GPU is performed through the RDMA network.
[0065] In an embodiment of the present application, the Bark scale subband filter bank adopts a non-uniform division method, and filter structures of different orders are set in different frequency bands; when calculating the MFCC-PLP hybrid features, overlapping frame processing is adopted, and the PLP coefficient order is dynamically adjusted according to the current signal-to-noise ratio.
[0066] The specific implementation process is as follows:
[0067] The Bark scale subband filter bank is designed according to the frequency perception characteristics of the human auditory system. It adopts a non-uniform partitioning method, with a narrower frequency band in the low-frequency part and a wider frequency band in the high-frequency part, to simulate the different sensitivity of the human ear to different frequency bands. For each subband, filters of different orders are set according to its specific frequency range. For example, in the low-frequency region (such as 20Hz-1kHz), since the signal changes relatively slowly, a lower-order filter can be used; in the high-frequency region (such as above 5kHz), since the signal changes rapidly, a higher-order filter is required to capture the rapidly changing details.
[0068] Through the above design, the Bark scale subband filter can effectively remove background noise while retaining the key features of the speech signal, providing high-quality input data for subsequent processing.
[0069] In addition, the input audio signal is divided into multiple short time frames (typically 20ms to 30ms), with a certain overlap between adjacent frames (e.g., 10ms). This processing method helps capture the temporal changes of the speech signal and improves the accuracy of feature extraction. When calculating MFCC (Mel-Frequency Cepstral Coefficients) and PLP (Perceptual Linear Prediction Coefficients), the order of the PLP coefficients is dynamically adjusted based on the current signal-to-noise ratio (SNR). When the SNR is high, a higher order can be used to capture more detailed information; in low SNR environments, the order is reduced to avoid overfitting the noise.
[0070] It's important to note that leveraging the powerful parallel computing capabilities of the GPU allows for simultaneous processing of multiple frames of data to rapidly generate the MFCC-PLP hybrid feature matrix. This approach not only improves computational efficiency but also ensures real-time processing. A Remote Direct Memory Access (RDMA) network is used to efficiently transfer feature data between the CPU and GPU. Traditionally, copying data from the CPU to the GPU requires multiple memory copy operations, which introduces significant latency. RDMA, on the other hand, allows data to be transferred directly from one device's memory to another's without CPU intervention, significantly reducing the time overhead of data transfer.
[0071] Through a carefully designed Bark-scale subband filter bank, specifically employing filter structures of varying orders for different frequency bands, noise can be effectively removed while maintaining the quality of the speech signal, a crucial step for subsequent feature extraction. By combining MFCC and PLP feature extraction methods and dynamically adjusting parameters based on the actual signal-to-noise ratio, the acoustic characteristics of the speech signal can be more comprehensively captured, improving the robustness and accuracy of the recognition system. Leveraging the parallel computing capabilities of GPUs and the zero-copy transmission technology of RDMA networks, the entire processing flow is significantly accelerated to meet real-time processing requirements. These optimization measures significantly reduce processing latency and ensure system responsiveness, especially in scenarios with large data volumes.
[0072] In summary, the S400 step has built an efficient three-level pipeline architecture through optimized design at the hardware level and fine-tuning at the algorithm level, enabling the system to not only maintain high-precision voiceprint recognition capabilities in complex environments, but also meet the requirements of real-time processing.
[0073] S500 generates a scene adaptation factor based on the MFCC-PLP hybrid feature matrix and real-time environmental parameters, combines the weighted feature combination with the scene adaptation factor for multi-feature joint matching, and when the confidence level of multiple consecutive matches is lower than the set threshold, calls the speech signal to activate the biological joint verification of the vocal cord vibration fundamental frequency and lip optical flow features.
[0074] In an embodiment of the present application, the biological joint verification includes dual verification conditions, which respectively verify the correlation between the fundamental frequency of vocal cord vibration and the fundamental frequency of speech, and the time domain alignment of lip movement and speech energy; when any verification condition is not met, the three-dimensional lip shape reconstruction verification based on optical detection is triggered.
[0075] Specifically, the implementation process is as follows: Based on the MFCC-PLP hybrid feature matrix calculated by the GPU cluster in step (S400), the MFCC-PLP hybrid feature matrix combines Mel-frequency cepstral coefficients (MFCC) and perceptual linear prediction coefficients (PLP), which can fully reflect the acoustic characteristics of the speech signal. Current environmental parameters, such as background noise level, temperature, humidity, etc., are collected through sensors or other means. These parameters help understand the impact of the current environment on the speech signal and provide a basis for adjusting the matching algorithm.
[0076] The MFCC-PLP hybrid feature matrix is combined with real-time environmental parameters, and a specific algorithm is used to generate a "scene adaptation factor". This scene adaptation factor reflects the most optimized matching conditions in the current environment, ensuring that the matching algorithm can adapt to changes in different scenarios.
[0077] The optimized weighted feature combination output by the Takagi-Sugeno inference mechanism is combined with the scene adaptation factor for the multi-feature joint matching process. This step aims to maximize the use of all available information to improve matching accuracy. During matching, the system continuously evaluates the match confidence. If the match confidence falls below the set threshold for multiple consecutive times, the current match result is considered unreliable and requires further verification.
[0078] When the matching confidence is lower than the threshold, the biometric joint verification process is started. This process includes two main verification conditions:
[0079] 1. Verification of the correlation between the fundamental frequency of vocal cord vibration and the fundamental frequency of speech: Analyze the vocal cord vibration data obtained from the laryngeal triaxial accelerometer, extract its fundamental frequency, and compare it with the fundamental frequency in the speech signal to confirm whether there is a high correlation between the two.
[0080] 2. Verify the temporal alignment of lip movements and speech energy: Use a camera or other visual sensor to capture the speaker's lip movements and analyze their temporal alignment with the energy changes in the speech signal. Good synchronization indicates a high degree of speaker authenticity.
[0081] If any of the above verification conditions are not met, the system will trigger 3D lip reconstruction verification based on optical detection. This process involves using a high-resolution camera to capture the speaker's facial image and reconstructing the 3D lip structure through computer vision technology to more accurately verify the speaker's identity.
[0082] By combining the MFCC-PLP hybrid feature matrix with the scene adaptation factor generated by real-time environmental parameters, the matching algorithm can be dynamically adjusted to adapt to different application scenarios, thereby improving the accuracy and robustness of the recognition system. By introducing a joint biometric verification mechanism, especially when the matching confidence is low, the dual verification conditions of vocal cord vibration fundamental frequency and lip optical flow features, as well as 3D lip shape reconstruction verification when necessary, greatly enhance the system's security protection capabilities and reduce the risk of misidentification.
[0083] Whether in a quiet or noisy environment, this method can ensure the stable operation of the system through adaptive mechanisms and multi-level verification strategies, ensuring that reliable voiceprint recognition services can be provided in various situations. This design is particularly suitable for occasions with extremely high security requirements, such as financial transaction authentication and access control.
[0084] It should be pointed out here that the calculation of the scene mutation index adopts composite judgment conditions, including: the divergence change of the acoustic environment fingerprint, the change of the semantic keyword distribution entropy value, and the change of the emotional impulse response energy variance; when the combined judgment criteria of multiple conditions are met, it is judged as a scene mutation event.
[0085] Specifically include:
[0086] 1. Calculation of scene mutation index:
[0087] Divergence changes in acoustic environment fingerprints. Acoustic environment fingerprints are extracted by analyzing the sound signals obtained from the MEMS microphone array to represent the characteristics of the current acoustic environment. These fingerprints may include spectral characteristics, noise levels, etc. Divergence changes: Monitor changes in acoustic environment fingerprints over time. For example, statistical distances (such as Kullback-Leibler divergence) are used to measure the difference between acoustic environment fingerprints in two time periods. If this difference exceeds a set threshold, the acoustic environment is considered to have changed significantly.
[0088] 2. Changes in semantic keyword distribution entropy:
[0089] Semantic keyword distribution: By performing natural language processing on the speech content, we identify and count the frequency of keywords or phrases, creating a keyword distribution chart. Entropy change: We calculate the entropy of keyword distribution to reflect the uncertainty of the information. If the entropy of keyword distribution changes significantly within a short period of time (such as a sudden increase or decrease), it indicates that the topic or context of the current conversation has changed significantly.
[0090] 3. Changes in the energy variance of emotional impulse response:
[0091] Emotional impulse response energy, based on vocal cord vibration information captured by a triaxial accelerometer and the speech signal itself, analyzes fundamental frequency jitter and other indicators to quantify the speaker's emotional state and converts it into emotional impulse response energy. Variance change: Monitors how the emotional impulse response energy changes over time. A significant increase in variance may indicate a significant fluctuation in the speaker's emotional state.
[0092] In addition, independent thresholds are set for each of the above indicators (divergence changes of acoustic environment fingerprints, changes in semantic keyword distribution entropy, and changes in emotional impulse response energy variance); when multiple conditions are met at the same time (for example, at least two indicators exceed their respective set thresholds), the system will determine that a sudden change has occurred in the scene. This composite judgment method can more accurately capture changes in the actual environment and avoid misjudgment of a single indicator.
[0093] That is to say, first of all, it is necessary to collect data in real time from devices such as MEMS microphone arrays and three-axis accelerometers, and perform necessary preprocessing (such as filtering, noise reduction, etc.). For each type of input data (acoustic, semantic, emotional), a specific algorithm is applied to extract the corresponding features. For example, for acoustic data, spectrum analysis technology can be used; for semantic data, natural language processing tools are used; and for emotional data, it is based on physiological signal analysis methods. Based on the extracted features, the divergence change of the acoustic environment fingerprint, the change in the entropy value of the semantic keyword distribution, and the change in the energy variance of the emotional impulse response are calculated respectively. The calculated change amount is compared with the pre-set threshold. If it is found that multiple indicators exceed their respective threshold ranges, the scene mutation detection mechanism is triggered.
[0094] The present application embodiment discloses a dynamic voiceprint recognition and updating device for multiple scenarios, referring to Figure 2 ,include:
[0095] 3D scene dataset construction module 001, which uses a circularly distributed 16-channel MEMS microphone array and a laryngeal three-axis accelerometer to synchronously collect voice signals to construct a 3D scene dataset containing acoustic scene fingerprints, semantic context vectors, and emotional biometrics;
[0096] Multidimensional feature set construction module 002 uses a scene-adaptive dual-mode processing architecture for feature extraction: in quiet scenes, RLS-lattice filtering is implemented in conjunction with dual-density complex wavelet transform to extract anti-aliasing acoustic features. In noisy scenes, a robust feature mode fused with Gammatone subband energy weighting and Teager energy operator is activated. Semantic features are simultaneously extracted through an attention-enhanced keyword matching tree. Emotional state coding vectors are generated by combining fundamental frequency jitter quantification. The acoustic features, semantic features, and emotion coding vectors are used to construct the corresponding multidimensional feature set.
[0097] Weighted feature combination output module 003 builds a dynamic topological structure of the voiceprint feature library based on a multidimensional feature set, calculates the scene parameter matching degree based on a Gaussian mixture model, triggers an online clustering reconstruction mechanism when it detects that the scene mutation index exceeds a set threshold, and uses the Takagi-Sugeno inference mechanism to dynamically adjust the fusion weights of acoustic features, semantic features, and emotion encoding vectors to output the optimized weighted feature combination;
[0098] Three-stage pipeline architecture deployment module 004, based on weighted feature combination, is used to deploy a three-stage pipeline architecture for real-time processing: the FPGA side is configured with a Bark-scale subband noise suppression filter for feature enhancement, the GPU cluster parallelly calculates the MFCC-PLP hybrid feature matrix for matching operations, and zero-copy transmission of feature data between the CPU and GPU is performed through the RDMA network;
[0099] The multi-feature joint matching module 005 generates a scene adaptation factor based on the MFCC-PLP mixed feature matrix and real-time environmental parameters, combines the weighted feature combination with the scene adaptation factor for multi-feature joint matching, and when the confidence of multiple consecutive matches is lower than the set threshold, calls the speech signal to activate the biological joint verification of the fundamental frequency of vocal cord vibration and lip optical flow features.
[0100] An embodiment of the present application further discloses a dynamic voiceprint recognition and updating device for multiple scenarios, comprising a processor, wherein the processor runs a program of any one of the above-mentioned dynamic voiceprint recognition and updating methods for multiple scenarios.
[0101] An embodiment of the present application further discloses a storage medium storing a program of any one of the above-mentioned dynamic voiceprint recognition and updating methods for multiple scenarios.
[0102] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A dynamic voiceprint recognition and updating method for multiple scenarios, characterized in that: include: The system uses a circularly distributed 16-channel MEMS microphone array and a three-axis throat accelerometer to synchronously collect speech signals, constructing a three-dimensional scene dataset containing acoustic scene fingerprints, semantic context vectors, and emotional biometrics. A scene-adaptive dual-mode processing architecture is used for feature extraction: In quiet scenes, RLS-lattice filtering combined with dual-density complex wavelet transform is implemented to extract anti-aliasing acoustic features. In noisy scenes, a robust feature mode fused with Gammatone subband energy weighting and Teager energy operator is activated. Semantic features are simultaneously extracted through an attention-enhanced keyword matching tree. Emotional state encoding vectors are generated by combining fundamental frequency jitter quantification. A corresponding multidimensional feature set is constructed based on acoustic features, semantic features, and emotion encoding vectors. The dynamic topological structure of the voiceprint feature library is constructed based on a multi-dimensional feature set. The scene parameter matching degree is calculated based on the Gaussian mixture model. When the scene mutation index exceeds the set threshold, the online clustering reconstruction mechanism is triggered. The Takagi-Sugeno inference mechanism is used to dynamically adjust the fusion weights of acoustic features, semantic features, and emotion coding vectors, and output the optimized weighted feature combination. A three-stage pipeline architecture based on weighted feature combination is deployed for real-time processing: a Bark-scale subband noise suppression filter is configured on the FPGA for feature enhancement, a GPU cluster parallelizes the MFCC-PLP hybrid feature matrix for matching, and zero-copy transmission of feature data between the CPU and GPU is achieved via the RDMA network. Based on the MFCC-PLP hybrid feature matrix and real-time environmental parameters, a scene adaptation factor is generated. The weighted feature combination is combined with the scene adaptation factor for multi-feature joint matching. When the confidence of multiple consecutive matches is lower than the set threshold, the speech signal is called to activate the biological joint verification of the vocal cord vibration fundamental frequency and lip optical flow features.
2. The multi-scenario dynamic voiceprint recognition update method according to claim 1 is characterized in that: The method also includes: The MEMS microphone array adopts a double-layer spiral topology structure, the inner ring microphone unit is equipped with an acoustic waveguide, and the outer ring microphone unit is covered with a nanoporous membrane; The laryngeal triaxial acceleration sensor retains the characteristic frequency band of biological vibration through a bandpass filter and realizes microsecond-level synchronous acquisition with the voice signal.
3. The multi-scenario dynamic voiceprint recognition update method according to claim 1, characterized in that: The method also includes: The RLS-lattice filter sets a multi-order adaptive structure and adopts a variable step size algorithm to update the coefficients; When the Teager energy operator fusion mode is activated, an auditory masking effect model of the critical frequency band is synchronously constructed, and the weight coefficient of each frequency band is dynamically adjusted according to the real-time signal-to-noise ratio.
4. The multi-scenario dynamic voiceprint recognition update method according to claim 1, characterized in that: The method also includes: The online clustering reconstruction mechanism sets a dynamic radius adjustment strategy, using different initialization radii in quiet scenes and noisy scenes; The Takagi-Sugeno reasoning mechanism sets up a multi-layer fuzzy rule base, and the input parameters include the scene parameter change rate and the feature vector similarity.
5. The multi-scenario dynamic voiceprint recognition update method according to claim 1, characterized in that: The method also includes: The Bark scale subband filter bank adopts a non-uniform partitioning method, and different frequency bands are set with filter structures of different orders; When calculating the MFCC-PLP hybrid features, overlapping frame processing is adopted, and the PLP coefficient order is dynamically adjusted according to the current signal-to-noise ratio.
6. The multi-scenario dynamic voiceprint recognition update method according to claim 1, characterized in that: The method also includes: The biological joint verification includes dual verification conditions, which respectively verify the correlation between the fundamental frequency of vocal cord vibration and the fundamental frequency of speech, and the time domain alignment between lip movement and speech energy; When any verification condition is not met, the 3D lip shape reconstruction verification based on optical detection is triggered.
7. The multi-scenario dynamic voiceprint recognition update method according to claim 1, characterized in that: The method also includes: The scene mutation index calculation adopts composite judgment conditions, including: the divergence change of the acoustic environment fingerprint, the change of the semantic keyword distribution entropy value, and the change of the emotional impulse response energy variance; when the combined judgment criteria of multiple conditions are met, it is determined to be a scene mutation event.
8. A dynamic voiceprint recognition and updating device for multiple scenarios, characterized in that: include: The 3D scene dataset construction module uses a circularly distributed 16-channel MEMS microphone array and a three-axis throat accelerometer to synchronously collect voice signals, and is used to construct a 3D scene dataset containing acoustic scene fingerprints, semantic context vectors, and emotional biometrics. The multidimensional feature set construction module uses a scene-adaptive dual-mode processing architecture for feature extraction: In quiet scenes, RLS-lattice filtering combined with dual-density complex wavelet transform is implemented to extract anti-aliasing acoustic features. In noisy scenes, a robust feature mode fused with Gammatone subband energy weighting and Teager energy operator is activated. Semantic features are simultaneously extracted through an attention-enhanced keyword matching tree. The emotional state coding vector is generated by combining fundamental frequency jitter quantification. The corresponding multidimensional feature set is constructed based on the acoustic features, semantic features, and emotional coding vector. The weighted feature combination output module builds a dynamic topological structure of the voiceprint feature library based on a multidimensional feature set, calculates the matching degree of scene parameters based on a Gaussian mixture model, triggers an online clustering reconstruction mechanism when the scene mutation index exceeds a set threshold, and uses the Takagi-Sugeno inference mechanism to dynamically adjust the fusion weights of acoustic features, semantic features, and emotion encoding vectors to output the optimized weighted feature combination; A three-stage pipeline architecture deployment module is used to deploy a three-stage pipeline architecture for real-time processing based on weighted feature combination. The FPGA side is configured with a Bark-scale subband noise suppression filter for feature enhancement. The GPU cluster parallelizes the calculation of the MFCC-PLP hybrid feature matrix for matching operations. Zero-copy transmission of feature data between the CPU and GPU is achieved through the RDMA network. The multi-feature joint matching module generates a scene adaptation factor based on the MFCC-PLP hybrid feature matrix and real-time environmental parameters, and combines the weighted feature combination with the scene adaptation factor for multi-feature joint matching. When the confidence of multiple consecutive matches is lower than the set threshold, the speech signal is called to activate the biological joint verification of the vocal cord vibration fundamental frequency and lip optical flow features.
9. A dynamic voiceprint recognition and updating device for multiple scenarios, characterized in that: The method comprises a processor running a program of the dynamic voiceprint recognition and updating method for multiple scenarios as claimed in any one of claims 1 to 7.
10. A storage medium, characterized in that: A program storing the multi-scenario dynamic voiceprint recognition and updating method according to any one of claims 1 to 7 is stored.
Citation Information
Cited By
Speech recognition authentication method and system based on multi-modal features and dynamic evaluation
CN120748413A