Finite element identity recognition system based on voiceprint comparison

By combining finite element method with modal and dual-unit estimation, combined with pre-balance processing and pseudo-measurement constraints, the anti-spoofing ability and environmental adaptability problems of the voice identity recognition system are solved, and stable and efficient identity recognition is achieved in different environments.

CN120708623AActive Publication Date: 2025-09-26POINT CONTROL CLOUD (BEIJING) INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511012471.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-09-26
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

Existing voice identity recognition systems have weak anti-spoofing capabilities, poor environmental adaptability, and high dependence on corpus, resulting in poor identity recognition accuracy; they are highly sensitive to noise and environment, have poor stability, and poor recognition effects.

Method used

Finite element combined with modal and dual-unit joint estimation is used to distinguish between real glottal excitation and playback simulation signals. Channel interference is corrected through pre-balancing processing, pseudo-measurement constraints are introduced to eliminate redundant components, excitation unit parameters are adaptively selected, and characterization unit correction is adaptively adjusted to reduce sensitivity to the environment and noise.

Benefits of technology

It improves the accuracy and stability of identity recognition, enhances the detection capability of anti-recording and anti-synthesis attacks, ensures stable extraction of key identity clues in different environments, and reduces performance fluctuations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708623A_ABST
    Figure CN120708623A_ABST
Patent Text Reader

Abstract

The invention discloses a finite element identity recognition system based on voiceprint comparison. The finite element identity recognition system comprises a voiceprint extraction module, a sound channel finite element modal analysis module, a front balance processing module, a double-unit construction module, a pseudo measurement constraint module, an excitation unit parameter adaptive selection module, a representation unit correction module and an identity recognition module. The invention belongs to the field of voice identity recognition, and particularly relates to a finite element identity recognition system based on voiceprint comparison. According to the scheme, real glottis excitation and playing simulation signals are distinguished through finite element combination modality and double-unit joint estimation, and the detection capacity of anti-recording and anti-synthesis attacks is enhanced; pseudo measurement constraints are introduced, redundant excitation components are eliminated, and the distinguishing degree of different speakers is improved; through the correction of the characterization unit, the real-time self-adaptive capability is provided for the environment sudden change aiming at the non-stable speaking and environment, dynamic scaling measurement noise and process noise, the performance fluctuation caused by assumed mismatch is reduced, and the identity recognition effect is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice identity recognition, and in particular to a finite element identity recognition system based on voiceprint comparison. Background Art

[0002] A voice ID recognition system is a technology that identifies speakers based on voice features. It extracts and analyzes unique physiological and behavioral characteristics in speech and compares them with pre-registered voice templates to accurately identify the speaker. However, typical voice ID recognition systems suffer from weak anti-spoofing capabilities, poor environmental adaptability, and a high reliance on corpus data, resulting in poor recognition accuracy. They are also highly sensitive to noise and environmental factors, exhibit poor stability, and are unable to adapt to environmental and speech non-stationarity, leading to poor recognition results. Summary of the Invention

[0003] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides a finite element identity recognition system based on voiceprint comparison. In view of the problems that general voice identity recognition systems have weak anti-spoofing ability, poor environmental adaptability, and high dependence on corpus, which lead to poor identity recognition accuracy, this solution uses finite elements combined with modal and dual-unit joint estimation to distinguish between real glottal excitations and playback simulation signals, thereby enhancing the detection capability of anti-recording and anti-synthesis attacks; through pre-balancing, channel interference correction is performed for call interference and recording equipment differences, ensuring that key identity clues can be stably extracted in mobile phones, conference microphones, and noisy backgrounds; the features based on physical modeling do not rely on a large number of specific spoken words and sentence training, and can directly capture individual characteristics through vocal tract and glottal parameters, reducing the dependence on massive corpus and Reliance on strong supervised training; thereby improving the accuracy of identity recognition; in view of the fact that general voice identity recognition systems are highly sensitive to noise and environment, have poor stability, and weak adaptability to environmental and speech non-stationarity, which leads to poor recognition effect, this scheme introduces pseudo-measurement constraints, eliminates redundant excitation components, and retains only a few non-zero dominant components to improve the discrimination between different speakers; based on the adaptive selection of excitation unit parameters, it can adaptively balance the system's tracking of excitations under different speech content and environments to ensure recognition stability; through representation unit correction, it dynamically scales measurement noise and process noise for speech and environmental non-stationarity, and has real-time adaptive capabilities to environmental mutations and changes in speaking methods, reducing performance fluctuations caused by hypothesis mismatch, thereby improving identity recognition results.

[0004] The technical solution adopted by the present invention is as follows: The present invention provides a finite element identity recognition system based on voiceprint comparison, including a voiceprint extraction module, a vocal tract finite element modal analysis module, a pre-balance processing module, a dual-unit construction module, a pseudo-measurement constraint module, an excitation unit parameter adaptive selection module, a characterization unit correction module and an identity recognition module;

[0005] The voiceprint extraction module performs feature extraction after preprocessing the voice signal;

[0006] The vocal tract finite element modal analysis module constructs a vocal tract dynamics equation to obtain a state space transfer form of a frame-level vocal tract representation vector;

[0007] The front-end balancing processing module constructs an observation matrix and a direct transfer matrix based on the voiceprint features and the state space transfer form;

[0008] The dual-unit construction module first predicts and updates the modal excitation vector through an optimal recursive state estimation unit in each frame of speech; then predicts and updates the vocal tract representation vector through another estimation unit using the modal excitation vector;

[0009] The pseudo-measurement constraint module constructs a sparse constraint equation using a sign function as a pseudo-measurement to eliminate redundant components in the excitation vector;

[0010] The excitation unit parameter adaptive selection module determines the excitation unit parameters based on the curve adaptively;

[0011] The characterization unit correction module adaptively adjusts the disturbance and the observed disturbance by self-adjusting the adjustment coefficient, and corrects the prediction uncertainty and the optimal update weight matrix accordingly;

[0012] The identity recognition module performs voiceprint comparison based on the frame-level laryngeal source excitation timing and the vocal tract modal representation to achieve identity recognition.

[0013] Furthermore, the voiceprint extraction module obtains the voice signal, pre-processes each frame of the voice signal, and extracts the vocal tract formant features. , spectral difference dynamic characteristics and pronunciation intensity contours ; The three types of features are regarded as multi-type sensor responses, expressed as: ; ; ;in, is the voiceprint observation vector of the k-th frame of speech; and are the vocal tract formant features of the k-th frame and the k-1-th frame speech respectively; is the total number of sampling points in each frame of speech; n is the sampling point index; is the speech amplitude of the nth sampling point in the kth frame.

[0014] Furthermore, the vocal tract finite element modal analysis module uses modal coordinates to describe the vocal tract vibration, which is expressed as: ;in, is the amplitude of the speech mode at time t; and They are The first and second derivatives of ; is the characteristic frequency vector of the speech mode; is the speech modal damping ratio vector; is the modal matrix of the vocal tract finite element model; T is the transpose operation; is the force load mapping operator, u is the laryngeal source excitation signal; discretized to each frame model, the vocal tract representation vector of the t-th frame speech Defined as: ; Through discretization, we get: ; ; ; ;in, and are the vocal tract representation vectors of the k+1th and kth frame speech respectively; is the frame interval; is the discrete time representation transfer matrix; B is the discrete time input matrix; is the modal excitation vector of the k-th frame speech; is the process disturbance; I is the identity matrix; is the natural angular frequency of each mode of the vocal tract.

[0015] Furthermore, the pre-balancing processing module constructs the observation matrix C and the direct transfer matrix D, which are expressed as: ; ; ;in, is the observation perturbation vector; is the response weight of the static spectrum to the modal displacement; is the response weight of the dynamic spectrum to the modal displacement; is the response weight of the pronunciation intensity profile to the modal excitation; the measurement matrix is ​​pre-balanced using the measurement disturbance uncertainty matrix R; it is expressed as: ; ;in, and They are the observation matrix and direct transfer matrix after pre-balancing processing.

[0016] Furthermore, in each frame of speech, the dual-unit construction module first uses an optimal recursive state estimation unit to predict and update the modal excitation vector as an excitation unit, and then uses another optimal recursive state estimation unit to predict and update the vocal tract representation vector with the latest excitation estimate as a representation unit; the excitation unit prediction is expressed as: ; ;in, It is the predicted value of the laryngeal source excitation before the observation update of the k-th frame speech; is the updated excitation posterior estimate of the k-1th frame speech observation; Is the predicted value The prediction uncertainty of , which measures the prediction uncertainty; is the posterior uncertainty of the excitation estimate after the k-1th frame speech update; is the uncertainty of the excitation unit disturbance; the excitation unit update is expressed as: ; ; ;in, is the optimal updated weight matrix of the k-th frame speech excitation unit; is the posterior estimate of the speech excitation of the kth frame; is the residual term, the difference between the observation and the prediction; is the posterior estimate of the representation of the k-1th frame of speech; is the a posteriori uncertainty of the stimulus; the representation unit prediction is expressed as: ; ;in, is the prior estimate of the representation of the k-th frame of speech before the measurement update; is the uncertainty matrix of the k-th frame speech representation prediction; is the uncertainty of the characterization unit disturbance; the characterization unit update is expressed as: ; ; ; ;in, is the optimal update weight matrix of the representation unit; is the posterior estimate of the representation of the kth frame; It is the voiceprint observation vector after pre-balancing processing.

[0017] Furthermore, the pseudo-measurement constraint module constructs a pseudo-measurement equation, which is expressed as: ; ;in, is the pseudo-measurement matrix, is a diagonal matrix, is the matrix element at row i and column i; is the estimated value of the excitation corresponding to the i-th mode in the k-th frame; sign(·) is the sign function; Is the tolerance vector, a very small positive number; iterative update, expressed as: ; ; ;in, is the pseudo-measurement optimal update weight matrix for the τth iteration; and are the uncertainty estimates of the excitation vector at the τth iteration and the τ+1th iteration, respectively. The uncertainty before the pseudo-measurement iteration is ; is a pseudo-measurement disturbance; and are the excitation vectors for the τ+1th iteration and the τth iteration respectively.

[0018] Furthermore, the excitation unit parameter adaptive selection module is sensitive to the excitation unit disturbance , calculate the mean square error , expressed as: ; For pseudo-measurement disturbances , calculate the excitation sparsity and observation fitting error , respectively expressed as: ; ; The inflection point is the disturbance uncertainty of the optimal excitation unit and pseudo-measurement disturbances ; is the L2 norm.

[0019] Furthermore, the characterization unit correction module adjusts the uncertainty of the characterization unit disturbance by adaptively adjusting the uncertainty of the characterization unit disturbance for the non-stationary speech and environment. and pseudo-measurement disturbances , and then realize the correction of the optimal update weight matrix of the representation unit; expressed as: ; ; ; ; ; ; ; ;in, is the self-regulating adjustment coefficient; is the trace of the matrix; is the estimate of the uncertainty of the actual innovation in the kth frame; is the uncertainty of the theoretical voiceprint residual; and are the kth and uth frame-level voiceprint prediction residuals respectively; is the residual uncertainty of the sample voiceprint; and It is adjusted and ; and are the prediction uncertainty and optimal update weight matrix after characterization unit correction, respectively.

[0020] Furthermore, the identification module selects the glottal excitation timing obtained after self-adjustment of the excitation unit parameters and correction of the characterization unit. and vocal tract modal response Combined as a bimodal feature; calculate the error metric with each user template in the database , select the smallest one and perform identity authentication, which is expressed as: ;like , the verification passes, otherwise the verification fails; is the error measure between the i-th user and the current speech feature; is the laryngeal source excitation timing vector of the current sentence to be recognized, and the frame-level glottal excitation timing vector Get it by splicing the head and tail together; is the excitation template vector stored in the database by user o; is the channel modal representation vector of the current sentence to be recognized, and the frame-level channel modal response Splicing obtained; is the representation template vector of user o; is the minimum error metric; T is the acceptance threshold; and is the measurement weight.

[0021] The beneficial effects achieved by the present invention using the above scheme are as follows:

[0022] (1) In view of the problems that general voice identity recognition systems have weak anti-spoofing capabilities, poor environmental adaptability, and high dependence on corpus, which lead to poor identity recognition accuracy, this scheme uses finite element combined with modal and dual-unit joint estimation to distinguish between real glottal excitations and playback simulation signals, thereby enhancing the detection capability of anti-recording and anti-synthesis attacks; through pre-balancing, channel interference correction is performed to address call interference and recording equipment differences, ensuring that key identity clues can be stably extracted in mobile phones, conference microphones, and noisy backgrounds; features based on physical modeling do not rely on a large number of specific spoken words and sentence training, and can directly capture individual characteristics through vocal tract and glottal parameters, reducing dependence on massive corpus and strong supervision training; thereby improving identity recognition accuracy.

[0023] (2) In view of the fact that general voice identity recognition systems are highly sensitive to noise and environment, have poor stability, and are weak in adapting to the non-stationarity of the environment and speech, which leads to poor recognition results, this scheme introduces pseudo-measurement constraints, eliminates redundant excitation components, and retains only a few non-zero dominant components, thereby improving the discrimination between different speakers; based on the adaptive selection of excitation unit parameters, the system can adaptively balance the tracking of excitations under different speech content and environments to ensure the stability of recognition; through representation unit correction, the system dynamically scales the measurement noise and process noise to address the non-stationarity of speech and environment, and has real-time adaptive capabilities to sudden changes in the environment and changes in speaking methods, reducing performance fluctuations caused by hypothesis mismatch, thereby improving the identity recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A schematic diagram of the process of a finite element identity recognition system based on voiceprint comparison provided by the present invention.

[0025] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION

[0026] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0027] In the description of the present invention, it should be understood that terms such as "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, they should not be understood as limiting the present invention.

[0028] Example 1, see Figure 1 The present invention provides a finite element identity recognition system based on voiceprint comparison, which includes a voiceprint extraction module, a vocal tract finite element modal analysis module, a pre-balance processing module, a dual-unit construction module, a pseudo-measurement constraint module, an excitation unit parameter adaptive selection module, a characterization unit correction module and an identity recognition module;

[0029] The voiceprint extraction module performs feature extraction after preprocessing the voice signal;

[0030] The vocal tract finite element modal analysis module constructs a vocal tract dynamics equation to obtain a state space transfer form of a frame-level vocal tract representation vector;

[0031] The front-end balancing processing module constructs an observation matrix and a direct transfer matrix based on the voiceprint features and the state space transfer form;

[0032] The dual-unit construction module first predicts and updates the modal excitation vector through an optimal recursive state estimation unit in each frame of speech; then predicts and updates the vocal tract representation vector through another estimation unit using the modal excitation vector;

[0033] The pseudo-measurement constraint module constructs a sparse constraint equation using a sign function as a pseudo-measurement to eliminate redundant components in the excitation vector;

[0034] The excitation unit parameter adaptive selection module determines the excitation unit parameters based on the curve adaptively;

[0035] The characterization unit correction module adaptively adjusts the disturbance and the observed disturbance by self-adjusting the adjustment coefficient, and corrects the prediction uncertainty and the optimal update weight matrix accordingly;

[0036] The identity recognition module performs voiceprint comparison based on the frame-level laryngeal source excitation timing and the vocal tract modal representation to achieve identity recognition.

[0037] Example 2, see Figure 1 This embodiment is based on the above embodiment. The voiceprint extraction module obtains the voice signal, pre-processes each frame of the voice signal, and extracts the vocal tract formant features. , spectral difference dynamic characteristics and pronunciation intensity contours , improving the ability to capture individual acoustic differences; preprocessing includes pre-emphasis, framing, windowing, and short-time Fourier transform; vocal tract formant features reflect the vocal tract formant, that is, the shape of the individual vocal tract; spectral difference dynamic features reflect the changing trend during pronunciation; pronunciation intensity contour reflects pronunciation intensity; the three types of features are regarded as multi-type sensor responses, expressed as: ; ; ;in, is the voiceprint observation vector of the k-th frame of speech; and are the vocal tract formant features of the k-th frame and the k-1-th frame speech respectively; is the total number of sampling points in each frame of speech; n is the sampling point index; is the speech amplitude of the nth sampling point in the kth frame.

[0038] Example 3, see Figure 1This embodiment is based on the above embodiment. The vocal tract finite element modal analysis module uses a finite element model to characterize the vocal tract dynamics. Combined with the physical generation mechanism of the speech signal, it retains the main vibration mode through modal dimension reduction, significantly reducing the computational overhead. The vocal tract vibration is described using modal coordinates, which are expressed as: ;in, is the amplitude of the speech mode at time t; and They are The first and second derivatives of ; is the characteristic frequency vector of the speech mode; is the speech modal damping ratio vector; is the modal matrix of the vocal tract finite element model; T is the transpose operation; It is a force load mapping operator that maps the laryngeal excitation signal into a linear operator of equivalent force load on the nodes of the vocal tract finite element model. u is the laryngeal excitation signal, which represents the physical driving force input from the glottis to the vocal tract. Discretized to each frame model, the vocal tract representation vector of the t-th frame speech Defined as: ; Through discretization, we get: ; ; ; ;in, and are the vocal tract representation vectors of the k+1th and kth frame speech respectively; is the frame interval; is the discrete time representation transfer matrix; B is the discrete time input matrix; is the modal excitation vector of the k-th frame speech; is the process disturbance; I is the identity matrix; is the natural angular frequency of each mode of the vocal tract; through glottal excitation and vocal tract modal estimation, more robust and physically consistent voiceprint features are extracted.

[0039] Example 4, see Figure 1 This embodiment is based on the above embodiment. The pre-balancing processing module maps the physical representation to the actual extracted voiceprint feature space to achieve finite element ↔ voiceprint bridging; different feature dimensions correspond to different observation channels; the observation matrix C and the direct transfer matrix D are constructed, which are expressed as: ; ; ;in, is the observed disturbance vector, which reflects the measurement disturbance mixed in the actual extracted voiceprint features. , R is the observation disturbance uncertainty matrix, and the disturbance levels of different characteristic components vary greatly; is the response weight of the static spectrum to the modal displacement; is the response weight of the dynamic spectrum to the modal displacement; is the response weight of the pronunciation intensity profile to the modal excitation; the disturbance levels of different voiceprint features vary greatly, and direct filtering is prone to numerical pathology. The measurement disturbance uncertainty matrix R is used to pre-balance the observation matrix; it can be expressed as: ; ;in, and They are the observation matrix and direct transfer matrix after pre-balancing processing; accurately recover the physical representation and glottal excitation from the disturbance; ensure that different feature channels can be reasonably weighed regardless of whether in a quiet environment or a disturbed environment, and the system will not fail due to excessive disturbance in one channel.

[0040] Example 5, see Figure 1 This embodiment is based on the above embodiment. In each frame of speech, the dual-unit construction module first uses an optimal recursive state estimation unit to predict and update the modal excitation vector as the excitation unit, and then uses another optimal recursive state estimation unit to predict and update the vocal tract representation vector with the latest excitation estimate as the representation unit. Through the alternating iteration of the excitation unit and the representation unit, the laryngeal source excitation and the vocal tract representation are jointly estimated to improve the physical consistency of the vocal tract dynamics and the pronunciation drive. The excitation unit prediction is expressed as: ; ;in, It is the predicted value of the laryngeal source excitation before the observation update of the k-th frame speech; is the updated excitation posterior estimate of the k-1th frame speech observation; Is the predicted value The prediction uncertainty of , which measures the prediction uncertainty; is the posterior uncertainty of the excitation estimate after the k-1th frame speech update; is the uncertainty of the excitation unit disturbance, is the random variation amplitude of the laryngeal source excitation between frames; the excitation unit update is expressed as: ; ; ;in, is the optimal updated weight matrix of the k-th frame speech excitation unit; is the posterior estimate of the speech excitation of the kth frame; is the residual term, the difference between the observation and the prediction; is the posterior estimate of the representation of the k-1th frame of speech; is the a posteriori uncertainty of the stimulus; the representation unit prediction is expressed as: ; ;in, is the prior estimate of the representation of the k-th frame of speech before the measurement update; is the uncertainty matrix of the k-th frame speech representation prediction; is the uncertainty of the characterization unit disturbance; the characterization unit update is expressed as: ; ; ; ;in, is the optimal update weight matrix of the representation unit; is the posterior estimate of the representation of the kth frame; It is the voiceprint observation vector after pre-balancing processing.

[0041] By performing the above operations, this solution addresses the problems of weak anti-spoofing ability, poor environmental adaptability, and high dependence on corpus in general voice identity recognition systems, which lead to poor identity recognition accuracy. By combining finite element analysis with modal and dual-unit joint estimation, this solution distinguishes between real glottal excitations and playback simulation signals, thereby enhancing the detection capability of anti-recording and anti-synthesis attacks. Through pre-balancing, channel interference correction is performed to address call interference and recording device differences, ensuring that key identity clues can be stably extracted even on mobile phones, conference microphones, and in noisy backgrounds. The features based on physical modeling do not rely on a large number of specific spoken words and sentence training, but can directly capture individual characteristics through vocal tract and glottal parameters, reducing dependence on massive corpus and strong supervised training, thereby improving identity recognition accuracy.

[0042] Example 6, see Figure 1 This embodiment is based on the above embodiment. The pseudo-measurement constraint module is based on the fact that when the speaker pronounces, only a few vocal cord muscle groups dominate the excitation. Therefore, pseudo-measurement is used to force L1 sparseness and eliminate redundant excitation components. Based on iterative convergence, only a few non-zero dominant components are retained. The pseudo-measurement equation is constructed and expressed as: ; ;in, is the pseudo-measurement matrix, is a diagonal matrix, is the matrix element at row i and column i; is the estimated value of the excitation corresponding to the i-th mode in the k-th frame; sign(·) is the sign function; Is the tolerance vector, a very small positive number; iterative update, expressed as: ; ; ;in, is the pseudo-measurement optimal update weight matrix for the τth iteration; and are the uncertainty estimates of the excitation vector at the τth iteration and the τ+1th iteration, respectively. The uncertainty before the pseudo-measurement iteration is ; is a pseudo-measurement disturbance; and are the excitation vectors for the τ+1th iteration and the τth iteration respectively.

[0043] Example 7, see Figure 1 This embodiment is based on the above embodiment, and the excitation unit parameter adaptive selection module is used to select the excitation unit disturbance uncertainty. and pseudo-measurement disturbances It is difficult to set a priori, and the two L-curve inflection points are adaptively selected; for the excitation unit disturbance , calculate the mean square error , expressed as: ; For pseudo-measurement disturbances , calculate the excitation sparsity and observation fitting error , respectively expressed as: ; ;for by As the horizontal axis, As the vertical axis, draw the L-curve; the inflection point is the disturbance uncertainty of the optimal excitation unit and pseudo-measurement disturbances ; is the L2 norm.

[0044] Example 8, see Figure 1 The representation unit correction module adjusts the uncertainty of the representation unit disturbance adaptively for the non-stationary speech and environment. and pseudo-measurement disturbances , and then realize the correction of the optimal update weight matrix of the representation unit; expressed as: ; ; ; ; ; ; ; ;in, is the self-regulating adjustment coefficient; is the trace of the matrix; is the estimate of the uncertainty of the actual innovation in the kth frame; is the uncertainty of the theoretical voiceprint residual; and are the kth and uth frame-level voiceprint prediction residuals respectively; is the residual uncertainty of the sample voiceprint; and It is adjusted and ; and are the prediction uncertainty and optimal update weight matrix after characterization unit correction, respectively.

[0045] Example 9, see Figure 1This embodiment is based on the above embodiment. The identification module selects the glottal excitation timing obtained after self-adjustment of the excitation unit parameters and correction of the characterization unit. and vocal tract modal response Combined as a bimodal feature; calculate the error metric with each user template in the database , select the smallest one and perform identity authentication, which is expressed as: ;like , the verification passes, otherwise the verification fails; is the error measure between the i-th user and the current speech feature; is the laryngeal source excitation timing vector of the current sentence to be recognized, and the frame-level glottal excitation timing vector Get it by splicing the head and tail together; is the excitation template vector stored in the database by user o; is the channel modal representation vector of the current sentence to be recognized, and the frame-level channel modal response Splicing obtained; is the representation template vector of user o; is the minimum error metric; T is the acceptance threshold; and is the measurement weight.

[0046] By performing the above operations, the general voice identity recognition system has the problems of high sensitivity to noise and environment, poor stability, and weak adaptability to environmental and speech non-stationarity, which leads to poor recognition effect. This scheme introduces pseudo-measurement constraints to eliminate redundant excitation components and retain only a few non-zero dominant components, thereby improving the discrimination between different speakers; based on the adaptive selection of excitation unit parameters, it can adaptively balance the system's tracking of excitations under different speech content and environments to ensure recognition stability; through representation unit correction, it dynamically scales measurement noise and process noise to address the non-stationarity of speech and environment, and has real-time adaptive capabilities to sudden environmental changes and changes in speaking methods, reducing performance fluctuations caused by hypothesis mismatch, thereby improving identity recognition effect.

[0047] While the embodiments of the present invention have been shown and described, it will be apparent to those skilled in the art that various changes, modifications, substitutions, and alterations can be made to the embodiments without departing from the principles and spirit of the invention.

[0048] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.

Claims

1. A finite element identity recognition system based on voiceprint comparison, characterized by: The system includes a voiceprint extraction module, a vocal tract finite element modal analysis module, a front-end balance processing module, a dual-unit construction module, a pseudo-measurement constraint module, an excitation unit parameter adaptive selection module, a characterization unit correction module, and an identity recognition module; The voiceprint extraction module performs feature extraction after preprocessing the voice signal; The vocal tract finite element modal analysis module constructs a vocal tract dynamics equation to obtain a state space transfer form of a frame-level vocal tract representation vector; The front-end balancing processing module constructs an observation matrix and a direct transfer matrix based on the voiceprint features and the state space transfer form; The dual-unit construction module first predicts and updates the modal excitation vector through an optimal recursive state estimation unit in each frame of speech; then predicts and updates the vocal tract representation vector through another estimation unit using the modal excitation vector; The pseudo-measurement constraint module constructs a sparse constraint equation using a sign function as a pseudo-measurement to eliminate redundant components in the excitation vector; The excitation unit parameter adaptive selection module determines the excitation unit parameters based on the curve adaptively; The characterization unit correction module adaptively adjusts the disturbance and the observed disturbance by self-adjusting the adjustment coefficient, and corrects the prediction uncertainty and the optimal update weight matrix accordingly; The identity recognition module performs voiceprint comparison based on the frame-level laryngeal source excitation timing and the vocal tract modal representation to achieve identity recognition.

2. The finite element identity recognition system based on voiceprint comparison according to claim 1, characterized in that: The voiceprint extraction module obtains the voice signal and extracts the vocal tract formant features after preprocessing each frame of the voice signal. , spectral difference dynamic characteristics and pronunciation intensity contours ; The three types of features are regarded as multi-type sensor responses, expressed as: ; ; ;in, is the voiceprint observation vector of the k-th frame of speech; and are the vocal tract formant features of the k-th frame and the k-1-th frame speech respectively; is the total number of sampling points in each frame of speech; n is the sampling point index; is the speech amplitude of the nth sampling point in the kth frame.

3. The finite element identity recognition system based on voiceprint comparison according to claim 2, characterized in that: The vocal tract finite element modal analysis module uses modal coordinates to describe the vocal tract vibration, which is expressed as: ;in, is the amplitude of the speech mode at time t; and They are The first and second derivatives of ; is the characteristic frequency vector of the speech mode; is the speech modal damping ratio vector; is the modal matrix of the vocal tract finite element model; T is the transpose operation; is the force load mapping operator, u is the laryngeal source excitation signal; discretized to each frame model, the vocal tract representation vector of the t-th frame speech Defined as: ; Through discretization, we get: ; ; ; ;in, and are the vocal tract representation vectors of the k+1th and kth frame speech respectively; is the frame interval; is the discrete time representation transfer matrix; B is the discrete time input matrix; is the modal excitation vector of the k-th frame speech; is the process disturbance; I is the identity matrix; is the natural angular frequency of each mode of the vocal tract.

4. The finite element identity recognition system based on voiceprint comparison according to claim 3 is characterized by: The pre-balancing processing module constructs the observation matrix C and the direct transfer matrix D, which are expressed as: ; ; ;in, is the observation perturbation vector; is the response weight of the static spectrum to the modal displacement; is the response weight of the dynamic spectrum to the modal displacement; is the response weight of the pronunciation intensity profile to the modal excitation; the measurement matrix is ​​pre-balanced using the measurement disturbance uncertainty matrix R; it is expressed as: ; ;in, and They are the observation matrix and direct transfer matrix after pre-balancing processing.

5. The finite element identity recognition system based on voiceprint comparison according to claim 4 is characterized in that: In each frame of speech, the dual-unit building block first uses an optimal recursive state estimation unit to predict and update the modal excitation vector as the excitation unit, and then uses another optimal recursive state estimation unit to predict and update the vocal tract representation vector with the latest excitation estimate as the representation unit; the excitation unit prediction is expressed as: ; ;in, It is the predicted value of the laryngeal source excitation before the observation update of the k-th frame speech; is the updated excitation posterior estimate of the k-1th frame speech observation; Is the predicted value The prediction uncertainty of , which measures the prediction uncertainty; is the posterior uncertainty of the excitation estimate after the k-1th frame speech update; is the uncertainty of the excitation unit disturbance; the excitation unit update is expressed as: ; ; ;in, is the optimal updated weight matrix of the k-th frame speech excitation unit; is the posterior estimate of the speech excitation of the kth frame; is the residual term, the difference between the observation and the prediction; is the posterior estimate of the representation of the k-1th frame of speech; is the a posteriori uncertainty of the stimulus; the representation unit prediction is expressed as: ; ;in, is the prior estimate of the representation of the k-th frame of speech before the measurement update; is the uncertainty matrix of the k-th frame speech representation prediction; is the uncertainty of the characterization unit disturbance; the characterization unit update is expressed as: ; ; ; ;in, is the optimal update weight matrix of the representation unit; is the posterior estimate of the representation of the kth frame; It is the voiceprint observation vector after pre-balancing processing.

6. The finite element identity recognition system based on voiceprint comparison according to claim 5, characterized in that: The pseudo-measurement constraint module constructs a pseudo-measurement equation, which is expressed as: ; ;in, is the pseudo-measurement matrix, is a diagonal matrix, is the matrix element at row i and column i; is the estimated value of the excitation corresponding to the i-th mode in the k-th frame; sign(·) is the sign function; is the tolerance vector, a very small positive number; iterative update is expressed as: ; ; ;in, is the pseudo-measurement optimal update weight matrix for the τth iteration; and are the uncertainty estimates of the excitation vector at the τth iteration and the τ+1th iteration, respectively. The uncertainty before the pseudo-measurement iteration is ; is a pseudo-measurement disturbance; and are the excitation vectors for the τ+1th iteration and the τth iteration respectively.

7. The finite element identity recognition system based on voiceprint comparison according to claim 6, characterized in that: The excitation unit parameter adaptive selection module is sensitive to the excitation unit disturbance , calculate the mean square error , expressed as: ; For pseudo-measurement disturbances , calculate the incentive sparsity and observation fitting error , respectively expressed as: ; ; The inflection point is the disturbance uncertainty of the optimal excitation unit and pseudo-measurement disturbances ; is the L2 norm.

8. The finite element identity recognition system based on voiceprint comparison according to claim 7, characterized in that: The characterization unit correction module adjusts the uncertainty of the characterization unit disturbance by adaptively adjusting the uncertainty of the characterization unit disturbance for the non-stationary speech and environment. and pseudo-measurement disturbances , and then realize the correction of the optimal update weight matrix of the representation unit; expressed as: ; ; ; ; ; ; ; ;in, is the self-regulating adjustment coefficient; is the trace of the matrix; is the estimate of the uncertainty of the actual innovation in the kth frame; is the uncertainty of the theoretical voiceprint residual; and are the kth and uth frame-level voiceprint prediction residuals respectively; is the residual uncertainty of the sample voiceprint; and It is adjusted and ; and are the prediction uncertainty and optimal update weight matrix after characterization unit correction, respectively.

9. The finite element identity recognition system based on voiceprint comparison according to claim 8, characterized in that: The identity recognition module selects the glottal excitation time sequence obtained after self-adjustment of the excitation unit parameters and correction of the characterization unit. and vocal tract modal response Combined as a bimodal feature; calculate the error metric with each user template in the database , select the smallest one and perform identity authentication, which is expressed as: ;like , the verification passes, otherwise the verification fails; is the error measure between the i-th user and the current speech feature; is the laryngeal source excitation timing vector of the current sentence to be recognized, and the frame-level glottal excitation timing vector Get it by splicing the head and tail together; is the excitation template vector stored in the database by user o; is the channel modal representation vector of the current sentence to be recognized, and the frame-level channel modal response Splicing obtained; is the representation template vector of user o; is the minimum error metric; T is the acceptance threshold; and is the measurement weight.

Citation Information

Patent Citations

  • Autonomous fitness for service assessment

    CA2912296A1

  • Determining method of numerical value of vibration noise of axle housing of drive axle

    CN107391816A

  • Multi-modal interactive structural mechanical analysis method and system, computer equipment and medium

    CN115756161A

  • Anti-disturbance voiceprint verification method and related device

    CN116994601A

  • Sound wave detection method for filtering vibration noise of transformer

    CN117690445A

Cited By

  • Magnetic flux regulation crosstalk calibration method for quantum bits

    CN121235140A