A finite element identity recognition system based on voiceprint comparison
By combining finite element method with modal and dual-element estimation, the system distinguishes between real glottal excitation and simulated playback signals, corrects channel interference, eliminates redundant components, and adaptively adjusts parameters. This solves the problems of weak anti-spoofing ability and poor environmental adaptability of voice identity recognition systems, achieving higher recognition accuracy and stability.
Patent Information
- Application Number
- CN202511012471.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-07-22
AI Technical Summary
Existing voice identity recognition systems have weak anti-spoofing capabilities, poor environmental adaptability, and high dependence on corpora, resulting in poor identity recognition accuracy; they are also highly sensitive to noise and environmental factors, have poor stability, and poor recognition performance.
By employing finite element method combined with modal and dual-element joint estimation, the system distinguishes between real glottal excitation and playback simulation signals. It corrects channel interference through pre-balance processing, introduces pseudo-measurement constraints to eliminate redundant components, adaptively adjusts excitation unit parameters, and performs characterization unit correction to achieve identity recognition.
It improves the accuracy and stability of identity recognition, enhances the detection capability against anti-recording and anti-synthesis attacks, reduces the dependence on massive corpora, has real-time adaptive capabilities, and reduces performance fluctuations.
Smart Images

Figure CN120708623B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of voice identity recognition, and particularly to a finite element identity recognition system based on voiceprint comparison. BACKGROUND
[0002] The voice identity recognition system is a technical system for identifying the identity of a speaker based on voice features. It extracts and analyzes unique physiological and behavioral features in the voice, compares them with pre-registered voice templates, and thus accurately identifies the identity of the speaker. However, the general voice identity recognition system has weak anti-fraud ability, poor environmental adaptability, high dependence on corpus, and thus poor identity recognition accuracy. The general voice identity recognition system has high sensitivity to noise and environment, poor stability, weak adaptability to non-stationary environment and speech, and thus poor recognition effect. SUMMARY
[0003] In view of the above problems, the present application provides a finite element identity recognition system based on voiceprint comparison, which can overcome the defects of the prior art. The present application can distinguish between real glottal excitation and playback simulation signals by combining modal and double-unit joint estimation, thereby enhancing the detection ability of anti-recording and anti-synthesis attacks. The pre-balance can correct the channel interference caused by the difference between the conversation interference and the recording device, ensuring that the key identity clues can be stably extracted in mobile phones, conference microphones, and noisy backgrounds. The features based on physical modeling are not dependent on a large number of specific speech words and sentence training, and can directly capture individual features through the vocal tract and glottal parameters, thereby reducing the dependence on massive corpus and strong supervision training. Thus, the identity recognition accuracy is improved. The present application can improve the discrimination of different speakers by introducing pseudo-measurement constraints, eliminating redundant excitation components, and retaining only a few non-zero dominant components. The excitation unit parameters are adaptively selected, which can adaptively balance the tracking of the excitation in different speech content and environments, ensuring the stability of the recognition. The representation unit correction can adaptively track the dynamic scaling of measurement noise and process noise in response to non-stationary speech and environment, and has real-time adaptive ability to environmental mutations and changes in speech methods, thereby reducing the performance fluctuations caused by hypothesis mismatch, and improving the identity recognition effect.
[0004] The technical scheme adopted by the present application is as follows: the present application provides a finite element identity recognition system based on voiceprint comparison, comprising a voiceprint extraction module, a sound channel finite element modal analysis module, a pre-balancing processing module, a double-unit construction module, a pseudo-measurement constraint module, an excitation unit parameter adaptive selection module, a representation unit correction module and an identity recognition module.
[0005] The voiceprint extraction module extracts features after pre-processing the voice signal.
[0006] The sound channel finite element modal analysis module constructs a sound channel dynamics equation to obtain a state space transfer form of a frame-level sound channel representation vector.
[0007] The pre-balancing processing module constructs an observation matrix and a direct transfer matrix based on voiceprint features and the state space transfer form.
[0008] The double-unit construction module first predicts and updates a modal excitation vector in each frame of voice through an optimal recursive state estimation unit; then predicts and updates a sound channel representation vector through another estimation unit with the modal excitation vector.
[0009] The pseudo-measurement constraint module constructs a sparse constraint equation with a sign function as a pseudo-measurement to eliminate redundant components in the excitation vector.
[0010] The excitation unit parameter adaptive selection module determines excitation unit parameters based on curve adaptation.
[0011] The representation unit correction module adjusts the disturbance and the observation disturbance adaptively by self-adjusting the adjustment coefficient, and corrects the prediction uncertainty and the optimal update weight matrix accordingly.
[0012] The identity recognition module compares voiceprints based on frame-level laryngeal source excitation timing and sound channel modal representation to achieve identity recognition.
[0013] Further, the voiceprint extraction module acquires a voice signal, pre-processes each frame of voice signal, and extracts sound channel formant features , spectral difference dynamic features and pronunciation intensity profiles ; the three types of features are regarded as responses of multiple types of sensors and represented as: ; ; ; wherein, is the voiceprint observation vector of the kth frame of voice; and are the sound channel formant features of the kth frame and the k-1th frame of voice, respectively; is the total number of sampling points of each frame of voice; n is the sampling point index; is the voice amplitude of the nth sampling point in the kth frame.
[0014] Further, the vocal tract finite element modal analysis module describes the vocal tract vibration with modal coordinates, denoted as: ; wherein, is the speech modal amplitude at time t; and are the first and second derivatives of , respectively; is the characteristic frequency vector of the speech modal; is the modal damping ratio vector of the speech; is the modal matrix of the vocal tract finite element model; T is the transpose operation; is the force load mapping operator, u is the glottal source excitation signal; discretized to each frame model, the vocal tract representation vector of the t-th frame speech is defined as: ; by discretization, we get: ; ; ; ; wherein, and are the vocal tract representation vectors of the k+1-th and k-th frame speech, respectively; is the frame interval; is the discrete-time representation transition matrix; B is the discrete-time input matrix; is the modal excitation vector of the k-th frame speech; is the process disturbance; I is the unit matrix; is the natural angular frequency of each modal of the vocal tract.
[0015] Further, the pre-whitening processing module constructs the observation matrix C and the direct transfer matrix D, denoted as: ; ; ; wherein, is the observation disturbance vector; is the static spectrum response weight to the modal displacement; is the dynamic spectrum response weight to the modal displacement; is the response weight of the pronunciation intensity profile to the modal excitation; the observation matrix is pre-whitened with the measurement disturbance uncertainty matrix R; denoted as: ; ; wherein, and are the observation matrix and the direct transfer matrix after pre-whitening, respectively.
[0016] Further, the dual-unit construction module predicts and updates the modal excitation vector as an excitation unit using one optimal recursive state estimation unit in each frame of speech, and predicts and updates the vocal tract representation vector as a representation unit using another optimal recursive state estimation unit with the latest excitation estimation; the excitation unit prediction is represented as: ; ; wherein, is the prediction of the glottal source excitation before the observation update in the kth frame of speech; is the posteriori excitation estimation after the observation update in the k-1th frame of speech; is the prediction uncertainty of the prediction ; and is the posteriori uncertainty of the excitation estimation after the update in the k-1th frame of speech; is the excitation unit disturbance uncertainty; the excitation unit update is represented as: ; ; ; wherein, is the optimal update weight matrix of the excitation unit in the kth frame of speech; is the posteriori excitation estimation in the kth frame of speech; is the residual term, the difference between the observation and the prediction; is the representation posteriori estimation in the k-1th frame of speech; is the posteriori uncertainty of the excitation; the representation unit prediction is represented as: ; ; wherein, is the representation priori estimation before the measurement update in the kth frame of speech; is the uncertainty matrix of the representation prediction in the kth frame of speech; is the representation unit disturbance uncertainty; the representation unit update is represented as: ; ; ; ; wherein, is the optimal update weight matrix of the representation unit; is the representation posteriori estimation in the kth frame of speech; is the voiceprint observation vector after the pre-emphasis processing.
[0017] Further, the pseudo-measurement constraint module is to construct a pseudo-measurement equation, represented as: ; ; wherein, is the pseudo-measurement matrix, which is a diagonal matrix, is the matrix element in the i-th row and the i-th column; is the excitation estimation value corresponding to the i-th mode in the kth frame of speech; sign(·) is the sign function; is the tolerance vector, a small positive number; the iterative update is denoted as: ; ; ; where, is the pseudo-measurement optimal update weight matrix at the τth iteration; and are the uncertainty estimates of the excitation vector at the τth iteration and the (τ+1)th iteration, respectively, and the uncertainty before the pseudo-measurement iteration is ; is the pseudo-measurement perturbation; and are the excitation vectors at the (τ+1)th iteration and the τth iteration, respectively.
[0018] Further, the excitation unit parameter adaptive selection module calculates the mean square error , denoted as: ; and calculates the excitation sparsity and the observation fitting error , denoted as: ; ; the turning point is the perturbation uncertainty of the optimal excitation unit and the pseudo-measurement perturbation ; is the L2 norm. Further, the characterization unit correction module, for speech and environmental non-stationarity, adaptively adjusts the characterization unit perturbation uncertainty and the pseudo-measurement perturbation
[0019] , thereby realizing correction of the optimal update weight matrix of the characterization unit; denoted as: ; ; ; ; ; ; ; ; ; ; where, is a self-adjusting adjustment coefficient; is the trace of a matrix; is the estimated value of the actual innovation uncertainty of the kth frame; is the theoretical voiceprint residual uncertainty; and are the kth and u-th frame level voiceprint prediction residuals, respectively; is the sample voiceprint residual uncertainty; and are the adjusted and ; and These are the prediction uncertainty after characterization unit correction and the optimal updated weight matrix, respectively.
[0020] Furthermore, the identity recognition module will select the glottal excitation timing based on the self-adjusting parameters of the excitation unit and the correction by the representation unit. With vocal tract modal response The combination is used as a bimodal feature; an error metric is calculated with each user template in the database. Select the smallest value for identity verification, as shown below: ;like If the result is positive, the verification passes; otherwise, the verification fails. It is the error measure between the i-th user and the current speech features; It is the laryngeal source excitation timing vector of the current sentence to be identified, and the frame-level glottal excitation timing... It is obtained by piecing the beginning and end together; It is the incentive template vector stored in the database for user number 0; It is the vocal tract modal representation vector of the current sentence to be identified, which will be the frame-level vocal tract modal response. obtained by piecing together; It is the representation template vector of user o; It is the minimum error metric; T is the acceptance threshold; and It measures weight.
[0021] The beneficial effects achieved by the present invention using the above solution are as follows:
[0022] (1) In view of the problems that general voice identity recognition systems have weak anti-spoofing ability, poor environmental adaptability, and high dependence on corpus, which leads to poor identity recognition accuracy, this solution uses finite element method combined with modal and dual-unit joint estimation to distinguish between real glottal excitation and playback simulation signal, thereby enhancing the detection capability of anti-recording and anti-synthesis attacks; through pre-balancing to correct channel interference for differences in call interference and recording equipment, it ensures that key identity clues can be stably extracted in mobile phones, conference microphones, and noisy backgrounds; the physical modeling-based features do not rely on a large number of specific words and sentences for training, and can directly capture individual features through vocal tract and glottal parameters, reducing the dependence on massive corpus and strong supervised training; thereby improving the accuracy of identity recognition.
[0023] (2) For the general speech identity recognition system, the system has high sensitivity to noise and environment, poor stability, weak adaptability to non-stationary environment and speaker, and poor recognition effect. The scheme introduces pseudo measurement constraint, removes redundant excitation components, and only retains a few non-zero dominant components to improve the discrimination of different speakers. Based on the adaptive selection of excitation unit parameters, the system can adaptively balance the tracking of excitation under different speaking content and environment, and ensure the stability of recognition. Through the correction of the representation unit, the dynamic scaling of the measurement noise and the process noise is adjusted according to the non-stationary environment and speaker, and the real-time adaptive ability to environmental changes and speaking methods is achieved, the performance fluctuation caused by hypothesis mismatch is reduced, and the identity recognition effect is improved. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 A flowchart of a finite element identity recognition system based on voiceprint comparison is provided.
[0025] The accompanying drawings are used to provide a further understanding of the present application, and constitute a part of the specification, and are used to explain the present application together with embodiments of the present application, and do not constitute a limitation on the present application. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present application will be described in detail below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0027] In the description of the present application, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore it cannot be understood as a limitation on the present application.
[0028] Embodiment one, refer to Figure 1 The present application provides a finite element identity recognition system based on voiceprint comparison, which comprises a voiceprint extraction module, a voice channel finite element modal analysis module, a pre-balance processing module, a double unit construction module, a pseudo measurement constraint module, an excitation unit parameter adaptive selection module, a representation unit correction module and an identity recognition module.
[0029] The voiceprint extraction module extracts features after pre-processing the speech signal;
[0030] The vocal tract finite element modal analysis module constructs a vocal tract dynamic equation to obtain a state space transfer form of a frame-level vocal tract characterization vector;
[0031] The front balance processing module constructs an observation matrix and a direct transfer matrix based on the voiceprint feature and the state space transfer form;
[0032] The double-unit construction module first predicts and updates a modal excitation vector in each frame of speech through an optimal recursive state estimation unit; then predicts and updates a vocal tract characterization vector through another estimation unit with the modal excitation vector;
[0033] The pseudo-measurement constraint module constructs a sparse constraint equation with a sign function as a pseudo-measurement to eliminate redundant components in the excitation vector;
[0034] The excitation unit parameter adaptive selection module determines the excitation unit parameters based on curve adaptation;
[0035] The characterization unit correction module adaptively adjusts the disturbance and the observation disturbance by self-adjusting the adjustment coefficient, and correspondingly corrects the prediction uncertainty and the optimal update weight matrix;
[0036] The identity recognition module performs voiceprint comparison based on the frame-level laryngeal source excitation timing and the vocal tract modal characterization to realize identity recognition.
[0037] Embodiment two, refer to Figure 1 Based on the above embodiment, the voiceprint extraction module acquires the speech signal, extracts the vocal tract formant feature , the spectral difference dynamic feature and the pronunciation intensity profile after pre-processing each frame of speech signal, improves the capture ability of individual acoustic differences; the pre-processing includes pre-emphasis, framing, windowing and short-time Fourier transform; the vocal tract formant feature reflects the vocal tract formant, i.e. the individual vocal tract shape; the spectral difference dynamic feature reflects the change trend in the pronunciation process; the pronunciation intensity profile reflects the pronunciation intensity; the three types of features are regarded as responses of multiple types of sensors and are expressed as: ; ; ; wherein, is the voiceprint observation vector of the kth frame of speech; and are the vocal tract formant features of the kth frame and the k-1th frame of speech, respectively; is the total number of sampling points of each frame of speech; n is the sampling point index; is the speech amplitude of the nth sampling point in the kth frame.
[0038] Embodiment three, refer to Figure 1, the embodiment is based on the above-mentioned embodiment, the vocal tract finite element modal analysis module uses a finite element model to depict the dynamics of the vocal tract, combines the physical generation mechanism of the speech signal, reduces the operation cost by retaining the main mode shape through modal dimension reduction; the modal coordinates are used to describe the vibration of the vocal tract, which is expressed as: ; wherein, is the modal amplitude of the speech at time t; and are the first derivative and the second derivative of , respectively; is the characteristic frequency vector of the speech mode; is the damping ratio vector of the speech mode; is the modal matrix of the vocal tract finite element model; T is the transpose operation; is a force load mapping operator, which is a linear operator that maps the glottal source excitation signal to the equivalent force load on the nodes of the vocal tract finite element model, u is the glottal source excitation signal, representing the physical driving force input to the vocal tract at the glottis; the vocal tract representation vector of the t-th frame of speech is defined as: ; by discretization, we get: ; ; ; ; wherein, and are the vocal tract representation vectors of the k+1-th and k-th frames of speech, respectively; is the frame interval; is the discrete-time representation transition matrix; B is the discrete-time input matrix; is the modal excitation vector of the k-th frame of speech; is the process disturbance; I is the identity matrix; is the natural angular frequency of each mode of the vocal tract; by estimating the glottis excitation and the vocal tract mode, more robust and physically consistent voiceprint features are extracted.
[0039] Embodiment four, see Figure 1 , the embodiment is based on the above-mentioned embodiment, the pre-balancing processing module maps the physical representation to the actual extracted voiceprint feature space, realizing the finite element ↔ voiceprint bridge; different feature dimensions correspond to different observation channels; the observation matrix C and the direct transfer matrix D are constructed, which are expressed as: ; ; ; wherein, is the observation disturbance vector, which reflects the measurement disturbance mixed in the actually extracted voiceprint features, and is taken as , R is the observation disturbance uncertainty matrix, and the disturbance levels of different feature components are very different; is the response weight of the static spectrum to the modal displacement; is the response weight of the dynamic spectrum to the modal displacement; is the response weight of the pronunciation intensity profile to the modal excitation; and different voiceprint feature perturbation levels differ greatly, direct filtering is numerically ill-conditioned, and a measurement perturbation uncertainty matrix R is used for pre-balancing processing of the observation matrix; is expressed as: ; ; wherein, and are the observation matrix and the direct transfer matrix after pre-balancing processing, respectively; physical representation and glottal excitation are accurately recovered from perturbation; and it is ensured that different feature channels can be reasonably balanced regardless of a quiet environment or a perturbed environment, and a system will not fail due to a certain road perturbation being too large.
[0040] Embodiment five, referring to Figure 1 , which is based on the above-mentioned embodiment, a double-unit construction module is used to predict and update the modal excitation vector as an excitation unit by using one optimal recursive state estimation unit in each frame of speech, and to predict and update the vocal tract representation vector as a representation unit by using another optimal recursive state estimation unit with the latest excitation estimation; through the alternating iteration of the excitation unit and the representation unit, the laryngeal source excitation and the vocal tract representation are jointly estimated, and the physical consistency of the vocal tract dynamics and the pronunciation driving is improved; the excitation unit prediction is expressed as: ; ; wherein, is the predicted value of the laryngeal source excitation before the observation update of the kth frame of speech; is the excitation posterior estimation value after the observation update of the k-1th frame of speech; is the prediction uncertainty of the predicted value ; is the posterior uncertainty of the excitation estimation after the update of the k-1th frame of speech; is the excitation unit perturbation uncertainty, which is the random fluctuation amplitude of the laryngeal source excitation between frames; the excitation unit update is expressed as: ; ; ; wherein, is the optimal update weight matrix of the excitation unit of the kth frame of speech; is the excitation posterior estimation of the kth frame of speech; is a residual term, the difference between the observation and the prediction; is the representation posterior estimation of the k-1th frame of speech; is the excitation posterior uncertainty; the representation unit prediction is expressed as: ; ; wherein, is the representation prior estimation of the kth frame of speech before the measurement update; is the uncertainty matrix of the representation prediction of the kth frame of speech; is the uncertainty of the unit perturbation; the update of the representation of the characterization unit is: ; ; ; ; wherein, is the optimal update weight matrix of the characterization unit; is the posteriori estimation of the characterization of the kth frame; is the voiceprint observation vector after pre-balancing processing.
[0041] By performing the above operations, for the general voice identity recognition system, the anti-cheating ability is weak, the environmental adaptability is poor, the dependence on corpus is high, and the identity recognition accuracy is poor. The scheme combines finite elements with modal and double-unit joint estimation to distinguish between real glottal excitation and playback simulation signals, enhance the detection ability of anti-recording and anti-synthesis attacks, correct channel interference for the differences between the recording device and the conversation interference, ensure that the key identity clues can be stably extracted in mobile phones, conference microphones, and noisy backgrounds, and reduce the dependence on massive corpus and strong supervision training by directly capturing individual features through the vocal tract and glottal parameters, thereby improving the identity recognition accuracy.
[0042] Embodiment six, refer to Figure 1 Based on the above embodiment, the pseudo-measurement constraint module is that only a few vocal cord muscle groups dominate the excitation when the speaker pronounces, so L1 sparsity is forced by using pseudo-measurement, redundant excitation components are removed, and based on iterative convergence, only a few non-zero dominant components are finally retained; the pseudo-measurement equation is constructed and expressed as: ; ; wherein, is the pseudo-measurement matrix, and is a diagonal matrix, is the matrix element of the i-th row and i-th column; is the excitation estimation value corresponding to the i-th mode of the kth frame; sign(·) is the sign function; is the tolerance vector, which is a very small positive number; the iterative update is expressed as: ; ; ; wherein, is the pseudo-measurement optimal update weight matrix of the τth iteration; and are the uncertainty estimates of the excitation vector for the τth iteration and the τ+1th iteration, respectively, and the uncertainty before the pseudo-measurement iteration is ; is the pseudo-measurement perturbation; and are the excitation vectors of the τ+1th iteration and the τth iteration, respectively.
[0043] Embodiment seven, refer to Figure 1 , which is based on the above embodiment, the excitation unit parameter adaptive selection module for excitation unit disturbance uncertainty and pseudo measurement disturbance It is difficult to set a priori, through two L-curve inflection point adaptive selection; for excitation unit disturbance , the mean square error , expressed as: ; for pseudo measurement disturbance , the excitation sparsity and observation fitting error , respectively expressed as: ; ; for with as the abscissa, with as the ordinate, draw the L-curve; the inflection point is the optimal excitation unit disturbance uncertainty and pseudo measurement disturbance ; is the L2 norm.
[0044] Embodiment eight, refer to Figure 1 , the characterization unit correction module for speech and environmental non-stationary, by adaptive adjustment of the characterization unit disturbance uncertainty and pseudo measurement disturbance , and then realize the correction of the optimal update weight matrix of the characterization unit; expressed as: ; ; ; ; ; ; ; ; wherein, is the self-regulating adjustment coefficient; is the trace of the matrix; is the estimated value of the actual innovation uncertainty of the kth frame; is the theoretical voiceprint residual uncertainty; and are the kth and uth frame level voiceprint prediction residual, respectively; is the sample voiceprint residual uncertainty; and are the adjusted and ; and are the prediction uncertainty and optimal update weight matrix of the characterization unit after correction, respectively.
[0045] Embodiment nine, refer to Figure 1This embodiment is based on the above embodiment. The identity recognition module selects the glottal excitation timing based on the self-adjustment of the excitation unit parameters and the correction of the characterization unit. With vocal tract modal response The combination is used as a bimodal feature; an error metric is calculated with each user template in the database. Select the smallest value for identity verification, as shown below: ;like If the result is positive, the verification passes; otherwise, the verification fails. It is the error measure between the i-th user and the current speech features; It is the laryngeal source excitation timing vector of the current sentence to be identified, and the frame-level glottal excitation timing... It is obtained by piecing the beginning and end together; It is the incentive template vector stored in the database for user number 0; It is the vocal tract modal representation vector of the current sentence to be identified, which will be the frame-level vocal tract modal response. obtained by piecing together; It is the representation template vector of user o; It is the minimum error metric; T is the acceptance threshold; and It measures weight.
[0046] By performing the above operations, this solution addresses the problems of general voice identity recognition systems, such as high sensitivity to noise and environment, poor stability, and weak adaptability to environmental and speech non-stationarity, leading to poor recognition performance. It introduces pseudo-measurement constraints to eliminate redundant excitation components, retaining only a few non-zero dominant components to improve the distinguishability between different speakers. Based on the adaptive selection of excitation unit parameters, the system can adaptively balance the tracking of excitations under different speech content and environments, ensuring recognition stability. Through representation unit correction, it dynamically scales measurement noise and process noise to address speech and environmental non-stationarity, providing real-time adaptive capability to environmental changes and speech patterns, reducing performance fluctuations caused by assumption mismatch, and thus improving identity recognition performance.
[0047] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.
[0048] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.
Claims
1. A finite element identity recognition system based on voiceprint comparison, characterized in that: The system comprises a voiceprint extraction module, a vocal tract finite element modal analysis module, a pre-balancing processing module, a double-unit construction module, a pseudo-measurement constraint module, an excitation unit parameter adaptive selection module, a representation unit correction module and an identity recognition module. The voiceprint extraction module extracts features after pre-processing the speech signal; The vocal tract finite element modal analysis module constructs a vocal tract dynamics equation to obtain a state space transfer form of a frame-level vocal tract representation vector; The pre-balancing processing module constructs an observation matrix and a direct transfer matrix based on the voiceprint features and the state space transfer form; The double-unit construction module first predicts and updates a modal excitation vector in each frame of speech through an optimal recursive state estimation unit, and then predicts and updates a vocal tract representation vector through another estimation unit using the modal excitation vector; The pseudo-measurement constraint module constructs a sparse constraint equation with a sign function as a pseudo-measurement to eliminate redundant components in the excitation vector; The excitation unit parameter adaptive selection module determines the excitation unit parameters based on curve adaptation; The representation unit correction module adaptively adjusts the disturbance and the observation disturbance by adjusting the adjustment coefficient, and correspondingly corrects the prediction uncertainty and the optimal update weight matrix; The identity recognition module compares the voiceprints based on the frame-level glottal excitation timing and the vocal tract modal representation to achieve identity recognition.
2. The finite element identity recognition system based on voiceprint comparison according to claim 1, characterized in that: The voiceprint extraction module acquires a speech signal, extracts a vocal tract formant feature after pre-processing each frame of the speech signal , a spectral difference dynamic feature , and a pronunciation intensity profile ; the three types of features are regarded as responses of multiple types of sensors and are expressed as: ; ; ; wherein, is a voiceprint observation vector of the kth frame of speech; and are vocal tract formant features of the kth frame and the (k-1)th frame of speech respectively; is the total number of sampling points of each frame of speech; n is a sampling point index; is a speech amplitude of the nth sampling point in the kth frame.
3. The finite element identity recognition system based on voiceprint comparison according to claim 2, characterized in that: The vocal tract finite element modal analysis module describes the vocal tract vibration with modal coordinates, expressed as: ; wherein, is the modal amplitude of speech at time t; and are the first and second derivatives of , respectively; is the characteristic frequency vector of the speech modal; is the damping ratio vector of the speech modal; is the modal matrix of the vocal tract finite element model; T is the transpose operation; is the force load mapping operator, u is the laryngeal source excitation signal; the vocal tract representation vector of the t-th frame of speech is discretized to each frame model is defined as: ; by discretization, we get: ; ; ; ; wherein, and are the vocal tract representation vectors of the k+1-th frame and the k-th frame of speech, respectively; is the frame interval; is the discrete-time representation transition matrix; B is the discrete-time input matrix; is the modal excitation vector of the k-th frame of speech; is the process disturbance; I is the unit matrix; is the natural angular frequency of each modal of the vocal tract.
4. The finite element identity recognition system based on voiceprint comparison according to claim 3, characterized in that: The pre-conditioning processing module constructs an observation matrix C and a direct transfer matrix D, expressed as: ; ; ; wherein, is an observation disturbance vector; is a static spectrum response weight to modal displacement; is a dynamic spectrum response weight to modal displacement; is a sound production intensity profile response weight to modal excitation; the observation matrix is pre-conditioned with a measurement disturbance uncertainty matrix R; expressed as: ; ; wherein, and are the observation matrix and the direct transfer matrix after pre-conditioning, respectively.
5. The finite element identity recognition system based on voiceprint comparison according to claim 4, characterized in that: The double-unit construction module uses one optimal recursive state estimation unit to predict and update the modal excitation vector as an excitation unit and uses another optimal recursive state estimation unit to predict and update the vocal tract representation vector as a representation unit in each frame of speech; the excitation unit prediction is represented as: ; ; wherein, is the prediction of the glottal source excitation before the observation update of the kth frame of speech; is the posterior excitation estimate after the observation update of the k-1th frame of speech; is the prediction uncertainty of the prediction ; is the posterior uncertainty of the excitation estimate after the update of the k-1th frame of speech; is the excitation unit disturbance uncertainty; the excitation unit update is represented as: ; ; ; wherein, is the optimal update weight matrix of the excitation unit of the kth frame of speech; is the posterior excitation estimate of the kth frame of speech; is the residual term, the difference between the observation and the prediction; is the representation posterior estimate of the k-1th frame of speech; is the posterior uncertainty of the excitation; the representation unit prediction is represented as: ; ; wherein, is the representation prior estimate before the measurement update of the kth frame of speech; is the uncertainty matrix of the representation prediction of the kth frame of speech; is the representation unit disturbance uncertainty; the representation unit update is represented as: ; ; ; ; wherein, is the optimal update weight matrix of the representation unit; is the representation posterior estimate of the kth frame; is the voiceprint observation vector after the pre-equalization processing.
6. The finite element identity recognition system based on voiceprint comparison according to claim 5, characterized in that: The pseudo-measurement constraint module is to construct a pseudo-measurement equation, denoted as: ; ; wherein, is a pseudo-measurement matrix, is a diagonal matrix, is the matrix element of the i-th row and i-th column; is the excitation estimation value corresponding to the i-th modality of the k-th frame; sign(·) is a sign function; is a tolerance vector, a very small positive number; iterative update, denoted as: ; ; ; wherein, is the pseudo-measurement optimal update weight matrix of the τ-th iteration; and are the uncertainty estimates of the excitation vector for the τ-th iteration and the τ+1-th iteration, respectively, and the uncertainty before the pseudo-measurement iteration is ; is a pseudo-measurement disturbance; and are the excitation vectors of the τ+1-th iteration and the τ-th iteration, respectively.
7. The finite element identity recognition system based on voiceprint comparison according to claim 6, characterized in that: The excitation unit parameter adaptive selection module is used for selecting the perturbation of the excitation unit , the mean square error is calculated , and is expressed as ; the pseudo-measurement perturbation is calculated , the excitation sparsity is calculated , and the observation fitting error is calculated , which are respectively expressed as ; ; the inflection point is the perturbation uncertainty of the optimal excitation unit and the pseudo-measurement perturbation ; is the L2 norm.
8. The finite element identity recognition system based on voiceprint comparison according to claim 7, characterized in that: The characterization unit correction module is non-stationary for speech and environment, and the characterization unit disturbance uncertainty is adaptively adjusted and pseudo-measurement disturbance , thereby realizing correction of the optimal update weight matrix of the characterization unit; expressed as: ; ; ; ; ; ; ; ; wherein, is a self-adjusting adjustment coefficient; is a trace of a matrix; is an estimated value of the kth frame actual innovation uncertainty; is a theoretical voiceprint residual uncertainty; and are the kth and uth frame level voiceprint prediction residuals, respectively; is a sample voiceprint residual uncertainty; and are adjusted and ; and are the prediction uncertainty and the optimal update weight matrix of the characterization unit after correction, respectively.
9. The finite element identity recognition system based on voiceprint comparison according to claim 8, characterized in that: The identity recognition module will be based on the excitation unit parameter self-adjusting selection and characterization unit correction after the glottal excitation timing and vocal tract modal response combined as a dual modal feature; calculate the error measure with each user template in the database , select the minimum one, perform identity verification, denoted as: ; if , the verification is passed, otherwise the verification fails; wherein, is the error measure of the i-th user and the current speech feature; is the laryngeal source excitation timing vector of the current sentence to be recognized, and the frame-level glottal excitation timing is spliced at the beginning and end to obtain; is the excitation template vector stored in the database of the o-th user; is the vocal tract modal representation vector of the current sentence to be recognized, and the frame-level vocal tract modal response is spliced to obtain; is the representation template vector of the o-th user; is the minimum error measure; T is the acceptance threshold; and are the measure weights.
Citation Information
Patent Citations
Determining method of numerical value of vibration noise of axle housing of drive axle
CN107391816A
Multi-modal interactive structural mechanical analysis method and system, computer equipment and medium
CN115756161A