Voiceprint recognition method and device under encircled signal domain and human-computer interaction device
By collecting voice and physiological signals through a wraparound wearable device, and combining dynamic beamforming and generative adversarial networks, the interference problem of voiceprint recognition in complex environments is solved, and high-precision user authentication and physiological anomaly detection are achieved.
Patent Information
- Application Number
- CN202510720835.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Existing voiceprint recognition technology is susceptible to interference in complex environments and cannot distinguish between recorded spoofing. Furthermore, the physiological monitoring function of wearable devices is independent of the identity authentication system, resulting in low data utilization and poor dynamic adaptability.
By collecting voice and physiological signals through a wraparound wearable device, and combining dynamic beamforming and generative adversarial networks, deep binding of voiceprint features and physiological parameters is achieved, generating a feature map containing user identity and physiological state parameters for event judgment.
It effectively eliminates environmental noise and equipment coupling vibration interference in motion scenarios, improves voiceprint recognition accuracy, and enables real-time warning of user identity verification and physiological abnormalities.
Smart Images

Figure CN120727017B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wearable device technology, and in particular to a voiceprint recognition method, device, and human-computer interaction device in an enveloping signal domain. Background Technology
[0002] Existing voiceprint recognition technologies typically rely on a single voice signal, making them susceptible to interference in complex environmental noise or when the user is in motion, and they cannot distinguish between security issues such as recording spoofing. Furthermore, the physiological monitoring functions and identity authentication systems of traditional wearable devices are independent, resulting in low data utilization and poor dynamic adaptability. For example, while smart bracelets can collect heart rate and activity data, they cannot be fused with voice features to achieve dynamic identity verification; and existing noise suppression methods do not consider the coupling noise generated by vibrations from device contact with skin. Summary of the Invention
[0003] The main objective of this invention is to provide a voiceprint recognition method, device, and human-computer interaction equipment in an enveloping signal domain, which can improve the accuracy of voiceprint recognition in complex environments while achieving real-time early warning of user identity verification and physiological abnormalities such as respiratory disorders.
[0004] To achieve the above objectives, the present invention provides a voiceprint recognition method in the surround signal domain, comprising the following steps:
[0005] Collect the target user's voice signal and simultaneously acquire physiological signals including at least wrist movement acceleration, skin surface vibration frequency, and blood oxygen saturation;
[0006] The speech signal is subjected to environmental noise separation processing to extract the voiceprint feature vectors of fundamental frequency jitter rate and formant migration trajectory contained in the speech signal;
[0007] The physiological signal is transformed by time and frequency and then fused with the voiceprint feature vector to generate a feature map containing user identity and physiological state parameters.
[0008] Event determination is performed on the speech feature segments and physiological feature segments in the feature map.
[0009] Furthermore, the steps for collecting the target user's voice signal include:
[0010] Multi-angle voice signals are acquired through a ring-shaped array of miniature microphones in a wraparound wearable device, and the directional beam of the microphone array dynamically adjusts the focusing area with the acceleration of wrist movement.
[0011] Furthermore, the step of simultaneously acquiring physiological signals including at least wrist motion acceleration, skin surface vibration frequency, and blood oxygen saturation includes:
[0012] The vibration frequency of the skin surface is obtained by using the piezoelectric sensor array embedded in the wraparound wearable device to capture the vibration frequency band distribution information of the skin surface in real time;
[0013] The wrist motion acceleration is obtained by dynamically correcting the time-frequency aliasing noise caused by limb displacement in the vibration frequency of the skin surface by the output value of the motion accelerometer.
[0014] Meanwhile, a dual-wavelength optical module is used to sample blood oxygen saturation.
[0015] Further, the step of performing environmental noise separation processing on the speech signal and extracting the voiceprint feature vectors of the fundamental frequency jitter rate and formant migration trajectory contained in the speech signal includes:
[0016] The sound field coverage of the microphone array is adjusted by using a dynamic beamforming algorithm combined with wrist motion acceleration parameters to suppress ambient noise components from directions other than the wearer.
[0017] The low-frequency modulation signal corresponding to the breathing rhythm in the speech signal is separated by a preset decomposition algorithm, and the fundamental frequency jitter rate of the fundamental frequency trajectory is extracted.
[0018] Capture changes in vocal tract morphology and combine them with the fluctuation trend of blood oxygen saturation to generate the physiologically relevant resonance peak migration trajectory;
[0019] The fundamental frequency jitter rate and the formant migration trajectory are vectorized to obtain the voiceprint feature vector.
[0020] Further, the step of fusing the physiological signal with the voiceprint feature vector after time-frequency transformation to generate a feature map containing user identification and physiological state parameters includes:
[0021] The physiological signals are processed by a two-stream convolutional network. The first-stream network extracts speech physiological features that are coupled with the speech signal, and the second-stream network extracts independent physiological features that are not related to speech. The speech physiological features are dynamically associated with the speaker feature vector in the frequency domain through a cross-modal attention mechanism.
[0022] A weight matrix is constructed based on blood oxygen saturation. The update period of the weight matrix is automatically adjusted according to the user's current wrist movement acceleration. When the movement acceleration exceeds a set threshold, the matrix update rate is reduced. The weight matrix is used to map and fuse the time-varying parameters of the formant migration trajectory, the frequency domain energy distribution of the skin vibration frequency, and the voiceprint feature vector.
[0023] The voiceprint feature vector is input into a generative adversarial network (GAN). The generator of the GAN dynamically enhances the identity-sensitive components based on physiological state parameters. The discriminator simultaneously detects physiological abnormalities and identity forgery features. By alternately optimizing the identity discrimination loss function and the physiological abnormality detection loss function, an encoding map is generated. The encoding basis of the encoding map is uniquely determined by the user's identity identifier. The mutation regions in the encoding map reflect the spatiotemporal distribution information of real-time physiological abnormalities. Furthermore, the dimension of the encoding map dynamically increases the number of feature channels based on the continuous wearing time of the device.
[0024] Furthermore, the step of performing event determination on the speech feature segments and physiological feature segments in the feature map includes:
[0025] The morphological abrupt change boundary of the physiological feature segment is detected and time-series matched with the fluctuation range of the fundamental frequency jitter rate in the voiceprint feature segment;
[0026] When the deviation between the resonance peak migration trajectory and the preset identity template exceeds the dynamic adjustment threshold, a secondary verification process is triggered in conjunction with the rate of change of blood oxygen saturation. The dynamic adjustment threshold is automatically updated based on the user's historical physiological parameter baseline.
[0027] Based on the correlation between skin vibration frequency energy distribution and wrist movement acceleration, the interference of voiceprint feature artifacts caused by limb movements is eliminated;
[0028] The vocal cord vibration spectrum characteristics are coupled with upper respiratory tract mechanical parameters for analysis, and the output results include the determination of event type and abnormal source path.
[0029] Furthermore, a weight matrix is constructed based on blood oxygen saturation. The update period of the weight matrix is automatically adjusted according to the user's current wrist motion acceleration. When the motion acceleration exceeds a set threshold, the matrix update rate is reduced. The step of mapping and fusing the time-varying parameters of the formant migration trajectory, the frequency domain energy distribution of skin vibration frequency, and the voiceprint feature vector through the weight matrix includes:
[0030] A weight matrix is constructed based on the trend of blood oxygen saturation change. The axial dimensions of the weight matrix correspond to the time-varying parameters of the resonance peak migration trajectory, the frequency domain energy distribution of the skin vibration frequency, and the amplitude change rate of the voiceprint feature vector, respectively.
[0031] The matrix update cycle is dynamically adjusted according to the amplitude range of the user's current wrist movement acceleration. When the acceleration exceeds the preset motion threshold, the matrix update rate is reduced to 1 / 3 to 1 / 2 of the original rate.
[0032] The instantaneous slope of the resonance peak migration trajectory is superimposed with the dominant frequency of the skin vibration frequency, and then the superposition result is convolved with the phase information of the voiceprint feature vector to generate the voiceprint feature vector with physiological state constraints.
[0033] Furthermore, the steps for constructing a weight matrix based on the trend of blood oxygen saturation changes include:
[0034] The basic weighting coefficient is set according to the direction of the average change of blood oxygen saturation within a preset time. An upward trend corresponds to increasing the fusion weight of skin vibration frequency, and a downward trend corresponds to increasing the discrimination weight of resonance peak trajectory.
[0035] The real-time value of wrist motion acceleration is linearly mapped to the weight update frequency. For every 1 m / s² increase in acceleration, the matrix update interval is extended by 0.5 seconds.
[0036] When the energy decay rate of the blood oxygen saturation and the skin vibration frequency fluctuates in opposite directions, a fixed ratio of attenuation is applied to the voiceprint feature vector.
[0037] The present invention proposes a voiceprint recognition device in an enveloping signal domain, comprising:
[0038] The acquisition unit is used to acquire the voice signal of the target user and simultaneously acquire physiological signals including at least wrist movement acceleration, skin surface vibration frequency and blood oxygen saturation.
[0039] The vector unit is used to perform environmental noise separation processing on the speech signal and extract the voiceprint feature vectors of the fundamental frequency jitter rate and formant migration trajectory contained in the speech signal.
[0040] The map unit is used to perform feature fusion with the voiceprint feature vector after the physiological signal is transformed by time and frequency to generate a feature map containing user identity and physiological state parameters.
[0041] The recognition unit is used to perform event determination on the speech feature segments and physiological feature segments in the feature map.
[0042] The present invention also provides a human-computer interaction device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described voiceprint recognition method in the surround signal domain.
[0043] The voiceprint recognition method, apparatus, and human-computer interaction device in the surround signal domain provided by this invention have the following beneficial effects:
[0044] (1) By using dynamic beamforming of the surround microphone array and adaptive filtering of skin vibration frequency, environmental noise and equipment coupling vibration interference are effectively eliminated, so that voiceprint feature extraction can still maintain high accuracy in motion scenarios.
[0045] (2) Using generative adversarial networks to construct a spatiotemporal joint coding map, deeply binding voiceprint features with physiological parameters such as blood oxygen and skin vibration;
[0046] (3) Based on the weight matrix of motion acceleration and the self-expanding feature map of wearing time, the device automatically optimizes the allocation of computing resources in different states such as stationary and motion. Attached Figure Description
[0047] Figure 1 This is a flowchart illustrating a voiceprint recognition method in an enveloping signal domain according to an embodiment of the present invention.
[0048] Figure 2 This is a structural block diagram of a voiceprint recognition device in an enveloping signal domain according to an embodiment of the present invention;
[0049] Figure 3 This is a schematic block diagram of the structure of a human-computer interaction device according to an embodiment of the present invention.
[0050] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0052] Reference Figure 1 This is a flowchart illustrating a voiceprint recognition method in the surround signal domain proposed in this invention. The method includes the following steps:
[0053] S1, collect the voice signal of the target user, and simultaneously acquire physiological signals including at least wrist movement acceleration, skin surface vibration frequency and blood oxygen saturation;
[0054] Step S1 includes the step of acquiring the target user's voice signal, which includes:
[0055] Multi-angle voice signals are acquired through a ring-shaped array of miniature microphones in a wraparound wearable device, and the directional beam of the microphone array dynamically adjusts the focusing area with the acceleration of wrist movement.
[0056] And, including the step of simultaneously acquiring physiological signals including at least wrist motion acceleration, skin surface vibration frequency, and blood oxygen saturation, the step including:
[0057] The vibration frequency of the skin surface is obtained by using the piezoelectric sensor array embedded in the wraparound wearable device to capture the vibration frequency band distribution information of the skin surface in real time;
[0058] The wrist motion acceleration is obtained by dynamically correcting the time-frequency aliasing noise caused by limb displacement in the vibration frequency of the skin surface by the output value of the motion accelerometer.
[0059] Meanwhile, a dual-wavelength optical module is used to sample blood oxygen saturation.
[0060] In the embodiment of step S1, the human-computer interaction device (specifically a smart bracelet in this embodiment) integrates a ring-shaped array of six miniature microphones on the inner side of the wristband, with an 8mm spacing between adjacent microphones, forming a 360° voice capture ring. When worn by the user, the directional beam of the microphone array is adjusted based on the real-time X / Y / Z axis acceleration values fed back by the wrist accelerometer, using the following formula:
[0061]
[0062] Among them, a x a y a z For the three-axis acceleration components, when the acceleration amplitude When the speed exceeds 2m / s², the beamwidth is automatically reduced to 60° to suppress ambient noise in the direction of movement (such as wind noise when running).
[0063] The skin surface vibration frequency is collected by arranging a 4×4 piezoelectric sensor array (specifically, a TDK PVDF thin-film sensor with a sampling rate of 1kHz) on the wristband contact surface, covering the pulsating area. In the process of correcting time-frequency aliasing noise, when the accelerometer detects wrist displacement (such as a waving motion), the skin surface vibration frequency is used to extract the frequency band signal and establish the acceleration amplitude. Vibration frequency of skin surface Linear regression model: The corrected wrist motion accelerations ax, ay, and az are obtained. For blood oxygen saturation sampling using a dual-wavelength optical module, red light AC (660nm) and infrared light DC (880nm) are emitted alternately at 200Hz. The reflected light intensity is received via a photodiode, and the ratio is calculated. According to the pre-calibrated curve S p O2 = 110−25R (obtain blood oxygen level S) p O2.
[0064] S2, perform environmental noise separation processing on the speech signal, and extract the voiceprint feature vectors of the fundamental frequency jitter rate and formant migration trajectory contained in the speech signal;
[0065] In step S2, a dynamic beamforming algorithm is used in conjunction with wrist motion acceleration parameters to adjust the sound field coverage of the microphone array and suppress environmental noise components in the direction of the wearer; a preset decomposition algorithm is used to separate the low-frequency modulation signal corresponding to the breathing rhythm in the speech signal and extract the fundamental frequency jitter rate of the fundamental frequency trajectory; changes in vocal tract morphology are captured and combined with the fluctuation trend of blood oxygen saturation to generate the physiologically relevant formant migration trajectory; the fundamental frequency jitter rate and the formant migration trajectory are vectorized to obtain the voiceprint feature vector.
[0066] In the embodiment of step S2, wrist triaxial acceleration data is input. Calculate the magnitude of the composite acceleration vector. Then, based on the magnitude of the acceleration vector... The beamwidth of the microphone array is dynamically adjusted, and the dynamic adjustment range of the beamwidth is as follows:
[0067]
[0068] Finally, a method is adopted to suppress out-of-beam noise by using a generalized sidelobe canceller (GSC) algorithm, with the main lobe always pointing towards the estimated position of the oral cavity, thereby suppressing environmental noise components in the direction of non-wearers.
[0069] In the process of separating the low-frequency modulation signal corresponding to the breathing rhythm in the speech signal by the decomposition algorithm, an improved variational mode decomposition (VMD) is performed on the denoised speech signal. With the mode number K=3 and penalty factor α=2000, the 50-300Hz fundamental frequency band, the breathing modulation band (0.2-2Hz), and the residual noise band are separated. The envelope signal of the breathing modulation band is extracted, and the peak offset of its cross-correlation function with the fundamental frequency band signal is calculated, defined as the fundamental frequency jitter rate J. F0 :
[0070]
[0071] in, The position of the cross-correlation peak in the i-th frame. This is the historical average.
[0072] In the process of capturing changes in vocal tract morphology and generating physiologically relevant formant migration trajectories by combining them with the fluctuation trends of blood oxygen saturation, time-varying linear predictive coding (TV-LPC) is used to analyze the speech signal. The formant parameters (F1-F3) are updated every 10ms. Then, a correlation is established between blood oxygen saturation (SpO2) and the formant migration trajectory. Association model:
[0073]
[0074] Where k=1,2,3 correspond to three resonance peaks, and α=0.15 and β=0.08 are physiological coupling coefficients obtained through user calibration. This is the resonance peak segment. For the resonance peak migration trajectory, The resonant peak segment is related to the fluctuation trend, d is the distance of the captured channel, and t is the time of the captured channel.
[0075] In the process of vectorizing the fundamental frequency jitter rate and formant migration trajectory to obtain the voiceprint feature vector, the fundamental frequency jitter rate J is... F0 Z-score standardization was performed, and the mean, variance, and skewness were calculated for each time window to determine the formant migration trajectory. PCA transformation is used to reduce the dimensionality to 10-dimensional principal components, and the two are concatenated to generate a 28-dimensional voiceprint feature vector.
[0076]
[0077] S3, after time-frequency transformation, the physiological signal is fused with the voiceprint feature vector to generate a feature map containing user identity and physiological state parameters.
[0078] In the embodiment of step S3, a two-stream convolutional network is used to process the physiological signals. The first-stream network extracts speech physiological features coupled with the speech signal, and the second-stream network extracts independent physiological features unrelated to speech. The speech physiological features are dynamically correlated with the voiceprint feature vector in the frequency domain through a cross-modal attention mechanism. A weight matrix is constructed based on blood oxygen saturation. The update period of the weight matrix is automatically adjusted according to the user's current wrist motion acceleration. When the motion acceleration exceeds a set threshold, the matrix update rate is reduced. The weight matrix is used to analyze the time-varying parameters of the formant migration trajectory and the skin vibration frequency. Frequency domain energy distribution and voiceprint feature vectors are mapped and fused; the voiceprint feature vectors are input into a generative adversarial network (GAN), wherein the generator of the GAN dynamically enhances the identity-sensitive components based on physiological state parameters, the discriminator simultaneously detects physiological abnormal events and identity forgery features, and an encoding map is generated by alternately optimizing the identity discrimination loss function and the physiological abnormality detection loss function. The encoding basis of the encoding map is uniquely determined by the user's identity identifier, the mutation regions in the encoding map reflect the spatiotemporal distribution information of real-time physiological abnormal events, and the dimension of the encoding map dynamically increases the number of feature channels according to the continuous wearing time of the device.
[0079] Specifically, a weight matrix is constructed based on the trend of blood oxygen saturation changes. The axial dimensions of the weight matrix correspond to the time-varying parameters of the formant migration trajectory, the frequency domain energy distribution of the skin vibration frequency, and the amplitude change rate of the voiceprint feature vector, respectively. The matrix update cycle is dynamically adjusted according to the amplitude range of the user's current wrist movement acceleration. When the acceleration exceeds a preset motion threshold, the matrix update rate is reduced to 1 / 3 to 1 / 2 of the original rate. The instantaneous slope of the formant migration trajectory is superimposed with the dominant frequency of the skin vibration frequency, and the superposition result is convolved with the phase information of the voiceprint feature vector to generate the voiceprint feature vector with physiological state constraints.
[0080] The above-mentioned basic weight coefficient is set according to the average change direction of blood oxygen saturation within a preset time. An upward trend corresponds to an increase in the fusion weight of skin vibration frequency, and a downward trend corresponds to an increase in the discrimination weight of resonance peak trajectory. The real-time value of wrist movement acceleration is linearly mapped to the weight update frequency. For every 1 m / s² increase in acceleration, the matrix update interval is extended by 0.5 seconds. When the energy decay rate of blood oxygen saturation and skin vibration frequency fluctuates in opposite directions, a fixed proportion of attenuation is applied to the voiceprint feature vector.
[0081] In the specific implementation of step S3, the dual-stream convolutional network consists of the following branches: a speech-related physiological feature stream and an independent physiological feature stream. The speech-related physiological feature stream is the input time-frequency transformed physiological signal (such as the skin vibration frequency STFT spectrum and blood oxygen time-series waveform). A 3-layer causal convolution (kernel size [3,5,3], number of channels [16,32,64]) is used to extract coupled features related to the speech fundamental frequency jitter. These features are then processed through a cross-modal attention layer and the speaker signature feature vector. Interaction:
[0082]
[0083] Where Q represents physiological characteristics, K and V represent voiceprint characteristics, and d k =64 is the scaling factor. Independent physiological feature streams, on the other hand, take the same physiological signal as input, but the convolutional layer introduces channel masks (such as masking speech-related frequency bands), and the output includes physiological indicators unrelated to identity, such as heart rate variability and skin impedance.
[0084] In the process of constructing and fusing the dynamic weight matrix described above, the weights driven by blood oxygenation are first initialized, and a three-dimensional weight matrix is defined. The T-axis represents the instantaneous slope of the corresponding resonance peak migration trajectory (sampled every 10ms), the F-axis represents the 1 / 3 octave band energy of the corresponding skin vibration frequency (center frequencies of 50Hz, 160Hz, and 500Hz), and the c-axis represents the amplitude change rate of the corresponding voiceprint feature vector (calculated separately for each of the 28 features). Regarding the response to blood oxygen trends, when the 5-second moving average of blood oxygen saturation increases by more than 1%, the F-axis weight increases by 1.8 times; when blood oxygen decreases continuously for 3 seconds, the T-axis weight coefficient is adjusted accordingly. Dynamic enhancement.
[0085] The update mechanism for motion acceleration constraints is the update cycle. With acceleration Mapping relationship:
[0086]
[0087] Then, the fusion process is performed, first by weighting and superimposing the instantaneous slope S(t) of the resonance peak with the skin's dominant frequency energy E(f): The first voiceprint feature was obtained. Then, the superposition result is convolved with the voiceprint phase information ϕ in the temporal domain to obtain the second voiceprint feature. .
[0088] The aforementioned Generative Adversarial Networks (GANs) include a generator and a discriminator, wherein...
[0089] The generator is the one that enhances the second voiceprint feature. Real-time physiological state parameters are used as input to the adversarial network. The structure of the adversarial network generated based on the input is a 5-layer fully connected network (hidden layer dimensions [256,128,64,32,16]), and the output is an identity-sensitive component. .
[0090] The discriminator creates a loss function through an identity discrimination branch (binary classification) and a physiological abnormality detection branch (multi-class classification), and then generates a coding map. The specific process of generating the coding map is as follows:
[0091] (1) Identity-sensitive components A 16-dimensional identity basis vector is generated through orthogonalization.
[0092] (2) The 128-dimensional mutation vector triggered by the abnormal event is superimposed on the base to obtain the abnormal event coding map.
[0093] S4, perform event judgment on the speech feature segments and physiological feature segments in the feature map.
[0094] In the embodiment of step S4, the morphological abrupt change boundary of the physiological feature segment is detected and time-series matched with the fluctuation range of the fundamental frequency jitter rate in the voiceprint feature segment; when the deviation between the formant migration trajectory and the preset identity template exceeds the dynamic adjustment threshold, a secondary verification process is triggered in conjunction with the blood oxygen saturation change rate, and the dynamic adjustment threshold is automatically updated according to the user's historical physiological parameter baseline; based on the correlation between skin vibration frequency energy distribution and wrist movement acceleration, voiceprint feature artifact interference caused by limb movements is eliminated; the vocal cord vibration spectrum features are coupled and analyzed with upper respiratory tract mechanical parameters to output a judgment result including event type and abnormal source tracing path.
[0095] In a specific embodiment of step S4, the process of morphological abrupt change boundary detection involves performing the Canny edge detection algorithm on the physiological feature segments (skin vibration energy, blood oxygen change gradient) to locate energy jump points, and then defining abrupt change boundary conditions such that the energy difference within adjacent 50ms windows exceeds three times the standard deviation of the historical baseline. During the time-series matching process with the fluctuation range of the fundamental frequency jitter rate in the voiceprint feature segments, the fundamental frequency jitter rate J in the voiceprint feature segments is extracted. F0 The fluctuation range (i.e., maximum value - minimum value) is used to align the time sequence deviation between the physiological mutation point and the fundamental frequency fluctuation peak using the Dynamic Time Warping (DTW) algorithm. When the alignment error is less than 50ms, it is marked as a valid associated event. The Canny edge detection algorithm and the Dynamic Time Warping (DTW) algorithm adopt existing technologies, which will not be described in detail in this embodiment.
[0096] When the deviation between the resonance peak migration trajectory and the preset identity template exceeds the dynamic adjustment threshold, in conjunction with the process of triggering the secondary verification process based on the rate of change of blood oxygen saturation, the preset identity template includes the mean value and fluctuation range of the resonance peak trajectory when the user registered; the cosine similarity between the current resonance peak migration trajectory and the template is calculated in real time, and the threshold is dynamically adjusted. When the cosine similarity is less than the dynamic adjustment threshold and the rate of decrease of blood oxygen is greater than 2% / s, secondary verification is triggered.
[0097] The process described above for eliminating voiceprint artifacts caused by limb movements involves establishing vibration energy E. v With acceleration a sum Compensation model:
[0098]
[0099] In the formula, It is the skin vibration frequency, when a sum >2m / s2 S At that time, a frequency domain mask is applied to the voiceprint feature vector (suppressing high-frequency components in the 200-500Hz range).
[0100] In the above-mentioned process of coupling the vocal cord vibration spectrum characteristics with upper respiratory tract mechanical parameters to output a judgment result including event type and abnormal source path, the parameters include vocal cord mass m=0.15g, stiffness k=1.2N / m, and upper respiratory tract airflow disturbance parameters are determined by blood oxygen saturation S. p O2 and skin vibration frequency f v Derivation:
[0101]
[0102] In the formula, These are parameters related to airflow disturbance in the upper respiratory tract.
[0103] Reference Appendix Figure 2 This is a structural block diagram of a voiceprint recognition device in an enveloping signal domain proposed in this invention. The device includes:
[0104] The acquisition unit is used to acquire the voice signal of the target user and simultaneously acquire physiological signals including at least wrist movement acceleration, skin surface vibration frequency and blood oxygen saturation.
[0105] The vector unit is used to perform environmental noise separation processing on the speech signal and extract the voiceprint feature vectors of the fundamental frequency jitter rate and formant migration trajectory contained in the speech signal.
[0106] The map unit is used to perform feature fusion with the voiceprint feature vector after the physiological signal is transformed by time and frequency to generate a feature map containing user identity and physiological state parameters.
[0107] The recognition unit is used to determine events from the speech feature segments and physiological feature segments in the feature map.
[0108] Reference Figure 3 This invention also provides a human-computer interaction device, which can be a server, and its internal structure can be as follows: Figure 3 As shown, the human-computer interaction device includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor, designed in this computer configuration, provides computing and control capabilities. The memory of the human-computer interaction device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the human-computer interaction device stores the data corresponding to this embodiment. The network interface of the human-computer interaction device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.
[0109] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the human-computer interaction device to which the present invention is applied.
[0110] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for voiceprint recognition under a ring-type signal domain, characterized in that, The method comprises the following steps: Collecting voice signals of a target user through a micro microphone array arranged in a ring shape in a ring-wearing device, and synchronously acquiring physiological signals including at least wrist movement acceleration, skin surface vibration frequency and blood oxygen saturation; Performing ambient noise separation processing on the voice signals to extract a voiceprint feature vector including a fundamental frequency jitter rate and a formant migration trajectory in the voice signals; using a dynamic beam forming algorithm to adjust the sound field coverage range of the microphone array in combination with the wrist movement acceleration parameter to suppress ambient noise components in the direction of non-wearers; The low-frequency modulation signal corresponding to the breathing rhythm in the speech signal is separated by a preset decomposition algorithm, and the fundamental frequency jitter rate of the fundamental frequency track is extracted, wherein in the process of separating the low-frequency modulation signal corresponding to the breathing rhythm in the speech signal by the decomposition algorithm, improved variational mode decomposition is performed on the noise-reduced speech signal, the mode number K is set to 3, the penalty factor α is set to 2000, the 50-300Hz fundamental frequency band, the 0.2-2Hz breathing modulation band and the residual noise band are separated, the envelope signal of the breathing modulation band is extracted, the peak value offset of the cross-correlation function of the envelope signal and the fundamental frequency band signal is calculated, and the fundamental frequency jitter rate J is defined as F0 ; wherein, is the i-th frame cross-correlation peak position, is the history mean value; Capture the vocal tract morphology changes, combined with the fluctuation trend of the blood oxygen saturation to generate the resonance peak migration trajectory with physiological correlation, wherein in the process of capturing the vocal tract morphology changes, combined with the fluctuation trend of the blood oxygen saturation to generate the resonance peak migration trajectory with physiological correlation, time-varying linear prediction coding is used to analyze the voice signal, the resonance peak parameters F1-F3 are updated every 10ms, and then the correlation model of the blood oxygen saturation SpO2 and the resonance peak migration trajectory is established: ; wherein k=1, 2, 3 correspond to three resonance peaks, and alpha=0.15 and beta=0.08 are physiological coupling coefficients obtained through user calibration, is a resonance peak segment, is a resonance peak migration track, is a resonance peak segment related to fluctuation trend, d is a distance of a capture sound channel, and t is a time of the capture sound channel. Vectorizing the fundamental frequency jitter rate and the formant migration trajectory to obtain the voiceprint feature vector; Performing feature fusion on the voiceprint feature vector after time-frequency transformation of the physiological signals to generate a feature map including a user identity and physiological state parameters; Performing event judgment on the voice feature segments and physiological feature segments in the feature map. 2.The voiceprint recognition method under the ring-type signal domain according to claim 1, wherein, The step of collecting voice signals of a target user comprises: Acquiring multi-angle voice signals through a micro microphone array arranged in a ring shape in a ring-wearing device, and dynamically adjusting the focusing area of the directional beam of the microphone array according to wrist movement acceleration. 3.The voiceprint recognition method under the ring-type signal domain according to claim 1, wherein, The step of synchronously acquiring physiological signals including at least wrist movement acceleration, skin surface vibration frequency and blood oxygen saturation comprises: Real-time capturing of skin surface vibration frequency band distribution information using a piezoelectric sensor array embedded in the ring-wearing device to obtain the skin surface vibration frequency; Dynamically correcting the time-frequency aliasing noise in the skin surface vibration frequency caused by limb displacement using the output value of a motion accelerometer to obtain the wrist movement acceleration; Meanwhile, using a dual-wavelength optical module to sample blood oxygen saturation. 4.The voiceprint recognition method under the ring-type signal domain according to claim 1, wherein, The step of performing feature fusion on the voiceprint feature vector after time-frequency transformation of the physiological signals to generate a feature map including a user identity and physiological state parameters comprises: Using a double-flow convolution network to process the physiological signals respectively, a first flow network extracts voice-physiological features having a coupling relationship with voice signals, and a second flow network extracts independent physiological features irrelevant to voice, wherein the voice-physiological features and the voiceprint feature vector establish dynamic correlation weights in the frequency domain through a cross-modal attention mechanism; Constructing a weight matrix based on blood oxygen saturation, the update period of the weight matrix is automatically adjusted according to the current wrist movement acceleration of the user, and the matrix update rate is reduced when the movement acceleration exceeds a set threshold, and the time-varying parameters of the formant migration trajectory, the frequency energy distribution of the skin vibration frequency and the voiceprint feature vector are mapped and fused through the weight matrix; Input the voiceprint feature vector into a generative adversarial network, wherein a generator of the generative adversarial network dynamically enhances the identity-sensitive component according to the physiological state parameter, a discriminator synchronously detects physiological abnormal event and identity forgery feature, generates an encoding atlas by alternately optimizing an identity discrimination loss function and a physiological abnormality detection loss function, an encoding base of the encoding atlas is uniquely determined by a user identity, and a mutation area in the encoding atlas reflects the spatiotemporal distribution information of real-time physiological abnormal events, and the dimension of the encoding atlas dynamically increases the number of feature channels according to the duration of continuous wearing of the device. 5.The voiceprint recognition method under the ring-type signal domain according to claim 4, characterized in that, The steps of event judgment on the voice feature segment and the physiological feature segment in the feature atlas include: detecting the morphological mutation boundary of the physiological feature segment and performing time sequence matching with the fluctuation interval of the jitter rate in the voiceprint feature segment; when the deviation of the formant migration trajectory from the preset identity template exceeds a dynamically adjusted threshold, triggering a secondary verification process in combination with the change rate of blood oxygen saturation, and the dynamically adjusted threshold is automatically updated according to the historical physiological parameter baseline of the user; based on the correlation between the skin vibration frequency energy distribution and the wrist motion acceleration, eliminating the voiceprint feature artifact interference caused by limb movement; coupling analysis of the vocal cord vibration spectrum feature and the upper respiratory tract mechanical parameter, outputting the judgment result containing the event type and the abnormal source path. 6.The voiceprint recognition method under the ring-type signal domain according to claim 4, characterized in that, constructing a weight matrix based on blood oxygen saturation, the update period of the weight matrix is automatically adjusted according to the current wrist motion acceleration of the user, and when the motion acceleration exceeds a set threshold, the matrix update rate is reduced, the steps of mapping and fusing the time-varying parameters of the formant migration trajectory, the frequency energy distribution of the skin vibration frequency, and the voiceprint feature vector through the weight matrix include: constructing a weight matrix based on the trend of blood oxygen saturation, the axial dimension of the weight matrix corresponds to the time-varying parameters of the formant migration trajectory, the frequency energy distribution of the skin vibration frequency, and the amplitude change rate of the voiceprint feature vector, respectively; dynamically adjusting the matrix update period according to the amplitude interval of the current wrist motion acceleration of the user, and when the acceleration exceeds a preset motion threshold, reducing the matrix update rate to 1 / 3 to 1 / 2 of the original rate; superimposing the instantaneous slope of the formant migration trajectory and the main frequency of the skin vibration frequency, and then performing convolution operation on the superimposed result and the phase information of the voiceprint feature vector to generate the voiceprint feature vector with physiological state constraint.
7. The voiceprint recognition method under the ring-type signal domain according to claim 6, characterized in that, The steps of constructing a weight matrix based on the trend of blood oxygen saturation include: setting a basic weight coefficient according to the mean change direction of the blood oxygen saturation within a preset time, increasing the fusion weight of the skin vibration frequency for the upward trend, and increasing the discrimination weight of the formant trajectory for the downward trend; linearly mapping the real-time value of the wrist motion acceleration to the weight update frequency, and extending the matrix update interval by 0.5 seconds for every 1 m / s² increase in acceleration; when the energy attenuation rates of the blood oxygen saturation and the skin vibration frequency present reverse fluctuations, applying a fixed proportional attenuation to the voiceprint feature vector.
8. A device for voiceprint recognition under a ring-type signal domain, characterized in that, The acquisition unit is configured to acquire a voice signal of a target user through a micro microphone array arranged in a ring shape in the ring-wearing device, and synchronously acquire physiological signals including at least wrist motion acceleration, skin surface vibration frequency and blood oxygen saturation; The vector unit is configured to perform ambient noise separation processing on the voice signal, extract a voiceprint feature vector including a pitch jitter rate and a formant shift trajectory in the voice signal, and adjust a sound field coverage range of the microphone array by using a dynamic beam forming algorithm combined with a wrist motion acceleration parameter to suppress ambient noise components in a non-wearer direction. The low-frequency modulation signal corresponding to the breathing rhythm in the speech signal is separated by a preset decomposition algorithm, and the fundamental frequency jitter rate of the fundamental frequency track is extracted, wherein in the process of separating the low-frequency modulation signal corresponding to the breathing rhythm in the speech signal by the decomposition algorithm, improved variational mode decomposition is performed on the noise-reduced speech signal, the mode number K is set to 3, the penalty factor α is set to 2000, the 50-300Hz fundamental frequency band, the 0.2-2Hz breathing modulation band and the residual noise band are separated, the envelope signal of the breathing modulation band is extracted, the peak value offset of the cross-correlation function of the envelope signal and the fundamental frequency band signal is calculated, and the fundamental frequency jitter rate J is defined as F0 ; wherein, is the i-th frame cross-correlation peak position, is the historical mean; Capture the vocal tract morphology changes, combined with the fluctuation trend of the blood oxygen saturation to generate the resonance peak migration trajectory with physiological correlation, wherein in the process of capturing the vocal tract morphology changes, combined with the fluctuation trend of the blood oxygen saturation to generate the resonance peak migration trajectory with physiological correlation, time-varying linear prediction coding is used to analyze the voice signal, the resonance peak parameters F1-F3 are updated every 10ms, and then the correlation model of the blood oxygen saturation SpO2 and the resonance peak migration trajectory is established: ; wherein k=1, 2, 3 correspond to three resonance peaks, and a=0.15, b=0.08 are physiological coupling coefficients obtained through user calibration, is a resonance peak segment, is a resonance peak migration track, is a resonance peak segment related to fluctuation trend, d is the distance of the capture sound channel, and t is the time of the capture sound channel. The pitch jitter rate and the formant shift trajectory are vectorized to obtain the voiceprint feature vector. The atlas unit is configured to perform feature fusion on the voiceprint feature vector and the physiological signals after time-frequency conversion to generate a feature atlas including a user identity and physiological state parameters. The recognition unit is configured to perform event judgment on a voice feature segment and a physiological feature segment in the feature atlas.
9. A human-machine interaction device comprising a memory and a processor, the memory having stored therein a computer program, characterized in that, The processor executes the computer program to implement the steps of the voiceprint recognition method in the ring-wearing signal domain according to any one of claims 1 to 7.