A User Authentication Method and System Based on Breath Sounds

Capturing intra-ear breathing signals through ear-mounted devices, using Gaussian hybrid model and triple neural network for identity authentication, solving the problem of relying on active behavior and vulnerability in the existing technology, and achieving high security and ease of use identity authentication.

CN120030518BActive Publication Date: 2025-08-01NANJING UNIV OF INFORMATION SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510503576.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-01
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Existing biometrics based on ear-wearing devices rely on the user's active behavior and are vulnerable to audio playback attacks, which have insufficient security and ease of use.

Method used

Intra-ear breathing signals are captured through ear-mounted devices, and difficult-to-false biological features are extracted using Gaussian hybrid models and triple neural networks for identity authentication.

Benefits of technology

Improves the security and accuracy of identity authentication, can resist spoof attacks, adapt to environmental noise and respiratory behavior changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030518B_ABST
    Figure CN120030518B_ABST
Patent Text Reader

Abstract

The present invention discloses a user identity authentication method and system based on breath sounds, including: capturing an inner-ear breath signal from the ear canal through an ear-worn device, and performing denoising processing on the inner-ear breath signal to obtain a breath denoised signal; using a Gaussian mixture model to calculate the probabilities of each frame of the breath denoised signal belonging to a breath event and a non-breath event respectively to obtain a breath prior probability and a non-breath prior probability, thereby segmenting the breath denoised signal to obtain a breath segment signal; extracting a biometric identifier from the breath segment signal; inputting the biometric identifier into a pre-trained triplet neural network to calculate the similarity between the biometric identifier and a pre-stored biometric template; performing user identity authentication according to the similarity; the present invention extracts biometric features that are difficult to forge from breath sounds, can highly resist spoofing attacks, and at the same time has strong resistance and adaptability to environmental noise and changes in breathing behavior.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of identity recognition, and particularly relates to a user identity authentication method and system based on breath sounds. Background Art

[0002] With the continuous expansion of the market scale of ear-worn devices, more and more sensors are integrated into headphones, thus giving rise to increasingly rich applications and providing a new direction for identity authentication technology. Existing biometric technologies based on ear-worn devices achieve identity authentication by extracting biometric features such as the geometric shape of the user's ear canal, gait, voice features, and tooth occlusion. However, these methods usually rely on the user's active behaviors (such as walking, speaking, or chewing) and require the combined use of transceiver sensors (such as speakers and microphones), which reduces the user experience.

[0003] Breathing, as a common and unconscious physiological behavior, provides new possibilities for continuous passive identity authentication. However, due to the weak breath sounds and the security risks of air conduction, such systems are vulnerable to audio replay attacks, and accurate identity authentication also requires the user to place the acquisition device close to the mouth or nose, which has obvious deficiencies in terms of security and usability. Therefore, there is an urgent need for a continuous passive user identity authentication system that maintains a high level of security while imposing no other behavioral or usage restrictions on the user. Summary of the Invention

[0004] The present invention provides a user identity authentication method and system based on breath sounds, which extracts biometric features that are difficult to forge from breath sounds, can highly resist spoofing attacks, and has strong resistance and adaptability to environmental noise and changes in breathing behavior.

[0005] To achieve the above object, the technical solution adopted by the present invention is:

[0006] In a first aspect of the present invention, there is provided a user identity authentication method based on breath sounds, including:

[0007] Capturing an in-ear breath signal from the ear canal through an ear-worn device, and performing denoising processing on the in-ear breath signal to obtain a breath-denoised signal;

[0008] Using a Gaussian mixture model to calculate the probabilities of each frame of the breath-denoised signal belonging to a breathing event and a non-breathing event respectively to obtain a breathing prior probability and a non-breathing prior probability, and calculating the likelihood ratio of each frame of the breath-denoised signal based on the breathing prior probability and the non-breathing prior probability; segmenting the breath-denoised signal according to the likelihood ratio to obtain a breath segment signal;

[0009] Extract a biometric identifier from the respiratory segment signal; input the biometric identifier into a pre-trained triplet neural network to calculate the similarity between the biometric identifier and a pre-stored biometric template; perform user authentication based on the similarity.

[0010] Further, perform denoising processing on the in-ear respiratory signal to obtain a denoised respiratory signal, specifically including:

[0011] Set an identification window on the in-ear respiratory signal; after moving the identification window on the in-ear respiratory signal according to a set jump length, calculate the power spectral density of the in-ear respiratory signal within the identification window, and discard the in-ear respiratory signal within the identification window when the power spectral density exceeds a preset density threshold to obtain a first-level denoised respiratory signal;

[0012] Input the first-level denoised respiratory signal into a band-pass filter to filter the first-level denoised respiratory signal within a set frequency range to obtain a second-level denoised respiratory signal;

[0013] Optimize the denoising parameters of the maximum overlap discrete wavelet transform model through a genetic algorithm; input the second-level denoised respiratory signal into the optimized maximum overlap discrete wavelet transform model, decompose the second-level denoised respiratory signal into multiple wavelet coefficients and scale coefficients, perform thresholding processing on the wavelet coefficients and scale coefficients, and generate a denoised respiratory signal through signal reconstruction.

[0014] Further, optimize the denoising parameters of the maximum overlap discrete wavelet transform model through a genetic algorithm, specifically including:

[0015] Define a set of denoising parameters for the maximum overlap discrete wavelet transform model, and the expression formula is:

[0016]

[0017] In the formula, is the set of denoising parameters in the iteration of the genetic algorithm; is the mother wavelet function in the iteration of the genetic algorithm; is the decomposition level in the iteration of the genetic algorithm; is the threshold function in the iteration of the genetic algorithm; is the selected threshold in the iteration of the genetic algorithm; is the threshold rescaling function in the iteration of the genetic algorithm;

[0018] Encode the set of denoising parameters into chromosomes, and randomly generate a set number of chromosomes to form an initial population;

[0019] The maximum overlap discrete wavelet transform model based on the chromosome-corresponding denoising parameter set converts the original noisy signal into a reconstructed signal and calculates the mean square error. The expression formula is as follows:

[0020]

[0021]

[0022] In the formula, is the maximum overlap discrete wavelet transform model; is the reconstructed signal; is the original noisy signal; is the mean square error; is the number of original noisy signals; n is the serial number of the original noisy signal;

[0023] Cross-process and mutate the chromosomes according to the mean square error and update the initial population of the chromosomes. Repeat the optimization process of the denoising parameters in the maximum overlap discrete wavelet transform model until the termination condition is met, and then output the optimized maximum overlap discrete wavelet transform model.

[0024] Furthermore, use the Gaussian mixture model to calculate the probabilities that each frame of respiratory denoised signal belongs to respiratory events and non-respiratory events respectively to obtain the respiratory prior probability and non-respiratory prior probability. Calculate the likelihood ratio of each frame of respiratory denoised signal through the respiratory prior probability and non-respiratory prior probability, specifically including:

[0025] Divide each frame of respiratory denoised signal into W sub-bands according to frequency;

[0026] Use the Gaussian mixture model to calculate the probability that each frame of respiratory denoised signal belongs to respiratory events respectively to obtain the respiratory prior probability. The expression formula is as follows:

[0027]

[0028] In the formula, is the respiratory prior probability; is the th logarithmic energy of the th sub-band in the th frame of respiratory denoised signal; is the th is the th is the mean of the th sub-band;

[0029] The probability that each frame of the respiratory denoised signal belongs to a non-respiratory event is calculated using a Gaussian mixture model to obtain the non-respiratory prior probability. The expression formula is:

[0030]

[0031] In the formula, is the non-respiratory prior probability;

[0032] The likelihood ratio of the sub-band in each frame of the respiratory denoised signal is calculated through the respiratory prior probability and the non-respiratory prior probability. The expression formula is:

[0033]

[0034] In the formula, is the likelihood ratio of the th sub-band in the th frame of the respiratory denoised signal;

[0035] Weights are assigned to the likelihood ratios of each sub-band according to the power cumulative distribution of each sub-band, and the weighted sum of the likelihood ratios of each sub-band is obtained to get the likelihood ratio of each frame of the respiratory denoised signal. The expression formula is:

[0036]

[0037] In the formula, is the weighting coefficient of the th sub-band; is the likelihood ratio of the th frame of the respiratory denoised signal.

[0038] [[ID=4३]]Furthermore, the respiratory denoised signal is segmented according to the likelihood ratio to obtain the respiratory segment signal, specifically including:

[0039] An adaptive threshold is calculated based on the mean and variance of the sub-bands in each frame of the respiratory denoised signal. The expression formula is:

[0040]

[0041] In the formula, is the adaptive threshold, represents the threshold weight of the th sub-band; represents the mean of the th sub-band in the non-respiratory mode; represents the mean of the th sub-band in the respiratory mode; represents the variance of the th sub-band in the non-respiratory mode; represents the variance of the th sub-band in the respiratory mode;

[0042] Compare the likelihood ratio of the frame respiratory denoised signal with the adaptive threshold to determine whether each frame of the respiratory denoised signal belongs to a respiratory event or a non-respiratory event, and delete the frames belonging to the non-respiratory event in the respiratory denoised signal to obtain a respiratory segment signal.

[0043] Further, the biometric identifier includes a body asymmetry identifier, an ear canal geometry identifier, and a respiratory tract descriptor; extracting the biometric identifier from the respiratory segment signal specifically includes:

[0044] Extract the start timestamp and end timestamp of the respiratory segment signal, and calculate the time difference of the respiratory segment signal. The expression formula is:

[0045]

[0046] In the formula, is the time difference of the respiratory segment signal; is the end timestamp of the respiratory segment signal; is the start timestamp of the respiratory segment signal;

[0047] If the time difference of the respiratory segment signal exceeds the preset respiratory time threshold , determine that the respiratory segment signal is a valid respiratory segment; otherwise, delete the respiratory segment signal;

[0048] Obtain multiple consecutive adjacent valid respiratory segments, and calculate the time interval between the start timestamp of the current valid respiratory segment and the end timestamp of the previous adjacent valid respiratory segment, denoted as the first time interval. The expression formula is:

[0049]

[0050] In the formula, is the start timestamp of the current valid respiratory segment; is the end timestamp of the previous adjacent valid respiratory segment; is the first time interval;

[0051] Calculate the time interval between the end timestamp of the current valid respiratory segment and the start timestamp of the next adjacent valid respiratory segment, denoted as the second time interval. The expression formula is:

[0052]

[0053] In the formula, is the second time interval; is the end timestamp of the current valid respiratory segment; is the start timestamp of the next adjacent valid respiratory segment;

[0054] If the first time interval is greater than the second time interval, it is determined that the effective breathing segment is an inhalation process; otherwise, it is determined that the effective breathing segment is an exhalation process; the effective breathing segments of the inhalation process and the effective breathing segments of the adjacent exhalation process are combined to form the breathing event characteristics of a single cycle.

[0055] The breathing event characteristics are divided into left-channel breathing event characteristics and right-channel breathing event characteristics; the cross-power spectral density is calculated based on the left-channel breathing event characteristics and the right-channel breathing event characteristics; the body asymmetry identifier is obtained by calculating from the cross-power spectral density and the auto-power spectrum of the right-channel breathing event characteristics, and the expression formula is:

[0056]

[0057] In the formula, is the body asymmetry identifier; is the cross-power spectral density; is the auto-power spectrum of the right-channel breathing event characteristics;

[0058] Set several target frequencies within the frequency range of the ear internal breathing sound; calculate the power cumulative distribution characteristics of a single breathing cycle according to the target frequencies.

[0059]

[0060] In the formula, is the time-frequency diagram of the breathing event characteristics; is the starting frequency of the breathing event characteristics; is the ending frequency of the breathing event characteristics; is the target frequency; is the power cumulative distribution characteristics of a single breathing cycle;

[0061] The ear canal geometry identifier is composed of the power cumulative distribution characteristics corresponding to each target frequency;

[0062] The Mel-frequency cepstral coefficients are extracted from the breathing event characteristics, and the Mel-frequency cepstral coefficients, their first-order derivatives and second-order derivatives are combined to form the respiratory tract descriptor.

[0063] Furthermore, the triplet neural network includes three sub-network models; three groups of convolutional units and three fully-connected layers are connected in sequence to form the sub-network model; a two-dimensional convolutional block and a max-pooling layer are configured in sequence within the three groups of convolutional units;

[0064] The training process of the triplet neural network includes:

[0065] Obtain the breathing audio training signals of multiple users from the database, multiply the audio waveforms of the breathing audio training signals by a random factor to adjust the breathing sound volume; change the duration and speed of the breathing sound in the breathing audio training signals; add Gaussian noise to the breathing audio training signals, and move the breathing audio training signals along the time domain to obtain breathing audio training samples;

[0066] Select a sample anchor point from the breathing audio training samples, set the breathing audio training samples with the same user breathing sound as the sample anchor point as positive samples, otherwise set the breathing audio training samples as negative samples;

[0067] Input the sample anchor point, positive samples, and negative samples into the triplet neural network, and the triplet neural network outputs the intra-class distance between the sample anchor point and the positive samples and the inter-class distance between the sample anchor point and the negative samples;

[0068] Calculate the training loss value according to the intra-class distance and the inter-class distance, and the expression formula is:

[0069]

[0070] In the formula, is the training loss value, is the sample anchor point, is the positive sample; is the negative sample; is the minimum distance between the positive sample and the negative sample; is the intra-class distance between the sample anchor point and the positive sample; is the inter-class distance between the sample anchor point and the negative sample;

[0071] Optimize the weight parameters of the triplet neural network according to the training loss value, and repeat the training process of the triplet neural network until the set number of iterations is reached, and output the trained triplet neural network.

[0072] Furthermore, input the biometric identifier into the pre-trained triplet neural network to calculate the similarity between the biometric identifier and the pre-stored biometric template; perform user identity authentication according to the similarity, specifically including:

[0073] Input the biometric identifier into the pre-trained triplet neural network to calculate the similarity between the biometric identifier and the pre-stored biometric template, and the expression formula is:

[0074]

[0075] In the formula, is the similarity between the biometric identifier and the k-th biometric template; is the sub-network model; is a body asymmetry identifier; is an ear canal geometry identifier; is a respiratory tract descriptor; is the body asymmetry identification feature in the k-th biometric template; is the ear canal geometry identification feature in the k-th biometric template; is the respiratory tract description feature in the k-th biometric template;

[0076] When it is determined that the biometric identifier belongs to the user identity corresponding to the k-th biometric template. is a similarity threshold;

[0077] When it is determined that the biometric identifier belongs to an illegal user.

[0078] The second aspect of the present invention provides a user identity authentication system based on breath sounds, including:

[0079] An acquisition module, configured to capture an in-ear breath signal from the ear canal through an ear-worn device, and perform denoising processing on the in-ear breath signal to obtain a denoised breath signal;

[0080] A segmentation module, configured to calculate the probabilities of each frame of the denoised breath signal belonging to a breath event and a non-breath event respectively using a Gaussian mixture model to obtain a breath prior probability and a non-breath prior probability, calculate the likelihood ratio of each frame of the denoised breath signal based on the breath prior probability and the non-breath prior probability; segment the denoised breath signal based on the likelihood ratio to obtain a breath segment signal;

[0081] An identification module, configured to extract a biometric identifier from the breath segment signal; input the biometric identifier into a pre-trained triplet neural network, calculate the similarity between the biometric identifier and a pre-stored biometric template; perform user identity authentication based on the similarity.

[0082] The third aspect of the present invention provides an electronic device, including a storage medium and a processor; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the user identity authentication method described in the first aspect of the present invention.

[0083] Compared with the prior art, the beneficial effects of the present invention:

[0084] The present invention uses a Gaussian mixture model to calculate the probabilities that each frame of the respiratory denoised signal belongs to a respiratory event and a non-respiratory event respectively, obtaining the respiratory prior probability and the non-respiratory prior probability, and calculating the likelihood ratio of each frame of the respiratory denoised signal based on the respiratory prior probability and the non-respiratory prior probability; segmenting the respiratory denoised signal according to the likelihood ratio to obtain a respiratory segment signal; the present invention can effectively filter out noise interference, reduce the influence of abnormal data, and enhance the robustness of the system in a complex environment; it can accurately distinguish respiratory and non-respiratory signals, adapt to the respiratory patterns of different users, and improve the accuracy of respiratory event detection.

[0085] The present invention extracts a biometric identifier from the respiratory segment signal; inputs the biometric identifier into a pre-trained triplet neural network to calculate the similarity between the biometric identifier and a pre-stored biometric template; performs user identity authentication according to the similarity; respiratory sounds are difficult to forge as biometric features, and combined with the triplet neural network of deep learning, it can accurately identify individual differences, effectively resist spoofing attacks, and significantly improve the security of identity authentication. Brief Description of the Drawings

[0086] Figure 1 It is a flowchart of the user identity authentication method based on respiratory sound provided in Embodiment 1.

[0087] Figure 2 It is a flowchart of the denoising process of the in-ear respiratory signal provided in Embodiment 1;

[0088] Figure 3 It is a flowchart of the segmentation of the respiratory denoised signal provided in Embodiment 1;

[0089] Figure 4 It is a flowchart of the training process of the triplet neural network provided in Embodiment 1;

[0090] Figure 5 It is a flowchart of the user identity authentication provided in Embodiment 1;

[0091] Figure 6 It is a structural diagram of the triplet neural network provided in Embodiment 1. Detailed Description of the Invention

[0092] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solutions of the present invention more clearly, and cannot be used to limit the protection scope of the present invention.

[0093] Embodiment 1

[0094] As Figure 1 shown, this embodiment provides a user identity authentication method based on respiratory sound, including:

[0095] As Figure 6As shown in the figure, three sets of convolutional units and three fully connected layers are connected in sequence to form the sub-network model; a triple neural network is constructed from three sub-network models; within the three sets of convolutional units, a two-dimensional convolutional block and a max pooling layer are configured in sequence; the two-dimensional convolutional block consists of a convolutional layer based on the ReLU activation function and a batch normalization layer; in this embodiment, the kernel size of the max pooling layer is 3x3 and the stride is 2.

[0096] As Figure 4 shown, the training process of the triple neural network includes:

[0097] Obtain the breathing audio training signals of multiple users from the database, multiply the audio waveforms of the breathing audio training signals by random factors to adjust the breathing sound volume; change the duration and speed of the breathing sound in the breathing audio training signals; add Gaussian noise to the breathing audio training signals to shift the breathing sound pitch up or down; move the breathing audio training signals along the time domain to obtain breathing audio training samples;

[0098] Select sample anchors from the breathing audio training samples, set the breathing audio training samples with the same user's breathing sound as the sample anchor as positive samples, otherwise set the breathing audio training samples as negative samples;

[0099] Input the sample anchor, positive sample, and negative sample into the triple neural network, and output the intra-class distance between the sample anchor and the positive sample and the inter-class distance between the sample anchor and the negative sample by the triple neural network;

[0100] Calculate the training loss value according to the intra-class distance and the inter-class distance, and the expression formula is: [[ID=S19]]

[0101]

[0102] In the formula, is the training loss value, is the sample anchor, is the positive sample; is the negative sample; is the minimum distance between the positive sample and the negative sample; is the intra-class distance between the sample anchor and the positive sample; is the inter-class distance between the sample anchor and the negative sample;

[0103] Use the Adam trainer to optimize the weight parameters of the triple neural network according to the training loss value, and repeat the training process of the triple neural network until the set number of iterations is reached, and output the trained triple neural network; in this embodiment, the learning rate, batch size, and number of iterations are set to 0.0001, 16, and 60 respectively.

[0104] In this embodiment, positive and negative sample pairs are constructed, and a specific loss function is designed to train the model. During the training process, the model adjusts the feature space so that the respiratory signal features of the same user are close to each other in the feature space (reducing the intra-class distance), while the respiratory signal features of different users are far from each other (increasing the inter-class distance). This optimization method enables the triplet neural network to more accurately identify the differences between different users, thereby improving the accuracy and reliability of identity authentication.

[0105] As Figure 2 shown, the in-ear respiratory signal is captured from the ear canal through an ear-worn device, and the in-ear respiratory signal is denoised to obtain a denoised respiratory signal, which specifically includes:

[0106] Set an identification window on the in-ear respiratory signal; after moving the identification window on the in-ear respiratory signal according to the set jump length, calculate the power spectral density of the in-ear respiratory signal within the identification window, and discard the in-ear respiratory signal within the identification window when the power spectral density exceeds the preset density threshold to obtain a first-level denoised respiratory signal; in this embodiment, the window length of the identification window is 10 s and the jump length is 1 s.

[0107] Input the first-level denoised respiratory signal into a band-pass filter, and filter the first-level denoised respiratory signal within the set frequency range to obtain a second-level denoised respiratory signal; since the noise generated by movement and heartbeat is mainly distributed below 150 Hz, while the main frequency of in-ear breath sounds is 0 to 2 KHz. Therefore, the filtering range of the band-pass filter in this embodiment is set to 150 to 3000 Hz.

[0108] Define the denoising parameter set of the maximum overlap discrete wavelet transform model, and the expression formula is:

[0109]

[0110] In the formula, is the denoising parameter set in the iteration of the genetic algorithm; is the mother wavelet function in the iteration of the genetic algorithm; is the decomposition level in the iteration of the genetic algorithm; is the threshold function in the iteration of the genetic algorithm; is the selected threshold in the iteration of the genetic algorithm; is the threshold rescaling function in the iteration of the genetic algorithm;

[0111] Encode the denoising parameter set into chromosomes, and randomly generate a set number of chromosomes to form an initial population;

[0112] The maximum overlap discrete wavelet transform model based on the chromosome-corresponding denoising parameter set converts the original noisy signal into a reconstructed signal and calculates the mean square error. The expression formula is as follows:

[0113]

[0114]

[0115] In the formula, is the maximum overlap discrete wavelet transform model, is the reconstructed signal; is the original noisy signal; is the mean square error; is the number of original noisy signals; n is the serial number of the original noisy signal;

[0116] Cross-process and mutate the chromosomes according to the mean square error and update the initial population of the chromosomes. Repeat the optimization process of the denoising parameters in the maximum overlap discrete wavelet transform model until the termination condition is met, and then output the optimized maximum overlap discrete wavelet transform model.

[0117] Input the second-level denoised respiratory signal into the optimized maximum overlap discrete wavelet transform model, decompose the second-level denoised respiratory signal into multiple wavelet coefficients and scale coefficients, perform thresholding processing on the wavelet coefficients and scale coefficients, and generate a respiratory denoised signal through signal reconstruction.

[0118] As Figure 3 shown, use the Gaussian mixture model to calculate the probabilities that each frame of the respiratory denoised signal belongs to the respiratory event and the non-respiratory event respectively to obtain the respiratory prior probability and the non-respiratory prior probability, and calculate the likelihood ratio of each frame of the respiratory denoised signal through the respiratory prior probability and the non-respiratory prior probability. Specifically, it includes:

[0119] Divide each frame of the respiratory denoised signal into W sub-bands according to the frequency; in this embodiment, W = 4; the ranges of the four sub-bands are 0 to 500 Hz, 500 to 1000 Hz, 1000 to 1500 Hz, and 1,500 to 3,000 Hz respectively.

[0120] Use the Gaussian mixture model to calculate the probability that each frame of the respiratory denoised signal belongs to the respiratory event to obtain the respiratory prior probability. The expression formula is as follows:

[0121]

[0122] In the formula, is the respiratory prior probability; is the th logarithmic energy of the is the parameter set for the th sub - frequency band; is the class label of the denoised respiration signal for each frame; is the th sub - frequency band mean value; is the th sub - frequency band variance; is pi; is the natural exponential function;

[0123] The probability that each frame of the denoised respiration signal belongs to a non - respiration event is calculated using a Gaussian mixture model to obtain the non - respiration prior probability. The expression formula is:

[0124]

[0125] In the formula, is the non - respiration prior probability;

[0126] The likelihood ratio of the sub - frequency band in each frame of the denoised respiration signal is calculated through the respiration prior probability and the non - respiration prior probability. The expression formula is:

[0127]

[0128] In the formula, is the th frame of the denoised respiration signal, and is the likelihood ratio of the

[0129] th sub - frequency band; [[ID=4t7]]

[0130]

[0131] In the formula, is the th sub - frequency band weighting coefficient; is the th frame of the denoised respiration signal likelihood ratio.

[0132] The respiration - denoised signal is segmented according to the likelihood ratio to obtain the respiration segment signal, specifically including:

[0133] The adaptive threshold is calculated according to the mean value and variance of the sub - frequency band in each frame of the denoised respiration signal. The expression formula is:

[0134]

[0135] In the formula, is the adaptive threshold, represents the It should be noted that there may be some inaccuracies in the translation due to the lack of complete context and the complexity of the technical content. You may need to adjust it according to the actual situation.Threshold weight of each sub - band; Denote the mean value of the th sub - band in the non - breathing mode; Denote the mean value of the th sub - band in the breathing mode; Denote the variance of the

[0136] th sub - band in the non - breathing mode; Compare the likelihood ratio of the

[0137] th frame of the breathing - denoised signal with the adaptive threshold to determine whether each frame of the breathing - denoised signal belongs to a breathing event or a non - breathing event, and delete the frames belonging to the non - breathing event in the breathing - denoised signal to obtain the breathing segment signal.

[0138] Extract the start timestamp and end timestamp of the breathing segment signal, and calculate the time difference of the breathing segment signal. The expression formula is:

[0139]

[0140] In the formula, is the time difference of the breathing segment signal; is the end timestamp of the breathing segment signal; is the start timestamp of the breathing segment signal;

[0141] If the time difference of the breathing segment signal exceeds the preset breathing time threshold , determine that the breathing segment signal is a valid breathing segment; otherwise, delete the breathing segment signal.

[0142] Obtain multiple consecutive adjacent valid breathing segments, and calculate the time interval between the start timestamp of the current valid breathing segment and the end timestamp of the previous adjacent valid breathing segment, denoted as the first time interval. The expression formula is:

[0143]

[0144] In the formula, is the start timestamp of the current valid breathing segment; is the end timestamp of the previous adjacent valid breathing segment; is the first time interval;

[0145] Calculate the time interval between the end timestamp of the current valid breathing segment and the start timestamp of the adjacent valid breathing segment behind, denoted as the second time interval, and the expression formula is:

[0146]

[0147] In the formula, is the second time interval; is the end timestamp of the current valid breathing segment; is the start timestamp of the adjacent valid breathing segment behind;

[0148] If the first time interval is greater than the second time interval, determine that the valid breathing segment is an inhalation process; otherwise, determine that the valid breathing segment is an exhalation process; form the breathing event characteristics of a single cycle by combining the valid breathing segments of the inhalation process and the valid breathing segments of the adjacent exhalation process;

[0149] As Figure 5 shown, the breathing event characteristics are divided into left-channel breathing event characteristics and right-channel breathing event characteristics; calculate the cross-power spectral density according to the left-channel breathing event characteristics and the right-channel breathing event characteristics; calculate and obtain the body asymmetry identifier from the cross-power spectral density and the auto-power spectrum of the right-channel breathing event characteristics, and the expression formula is:

[0150]

[0151] In the formula, is the body asymmetry identifier; is the cross-power spectral density; is the auto-power spectrum of the right-channel breathing event characteristics;

[0152] In this embodiment, taking advantage of the individual differences and asymmetries of the human skeleton, the vibrations caused by breathing are transmitted to the left and right ear canals through the bones in the body, and the left and right channel signals captured by the microphone are regarded as two independent signals with different bone conduction paths to calculate the cross-power spectral density (CPSD) respectively, obtaining the correlation within the frequency spectrum range, and further obtaining the body asymmetry identifier.

[0153] Set a number of target frequencies within the frequency range of the ear breathing sound; calculate the power cumulative distribution characteristics of a single breathing cycle according to the target frequencies;

[0154]

[0155] In the formula, is the time-frequency diagram of the breathing event characteristics; is the starting frequency of the breathing event characteristics; is the termination frequency of the breathing event characteristics; is the target frequency; The power cumulative distribution feature for a single respiratory cycle; in this embodiment, the target frequencies are set to 500 Hz, 1000 Hz, 1500 Hz, and 3000 Hz;

[0156] The ear canal geometric identifier is composed of the power cumulative distribution features corresponding to each target frequency; in this embodiment, taking advantage of the differences in the ear canal structures among individuals, the impedance effect of the occlusion effect on respiratory sounds varies from person to person. In this embodiment, to characterize the differences caused by the occlusion effect among users, the power cumulative distribution of a single respiratory cycle is calculated, and then the ear canal geometric identifier is obtained.

[0157] Mel-frequency cepstral coefficients are extracted from the respiratory event features, and the mel-frequency cepstral coefficients, their first-order derivatives, and second-order derivatives are combined to form a respiratory tract descriptor. In this embodiment, during the breathing process, the high-speed airflow collides with the respiratory tract, causing resonance. Based on the significant individual differences in the resonance frequencies, the respiratory tract descriptor is extracted.

[0158] The biometric identifier is input into a pre-trained triplet neural network to calculate the similarity between the biometric identifier and the pre-stored biometric template; user authentication is performed according to the similarity, specifically including:

[0159] The biometric identifier is input into a pre-trained triplet neural network to calculate the similarity between the biometric identifier and the pre-stored biometric template. The expression formula is:

[0160]

[0161] In the formula, is the similarity between the biometric identifier and the k-th biometric template; is the sub-network model; is the body asymmetry identifier; is the ear canal geometric identifier; is the respiratory tract descriptor; is the body asymmetry identification feature in the k-th biometric template; is the ear canal geometric identification feature in the k-th biometric template; is the respiratory tract description feature in the k-th biometric template;

[0162] When , it is determined that the biometric identifier belongs to the user identity corresponding to the k-th biometric template; is the similarity threshold;

[0163] When , it is determined that the biometric identifier belongs to a non-legitimate user.

[0164] In this embodiment, the ear-internal respiration signal is first denoised to obtain a respiration-denoised signal. Then, the Gaussian mixture model is used to calculate the probabilities that each frame of the respiration-denoised signal belongs to a respiration event and a non-respiration event respectively, so as to obtain a respiration prior probability and a non-respiration prior probability. Furthermore, the respiration-denoised signal is segmented to obtain a respiration segment signal, making this embodiment resistant to environmental noise and respiration behavior changes and having strong adaptability; by extracting biometric features that are difficult to forge, it can highly resist spoofing attacks.

[0165] Embodiment 2

[0166] This embodiment provides a user identity authentication system based on breath sounds. The system described in this embodiment can be applied to the method described in Embodiment 1. The user identity authentication system includes:

[0167] An acquisition module, configured to capture an ear-internal respiration signal from the ear canal through an ear-worn device, and denoise the ear-internal respiration signal to obtain a respiration-denoised signal;

[0168] A segmentation module, configured to use the Gaussian mixture model to calculate the probabilities that each frame of the respiration-denoised signal belongs to a respiration event and a non-respiration event respectively to obtain a respiration prior probability and a non-respiration prior probability, and calculate the likelihood ratio of each frame of the respiration-denoised signal through the respiration prior probability and the non-respiration prior probability; segment the respiration-denoised signal according to the likelihood ratio to obtain a respiration segment signal;

[0169] An identification module, configured to extract a biometric identifier from the respiration segment signal; input the biometric identifier into a pre-trained triplet neural network, and calculate the similarity between the biometric identifier and a pre-stored biometric template; perform user identity authentication according to the similarity.

[0170] The acquisition module denoises the ear-internal respiration signal to obtain a respiration-denoised signal, which specifically includes:

[0171] Set an identification window on the ear-internal respiration signal; after moving the identification window on the ear-internal respiration signal according to a set jump length, calculate the power spectral density of the ear-internal respiration signal within the identification window, and discard the ear-internal respiration signal within the identification window when the power spectral density exceeds a preset density threshold to obtain a first-stage noise-reduced respiration signal;

[0172] Input the first-stage noise-reduced respiration signal into a band-pass filter, and filter the first-stage noise-reduced respiration signal within a set frequency range to obtain a second-stage noise-reduced respiration signal;

[0173] Optimize the denoising parameters of the maximum overlap discrete wavelet transform model through a genetic algorithm; input the second-level denoised respiratory signal into the optimized maximum overlap discrete wavelet transform model, decompose the second-level denoised respiratory signal into multiple wavelet coefficients and scale coefficients, perform thresholding processing on the wavelet coefficients and scale coefficients, and generate a respiratory denoised signal through signal reconstruction.

[0174] The biometric identifier includes a body asymmetry identifier, an ear canal geometry identifier, and a respiratory tract descriptor; the recognition module extracts the biometric identifier from the respiratory segment signal, specifically including:

[0175] Extract the start timestamp and end timestamp of the respiratory segment signal, and calculate the time difference of the respiratory segment signal. The expression formula is:

[0176]

[0177] In the formula, is the time difference of the respiratory segment signal; is the end timestamp of the respiratory segment signal; is the start timestamp of the respiratory segment signal;

[0178] If the time difference of the respiratory segment signal exceeds the preset respiratory time threshold , determine that the respiratory segment signal is a valid respiratory segment; otherwise, delete the respiratory segment signal.

[0179] Obtain multiple consecutive adjacent valid respiratory segments, calculate the time interval between the start timestamp of the current valid respiratory segment and the end timestamp of the previous adjacent valid respiratory segment, denoted as the first time interval. The expression formula is:

[0180]

[0181] In the formula, is the start timestamp of the current valid respiratory segment; is the end timestamp of the previous adjacent valid respiratory segment; is the first time interval;

[0182] Calculate the time interval between the end timestamp of the current valid respiratory segment and the start timestamp of the next adjacent valid respiratory segment, denoted as the second time interval. The expression formula is:

[0183]

[0184] In the formula, is the second time interval; is the end timestamp of the current valid respiratory segment; is the start timestamp of the next adjacent valid respiratory segment;

[0185] If the first time interval is greater than the second time interval, determine that the effective breathing segment is an inhalation process; otherwise, determine that the effective breathing segment is an exhalation process; combine the effective breathing segments of the inhalation process with the effective breathing segments of the adjacent exhalation process to form the breathing event characteristics of a single cycle.

[0186] The breathing event characteristics are divided into left-channel breathing event characteristics and right-channel breathing event characteristics; calculate the cross-power spectral density based on the left-channel breathing event characteristics and the right-channel breathing event characteristics; obtain the body asymmetry identifier from the cross-power spectral density and the auto-power spectrum of the right-channel breathing event characteristics, and the expression formula is:

[0187]

[0188] In the formula, is the body asymmetry identifier; is the cross-power spectral density; is the auto-power spectrum of the right-channel breathing event characteristics;

[0189] Set several target frequencies within the frequency range of the breathing sound in the ear; calculate the power cumulative distribution characteristics of a single breathing cycle according to the target frequencies.

[0190]

[0191] In the formula, is the time-frequency diagram of the breathing event characteristics; is the starting frequency of the breathing event characteristics; is the ending frequency of the breathing event characteristics; is the target frequency; is the power cumulative distribution characteristics of a single breathing cycle;

[0192] The ear canal geometry identifier is composed of the power cumulative distribution characteristics corresponding to each target frequency;

[0193] Extract the Mel-frequency cepstral coefficients from the breathing event characteristics, and form the respiratory tract descriptor by the Mel-frequency cepstral coefficients and their first and second derivatives.

[0194] Embodiment 3

[0195] This embodiment provides an electronic device, including a storage medium and a processor; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the user identity authentication method described in Embodiment 1.

[0196] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0197] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0198] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0199] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0200] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A user authentication method based on breath sounds, characterized in that, Including: Capturing an in-ear respiration signal from the ear canal through an ear-worn device, and performing denoising processing on the in-ear respiration signal to obtain a respiration denoised signal; Using a Gaussian mixture model to calculate the probabilities that each frame of the respiration denoised signal belongs to a respiration event and a non-respiration event respectively to obtain a respiration prior probability and a non-respiration prior probability, and calculating the likelihood ratio of each frame of the respiration denoised signal through the respiration prior probability and the non-respiration prior probability; Segmenting the respiration denoised signal according to the likelihood ratio to obtain a respiration segment signal; specifically including: Calculating an adaptive threshold according to the mean and variance of the sub-bands in each frame of the respiration denoised signal, and the expression formula is: ; In the formula, is the adaptive threshold, represents the threshold weight of the th sub-band; represents the mean value of the th sub-band in the non-breathing mode; represents the mean value of the th sub-band in the breathing mode; represents the variance of the th sub-band in the non-breathing mode; represents the variance of the th sub-band in the breathing mode; W is the number of divided sub-bands; Compare the likelihood ratio of the frame respiratory denoised signal with the adaptive threshold to determine whether each frame of the respiratory denoised signal belongs to a respiratory event or a non-respiratory event, and delete the frames belonging to the non-respiratory event in the respiratory denoised signal to obtain a respiratory segment signal; Extracting a biometric identifier from the respiration segment signal, and the biometric identifier includes a body asymmetry identifier; specifically including: Extracting the start timestamp and the end timestamp of the respiration segment signal, and calculating the time difference of the respiration segment signal, and the expression formula is: ; In the formula, is the time difference of the respiration segment signal; is the end timestamp of the respiration segment signal; is the start timestamp of the respiration segment signal; If the time difference of the respiration segment signal exceeds a preset respiration time threshold , determine that the respiration segment signal is a valid respiration segment; otherwise, delete the respiration segment signal. Obtaining a plurality of consecutive adjacent valid respiration segments, and calculating the time interval between the start timestamp of the current valid respiration segment and the end timestamp of the adjacent valid respiration segment in front, denoted as the first time interval, and the expression formula is: ; In the formula, is the start timestamp of the currently valid breathing segment; is the end timestamp of the adjacent previous valid breathing segment; is the first time interval; Calculating the time interval between the end timestamp of the current valid respiration segment and the start timestamp of the adjacent valid respiration segment behind, denoted as the second time interval, and the expression formula is: ; In the formula, is the second time interval; is the end timestamp of the current valid breathing segment; is the start timestamp of the adjacent subsequent valid breathing segment; If the first time interval is greater than the second time interval, determining that the valid respiration segment is an inhalation process; otherwise, determining that the valid respiration segment is an exhalation process; and forming the valid respiration segments of the inhalation process and the valid respiration segments of the adjacent exhalation process into the respiration event characteristics of a single cycle; The respiration event characteristics are divided into left-channel respiration event characteristics and right-channel respiration event characteristics; calculating the cross-power spectral density according to the left-channel respiration event characteristics and the right-channel respiration event characteristics; and calculating and obtaining the body asymmetry identifier from the cross-power spectral density and the auto-power spectral of the right-channel respiration event characteristics, and the expression formula is: ; In the formula, is the body asymmetry identifier; is the cross power spectral density; is the auto power spectrum of the right channel respiratory event feature; Inputting the biometric identifier into a pre-trained triplet neural network, and calculating the similarity between the biometric identifier and a pre-stored biometric template; and performing user identity authentication according to the similarity.

2. The method for user identity authentication based on breath sounds according to claim 1, wherein Performing denoising processing on the in-ear respiration signal to obtain a respiration denoised signal, specifically including: Setting an identification window on the in-ear respiration signal; after moving the identification window on the in-ear respiration signal according to a set jump length, calculating the power spectral density of the in-ear respiration signal within the identification window, and discarding the in-ear respiration signal within the identification window when the power spectral density exceeds a preset density threshold to obtain a first-stage denoised respiration signal; Inputting the first-stage denoised respiration signal into a band-pass filter, and filtering the first-stage denoised respiration signal within a set frequency range to obtain a second-stage denoised respiration signal; Optimizing the denoising parameters of the maximum overlap discrete wavelet transform model through a genetic algorithm; inputting the second-stage denoised respiration signal into the optimized maximum overlap discrete wavelet transform model, decomposing the second-stage denoised respiration signal into a plurality of wavelet coefficients and scale coefficients, performing thresholding processing on the wavelet coefficients and the scale coefficients, and generating a respiration denoised signal through signal reconstruction.

3. The method for user authentication based on breath sounds according to claim 2, wherein Optimizing the denoising parameters of the maximum overlap discrete wavelet transform model through a genetic algorithm, specifically including: Define the denoising parameter set of the maximum overlap discrete wavelet transform model, and the expression formula is: ; In the formula, is the set of denoising parameters in the ith iteration of the genetic algorithm; is the mother wavelet function in the ith iteration of the genetic algorithm; is the decomposition level in the ith iteration of the genetic algorithm; is the threshold function in the ith iteration of the genetic algorithm; is the selected threshold in the ith iteration of the genetic algorithm; is the threshold rescaling function in the ith iteration of the genetic algorithm; Encode the denoising parameter set into chromosomes, and randomly generate a set number of chromosomes to form an initial population; Based on the maximum overlap discrete wavelet transform model corresponding to the denoising parameter set of the chromosome, convert the original noisy signal into a reconstructed signal and calculate the mean square error. The expression formula is: ; ; In the formula, is the maximum overlap discrete wavelet transform model, is the reconstructed signal; is the original noisy signal; is the mean square error; is the number of original noisy signals; n is the serial number of the original noisy signal; Perform crossover processing and mutation processing on the chromosomes according to the mean square error and update the initial population of the chromosomes. Repeat the optimization process of the denoising parameters in the maximum overlap discrete wavelet transform model until the termination condition is met, and output the optimized maximum overlap discrete wavelet transform model.

4. The method for user identity authentication based on breath sounds according to claim 1, wherein Use the Gaussian mixture model to calculate the probabilities that each frame of the respiratory denoised signal belongs to respiratory events and non-respiratory events respectively to obtain the respiratory prior probability and non-respiratory prior probability. Calculate the likelihood ratio of each frame of the respiratory denoised signal through the respiratory prior probability and non-respiratory prior probability, specifically including: Divide each frame of the respiratory denoised signal into W sub-bands according to frequency; Use the Gaussian mixture model to calculate the probability that each frame of the respiratory denoised signal belongs to respiratory events respectively to obtain the respiratory prior probability. The expression formula is: ; In the formula, is the prior probability of respiration; is the -th frame of the respiration-denoised signal, and is the logarithmic energy of the -th sub-band; is the parameter set of the -th sub-band; is the class label of each frame of the respiration-denoised signal; is the mean value of the -th sub-band; is the variance of the -th sub-band; is the natural exponential function; Use the Gaussian mixture model to calculate the probability that each frame of the respiratory denoised signal belongs to non-respiratory events respectively to obtain the non-respiratory prior probability. The expression formula is: ; In the formula, is the non-breathing prior probability; Calculate the likelihood ratio of the sub-bands in each frame of the respiratory denoised signal through the respiratory prior probability and non-respiratory prior probability. The expression formula is: ; In the formula, is the likelihood ratio of the sub - frequency band in the frame - based respiration - denoised signal; Assign weights to the likelihood ratios of each sub-band according to the power cumulative distribution of each sub-band, and perform weighted summation on the likelihood ratios of each sub-band to obtain the likelihood ratio of each frame of the respiratory denoised signal. The expression formula is: ; In the formula, is the weighting coefficient of the th sub-band; is the likelihood ratio of the th frame of respiratory denoised signal; W is the number of divided sub-bands.

5. The method for user authentication based on breath sounds according to claim 1, characterized in that, The biometric identifier further includes an ear canal geometry identifier and a respiratory tract descriptor; extract the ear canal geometry identifier and the respiratory tract descriptor from the respiratory segment signal, specifically including: Set a number of target frequencies within the frequency range of the ear respiratory sound; calculate the power cumulative distribution characteristics of a single respiratory cycle according to the target frequencies. The expression formula is: ; In the formula, is the time-frequency diagram of the respiratory event feature; is the starting frequency of the respiratory event feature; is the ending frequency of the respiratory event feature; is the target frequency; is the power cumulative distribution feature of a single respiratory cycle; The ear canal geometry identifier is composed of the power cumulative distribution characteristics corresponding to each target frequency; Extract the Mel frequency cepstral coefficients from the respiratory event characteristics, and form the respiratory tract descriptor by the Mel frequency cepstral coefficients and their first-order and second-order derivatives.

6. The method for user authentication based on breath sounds according to claim 5, wherein, The triple neural network includes three sub-network models; three groups of convolutional units and three fully connected layers are connected in sequence to form the sub-network model; within the three groups of convolutional units, a two-dimensional convolutional block and a max pooling layer are configured to be connected in sequence; The training process of the triple neural network includes: Obtain the respiratory audio training signals of multiple users from the database, and multiply the audio waveforms of the respiratory audio training signals by random factors to adjust the volume of the respiratory sound; Change the duration and speed of the respiratory sound in the respiratory audio training signal; add Gaussian noise to the respiratory audio training signal, and move the respiratory audio training signal along the time domain to obtain the respiratory audio training samples; Select a sample anchor point from the respiratory audio training samples, set the respiratory audio training samples with the same user's respiratory sound as the sample anchor point as positive samples, otherwise set the respiratory audio training samples as negative samples; Input the sample anchor, positive sample, and negative sample into the triple neural network, and the triple neural network outputs the intra-class distance between the sample anchor and the positive sample and the inter-class distance between the sample anchor and the negative sample; Calculate the training loss value according to the intra-class distance and the inter-class distance, and the expression formula is: ; In the formula, is the training loss value, is the sample anchor point, is the positive sample; is the negative sample; is the minimum distance between the positive sample and the negative sample; is the within-class distance between the sample anchor point and the positive sample; is the between-class distance between the sample anchor point and the negative sample; Optimize the weight parameters of the triple neural network according to the training loss value, and repeat the training process of the triple neural network until the set number of iterations is reached, and output the trained triple neural network.

7. The method for user authentication based on breath sounds according to claim 6, wherein Input the biometric identifier into the pre-trained triple neural network, and calculate the similarity between the biometric identifier and the pre-stored biometric template; Perform user identity authentication according to the similarity, specifically including: Input the biometric identifier into the pre-trained triple neural network, and calculate the similarity between the biometric identifier and the pre-stored biometric template, and the expression formula is: ; In the formula, is the similarity between the biometric identifier and the k-th biometric template; is the sub-network model; is the body asymmetry identifier; is the ear canal geometry identifier; is the respiratory tract descriptor; is the body asymmetry identification feature in the k-th biometric template; is the ear canal geometry identification feature in the k-th biometric template; is the respiratory tract description feature in the k-th biometric template; Judge the user identity corresponding to the k-th biometric template to which the biometric identifier belongs according to the similarity; judge that the biometric identifier belongs to a non-legitimate user according to the similarity.

8. A user identity authentication system based on breath sounds, characterized in that, Include: An acquisition module, configured to capture the ear respiration signal from the ear canal through an ear-worn device, and perform denoising processing on the ear respiration signal to obtain a respiration denoised signal; A segmentation module, configured to calculate the probability that each frame of the respiration denoised signal belongs to a respiration event and a non-respiration event respectively by using a Gaussian mixture model to obtain a respiration prior probability and a non-respiration prior probability, and calculate the likelihood ratio of each frame of the respiration denoised signal through the respiration prior probability and the non-respiration prior probability; segment the respiration denoised signal according to the likelihood ratio to obtain a respiration segment signal; An identification module, configured to extract a biometric identifier from the respiration segment signal; input the biometric identifier into the pre-trained triple neural network, and calculate the similarity between the biometric identifier and the pre-stored biometric template; perform user identity authentication according to the similarity; The biometric identifier includes a body asymmetry identifier, and the identification module extracts the body asymmetry identifier from the respiration segment signal; specifically including: Extract the start timestamp and end timestamp of the respiration segment signal, and calculate the time difference of the respiration segment signal, and the expression formula is: ; In the formula, is the time difference of the respiration segment signal; is the end timestamp of the respiration segment signal; is the start timestamp of the respiration segment signal; If the time difference of the respiratory segment signal exceeds a preset respiratory time threshold , determine that the respiratory segment signal is a valid respiratory segment; otherwise, delete the respiratory segment signal; Obtain a plurality of consecutive adjacent valid respiration segments, and calculate the time interval between the start timestamp of the current valid respiration segment and the end timestamp of the previous adjacent valid respiration segment, denoted as the first time interval, and the expression formula is: ; In the formula, is the start timestamp of the currently valid breathing segment; is the end timestamp of the adjacent previous valid breathing segment; is the first time interval; Calculate the time interval between the end timestamp of the current valid respiration segment and the start timestamp of the next adjacent valid respiration segment, denoted as the second time interval, and the expression formula is: ; In the formula, is the second time interval; is the end timestamp of the current valid breathing segment; is the end-start timestamp of the adjacent valid breathing segment behind; If the first time interval is greater than the second time interval, determine that the valid respiration segment is an inhalation process; otherwise, determine that the valid respiration segment is an exhalation process; form the valid respiration segments of the inhalation process and the valid respiration segments of the adjacent exhalation process into the respiratory event characteristics of a single cycle; The respiratory event features are divided into left-channel respiratory event features and right-channel respiratory event features; the cross-power spectral density is calculated based on the left-channel respiratory event features and the right-channel respiratory event features; a body asymmetry identifier is obtained by calculating from the cross-power spectral density and the auto-power spectral of the right-channel respiratory event features, and the expression formula is: ; In the formula, is the body asymmetry identifier; is the cross-power spectral density; is the auto-power spectrum of the right-channel respiratory event feature; The segmentation module segments the respiratory denoised signal according to the likelihood ratio to obtain a respiratory segment signal, specifically including: An adaptive threshold is calculated according to the mean and variance of the sub-bands in each frame of the respiratory denoised signal, and the expression formula is: ; In the formula, is the adaptive threshold, represents the threshold weight of the th sub - frequency band; represents the mean value of the th sub - frequency band in the non - breathing mode; represents the mean value of the th sub - frequency band in the breathing mode; represents the variance of the th sub - frequency band in the non - breathing mode; represents the variance of the th sub - frequency band in the breathing mode; W is the number of divided sub - frequency bands; Compare the likelihood ratio of the frame respiratory denoised signal with the adaptive threshold to determine whether each frame of the respiratory denoised signal belongs to a respiratory event or a non-respiratory event, and delete the frames belonging to the non-respiratory event in the respiratory denoised signal to obtain the respiratory segment signal.

9. An electronic device, comprising a storage medium and a processor; the storage medium is used for storing instructions; characterized in that, The processor is configured to operate according to the instruction to execute the user identity authentication method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • User authentication method and related equipment

    CN115982686A