A biometric-based intelligent wearable device identity authentication method and system
The noise is removed through spectral subtraction and Wiener filter, combined with convolutional neural network and gated loop unit to extract speech features, solve the problem of verification failure of smart wearable devices under specific conditions, and achieve more accurate speech pattern recognition and data output accuracy.
Patent Information
- Application Number
- CN202510317804.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-18
AI Technical Summary
Existing smart wearable devices rely on fingerprint or facial recognition to fail under specific conditions to verify, and it is difficult to dynamically adjust abnormal data, resulting in limited processing capabilities when data quality is impaired.
Spectral subtraction and Wiener filter are used to remove background noise of speech samples, and time spectrum features are extracted using convolutional neural network, combined with the gated loop unit to learn the intonation and rhythm of speech data, and analyze sensor data through an isolated forest algorithm, and optimize data processing is applied to adaptive filtering and cross-validation.
It improves the user's voice capture and parsing ability, ensures the accuracy of data output and equipment health assessment, and optimizes the accuracy of identity authentication and data processing accuracy.
Smart Images

Figure CN119851672B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biometric technologies, and particularly to an identity authentication method and system for intelligent wearable devices based on biometric features. Background Art
[0002] Biometric technology is a technical field that uses an individual's physiological or behavioral characteristics for identity verification. These characteristics include, but are not limited to, fingerprints, faces, irises, voices, gestures, and gaits, etc. Due to the uniqueness and non-replicability of these characteristics, biometric technology provides a highly secure and convenient user identity verification method. However, existing intelligent wearable devices rely on fingerprint or face recognition, which may lead to verification failures under specific conditions, such as insufficient light. In addition, it is difficult to dynamically adjust abnormal data, resulting in limited processing capabilities when data quality is damaged. Therefore, improvements are needed. Summary of the Invention
[0003] The purpose of the present invention is to solve the drawbacks existing in the prior art, and to propose an identity authentication method and system for intelligent wearable devices based on biometric features.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions. An identity authentication method for intelligent wearable devices based on biometric features includes the following steps:
[0005] Collect a user's voice sample, remove background noise from the voice sample using spectral subtraction and Wiener filter technology to generate a denoised voice sample; apply a convolutional neural network to the denoised voice sample to separate the time series and spectral information of the voice signal, and extract a time-frequency feature vector;
[0006] Input the time-frequency feature vector into a gated recurrent unit, strengthen the learning of intonation and rhythm in the voice data through sequence analysis to generate a voice dynamic parsing model; verify the user's identity through the voice dynamic parsing model to obtain a user identity verification result;
[0007] Based on the user identity verification result, collect the sensor data stream of the intelligent wearable device, use the isolation forest algorithm to perform an independence analysis on each data point in the data stream, identify the data points that deviate from the normal range, and generate an abnormal data identification result; based on the abnormal data identification result, adjust the data output according to the normal data, correct the abnormal data, and output an automatically corrected data result;
[0008] Apply adaptive filtering to the automatically corrected data result, and use cross-validation to verify the automatically corrected data result to obtain a device health status result.
[0009] Preferably, the step of obtaining the denoised voice sample is as follows:
[0010] Collect the user's voice samples, apply frequency domain analysis through spectral subtraction to identify the interference of non-speech background noise, and generate preliminarily denoised voice samples;
[0011] Based on the preliminarily denoised voice samples, use a Wiener filter for processing, optimize the voice signal through spectrum reconstruction, and obtain further denoised voice samples;
[0012] Based on the further denoised voice samples, perform voice quality detection and adjustment, conduct signal quality assessment, and obtain denoised voice samples.
[0013] Preferably, the steps for obtaining the time-frequency spectrum feature vector are as follows:
[0014] Perform in-depth feature analysis on the denoised voice samples, use a convolutional neural network to process the voice signal, and separate the time series and spectrum information in the voice signal through convolutional layers and pooling layers to obtain time series and spectrum separation data;
[0015] Based on the time series and spectrum separation data, use a multi-layer perceptron to analyze and identify the relationship between each time point and frequency component, extract key time-frequency spectrum features, and generate a time-frequency spectrum feature vector.
[0016] Preferably, the steps for obtaining the voice dynamic parsing model are as follows:
[0017] Input the time-frequency spectrum feature vector into a gated recurrent unit, perform initial sequence analysis on the input data, convert it into a serialized format, and generate serialized feature data;
[0018] Based on the serialized feature data, learn and analyze the intonation and rhythm in the voice data, repeatedly iterate the feature processing process through cyclic connections, adjust the data representation, and construct a voice dynamic parsing model.
[0019] Preferably, the steps for obtaining the user authentication result are as follows:
[0020] Use the voice dynamic parsing model to process the time-frequency spectrum feature vector, identify the user's voice pattern, and generate preliminary authentication data;
[0021] Based on the preliminary authentication data, conduct statistical analysis, calculate the identity recognition rate, and the calculation formula is:
[0022] ;
[0023] Wherein, represents the identity recognition rate, is the th element of the user feature vector, is the average value of the th element in the feature vector, is the standard deviation of the th element, represents the contribution rate of each element to identity recognition, and m represents the total number of elements in the feature vector;
[0024] Based on the identity recognition rate, evaluate the matching degree with the preset standard. If the identity recognition rate exceeds the preset standard, confirm the user's identity; otherwise, reject the access to obtain the user identity authentication result.
[0025] Preferably, the steps for obtaining the abnormal data identification result are:
[0026] Based on the user identity authentication result, activate the sensors of the smart wearable device to collect physiological and environmental data and generate a real-time sensor data stream;
[0027] Based on the real-time sensor data stream, apply the Isolation Forest algorithm for anomaly detection and calculate the anomaly index of each data point. The calculation formula is:
[0028] ;
[0029] where represents the anomaly index of the data point , is the eigenvalue of the data point in the th dimension, is the importance coefficient of this dimension feature, is the feature dimension;
[0030] Based on the anomaly index, evaluate the degree of deviation from the normal range for each data point, and identify the data points with anomaly index higher than the threshold as abnormal to generate the abnormal data identification result.
[0031] Preferably, the steps for obtaining the automatically corrected data result are:
[0032] Based on the abnormal data identification result, apply the Kalman filter algorithm to identify the abnormal data points and separate them from the data stream of the smart wearable device to generate a preliminarily cleaned data stream;
[0033] Based on the preliminarily cleaned data stream, calculate the smoothed value of each data point. The calculation formula is:
[0034] ;
[0035] where represents the smoothed value of the data point , and respectively represent the data points before and after the data point , and are weighting factors, representing the data point range;
[0036] Based on the smoothed value, the data stream is readjusted to obtain an automatically corrected data result.
[0037] Preferably, the step of obtaining the device health status result is as follows:
[0038] Apply adaptive filtering to process the automatically corrected data result to obtain an optimized data stream;
[0039] Based on the optimized data stream, by comparing the performance between data segmentation sets, analyze and evaluate the operating status of the smart wearable device, determine the health status of the device, and generate a device health status result.
[0040] The present invention provides an identity authentication system, including:
[0041] A voiceprint acquisition module, which receives a voice sample from a user, extracts time series and spectral information therefrom, and generates a time-frequency spectral feature vector;
[0042] A voiceprint analysis module, which inputs the time-frequency spectral feature vector into a sequence processing unit, analyzes intonation and rhythm, and generates a voice dynamic analysis model;
[0043] An identity verification module, which uses the voice dynamic analysis model to compare with a stored user model, verifies voice consistency, and obtains a user identity verification result;
[0044] A data anomaly detection module, which analyzes the sensor data stream of the smart wearable device according to the user identity verification result, identifies abnormal data points, and generates an abnormal data identification result;
[0045] A data correction and health assessment module, based on the abnormal data identification result, adjusts and corrects the data output, applies a verification process to the automatically corrected data, and obtains a device health status result.
[0046] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0047] The present invention uses spectral subtraction and Wiener filter to remove background noise from voice samples, and combines convolutional neural network to extract time-frequency spectral features, improving the ability to capture and analyze user voice. In addition, the gated recurrent unit deeply learns the intonation and rhythm of voice data, making the recognition of voice patterns more accurate. The isolation forest algorithm performs independence analysis on sensor data, can identify and correct abnormal data, and ensure the accuracy of data output. Through adaptive filtering and cross-validation techniques, the accuracy of data processing is improved, and the health status assessment of the device is optimized. Brief Description of the Drawings
[0048] Figure 1 It is a schematic diagram of the steps of the present invention;
[0049] Figure 2 It is a schematic diagram of the steps for obtaining the denoised speech sample in the present invention;
[0050] Figure 3 It is a schematic diagram of the steps for obtaining the time-frequency spectrum feature vector in the present invention. Detailed implementation manners
[0051] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] Please refer to Figures 1 - 3 , the present invention provides a technical solution, an identity authentication method for an intelligent wearable device based on biometric features, including the following steps:
[0053] Collect the user's speech sample, remove background noise from the speech sample using spectral subtraction and Wiener filter technology to generate a denoised speech sample; apply a convolutional neural network to the denoised speech sample to separate the time series and spectral information of the speech signal and extract the time-frequency spectrum feature vector;
[0054] Input the time-frequency spectrum feature vector into a gated recurrent unit, strengthen the learning of the intonation and rhythm in the speech data through sequence analysis to generate a speech dynamic parsing model; verify the user's identity through the speech dynamic parsing model to obtain the user identity verification result;
[0055] Based on the user identity verification result, collect the sensor data stream of the intelligent wearable device, use the isolation forest algorithm to perform independence analysis on each data point in the data stream, identify the data points that deviate from the normal range, and generate an abnormal data identification result; based on the abnormal data identification result, adjust the data output according to the normal data, correct the abnormal data, and output the automatically corrected data result;
[0056] Apply adaptive filtering to the automatically corrected data result and use cross-validation to verify the automatically corrected data result to obtain the device health status result.
[0057] The steps for obtaining the denoised speech sample are as follows:
[0058] Collect the user's speech sample, apply frequency domain analysis through spectral subtraction to identify the interference of non-speech background noise, and generate a preliminarily denoised speech sample;
[0059] Based on the preliminarily denoised speech samples, the Wiener filter is used for processing, and the speech signal is optimized through spectrum reconstruction to obtain further denoised speech samples;
[0060] Based on the further denoised speech samples, the sound quality is detected and adjusted, and the signal quality is evaluated to obtain the denoised speech samples.
[0061] Specifically, first refer to the audio sampling and background noise reference data, set a noise energy threshold according to the average noise energy level continuously recorded in the actual sampling environment. This threshold is obtained by performing spectrum statistics on the noise signals collected in different scenarios and taking their average value. Then, this threshold is used as a basic parameter for determining non-speech components during the calculation process. Next, the collected original speech signal is segmented using a fixed frame length slicing method. For example, each frame takes 1024 samples and there is an overlap area of 512 samples between adjacent frames. The segmented signals are subjected to a fast Fourier transform to obtain the energy distribution in each frequency band. The estimated noise energy is subtracted from these frequency band energies through spectral subtraction to achieve the process of weakening the noise components. At the same time, compare the relationship between the remaining energy after denoising in each frequency band and the threshold. If the remaining energy is lower than the threshold, it is determined as invalid non-speech interference and this component is excluded during the inverse transform, thereby reducing the proportion of unnecessary background noise. And further check the abnormal frames during the analysis process according to the preset valid range. For example, mark the frames with the average energy deviating from the 0 to 100 unit interval as abnormal. Then summarize all the processed segmented data to reconstruct the time-domain signal and obtain the preliminarily denoised speech samples.
[0062] After obtaining the preliminarily denoised speech samples, calculate the gain coefficient of the Wiener filter based on the previous statistical results of the noise power spectrum and the speech power spectrum , where can be obtained by comparing the energy distributions of speech and noise at different frequencies. The specific method is to first record the energy mean of speech in each frequency band and the energy mean of noise in the same frequency band over a period of time, and set the ratio of these two energy means as the reference for the filter gain. Obtain a filter step size according to experience or experiments , when is larger, it indicates that the speech component dominates in the corresponding frequency band. If is too small, it means that the noise dominates in this frequency band. Continuously adjust according to the previously obtained energy ratio and observe whether gain oscillation occurs. If it is still in a frequent fluctuation state after repeated iterations, then appropriately reduce Perform smoother adjustments. After the parameters are stabilized, obtain the filtered speech signal through spectrum reconstruction. During this period, if it is found that the energy of certain frequency bands still exceeds the previously measured noise mean plus a 10% redundancy interval, it is regarded as potential noise residue and the gain coefficient of the corresponding frequency band is adjusted again. Finally, after multiple iterations, an ideal spectrum balance is achieved, and a further denoised speech sample is obtained.
[0063] After obtaining the further denoised speech sample, first detect the distortion degree contained in each frequency band with reference to the previously reserved sound quality measurement rules. Express the sound quality distortion degree by dividing several levels within the range of 0 to 1. For example, a value less than 0.2 is regarded as a very low distortion range, 0.2 to 0.5 is a medium distortion range, and a value greater than 0.5 is determined as a high distortion range. If the distortion index corresponding to certain time periods exceeds 0.5, fine-tuning is performed in combination with the waveform change trend of the current speech signal in the time domain. During this adjustment process, linear interpolation can be used to interpolate and trim the frames with mutations. Subsequently, signal quality evaluation is carried out according to the average energy distribution, distortion metric, and frame-to-frame smoothness obtained previously. During evaluation, the distortion indices of each segmented speech are compared with the pre-set quality determination threshold. If most of the distortion indices are less than 0.3, it is determined that the speech sample meets the quality standard. If there are still many frames falling within the range of 0.3 to 0.5, these frames are further adjusted and the distortion indices are recalculated. After all evaluations are completed, the trimming and detection results are integrated together to obtain the denoised speech sample.
[0064] The steps for obtaining the time-frequency spectrum feature vector are as follows:
[0065] Perform in-depth feature analysis on the denoised speech sample, use a convolutional neural network to process the speech signal, and separate the time series and spectrum information in the speech signal through the convolutional layer and pooling layer to obtain time series and spectrum separation data;
[0066] Based on the time series and spectrum separation data, use a multi-layer perceptron to analyze and identify the relationship between each time point and frequency component, extract key time-frequency spectrum features, and generate a time-frequency spectrum feature vector.
[0067] Specifically, for in-depth feature analysis of the denoised speech samples, first set the batch size to 32 according to the previously recorded speech data scale and sampling frame length, select a learning rate of 0.001, prepare a training set containing a number of labeled samples and distinguish the validation set and the test set. Then, by reading the speech data in the training set and organizing it into a multi-channel input format, when using a convolutional neural network to process the speech signal, divide several convolutional kernel sizes to cover different time windows and frequency ranges. When defining the number of convolutional kernels, refer to the previously accumulated sample diversity statistics. Initially, 64 convolutional kernels can be set as the initial value. If it is observed that the convergence process of the model is relatively slow after several training rounds, moderately increase the number of convolutional kernels and record the change in training error. Each convolution operation will output the corresponding feature map. When performing downsampling on the feature map through the pooling layer, a 2×2 pooling window can be selected and the stride can be set to 2. If insufficient local time resolution is found in the validation set, try reducing the pooling window to 1×2 to balance the computational cost and the degree of feature retention. During this period, monitor the loss value of each training round and compare it with the pre-given target loss threshold, which is empirically determined after compromising between the speech recognition accuracy and the model complexity. If the loss value remains higher than this threshold in 10 consecutive training rounds, increase the upper limit of the training rounds and repeat the parameter update. When the loss value is lower than this threshold and remains stable in several additional trainings, it is considered that the model training is completed. After the last training, perform the recognition accuracy verification on the test set and distinguish the time series information and spectral information output by the convolutional layer and the pooling layer, respectively summarize and record the feature data at each moment and the corresponding frequency band, and form the time series and spectral separation data that can be used for subsequent analysis.
[0068] Based on time series and spectrum-separated data, first determine the dimensional range of the input layer and combine the time dimension index and frequency component labels into a multi-dimensional vector. Then, when using a multi-layer perceptron, it is necessary to specify the number of hidden layers and the number of neurons. For example, arrange 128 neurons in the first hidden layer and 64 neurons in the second hidden layer. If signs of overfitting are observed during the training process, several samples can be taken from the previously sorted validation set for cross-checking and a dropout rate between 0.2 and 0.3 can be introduced in the corresponding hidden layer. Subsequently, set the loss function to cross-entropy and use the backpropagation method based on stochastic gradient descent to update the weights and biases. After each update, compare the difference between the training loss and the validation loss. If the training loss shows a downward trend while the validation loss continues to rise in multiple iterations, reduce the learning rate to 1 / 10 of the original value and observe whether the overfitting tendency is alleviated in subsequent iterations. To confirm the recognition accuracy of the model at different time periods and different frequency bands, segmented statistics will be performed on several batches of data and compared with the pre-determined accuracy reference value. For example, the range from 0 to 0.8 is determined as the acceptable range, and if it is greater than 0.8, it is considered too high and the hidden layer capacity needs to be fine-tuned additionally. Finally, when the training and validation processes are completed, the important correlation weights between each time point and the corresponding frequency components will be analyzed and recorded. Through this multi-layer perceptron, key time-frequency spectrum features are extracted to generate a time-frequency spectrum feature vector.
[0069] The steps to obtain the voice dynamic analysis model are as follows:
[0070] Input the time-frequency spectrum feature vector into a gated recurrent unit to perform an initial sequence analysis on the input data, convert it into a serialized format, and generate serialized feature data;
[0071] Based on the serialized feature data, learn and analyze the intonation and rhythm in the voice data, repeatedly iterate the feature processing process through cyclic connections, adjust the data representation, and construct the voice dynamic analysis model.
[0072] Specifically, when inputting the time-frequency spectrum feature vector into the gated recurrent unit, first set the length and number of channels of the input sequence according to the previously obtained feature vector dimension range. Subsequently, define the respective thresholds of the update gate and the reset gate during the gating process, and determine whether to adjust these thresholds by performing multiple sequence scans on the training data. For example, initially set the threshold of the update gate to 0.5 and observe the difference between the output of the update gate and the historical hidden state in subsequent iterations. If there is still a large error after 10 iterations, lower the threshold to 0.45 and retrain. If the error shows a downward trend, continue to maintain the threshold and record the trend of the loss value. To ensure that the feature information of some key time periods is not missed when performing the initial sequence analysis on the input data, compare the output of each time step with the input state of the next time step and calculate the sequence consistency index, and compare this index with the pre-established reference range. For example, if the consistency index is in the range of 0 to 0.6, it is determined as a medium range. If it is higher than 0.6, moderately increase the memory factor weight inside the gated unit in subsequent processing. At the same time, to make the sequence analysis produce a stable serialized format, check the stability of the time step output under different speech segments after each round of iterative training. Consider the difference degree between time steps greater than 0.2 as not stable enough and then re-calculate retrospectively. Finally, obtain a stable sequence output format through multiple rounds of iteration and continuous comparison, record it, and generate serialized feature data.
[0073] When learning and analyzing the intonation and rhythm in the speech data based on the serialized feature data, the number of layers of the recurrent connection can be set to 2 to 3 layers and the dimension of the hidden state of each layer can be distinguished. For example, set the hidden state size of the first recurrent layer to 64 and the second recurrent layer to 32. Confirm whether to increase or decrease the recurrent layer by gradually iterating the feature processing on the training set and recording the change of the loss value at different time steps. If it is observed that the training loss fluctuates and cannot drop to the predetermined range in multiple iterations, check whether the parameter update rate between each layer deviates from the previously recorded numerical range. If the deviation is large, reduce the learning step inside the recurrent unit and compare the change trend of the loss value again. To better reflect the differences in intonation and rhythm in the feature representation, the serialized feature data can be additionally divided into intonation components and rhythm components inside each layer for separate tracking. If it is found that the rhythm component increases abnormally in the later stage of training, increase the number of training rounds by 10% and re-verify in combination with the previously obtained update gate and reset gate thresholds. During this period, fine-tune each gating threshold at a step size of 0.1. If the error curve is stable after several consecutive iterations, retain this state for generating the final output data. After completing all the recurrent connection and iterative feature processing processes, adjust the data representation to construct a speech dynamic parsing model.
[0074] The steps to obtain the user authentication result are as follows:
[0075] Process the time-frequency spectrum feature vector using a voice dynamic parsing model, identify the user's voice pattern, and generate preliminary authentication data;
[0076] Based on the preliminary authentication data, perform statistical analysis and calculate the identity recognition rate. The calculation formula is:
[0077] ;
[0078] Where, represents the identity recognition rate, is the th element of the user feature vector, is the average value of the th element in the feature vector, is the standard deviation of the th element, represents the contribution rate of each element to identity recognition, and m represents the total number of elements in the feature vector;
[0079] Based on the identity recognition rate, evaluate the matching degree with the preset standard. If the identity recognition rate exceeds the preset standard, confirm the user's identity; otherwise, reject the access to obtain the user identity authentication result.
[0080] Specifically, when using the voice dynamic parsing model to process the time-frequency spectrum feature vector and identify the user's voice pattern, it is necessary to first read the previously obtained time-frequency spectrum feature vector and arrange it in the form of consecutive frames. For each frame of data points, use the pre-statistical user pitch range as a reference, compare the acoustic energy concentration with the previously obtained voice feature energy interval. During the comparison, frames with energy lower than 0.01 or higher than 2.5 are classified as the abnormal range and checked. During the process, continuously record the change frequency of the pitch and compare the recording result with the historical mean. For example, compare the pitch with the range of 80 Hz to 300 Hz, compare the speech rate with the range of 2 to 7 syllables per second, map the pronunciation features to the previously obtained timbre parameters, confirm the continuity and stability of the voice pattern through multiple loops and mark the time periods with large deviations. When the cumulative deviation of the recorded values exceeds the reference threshold, re-check the voiceprint components of the specific time period and observe the integrity of the vocal tract spectrum. This reference threshold can be determined according to the pitch and speech rate change ranges in the common voice segments of the user collected previously. After multiple rounds of comparison and cumulative analysis, obtain the judgment data for distinguishing abnormal pronunciation from normal pronunciation. Through these judgment data, identify the basic pattern of the user's voice features and classify them as an overall feature, and finally output the feature as the preliminary data format for identity authentication and record it to generate preliminary authentication data.
[0081] The advantage of the formula is that it combines the user feature vector with the average value and the standard deviation, and introduces the Item, to achieve a comprehensive consideration of the degree of difference and the weight of the difference impact; The acquisition step of is to read the value from the th element in the time-frequency spectrum feature vector determined previously and normalize it according to the acoustic parameters recorded in the training set. The acquisition step of is to calculate the long-term average value of the corresponding elements in all homologous speech feature vectors and perform floating-point precision correction.
[0082] Calculation process: When is , , , , first calculate the numerator:
[0083] ;
[0084] Then calculate the denominator part:
[0085] ;
[0086] Therefore, the first half of the fraction is:
[0087] ;
[0088] Next, calculate the cumulative product, where , , , so:
[0089] ;
[0090] ;
[0091] ;
[0092] Multiply the three terms:
[0093] ;
[0094] The final overall result is This result shows that the identity recognition rate is 1.08. The value 1.08 can be compared with the recognition threshold defined previously. The larger the value, the higher the degree of coincidence with the user's original features.
[0095] When evaluating the matching degree with a preset standard based on the identity recognition rate, it is necessary to first organize the sequence of identity recognition rates obtained during continuous monitoring and compare it with the historical standard sequence. By recording and summarizing the recognition rates of daily or each voice interaction, a determination range is set according to the previously collected average recognition rate and fluctuation range of the user. For example, a recognition rate lower than 0.5 is regarded as having an obvious anomaly and further comparison of component features is required. A recognition rate in the range of 0.5 to 1.0 is regarded as an acceptable range, and a recognition rate exceeding 1.0 belongs to the high-matching group. If the recognition rate is higher than 1.0 for multiple consecutive times, it is marked as an extremely high-matching situation. Each detection is compared with the previously calculated user baseline distribution. If the cumulative recognition rate is lower than 0.5 for multiple times, the sound source is read again and the pitch, energy concentration, and speech rate are checked to observe whether there is a change in the voice state or other voice anomalies of the current user. The process includes arranging the daily average and peak values of the aforementioned recognition rate records and matching them within the corresponding numerical range. If there is a deviation, the moment when the deviation value exceeds 0.1 is marked and rechecked during the next interaction. Finally, after these comparison operations are completed, the stability degree of the user recognition result is obtained and the final flag is confirmed. Through this process, it can be determined whether the preset standard is met, and when the value exceeds the standard, the user is immediately passed, otherwise access is refused to obtain the user identity verification result.
[0096] The steps for obtaining the abnormal data identification result are as follows:
[0097] Based on the user identity verification result, activate the sensors of the smart wearable device, collect physiological and environmental data, and generate a real-time sensor data stream;
[0098] Based on the real-time sensor data stream, apply the Isolation Forest algorithm for anomaly detection, calculate the anomaly index of each data point, and the calculation formula is:
[0099] ;
[0100] Among them, represents the anomaly index of the data point , is the eigenvalue of the data point in the th dimension, is the importance coefficient of this dimension of the feature, is the feature dimension;
[0101] Based on the anomaly index, evaluate the degree of deviation from the normal range for each data point, mark the data points with an anomaly index higher than the threshold as anomalies, and generate an abnormal data identification result.
[0102] Specifically, when activating the sensors of the smart wearable device based on the user authentication result, it is necessary to first read whether there is valid authorization information in the previously recorded authentication result. Subsequently, before turning on the sensors, check the preset sampling frequency and the number of sampling channels. The sampling frequency can be set within the range of 5 to 10 times per second according to the detection statistics of environmental noise and the device operation power consumption. If it is detected that the sampling frequency still deviates from the set range after multiple tests, inspect the working mode inside the sensors and record the inspection information. Then, select appropriate physiological data and environmental data monitoring categories. For example, in terms of physiological data, divide it into multiple items such as heart rate, body temperature, blood oxygen saturation, etc. for synchronous acquisition. In terms of environmental data, divide it into items such as temperature, humidity, and the concentration of certain chemical components in the air for parallel acquisition. Arrange the above-acquired physiological and environmental data in chronological order and perform index marking. If there is data missing or abnormal jump during the process, check the current power consumption and connection stability of the sensors and collect the data for the corresponding period again. At the same time, record the self-check log of the sensors during the entire acquisition process and compare it with the established sensor accuracy reference values. For example, compare the currently real-time acquired heart rate with the previously set range of 40 to 180 beats per minute, compare the body temperature with the range of 30°C to 45°C, compare the air humidity with the range of 10% to 90%, and compare the harmful gas concentration with the range of 0 ppm to 10 ppm. If a certain monitored value continuously exceeds these ranges, conduct further fault troubleshooting or observe the results of multiple rounds of repeated sampling. After all are satisfied, pack and summarize the multi-channel acquisition results and update them to the real-time sensor data stream to generate the real-time sensor data stream.
[0103] The advantage of the formula is that by accumulating the comprehensive weights of multi-dimensional features and using function to map the result to the interval (0, 1), so as to quantify the abnormality degree of a single data point; The acquisition steps of are to respectively extract the measured values corresponding to the data point under the dimension according to the previously collected physiological data and environmental data, and perform unit conversion and normalization. The acquisition steps of
[0104] Calculation process: If , let , where 0.25 can come from the normalized value of a physiological index (such as body temperature), 0.60 can come from the normalization result of an environmental index (such as humidity), and 0.10 can come from the normalization result of another physiological or environmental index. , and these parameters are obtained by calculating and weighted averaging the abnormal distribution probabilities of multiple groups of sensor data in the early stage. Each weight will float within the range of 0.1 to 0.6 with the distribution of monitoring data and device differences. First, calculate the exponential part:
[0105] ;
[0106] Then substitute it into the exponent:
[0107] ;
[0108] ;
[0109] Then add 1 to get the denominator:
[0110] ;
[0111] Thus:
[0112] ;
[0113] This result indicates that the abnormal index of a certain data point is 0.585. If judged in combination with the predetermined standard, when this value is greater than 0.50, it is listed as abnormal and waiting to be observed. If it is less than 0.50, it is classified into the normal range.
[0114] When evaluating the degree of deviation from the normal range for each data point based on abnormal indicators, it is necessary to first compare the sampled values of each dimension with the benchmark interval. Compare the numerical values of different indicators such as heart rate, body temperature, and humidity obtained from the previous statistics with their corresponding reasonable intervals frame by frame. Then, summarize the abnormal indicators corresponding to each frame and record them in a time series. To ensure the continuity of the entire judgment process, it is also necessary to update the abnormal mark of the current frame after each iterative comparison. By superimposing the abnormal mark on the time series and retrieving the periods where abnormalities continuously occur in multiple frames to observe whether it is a short-term fluctuation or a persistent deviation, track those mark values that are continuously higher than the specified reference threshold for multiple times. For example, set the reference threshold in the range around 0.50, and set this threshold by collecting and statistically analyzing the sensor data deviations in the same type of environmental scenarios in the past. If the abnormal indicator quickly rises above 0.80 within a short period of time, it is regarded as an obvious abnormal situation. Then, compare the heart rate or the component of harmful gas concentration in this period to see if they both exceed their respective normal intervals to corroborate the validity of the sensor readings. If multiple indicators are near the high threshold or have crossed this threshold, label the corresponding period as an abnormal point and attach a timestamp. Identify these abnormal points and list them uniformly during subsequent parsing. At the end of the entire process, finally obtain a judgment list indicating whether the data points are abnormal, and generate the abnormal data identification result.
[0115] The steps to obtain the automatically corrected data result are as follows:
[0116] Based on the abnormal data identification result, apply the Kalman filter algorithm to identify the abnormal data points and separate them from the data stream of the smart wearable device to generate a preliminarily cleaned data stream;
[0117] Based on the preliminarily cleaned data stream, calculate the smoothed value of each data point. The calculation formula is:
[0118] ;
[0119] Where, represents the smoothed value of the data point , and respectively represent the data points before and after the data point , and are the weight factors, represents the data point range;
[0120] Based on the smoothed value, readjust the data stream to obtain the automatically corrected data result.
[0121] Specifically, when identifying abnormal data points based on the abnormal data identification results and applying the Kalman filter algorithm, it is necessary to first read from the previously obtained abnormal data identification results which data points are marked as abnormal, and then locate these abnormal data points and their corresponding timestamps during the Kalman filter processing stage. During this process, the variance of the sensor measurement noise needs to be compared with the historical measurement error data, and the prediction and update steps in the filtering process are derived through the sensor state transition relationships recorded multiple times. At the same time, each data point needs to be analyzed. If it deviates significantly from the statistical mean interval or still maintains a deviation trend after multiple repeated measurements, it is compared again to confirm whether it is truly abnormal, and the confirmed abnormal points are removed. During this process, attention should also be paid to whether the sampling accuracy and sampling frequency of the device itself are within the established range. For example, the sampling frequency should be maintained between 5 and 10 times per second, and if it deviates from this range, the credibility of the filtering needs to be re-evaluated. After confirmation, the normal points and abnormal points are distinguished, and the abnormal points are removed. All normal data is re-aggregated in chronological order to obtain a preliminarily cleaned data stream.
[0122] The benefit of the formula is to use the weighted values of the adjacent data points on both the left and right sides to balance the interference factors of the data points and achieve the balance of the weighted results with the help of the denominator ; The acquisition step of is to first determine the index of the data point in the time series and extract the data points before and after it and The acquisition step of and is to read the true sampling values at adjacent positions in the preliminarily cleaned data stream. The acquisition step of
[0123] Calculation process:
[0124] Let , select the adjacent points of the data point at the sequence position index. Assume that through device observation and recording, it is obtained:
[0125] ;
[0126] ;
[0127] The weight factor is determined based on the mean square error of the most recent 100 samples and the statistical range of adjacent value fluctuations. If the following weights are obtained:
[0128] ;
[0129] where each may be slightly adjusted with data fluctuations within the range of 0.2 to 0.5, but fixed values are taken here for example. Then, the left and right summations are calculated in sequence:
[0130] ;
[0131] ;
[0132] Add these two together:
[0133] ;
[0134] Then divide by :
[0135] ;
[0136] This result indicates that the smoothed value of the data point at this moment is 0.3503. If the comparison with the previously collected average water level value of the device is around 1.00, it means that the data point deviates at this time, and in subsequent smoothing correction, this deviation will return to the nearby reference value in a gradually transitional manner.
[0137] When re - adjusting the data stream based on the smoothed value, it is necessary to first continuously retrieve the smoothed value of each data point, arrange the smoothed values corresponding to all data points in chronological order and compare them with the original data point values. If the difference in the smoothed values between some data points in adjacent frames exceeds the set uniform transition threshold, re - segment and further observe the amplitude change of this segment of data. If the amplitude continuously deviates from the reference range, compare it with the previously collected sensor noise characteristics to see if there is a reason for periodic jitter. For each analysis, the final mean and variance of the difference can be statistically obtained through multiple iterative smoothing operations and compared with thresholds such as 0.1 or 0.2. This threshold is set by the device recording the sensor drift rate and output instability during actual operation. If multiple observations are within the range, this batch of data is included in the available list. If it still exceeds, it is prompted to check whether it is caused by sensor failure or a sudden change in the environment. Finally, after globally numbering the smoothing correction results of each time period, they are summarized into a complete and smoothly connected data sequence to obtain the automatically corrected data result.
[0138] The steps to obtain the device health status result are as follows:
[0139] Apply adaptive filtering processing to automatically correct the data results and obtain an optimized data stream;
[0140] Based on the optimized data stream, analyze and evaluate the operating conditions of the smart wearable device by comparing the performance between data segmentation sets, determine the health status of the device, and generate the device health status result.
[0141] Specifically, when applying adaptive filtering processing to automatically correct the data results, it is necessary to first distinguish the time series of continuous sampling in the previously obtained automatically corrected data results, and select to track the dynamic characteristics of the noise as the adjustment basis for adaptive filtering. During this period, an error reference value for measuring the noise change amplitude can be set for each observation window. For example, the mean square error range from the recent 50 samplings is included in the calculation and this reference value is updated in real time during the comparison. If the mean square error in several windows deviates from this reference value by more than a specified amplitude, the filter parameters are adjusted. In specific implementation, each data point is compared with the historical observation value and the filter gain is increased or decreased in combination with the error signal. If the trend of the error signal still deviates from the preset threshold interval after multiple increases, it is necessary to re-evaluate the sampling environment and observe the working mode of the sensor. The reference threshold interval usually comes from the error distribution situation formed by multiple samplings of the same device under normal temperature and pressure. By calculating the distribution mean and dispersion, an upper and lower limit is obtained and recorded for real-time adjustment. When the error signal is lower than this upper and lower limit for multiple consecutive frames, it indicates that the current filter gain is appropriate. If not, the filter weight coefficient is updated and the iteration is performed again. During the whole process, the time series generated after filtering is also compared segment by segment with the unfiltered sequence before. For example, each data segment is split into several groups and whether there is a significant deviation at the same time point is compared. If a single-point deviation exceeds the predetermined standard, further track whether there is sensor noise drift or sudden abnormality in the original data. After completing the filtering and correction of all data segments, the filtered sequences are spliced together to obtain a time series data with higher smoothness and lower noise influence, and the optimized data stream is obtained.
[0142] When optimizing the data stream and comparing the performance between data subsets, it is necessary to first divide the optimized data stream into several subsets according to different time periods or different load states. Then, statistical features that match the normal operating indicators of the device are recorded for each subset. For example, in physiological sensor measurements, data such as heart rate or body temperature can be grouped and the range of their changes over time can be statistically analyzed. The average value of each group is compared with the historical normal value. If it is found that the average value of multiple samples deviates from the set range within a short period of time, this phenomenon is recorded and the data details are traced further. The performance comparison between data subsets can adopt multiple comparison methods. For example, quantitative analysis is carried out on the peak values, valley values, and change rates within each subset, and the cases where the change rate exceeds a reference value such as 0.2 or 0.3 are marked. The reference value is statistically obtained from the historical volatility distribution collected previously and will be increased or decreased by a certain proportion according to the usage scenario of the device. Then, the marked results of each subset are integrated and arranged. If it is found that the abnormal frequency of some subsets is too high or the vibration amplitude increases significantly, further queries are made to check whether there is an association with the external environment. A similar method can also be adopted when cross-comparing indicators such as pressure, temperature, current, and voltage. Each indicator has a specific reasonable range, such as pressure comparison from 0 MPa to 2 MPa, current comparison from 0 A to 5 A, etc. Each cycle of comparison needs to monitor whether it crosses this range and the number of crossings. Finally, the overall operating condition of the smart wearable device is analyzed based on the distribution and fluctuation characteristics of different indicators between subsets, and corresponding judgments are made to generate the device health status result.
[0143] An identity authentication system, comprising:
[0144] A voiceprint acquisition module, which receives a voice sample from a user, extracts time series and spectrum information therefrom, and generates a time-frequency spectrum feature vector;
[0145] A voiceprint analysis module, which inputs the time-frequency spectrum feature vector into a sequence processing unit, analyzes intonation and rhythm, and generates a voice dynamic analysis model;
[0146] An identity verification module, which uses the voice dynamic analysis model to compare with the stored user model, verifies voice consistency, and obtains a user identity verification result;
[0147] A data anomaly detection module, which analyzes the sensor data stream of the smart wearable device according to the user identity verification result, identifies abnormal data points, and generates an abnormal data identification result;
[0148] A data correction and health assessment module, which adjusts and corrects the data output based on the abnormal data identification result, applies a verification process to the automatically corrected data, and obtains a device health status result.
[0149] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the relevant art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. An identity authentication method for intelligent wearable devices based on biometrics, characterized in that It includes the following steps: Collect user voice samples, remove background noise from the voice samples using spectral subtraction and Wiener filter techniques to generate denoised voice samples; apply a convolutional neural network to the denoised voice samples to separate the time series and spectral information of the voice signal and extract time-frequency feature vectors; Read the obtained time-frequency feature vectors, verify the user's identity, and obtain the user identity verification result; Based on the user identity verification result, collect the sensor data stream of the intelligent wearable device, use the Isolation Forest algorithm to perform independence analysis on each data point in the data stream, identify the data points that deviate from the normal range, and generate an abnormal data identification result; Based on the abnormal data identification result, adjust the data output according to the normal data, correct the abnormal data, and output the automatically corrected data result; Apply adaptive filtering to the automatically corrected data result and use cross-validation to verify the automatically corrected data result to obtain the device health status result; The steps for obtaining the user identity verification result are as follows: Read the time-frequency feature vectors obtained previously and arrange them in the form of consecutive frames. For each frame of data points, use the pre-statistical user pitch range as a reference, compare the acoustic energy concentration with the voice feature energy interval. During the comparison, classify the frames with energy below 0.01 or above 2.5 as the abnormal range and conduct a check. Continuously record the change frequency of the pitch during the process and compare the recording result with the historical mean. Compare the speech rate with the interval of 2 to 7 syllables per second, map the pronunciation features to the timbre parameters, cycle through to confirm the continuity and stability of the speech pattern and mark the time periods with deviations. When the cumulative deviation of the recorded data exceeds the reference threshold, re-check the voiceprint components of the specific time period and observe the integrity of the vocal tract spectrum, identify the basic pattern of the user's voice features and classify it as an overall feature to generate preliminary identity verification data; Based on the preliminary identity verification data, conduct statistical analysis, calculate the identity recognition rate, and the calculation formula is: ; Among them, represents the identity recognition rate, is the -th element of the user feature vector, is the long-term average of the corresponding elements in all homologous speech feature vectors, is the standard deviation obtained by taking the square root of the variance calculated for the corresponding elements in the same set of speech feature vectors, represents the contribution rate of each element to identity recognition, and m represents the total number of elements in the feature vector; Based on the identity recognition rate, evaluate the matching degree with the preset standard. If the identity recognition rate exceeds the preset standard, confirm the user's identity; otherwise, reject the access to obtain the user identity verification result.
2. The biometric-based intelligent wearable device identity authentication method according to claim 1, wherein, The steps for obtaining the denoised voice samples are as follows: Collect user voice samples, apply frequency domain analysis through spectral subtraction to identify the interference of non-speech background noise and generate a preliminarily denoised voice sample; Based on the preliminarily denoised voice sample, use a Wiener filter for processing, optimize the voice signal through spectral reconstruction to obtain a further denoised voice sample; Based on the further denoised voice sample, conduct voice quality detection and adjustment, and perform signal quality assessment to obtain the denoised voice sample.
3. The biometric-based intelligent wearable device identity authentication method according to claim 1, characterized in that, The steps for obtaining the time-frequency feature vectors are as follows: Perform in-depth feature analysis on the denoised voice samples, use a convolutional neural network to process the voice signal, and separate the time series and spectral information in the voice signal through the convolutional layer and pooling layer to obtain the time series and spectral separation data; Based on the time series and spectrum separation data, a multi-layer perceptron is used to analyze and identify the relationship between each time point and frequency component, extract key time-frequency spectrum features, and generate time-frequency spectrum feature vectors.
4. The biometric-based intelligent wearable device identity authentication method according to claim 1, characterized in that, The steps for obtaining the abnormal data identification result are as follows: Based on the user authentication result, activate the sensors of the smart wearable device, collect physiological and environmental data, and generate a real-time sensor data stream; Based on the real-time sensor data stream, apply the Isolation Forest algorithm for anomaly detection, calculate the anomaly index for each data point, and the calculation formula is: ; Among them, represents the abnormal index of the data point, is the eigenvalue of the data point in the dimensional feature value, is the importance coefficient of the feature in this dimension, is the feature dimension; Based on the anomaly index, evaluate the degree of deviation from the normal range for each data point, identify the data points with anomaly indexes higher than the threshold as anomalies, and generate the abnormal data identification result.
5. The biometric-based intelligent wearable device identity authentication method according to claim 1, characterized in that The steps for obtaining the automatically corrected data result are as follows: Based on the abnormal data identification result, apply the Kalman filtering algorithm to identify the abnormal data points and separate them from the data stream of the smart wearable device to generate a preliminarily cleaned data stream; Based on the preliminarily cleaned data stream, calculate the smoothed value for each data point, and the calculation formula is: ; Among them, represents the smoothed value of the data point, and respectively represent the data points before and after the data point, and are weight factors, representing the data point range; Based on the smoothed value, readjust the data stream to obtain the automatically corrected data result.
6. The biometric-based intelligent wearable device identity authentication method according to claim 1, characterized in that, The steps for obtaining the device health status result are as follows: Apply adaptive filtering to process the automatically corrected data result to obtain an optimized data stream; Based on the optimized data stream, analyze and evaluate the operating status of the smart wearable device by comparing the performance between data segmentation sets, determine the health status of the device, and generate the device health status result.
Citation Information
Patent Citations
Mobile internet intelligent wearable device
CN107049280A
Human voice recognition system
CN113077794A
Voiceprint recognition method and device, terminal equipment and computer readable storage medium
CN114765028A
Target behavior monitoring method and system based on electromagnetic wave reflection
CN117214899A
Information monitoring method and system for medical equipment data
CN118861953A