Biological feature recognition method driven by PPG big data

By preprocessing the PPG signal with high-pass and low-pass filtering, fusing the Mel-frequency cepstral coefficients and Gammatone-frequency cepstral coefficients, constructing a Gaussian mixture model and i-vector framework, and using a long short-term memory network for classification, the problem that the PPG signal is susceptible to motion artifacts is solved, the signal quality and recognition accuracy are improved, and the robustness of the system is enhanced.

CN120611260AActive Publication Date: 2025-09-09HUBEI UNIV OF ECONOMICS

Patent Information

Application Number
CN202511103204.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-09-09
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

PPG signals are easily affected by motion artifacts, which degrades the signal quality. Traditional feature extraction methods are difficult to capture individual difference characteristics, and existing recognition models are not robust enough.

Method used

A biometric recognition method driven by PPG big data is adopted, including high-pass filtering and low-pass filtering preprocessing, fusion of Mel-frequency cepstral coefficients and Gammatone frequency cepstral coefficients, construction of the universal background model (UBM) of the Gaussian mixture model, calculation of the Baum-Welch statistic, extraction of the i-vector identity authentication vector, and training and classification through the long short-term memory network.

Benefits of technology

It effectively eliminates baseline drift, motion artifacts and ambient light interference, improves signal quality and feature differentiation capabilities, enhances recognition accuracy and robustness, and adapts to the needs of different individuals and scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611260A_ABST
    Figure CN120611260A_ABST
Patent Text Reader

Abstract

The invention provides a PPG big data driven biological feature recognition method, which comprises the following steps: acquiring PPG signal data of trainees, and performing high-pass filtering and low-pass filtering preprocessing on the PPG signal data to obtain preprocessed PPG signals; framing is carried out on the preprocessed PPG signal, a Mel frequency cepstral coefficient feature and a Gammatone frequency cepstral coefficient feature are extracted respectively, the Mel frequency cepstral coefficient feature and the Gammatone frequency cepstral coefficient feature are fused, then principal component analysis dimension reduction is carried out, and a fusion feature is obtained; a universal background model UBM of a Gaussian mixture model is constructed based on the fusion features, zero-order, first-order and second-order Baum-Welch statistics are calculated, a global difference space matrix is estimated from the Baum-Welch statistics, and i-vector identity authentication vectors are extracted; and inputting the i-vector identity authentication vector into a long short-term memory network for training and classification to obtain an identity recognition result of the trainees. According to the invention, the influence of motion artifacts can be eliminated, the overall variability of PPG signals is captured, and an identity authentication scheme is provided for wearable equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biometric identification technology, and in particular to a biometric identification method driven by PPG big data. Background Art

[0002] Photoplethysmography (PPG) data can accurately reflect vital signs and provide an important reference for identification during training. However, due to the diverse training scenarios, it is susceptible to motion artifacts (MA), which can distort signal fidelity and inhibit the reliable estimation of important parameters.

[0003] At present, the main problems in the field of PPG signal processing are as follows: first, PPG signals are easily interfered by motion artifacts, resulting in a decrease in signal quality; second, traditional feature extraction methods are difficult to effectively capture individual difference characteristics in PPG signals; third, existing recognition models are not robust enough in complex training environments.

[0004] Therefore, there is an urgent need for a biometric recognition method that can effectively eliminate the influence of motion artifacts, accurately capture the individual characteristics of PPG signals and has strong robustness to meet the needs of identity recognition of trainees during training. Summary of the Invention

[0005] The purpose of the present invention is to provide a biometric feature recognition method driven by PPG big data, aiming to solve the problem in the prior art that PPG signals are easily interfered by motion artifacts, resulting in reduced reliability of important parameter estimation.

[0006] To achieve the above objectives, the present invention provides a biometric feature recognition method driven by PPG big data, comprising: obtaining PPG signal data of trainees, performing high-pass filtering and low-pass filtering preprocessing on the PPG signal data to obtain a preprocessed PPG signal; framing the preprocessed PPG signal, extracting Mel-frequency cepstral coefficient features and Gammatone-frequency cepstral coefficient features respectively, fusing the Mel-frequency cepstral coefficient features and Gammatone-frequency cepstral coefficient features, and then performing principal component analysis dimensionality reduction to obtain fused features; constructing a universal background model (UBM) of a Gaussian mixture model based on the fused features, calculating zero-order, first-order, and second-order Baum-Welch statistics, estimating a global difference space matrix from the Baum-Welch statistics, and extracting an i-vector identity authentication vector; inputting the i-vector identity authentication vector into a long short-term memory network for training and classification to obtain an identity recognition result of the trainee.

[0007] Preferably, the performing high-pass filtering and low-pass filtering preprocessing on the PPG signal data includes: configuring filter parameters on the PPG signal data, setting the cutoff frequency of the fourth-order Chebyshev type I high-pass filter to 0.5 Hz, setting the passband fluctuation to within ±0.5 decibels, and the stopband attenuation slope to negative 80 decibels per decade, to obtain high-pass filter parameters; constructing the fourth-order Chebyshev type I high-pass filter based on the high-pass filter parameters, applying the fourth-order Chebyshev type I high-pass filter to the PPG signal data to eliminate baseline drift and obtain a first filtered signal; and evaluating the filtering effect according to the time domain and frequency domain characteristics of the first filtered signal, calculating the signal-to-noise ratio and the degree of baseline drift suppression, to ensure that preset signal quality indicators are met.

[0008] Preferably, the high-pass filtering and low-pass filtering preprocessing of the PPG signal data includes: configuring filter parameters on the first filtered signal, setting the cutoff frequency of the fourth-order elliptical low-pass filter to 18 Hz, setting the passband to 0 Hz to 18 Hz with a fluctuation of less than or equal to 0.1 decibel, setting the stopband to start at 22 Hz and attenuating greater than or equal to 40 decibels at 22 Hz and greater than or equal to 60 decibels at 50 Hz, to obtain low-pass filter parameters; constructing the fourth-order elliptical low-pass filter based on the low-pass filter parameters, applying the fourth-order elliptical low-pass filter to the first filtered signal, eliminating motion artifacts, ambient light interference and muscle noise, and obtaining the preprocessed PPG signal; and evaluating the noise suppression effect and calculating the degree of high-frequency interference suppression based on the spectral characteristics of the preprocessed PPG signal to ensure that preset signal quality requirements are met.

[0009] Preferably, the framing of the preprocessed PPG signal and extracting Mel-frequency cepstral coefficient features and Gammatone frequency cepstral coefficient features respectively include: performing pre-emphasis processing on the preprocessed PPG signal to enhance the high-frequency component to obtain a pre-emphasized signal; framing the pre-emphasized signal according to a preset frame length, adding a Hamming window to each frame signal for smooth transition to obtain a first windowed signal; performing fast Fourier transform on the first windowed signal to obtain a spectrum, inputting the spectrum into a Mel filter bank for frequency envelope extraction to obtain a Mel spectrum; performing discrete cosine transform on the Mel spectrum to compress the feature dimension to obtain the Mel-frequency cepstral coefficient features.

[0010] Preferably, the framing of the preprocessed PPG signal and extracting Mel-frequency cepstral coefficient features and Gammatone frequency cepstral coefficient features respectively include: framing the preprocessed PPG signal and applying a Hamming window to each frame signal to obtain a second windowed signal; performing fast Fourier transform on the second windowed signal to obtain a frequency domain signal; inputting the frequency domain signal into a Gammatone filter group, performing frequency band division according to the ERB frequency scale, and obtaining filtered outputs of each channel; performing energy calculation and logarithmic compression on the filtered outputs of each channel, and performing discrete cosine transform to obtain the Gammatone frequency cepstral coefficient features.

[0011] Preferably, the universal background model UBM of the Gaussian mixture model is constructed based on the fusion feature, the zero-order, first-order and second-order Baum-Welch statistics are calculated, the global difference space matrix is ​​estimated from the Baum-Welch statistics, and the i-vector identity authentication vector is extracted, including: based on the fusion feature, the number of Gaussian mixture components is set to 32, a universal background model UBM is constructed, and its covariance matrix is ​​set to a diagonal matrix to obtain initial model parameters; according to the initial model parameters, the EM algorithm is used to calculate the zero-order, first-order and second-order Baum-Welch statistics, and the global difference space matrix is ​​iteratively estimated until the change in the log-likelihood value is less than a preset threshold, thereby stopping the iteration and obtaining the global difference space matrix; based on the global difference space matrix and the Baum-Welch statistics, the i-vector identity authentication vector with a dimension of 200 is extracted by the maximum a posteriori probability estimation method.

[0012] Preferably, it also includes: obtaining the high-pass filter parameters, the low-pass filter parameters, and the extraction parameters of the Mel-frequency cepstral coefficient features and the Gammatone frequency cepstral coefficient features, constructing a parameter optimization framework based on a genetic algorithm, encoding the high-pass filter parameters, the low-pass filter parameters and the extraction parameters into chromosomes of an initial population; based on the chromosomes of the initial population, calculating the fitness function value of each chromosome, wherein the fitness function value comprehensively evaluates the three indicators of signal-to-noise ratio gain, signal quality index and feature discrimination, and obtains the population fitness distribution; according to the population fitness distribution, retaining the 10% individuals with the highest fitness to enter the next generation, and adopting the brocade method for the remaining 90% individuals. The competition selects crossover and mutation, where the crossover rate is initially set to 0.8 and dynamically adjusted according to the population diversity, and the mutation rate is initially set to 0.1 and gradually decreases to 0.01 as the number of iterations increases to obtain a new generation of population; every 10 generations, the 10% individuals with the highest fitness in the new generation of population are subjected to simulated annealing local optimization until the preset number of iterations is reached to obtain the optimal parameter configuration; based on the optimal parameter configuration, the PPG signal of each trainee is personalized preprocessed and feature extracted to generate optimized PPG signals and features; an incremental optimization mechanism is set up, and when the recognition accuracy drops by more than a preset threshold or reaches a fixed optimization period, the parameter re-optimization is triggered, and the optimization process is accelerated by parallel computing.

[0013] Preferably, it also includes: obtaining the initial fusion weights and dimensionality reduction matrix of the Mel frequency cepstral coefficient features and the Gammatone frequency cepstral coefficient features, constructing a dual-objective optimization problem for feature fusion, setting the first objective to maximize the Fisher discriminant ratio, and the second objective to minimize the feature dimension; based on the dual-objective optimization problem, using the NSGA-II algorithm to generate the initial population, each individual contains two parts: feature fusion weight and dimensionality reduction, wherein the fusion weight range is [0,1] and the sum is 1, and the dimensionality reduction range is [8,20]; performing non-dominated sorting and crowding calculation on the initial population, generating a parent population through binary tournament selection, and using simulation Binary crossover and polynomial mutation generate a population of offspring; a local search is performed on the solutions on the non-dominated front every 10 generations, and the search step size is dynamically adjusted with the number of iterations to accelerate the convergence of the algorithm and obtain a Pareto optimal solution set; based on the Pareto optimal solution set, the optimal solution is selected using a knee point detection method, which first arranges the solution set in ascending order according to the first objective, then calculates the curvature of adjacent solutions, selects the point with the largest curvature as the knee point, and generates a personalized feature fusion model; an adaptive adjustment mechanism is set to monitor recognition performance. When the cumulative error rate exceeds 5% or after 30 days, the feature fusion model is automatically updated to ensure stable performance.

[0014] Preferably, it also includes: constructing an i-vector framework parameter space and a decision model parameter space, wherein the i-vector framework parameter space includes the number of Gaussian mixture components, the rank of the T matrix, the dimension of the eigenvector and the number of iterations, and the decision model parameter space includes the number of layers, the number of hidden units, the dropout rate and the learning rate of the LSTM network; in the i-vector framework parameter space and the decision model parameter space, respectively, the covariance matrix adaptive evolution strategy CMA-ES is used for optimization, and the interactive optimization between subspaces is realized through the three-layer collaborative mechanism of parameter cross-evaluation, fitness sharing and search direction coordination; according to the progress speed and performance improvement of the interactive optimization of the subspace, the calculation time allocated to each subspace is dynamically adjusted. Computing resources are utilized, and multi-threading technology is used to simultaneously evaluate different parameter combinations. For different scenarios in training, typical scenario configurations are pre-defined and the optimization process is executed separately to generate scenario-specific parameter configuration schemes, and parameter configurations are automatically switched according to the training scenarios detected in real time. A regularization control strategy is adopted to add L2 regularization constraints in the i-vector subspace to control the complexity of the T matrix, and batch normalization and early stopping strategies are introduced in the decision model subspace to improve the generalization ability of the model. A two-level strategy of lightweight update strategy and comprehensive update strategy is adopted to achieve incremental optimization, wherein the lightweight update strategy fine-tunes parameters in a local range, and the comprehensive update strategy re-executes the complete co-evolutionary optimization after the accumulated new data exceeds the preset data value.

[0015] Preferably, the i-vector identity authentication vector is input into the long short-term memory network for training and classification to obtain the identity recognition result of the trainee, including: normalizing and time-rearranging the i-vector identity authentication vector to generate a training sample sequence and a corresponding identity label; based on the training sample sequence and the corresponding identity label, a network structure is constructed including an input layer, two LSTM layers and a fully connected layer, wherein the number of LSTM units in the first layer is 128, the number of LSTM units in the second layer is 64, and the dropout rate is set to 0.3 to obtain an initial network model; cross entropy is used as the loss function, and the initial network model is iteratively trained using the Adam optimizer, the initial value of the learning rate is set to 0.001 and dynamically adjusted using the cosine annealing strategy to obtain a trained LSTM network model; the i-vector identity authentication vector to be identified is input into the trained LSTM network model, the probability distribution of each identity category is calculated by the Softmax classifier, and the category with the highest probability is selected as the recognition result.

[0016] The beneficial effects of the present invention are: 1. This invention pre-processes the PPG signal through high-pass and low-pass filtering, effectively eliminating interference such as baseline drift, motion artifacts, ambient light interference, and muscle noise, thereby improving signal quality; 2. This paper adopts the fusion feature of MFCC and GFCC, making full use of the complementarity of the two features. MFCC can capture the overall spectral envelope of the signal, while GFCC emphasizes the finer spectral details of the low-frequency part through the Gammatone filter bank, thereby improving the distinguishing ability of the features. 3. This paper uses the i-vector framework to model PPG signals, compressing feature vector sequences into compact low-dimensional representations, effectively capturing the overall variability of PPG signals and improving recognition accuracy and efficiency. 4. This invention uses a long short-term memory network (LSTM) to make decisions and discriminate i-vectors, which can effectively capture the temporal characteristics of PPG signals and further improve recognition accuracy. 5. The present invention provides a parameter optimization framework based on evolutionary algorithms, a feature fusion strategy based on multi-objective optimization, and an i-vector and model parameter optimization method based on collaborative evolution, enabling the system to adapt to the needs of different individuals, different motion states, and different scenarios, thereby improving the robustness and adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 Flowchart of a biometric feature recognition method driven by PPG big data in an embodiment of the present invention; Figure 2 2 is a comparison diagram of the PPG signal before and after filtering in an embodiment of the present invention; Figure 3 MFCC feature extraction flow chart in an embodiment of the present invention; Figure 4 This is a block diagram of GFCC feature extraction in an embodiment of the present invention; Figure 5 : This is the average energy distribution diagram of MFCC and GFCC in an embodiment of the present invention; Figure 6 This is a graph showing the cumulative contribution rate of each dimension of the MFCC eigenvalues ​​in an embodiment of the present invention; Figure 7 This is a graph showing the cumulative contribution rate of each dimension of the GFCC eigenvalue in an embodiment of the present invention; Figure 8 This is a flow chart of estimating the T matrix using the EM algorithm according to an embodiment of the present invention; Figure 9 Schematic diagram of stratified data sampling in an embodiment of the present invention. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0020] This invention provides a biometric recognition method driven by PPG big data. This method, designed for wristband sensor acquisition scenarios, utilizes the i-vector framework to model the inherent time series of a trainee's heartbeat signal for efficient recognition. The invention is described in detail below with reference to specific embodiments.

[0021] Example 1 like Figure 1 As shown, this embodiment provides a biometric feature recognition method driven by PPG big data, including the following steps: Step S1: Obtain the PPG signal data of the trainee, and perform high-pass filtering and low-pass filtering preprocessing on the PPG signal data to obtain a preprocessed PPG signal.

[0022] Specifically, 30 healthy participants aged 18 to 58 were selected and asked to complete three types of activities in a simulated training scenario. For each participant, PPG signals were obtained from a wrist-worn oximeter equipped with a green LED (wavelength: 609 nm). Each recording lasted approximately 20 minutes and included physiological signals sampled at a frequency of 125 Hz. Real-world data from the simulated training scenario was obtained from the 30 participants using the following methods: Type 1 (T1): 1 to 10 participants participated and completed a certain length and speed of walking or running on a treadmill in the following order: walking at 1 to 2 km / h for 0.5 minutes, jogging at 6 to 8 km / h for 1 minute, running at 15 km / h for 12 to 15 minutes, jogging at 6 to 8 km / h for 1 minute, running at 12 to 15 km / h for 1 minute, and walking at 1 to 2 km / h for 0.5 minutes.

[0023] Type 2 (T2): There are 9 to 22 participants who mainly complete general upper limb exercises, such as handshakes, stretching, pushing, running, jumping and push-ups.

[0024] Type 3 (T3): 21 to 30 participants participate in the training, mainly completing upper limb violent exercises, such as boxing, dagger fighting and pull-ups.

[0025] Since participants 9 and 10 participated in T1 and T2 at the same time, and participants 21 and 22 participated in T2 and T3 at the same time, a total of 34 data records were obtained from this simulation test.

[0026] The PPG signal frequency band ranges from 0.5 to 18 Hz. The main noise sources in PPG recordings include baseline drift (respiration, temperature changes), motion artifact (MA), ambient light interference (power frequency interference), and muscle noise (EMG). To suppress these interferences, multi-stage filtering is implemented: first, a Chebyshev type I high-pass filter with a cutoff frequency of 0.5 Hz is used to eliminate baseline drift.

[0027] The implementation form of a 4th-order Chebyshev type I high-pass filter with a cutoff frequency of 0.5Hz is:

[0028] The coefficient is the filter coefficient, and its specific value is: b=[0.9391,-3.7564,5.6346,-3.7564,0.9391] a=[1.0000,-3.8746,5.6280,-3.6407,0.8873] The corresponding difference equation is:

[0029] The filter's characteristics are: Passband (>0.5Hz): Fluctuation within ±0.5 dB.

[0030] Stop band (<0.5Hz): Attenuation slope ≈ -80 dB / decade (≥40 dB at 0.1Hz).

[0031] The images were then filtered through an elliptical low-pass filter with a cutoff frequency of 18 Hz to remove motion artifacts, ambient light interference (power frequency interference), and muscle noise.

[0032] The implementation form of the 4th-order elliptic low-pass filter with a cutoff frequency of 18 Hz is:

[0033] The coefficient is the filter coefficient, and its specific value is: b = [0.0126, 0.0344, 0.0422, 0.0344, 0.0126] a = [1.0000, -2.6236, 2.7247, -1.3226, 0.2406] The corresponding difference equation is:

[0034] The filter's characteristics are: Passband: 0-18Hz signal retention, fluctuation ≤0.1 dB.

[0035] Stopband: Signals ≥22 Hz are strongly suppressed (≥40 dB at 22 Hz, ≥60 dB at 50 Hz).

[0036] Extremely narrow transition band: Attenuation drops rapidly from -3dB to -40dB in just 4Hz (18Hz→22Hz).

[0037] The final stage of preprocessing includes smoothing to further improve the signal quality, such as Figure 2 shown.

[0038] Step S2: Frame the preprocessed PPG signal, extract Mel-frequency cepstral coefficient features and Gammatone-frequency cepstral coefficient features respectively, fuse the Mel-frequency cepstral coefficient features and Gammatone-frequency cepstral coefficient features, and perform principal component analysis and dimensionality reduction to obtain fused features.

[0039] Specifically, after noise reduction processing, MFCC and GFCC are used for feature extraction as the main analysis features. MFCC is a mature algorithm widely used in the field of signal processing (especially speech recognition). The cepstrum is a homomorphic transformation from convolution to addition, and is scaled according to the Mel scale (the perceptual auditory scale that simulates the human ear's response to audio frequency). The present invention selects the MFCC calculation method based on discrete cosine transform (DCT) to reduce the computational cost. The calculation process is as follows: Figure 3 shown.

[0040] The first step is to input the signal (Indicates the sample point The original PPG signal is pre-emphasized to enhance the high-frequency components. The calculation formula is:

[0041] Where 0.97 is the pre-emphasis coefficient, which is often used in signal processing to enhance high-frequency components and reduce low-frequency dominance.

[0042] Then the pre-emphasized signal Split into short time frames of equal length (e.g. 25 ms), and add a window to each frame to reduce spectral leakage and ensure smooth transition. The calculation formula is:

[0043] Where N is the number of samples per frame, and the expression in parentheses defines the Hamming window applied to each frame.

[0044] Finally, each frame of windowed signal h[n] is converted to the frequency domain through fast Fourier transform (FFT):

[0045] In the formula Indicates the The complex spectrum of frequency bins.

[0046] This step calculates the Mel spectrum by inputting the Fourier transform signal amplitude into a Mel filter bank (a set of triangular bandpass filters). The expression is:

[0047] It is A triangular filter, and Respectively represent The lower cutoff frequency and upper cutoff frequency of the filter. Apply discrete cosine transform (DCT) to generate Mel-frequency cepstral coefficients (MFCC):

[0048] In the formula Indicates the MFCC coefficients, The total number of coefficients selected for the eigenvector.

[0049] Traditional MFCCs (Mel-Frequency Cepstral Coefficients) are designed based on the Mel-scale of human hearing. However, the frequency domain characteristics of PPG signals (with energy concentrated in the 0.5-5 Hz range) are significantly different from speech signals. This paper introduces GFCCs (Gammatone Frequency Cepstral Coefficients), the core of which is the simulation of the frequency selectivity of the cochlear basilar membrane by the Gammatone filter bank. The Gammatone filter has higher frequency resolution in the low-frequency band, and its transfer function is The pole-zero distribution characteristics of the GFCC are more suitable for the narrowband spectral structure of the PPG signal. By calculating the energy of each filter channel and performing logarithmic compression, the DCT transform further extracts a compact representation in the cepstral domain. The resulting GFCC coefficients simultaneously retain the key time-frequency domain information of the PPG signal.

[0050] The impulse response of the gammatone filter is defined as:

[0051] in is the amplitude gain, is the filter order, represents the filter bandwidth, is the center frequency in Hertz (Hz), is the phase offset. The center frequencies are equally spaced within the filter bank boundaries according to the equivalent rectangular bandwidth (ERB) scale. The gammatone filter bandwidth B is given by the following formula:

[0052] Among them, any frequency (unit ) The calculation formula is:

[0053] here represents the high-frequency asymptotic filter quality factor, is the minimum bandwidth at low frequencies, Usually it is 1 or 2. The Glasberg-Moore parameter is used in the present invention, that is Parameters; EarQ and , this parameter can better improve the resolution of the spectrum.

[0054] like Figure 4 As shown in the figure, GFCC can be considered an improved algorithm of MFCC: first, the audio signal is framed and windowed, and the fast Fourier transform (FFT) is performed frame by frame to obtain the spectrum. Then, the FFT signal is processed with a gammatone filter to enhance the perceptually significant frequency bands, and finally a logarithmic operation and discrete cosine transform are applied. The GFCC calculation formula is:

[0055] In the formula Indicates the The energy of the spectrum band, is the number of Gammatone filters, For the GFCC coefficients, is the total number of selected coefficients of the eigenvector.

[0056] In addition, if Figure 5 The average energy distribution of MFCC and GFCC for the same PPG signal segment shows that the Mel filter approach produces a more uniform energy distribution, while the Gammatone filter framework generates significant energy peaks coinciding with key regions such as the PPG low-frequency complex. This characteristic reinforces the advantage of cubic compression in preserving subtle amplitude variations. This difference is also reflected in histogram analysis: the Mel filter output exhibits a broader distribution, while the Gammatone filter output is concentrated in a specific (lower) frequency range. Therefore, combining the two features ensures a richer preservation of subject-specific characteristics. This approach, while maintaining resolution in mid- and high-frequency regions, while optimizing low-frequency representation, captures more individual PPG signal variability, thereby enhancing biometric recognition.

[0057] Principal component analysis (PCA) is a variance-based dimensionality reduction technique used to transform potentially correlated variables into a set of linearly independent principal components. This method effectively reduces data dimensionality and redundancy while preserving key feature information.

[0058] The PCA dimensionality reduction process includes: 1. Zero mean: Subtract the mean of each row of the data matrix to make the data distribution around the origin and eliminate the offset.

[0059] 2. Calculate the covariance matrix: measure the correlation and degree of variation between different attributes.

[0060] 3. Eigendecomposition and dimensionality reduction: Calculate the eigenvalues ​​and eigenvectors of the covariance matrix; sort them from largest to smallest eigenvalues, and select the first k eigenvectors to form the dimensionality reduction matrix.

[0061] In order to fully retain key information while reducing computational complexity, the cumulative contribution rate of the eigenvector needs to be calculated. The formula is as follows:

[0062] Where, is the corresponding eigenvalue in the matrix; is the number of eigenvalues. When the cumulative contribution rate is close to 100%, it indicates that the selected principal component has covered most of the information.

[0063] For PPG signal feature dimensionality reduction: 1. Data source: PPG data of 30 trainees, 15 groups of which were selected as test samples.

[0064] 2. Feature fusion: MFCC (Mel-frequency cepstral coefficients) + GFCC (Gammatone frequency cepstral coefficients) → fused into MGCC features (higher dimension).

[0065] 3. Dimensionality reduction analysis: MFCC: The cumulative contribution rate of the first 8 dimensions reaches 100%, such as Figure 6 The figure shows the cumulative contribution rate of each dimension of MFCC eigenvalue, so 8 dimensions are retained; GFCC: the first 4 dimensions already contain the main information, such as Figure 7 The figure shows the cumulative contribution rate of each dimension of the GFCC eigenvalue, so 4 dimensions are retained.

[0066] 4. Final features: After dimensionality reduction, 12-dimensional MFCC+GFCC fusion features are obtained, which effectively improves model training speed and recognition accuracy.

[0067] Step S3: constructing a universal background model (UBM) of a Gaussian mixture model based on the fusion features, calculating zero-order, first-order, and second-order Baum-Welch statistics, estimating a global difference space matrix from the Baum-Welch statistics, and extracting an i-vector identity authentication vector.

[0068] Specifically, in real-world PPG signal recognition scenarios, PPG signals contain not only target physiological characteristics (such as heart rate and blood pressure) but also various noise and motion artifacts. Different acquisition devices and environmental conditions (such as skin contact pressure and ambient lighting) introduce signal variation, which interferes with the collected PPG data and thus affects the accuracy of physiological parameter recognition. Traditional GMM-UBM methods can encounter challenges in handling this signal variation, leading to unstable recognition system performance. To address this issue, the i-vector method models the global variance in PPG data, incorporating both physiological feature variations and acquisition interference into the model. This approach reduces reliance on training data quality, simplifies the computational process, and improves overall system performance. The i-vector method boasts high computational efficiency and is suitable for processing large-scale PPG datasets. Furthermore, thanks to i-vector's inherent cross-device stability and the incorporation of signal compensation techniques such as linear discriminant analysis (LDA), i-vector demonstrates enhanced robustness in cross-device recognition tasks.

[0069] The i-vector is modeled in a global difference space, which includes sound feature differences and channel differences. The model can be expressed as: M = m + Tw Among them, M represents the mean supervector of the GMM generated by the current training sample; m represents the mean supervector that is independent of the sound feature information and channel information of the current training sample, that is, the mean supervector of the UBM; T represents the global difference space matrix, which is used to describe all variability in the sound data that is related to and unrelated to the current fault type. It is learned from a large number of training data of different fault types; w represents a full variable difference factor that obeys a standard Gaussian distribution, that is, i-vector. Among these four quantities, M and m are known and can be calculated by GMM and UBM. Then, in order to extract the i-vector, the global difference space matrix T must be estimated first. The estimation of this matrix is ​​done through statistical modeling methods, usually using the EM algorithm on a large amount of training data. The process is as follows Figure 8 The detailed steps are as follows: (1) First, we need to calculate the Baum-Welch statistic corresponding to each fault in the training data. Its zero-order, first-order, and second-order statistics are:

[0070]

[0071]

[0072] in is the mean supervector in the UBM model The first The mean of the Gaussian components, is the dimension of the input PPG signal feature, That is the Frame feature vector.

[0073] (2) Next, the global difference space matrix T is estimated using the EM algorithm. First, the T matrix is ​​initialized and the current zero-order, first-order, and second-order Baum-Welch statistics are calculated.

[0074] (3) Execute the E-step to calculate the first-order and second-order expectations of the total variable difference factor w:

[0075]

[0076]

[0077] in is the precision matrix, which is the inverse covariance matrix of the posterior distribution of i-vector, is the identity matrix, is the diagonal covariance matrix containing the variances of the mixture components of the UBM on the diagonal.

[0078] (4) Next, execute the M-step to update the global difference space matrix T and the covariance matrix ∑ respectively. The update formula is as follows:

[0079] (5) Repeat steps 3 and 4. After each iteration, calculate the log-likelihood value under the current parameters. If the change is less than the threshold, it is considered converged. Finally, the global difference space matrix T is obtained.

[0080] Log-likelihood The calculation formula is as follows:

[0081] in 、 are the second-order and first-order statistics, and is the expectation of i-vector.

[0082] The convergence condition is that the difference in log-likelihood between two adjacent iterations satisfies:

[0083] After the global difference space matrix T is estimated, the i-vector is the maximum a posteriori estimate of w, and the formula for extracting the i-vector is:

[0084] Where I is the identity matrix, N is the zero-order statistic, that is, the frame weight sum of each Gaussian component, the supervector F is the first-order central statistic, ∑ represents the covariance matrix (diagonal matrix) of the UBM, and T is the trained global difference space matrix.

[0085] (6) Selection of the number of Gaussian mixture components: The number of Gaussian mixture components is the number of independent Gaussian distributions in the model. Theoretically, the higher the number of Gaussian mixture components, the stronger the model's ability to fit the category. However, as the degree of mixture increases, the model becomes too complex, significantly increasing the computational cost, leading to a significant increase in model training and testing time, and also leading to the risk of overfitting. Therefore, the degree of mixture has a direct impact on the model's recognition accuracy. The number of Gaussian mixture components was set to 8, 16, 32, and 64, respectively, leaving other parameters unchanged by default. The fusion feature consisting of the 12 extracted feature vectors was input for testing. The results are shown in Table 1: Table 1 Effect of the number of Gaussian mixture components on model performance

[0086] As can be seen from the table, as the number of Gaussian mixture components increases, the EER becomes smaller and smaller. When it increases from 8 to 32, the EER decreases significantly. However, when it continues to increase to 64, the EER also decreases, but the amplitude is very small. Considering that the increase in the number of Gaussian mixture components will greatly increase the computational cost, reduce the efficiency of the model, and deteriorate the real-time performance, the number of Gaussian mixture components used in the final model is 32.

[0087] (7) Selection of i-vector dimension: EER is a commonly used performance evaluation metric in biometrics (such as speaker recognition and fingerprint recognition) and classification systems. It represents the error rate when the false acceptance rate (FAR) and false rejection rate (FRR) are equal. Its core meaning is the overall error probability of the system when the two types of errors (false positives and false negatives) are balanced. The lower the value, the better the system performance.

[0088] In methods using i-vectors, the i-vector dimension has a significant impact on model accuracy. A low-dimensional i-vector may not fully capture relevant information about the category, resulting in poor model recognition accuracy. An excessively high i-vector dimension may contain more information, but when training samples are limited, it may lead to model overfitting and increased computational costs. The following experiments determine the optimal i-vector dimension for this model. The i-vector dimensions of the model output are set to 100, 200, 300, and 400, respectively. Other parameters remain unchanged by default. The extracted 12-dimensional fusion features are input for testing. The results are shown in Table 2: Table 2 The impact of i-vector dimension on model performance

[0089] As can be seen from the table, as the i-vector dimension increases, the model's EER gradually decreases. When the dimension reaches 200, the model accuracy improves significantly. When the dimension reaches 300 and 400, the EER decreases slightly, but the decrease is not significant. Considering that i-vector feature vectors with a large dimension increase the risk of overfitting in subsequent classification decisions, 200 dimensions are selected as the i-vector dimension of the final model output.

[0090] (8) Selection of input feature duration: In the method using i-vector, the input feature duration has a greater impact on the recognition accuracy of the model. A shorter duration means less effective information for training and testing, which may lead to insufficient representation of the i-vector and reduce the accuracy of the model. A longer duration helps to generate a more stable and discriminative i-vector. In practical applications, it is necessary to balance the model performance and the amount of computation and select an appropriate feature duration. The feature duration of the input model is set to 2s, 3s, 4s, and 5s respectively. The feature duration here depends on the number of frames of the fused feature. This is because the feature of the input model is based on the frame feature. However, regardless of the feature duration of the input model, the i-vector model can output an i-vector of uniform length, which is one of its advantages. Other parameters remain unchanged by default, and the 12-dimensional fused feature extracted is input for testing. The results are shown in Table 3: Table 3 Impact of feature duration on model performance

[0091] As can be seen from the table, as the feature duration of the input model increases, the model performance gradually improves and the EER gradually decreases. When the feature duration increases from 2s to 4s, the model accuracy improves significantly. When the feature duration continues to increase to 5s, the EER decreases but the magnitude is not large. Since the PPG signal is periodic, the continued increase in the feature duration of the input model does not have a significant impact on the model performance. Considering the issue of computational efficiency, this paper selects 4s as the final input feature duration.

[0092] The final model input MGCC fusion features are 12-dimensional, the number of Gaussian mixture components is 32, the output i-vector dimension is 200-dimensional, and the feature duration of the input model is 4 seconds. At this time, the model takes into account the quality of the output i-vector and computational efficiency, preparing for the next classification decision.

[0093] Based on the fusion features, the number of Gaussian mixture components was set to 32, and a universal background model (UBM) was constructed. Its covariance matrix was set to a diagonal matrix to obtain the initial model parameters. The initialization process used a modified K-means clustering method to cluster all training data into 32 clusters. The mean and variance of each cluster were used as the initial mean and covariance of the GMM, and the initial weights of each component were set equal. This clustering-based initialization method can accelerate the convergence of the EM algorithm and improve the quality of the final model.

[0094] UBM training uses the expectation-maximization (EM) algorithm to iteratively optimize the parameters of the GMM. Each iteration consists of two steps: the expectation step (E-step), which calculates the posterior probability of each sample belonging to each Gaussian component; and the maximization step (M-step), which updates the model parameters (weights, mean, and covariance) based on these posterior probabilities. Specifically, in iteration t: E-step: For each training sample x, calculate the posterior probability that it belongs to the kth Gaussian component M-step: Update the weight, mean and covariance matrix of each Gaussian component To improve the robustness of the UBM, a variance floor constraint is applied during training to ensure that the diagonal elements of the covariance matrix are no less than a preset threshold (usually 0.01) to prevent the occurrence of singular covariance matrices. At the same time, weight regularization is implemented to prevent certain component weights from being too small or zero.

[0095] Based on the initial model parameters, the EM algorithm is used to calculate the zero-order, first-order, and second-order Baum-Welch statistics. The global difference space matrix is ​​iteratively estimated until the change in the log-likelihood value is less than a preset threshold. The global difference space matrix is ​​then obtained. The convergence of the EM algorithm is determined by the change in the log-likelihood function between two consecutive iterations being less than a preset threshold (typically set to 1e-5) or by the maximum number of iterations (typically set to 100).

[0096] Estimating the global discrepancy space matrix T is a key step in i-vector extraction. The initial T matrix is ​​obtained through principal component analysis (PCA) of the covariance matrix of the training data, and the final result is obtained through iterative optimization. In each iteration, the posterior distribution parameters (mean and covariance) are first calculated based on the current T matrix. The T matrix is ​​then updated based on these posterior statistics until the algorithm converges.

[0097] Based on the global difference space matrix and the Baum-Welch statistic, a maximum a posteriori probability estimation method is used to extract the i-vector identity authentication vector with a dimension of 200. i-vector extraction uses a closed-form solution, which is computationally efficient and stable. For each utterance, the i-vector is a point estimate (maximum a posteriori estimate) extracted from the posterior probability distribution, capturing the utterance's offset from the UBM.

[0098] Step S4: Input the i-vector identity authentication vector into the long short-term memory network for training and classification to obtain the identity recognition result of the trainee.

[0099] Specifically, in the overall variation space, the two most commonly used compensation techniques for i-vectors are within-class covariance normalization (WCCN) and linear discriminant analysis (LDA), which are often used to reduce variability. WCCN uses the inverse of the within-class covariance matrix to normalize the cosine kernel. LDA minimizes the within-class variance caused by the variation effect through coordinate axis transformation while maximizing the between-class variance. Experiments show that the cosine kernel function is an effective classifier for classifying i-vectors. The cosine kernel between two i-vectors: w1 and w2 is defined as follows:

[0100] To further optimize the performance of this method, a variety of commonly used machine learning and deep learning methods can be used to select effective classifiers for i-vector classification, including support vector machines (SVMs), multilayer perceptrons (MLPs), random forests (RFs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs). These methods can be authenticated by capturing the time and frequency domain features of PPG signals. Support vector machines (SVMs) excel in nonlinear classification problems by determining the optimal separating hyperplane between data points and, when combined with kernel functions, are particularly well-known for their high accuracy with limited datasets and robustness in noisy environments. Multilayer perceptrons (MLPs) model nonlinear relationships in data using one or more hidden layers, revealing signal characteristics through sequential weighting and activation functions. Random forests (RFs) capture different sub-features of the data through an ensemble structure of multiple decision trees, achieving robust classification. Convolutional neural networks (CNNs) automatically extract PPG signal features hierarchically through convolutional filters and pooling layers, capturing local patterns and building noise-resistant models. Recurrent neural networks (RNNs) use hidden state updates to model temporal dependencies. In particular, the addition of gating structures such as LSTM and GRU helps alleviate long-range dependency problems and effectively capture continuous heartbeat patterns in PPG signals.

[0101] The i-vector authentication vectors are normalized and reordered to generate a sequence of training samples and corresponding identity labels. Normalization is achieved using the Z-score method, which sets the data to a mean of 0 and a standard deviation of 1. Reordering organizes the training samples in chronological order, allowing the LSTM network to capture temporal features. The training samples and labels are divided into a training set and a validation set with an 80%:20% ratio.

[0102] Based on the training sample sequences and corresponding identity labels, a network structure was constructed consisting of an input layer, two LSTM layers, and a fully connected layer. The number of LSTM units in the first layer was 128, the number of units in the second layer was 64, and the dropout rate was set to 0.3 to obtain the initial network model. The LSTM layer used the tanh activation function, and the fully connected layer used a Softmax activation function with the number of output nodes equal to the number of trainees. The detailed structure of the network is as follows: Input layer: shape is (batch size , sequence length , feature dim ), where feature dim is the i-vector dimension; First layer LSTM: 128 units, return sequences =True, use tanh activation function; Dropout layer: the ratio is 0.3 to prevent overfitting; Second layer LSTM: 64 units, return sequences =False, use tanh activation function; Dropout layer: the ratio is 0.3 to prevent overfitting; Fully connected layer: The number of output nodes is equal to the number of trainees; Softmax layer: Converts the output into a probability distribution.

[0103] The initial network model was iteratively trained using the Adam optimizer, using cross-entropy as the loss function. The learning rate was initially set to 0.001 and dynamically adjusted using a cosine annealing strategy to obtain the trained LSTM network model. Training was performed using a mini-batch approach with a batch size of 32 and a maximum number of iterations of 100. To prevent overfitting, an early stopping strategy was implemented, stopping training after five consecutive epochs of no improvement in the validation set loss. Furthermore, the validation set accuracy was calculated after each epoch, and the optimal model parameters were recorded.

[0104] The i-vector identity authentication vector to be identified is input into the trained LSTM network model. A Softmax classifier is used to calculate the probability distribution of each identity category, and the category with the highest probability is selected as the identification result. In practice, a confidence threshold (typically 0.7) is set. When the highest probability falls below the threshold, the system marks the identity as "unknown," improving recognition reliability.

[0105] Example 2 This embodiment provides a biometric feature recognition method driven by PPG big data, including the following steps: Step S1: Obtain the PPG signal data of the trainee, and perform high-pass filtering and low-pass filtering preprocessing on the PPG signal data to obtain a preprocessed PPG signal.

[0106] Specifically, filter parameters are configured for the PPG signal data, the cutoff frequency of the fourth-order Chebyshev type I high-pass filter is set to 0.5 Hz, the passband fluctuation is set to within ±0.5 decibels, and the stopband attenuation slope is set to negative 80 decibels per decade, to obtain high-pass filter parameters; based on the high-pass filter parameters, the fourth-order Chebyshev type I high-pass filter is constructed, and the fourth-order Chebyshev type I high-pass filter is applied to the PPG signal data to eliminate baseline drift and obtain a first filtered signal; based on the time domain and frequency domain characteristics of the first filtered signal, the filtering effect is evaluated, and the signal-to-noise ratio and the degree of baseline drift suppression are calculated to ensure that preset signal quality indicators are met.

[0107] Filter parameters are configured for the first filtered signal, with the cutoff frequency of the fourth-order elliptical low-pass filter set to 18 Hz, the passband set to 0 Hz to 18 Hz with a fluctuation of less than or equal to 0.1 decibel, the stopband starting at 22 Hz with an attenuation of greater than or equal to 40 decibels at 22 Hz, and greater than or equal to 60 decibels at 50 Hz, to obtain low-pass filter parameters. A fourth-order elliptical low-pass filter is constructed based on the low-pass filter parameters, and the fourth-order elliptical low-pass filter is applied to the first filtered signal to eliminate motion artifacts, ambient light interference, and muscle noise, thereby obtaining the preprocessed PPG signal. Based on the spectral characteristics of the preprocessed PPG signal, the noise suppression effect is evaluated, and the degree of high-frequency interference suppression is calculated to ensure that preset signal quality requirements are met.

[0108] Step S2: Frame the preprocessed PPG signal, extract Mel-frequency cepstral coefficient features and Gammatone-frequency cepstral coefficient features respectively, fuse the Mel-frequency cepstral coefficient features and Gammatone-frequency cepstral coefficient features, and perform principal component analysis and dimensionality reduction to obtain fused features.

[0109] Specifically, the pre-processed PPG signal is pre-emphasized to enhance the high-frequency component and obtain a pre-emphasized signal; the pre-emphasized signal is framed according to a preset frame length, and a Hamming window is added to each frame signal for smooth transition to obtain a first windowed signal; the first windowed signal is fast Fourier transformed to obtain a spectrum, and the spectrum is input into a Mel filter bank for frequency envelope extraction to obtain a Mel spectrum; the Mel spectrum is discrete cosine transformed to compress the feature dimension to obtain the Mel-frequency cepstral coefficient feature.

[0110] The preprocessed PPG signal is framed, and a Hamming window is applied to each frame signal to obtain a second windowed signal; the second windowed signal is fast Fourier transformed to obtain a frequency domain signal; the frequency domain signal is input into a Gammatone filter bank, and frequency bands are divided according to an ERB frequency scale to obtain filtered outputs of each channel; energy calculation and logarithmic compression are performed on the filtered outputs of each channel, and discrete cosine transform is performed to obtain the Gammatone frequency cepstral coefficient characteristics.

[0111] Step S3: constructing a universal background model (UBM) of a Gaussian mixture model based on the fusion features, calculating zero-order, first-order, and second-order Baum-Welch statistics, estimating a global difference space matrix from the Baum-Welch statistics, and extracting an i-vector identity authentication vector.

[0112] Specifically, based on the fusion features, the number of Gaussian mixture components is set to 32, a universal background model (UBM) is constructed, and its covariance matrix is ​​set to a diagonal matrix to obtain initial model parameters; according to the initial model parameters, the EM algorithm is used to calculate the zero-order, first-order, and second-order Baum-Welch statistics, and the global difference space matrix is ​​iteratively estimated until the change in the log-likelihood value is less than a preset threshold, thereby obtaining the global difference space matrix; based on the global difference space matrix and the Baum-Welch statistic, the i-vector identity authentication vector with a dimension of 200 is extracted through the maximum a posteriori probability estimation method.

[0113] Step S4: Input the i-vector identity authentication vector into the long short-term memory network for training and classification to obtain the identity recognition result of the trainee.

[0114] Specifically, the i-vector identity authentication vector is normalized and time-rearranged to generate a training sample sequence and a corresponding identity label; based on the training sample sequence and the corresponding identity label, a network structure is constructed including an input layer, two LSTM layers and a fully connected layer, wherein the number of LSTM units in the first layer is 128, the number of LSTM units in the second layer is 64, and the dropout rate is set to 0.3 to obtain an initial network model; cross entropy is used as the loss function, and the initial network model is iteratively trained using the Adam optimizer, the initial value of the learning rate is set to 0.001 and dynamically adjusted using the cosine annealing strategy to obtain a trained LSTM network model; the i-vector identity authentication vector to be identified is input into the trained LSTM network model, the probability distribution of each identity category is calculated using the Softmax classifier, and the category with the highest probability is selected as the recognition result.

[0115] Example 3 This embodiment provides a biometric feature recognition method driven by PPG big data, including the following steps: Step S1: Obtain the PPG signal data of the trainee, and perform high-pass filtering and low-pass filtering preprocessing on the PPG signal data to obtain a preprocessed PPG signal.

[0116] Step S2: Frame the preprocessed PPG signal, extract Mel-frequency cepstral coefficient features and Gammatone-frequency cepstral coefficient features respectively, fuse the Mel-frequency cepstral coefficient features and Gammatone-frequency cepstral coefficient features, and perform principal component analysis and dimensionality reduction to obtain fused features.

[0117] Step S3: constructing a universal background model (UBM) of a Gaussian mixture model based on the fusion features, calculating zero-order, first-order, and second-order Baum-Welch statistics, estimating a global difference space matrix from the Baum-Welch statistics, and extracting an i-vector identity authentication vector.

[0118] Step S4: Input the i-vector identity authentication vector into the long short-term memory network for training and classification to obtain the identity recognition result of the trainee.

[0119] Step S5: Obtain the high-pass filter parameters, the low-pass filter parameters, and the extraction parameters of the Mel-frequency cepstral coefficient features and the Gammatone frequency cepstral coefficient features, construct a parameter optimization framework based on a genetic algorithm, encode the high-pass filter parameters, the low-pass filter parameters and the extraction parameters into chromosomes of the initial population; based on the chromosomes of the initial population, calculate the fitness function value of each chromosome, wherein the fitness function value comprehensively evaluates the three indicators of signal-to-noise ratio gain, signal quality index and feature discrimination, and obtains the population fitness distribution; according to the population fitness distribution, retain the 10% individuals with the highest fitness to enter the next generation, and adopt the trophy method for the remaining 90% individuals The competition is selected to perform crossover and mutation, where the crossover rate is initially set to 0.8 and dynamically adjusted according to the population diversity, and the mutation rate is initially set to 0.1 and gradually decreases to 0.01 as the number of iterations increases to obtain a new generation of population; every 10 generations, the 10% individuals with the highest fitness in the new generation of population are subjected to simulated annealing local optimization until the preset number of iterations is reached to obtain the optimal parameter configuration; based on the optimal parameter configuration, the PPG signal of each trainee is personalized preprocessed and feature extracted to generate optimized PPG signals and features; an incremental optimization mechanism is set to trigger parameter re-optimization when the recognition accuracy drops by more than a preset threshold or reaches a fixed optimization period, and the optimization process is accelerated by parallel computing.

[0120] Step S6: Obtain the initial fusion weights and dimensionality reduction matrix of the Mel frequency cepstral coefficient features and the Gammatone frequency cepstral coefficient features, construct a dual-objective optimization problem for feature fusion, set the first objective to maximize the Fisher discriminant ratio, and the second objective to minimize the feature dimension; Based on the dual-objective optimization problem, use the NSGA-II algorithm to generate the initial population, each individual contains two parts: feature fusion weight and dimensionality reduction, where the fusion weight range is [0,1] and the sum is 1, and the dimensionality reduction range is [8,20]; perform non-dominated sorting and crowding calculation on the initial population, generate the parent population through binary tournament selection, and use simulated binary Binary crossover and polynomial mutation generate a population of offspring; a local search is performed on the solutions on the non-dominated frontier every 10 generations, and the search step size is dynamically adjusted with the number of iterations to accelerate the convergence of the algorithm and obtain the Pareto optimal solution set; based on the Pareto optimal solution set, the optimal solution is selected using the knee point detection method, which first arranges the solution set in ascending order according to the first objective, then calculates the curvature of adjacent solutions, selects the point with the largest curvature as the knee point, and generates a personalized feature fusion model; an adaptive adjustment mechanism is set to monitor the recognition performance. When the cumulative error rate exceeds 5% or after 30 days, the feature fusion model is automatically triggered to update to ensure stable performance.

[0121] Step S7: Construct an i-vector framework parameter space and a decision model parameter space, wherein the i-vector framework parameter space includes the number of Gaussian mixture components, the rank of the T matrix, the dimension of the eigenvector, and the number of iterations, and the decision model parameter space includes the number of layers, the number of hidden units, the dropout rate, and the learning rate of the LSTM network; the covariance matrix adaptive evolution strategy CMA-ES is used to optimize the i-vector framework parameter space and the decision model parameter space respectively, and the interactive optimization between subspaces is realized through the three-layer collaborative mechanism of parameter cross-evaluation, fitness sharing, and search direction coordination; the calculations allocated to each subspace are dynamically adjusted according to the progress speed and performance improvement of the interactive optimization of the subspace. Resources, use multi-threading technology to simultaneously evaluate different parameter combinations; for different scenarios in training, pre-define typical scenario configurations and execute the optimization process separately to generate scenario-specific parameter configuration schemes, and automatically switch parameter configurations according to the training scenarios detected in real time; adopt a regularization control strategy, add L2 regularization constraints in the i-vector subspace to control the complexity of the T matrix, introduce batch normalization and early stopping strategies in the decision model subspace to improve the generalization ability of the model; adopt a two-level strategy of lightweight update strategy and comprehensive update strategy to achieve incremental optimization, wherein the lightweight update strategy fine-tunes parameters in a local range, and the comprehensive update strategy re-executes the complete co-evolutionary optimization after the accumulated new data exceeds the preset data value.

[0122] Example 4 This embodiment provides a parameter optimization framework based on an evolutionary algorithm to enhance the effects of PPG signal preprocessing and feature extraction.

[0123] Specifically, after obtaining the participants' PPG signal data, the high-pass and low-pass filter parameters, as well as the extraction parameters for the Mel-frequency cepstral coefficient features and the gammatone-frequency cepstral coefficient features, were first obtained. A genetic algorithm-based parameter optimization framework was then constructed, encoding the high-pass and low-pass filter parameters and the extraction parameters into chromosomes for the initial population. For the high-pass filter, the encoding details included the cutoff frequency (0.3-0.7 Hz), filter order (2-6), and filter type (Chebyshev, Butterworth, or elliptic); for the low-pass filter, the encoding details included the cutoff frequency (15-25 Hz), filter order (3-5), and filter type. Parameters used in the MFCC and GFCC feature extraction processes were also encoded, including the number of Mel-frequency filters (20-40), the number of cepstral coefficients (8-16), the frame length (20-30 ms), and the frame shift (10-15 ms). 100 chromosomes were randomly generated from the initial population, covering different regions of the parameter space to ensure diversity.

[0124] Based on the chromosomes of the initial population, the fitness function value of each chromosome is calculated, wherein the fitness function value comprehensively evaluates the three indicators of signal-to-noise ratio gain, signal quality index and feature discrimination to obtain the population fitness distribution. The signal-to-noise ratio gain is calculated by comparing the signal-to-noise power ratio before and after filtering; the signal quality index is based on the periodicity and morphological characteristics of the PPG waveform, including peak detection rate, waveform symmetry and baseline stability; the feature discrimination is measured by calculating the ratio of the intra-class distance to the inter-class distance, using the Fisher discriminant ratio as a quantitative indicator. The overall fitness function value is defined as:

[0125] The weight coefficient 、 、 Adjust according to the application scenario. In this solution, it is set to 0.3, 0.3 and 0.4.

[0126] Based on the population's fitness distribution, the 10% individuals with the highest fitness are retained for the next generation. The remaining 90% undergo crossover and mutation using tournament selection. The crossover rate is initially set to 0.8 and dynamically adjusted based on population diversity. The mutation rate is initially set to 0.1 and gradually decreases to 0.01 with increasing iterations, resulting in the next generation of populations. The crossover operation uses adaptive arithmetic crossover, with the crossover rate initially set to 0.8 and dynamically adjusted based on population diversity. The mutation operation uses Gaussian mutation, with the mutation rate initially set to 0.1 and gradually decreasing to 0.01 with increasing iterations, to balance global exploration and local exploitation capabilities.

[0127] Every 10 generations, simulated annealing is performed on the 10% of individuals with the highest fitness in the new generation population until the preset number of iterations is reached to obtain the optimal parameter configuration. Local optimization helps accelerate convergence, carefully explore the neighborhood space of high-quality solutions, and improve the quality of the final solution.

[0128] Based on the optimal parameter configuration, each participant's PPG signal is personalized preprocessed and feature extracted to generate optimized PPG signals and features. This personalized processing can adapt to the physiological characteristics of different participants and improve the pertinence and effectiveness of feature extraction.

[0129] Finally, an incremental optimization mechanism is implemented. When the recognition accuracy drops below a preset threshold or reaches a fixed optimization period, parameter re-optimization is triggered, and the optimization process is accelerated through parallel computing. In this way, the system can continuously adapt to changes in the trainee's physiological state and environmental conditions, maintaining long-term stable recognition performance.

[0130] Experimental results show that compared to a fixed-parameter solution, the PPG signal processing solution optimized by the evolutionary algorithm improves noise immunity by 32%, feature discrimination by 28%, and overall recognition accuracy by 3.5 percentage points. In particular, in high-intensity exercise scenarios (T3 type), recognition accuracy increased from 92.7% to 97.8%, validating the solution's superior performance in complex training environments.

[0131] Example 5 This embodiment provides a feature fusion strategy based on multi-objective optimization to enhance the fusion effect of MFCC and GFCC features.

[0132] Specifically, we first obtain the initial fusion weights and dimensionality reduction matrix of the Mel frequency cepstral coefficient features and the Gammatone frequency cepstral coefficient features, and then construct a dual-objective optimization problem for feature fusion. The first objective is to maximize the Fisher discriminant ratio, and the second objective is to minimize the feature dimension. Let the MFCC feature vector be X M , the GFCC eigenvector is X G, the fusion weight vector is W=[w M , w G ], the goal is to find the optimal weight The dimensionality reduction matrix P is used to maximize the inter-class distance and minimize the intra-class distance in the fused feature Z. The Fisher discriminant ratio is used as a measure of the feature's discriminative ability, while the feature dimension is considered as a complexity constraint, forming a dual-objective optimization problem: maximizing the discriminative ability while minimizing the feature dimension.

[0133] Before implementing multi-objective optimization, Z-score normalization is performed on each MFCC and GFCC feature to ensure comparability during the fusion process. Outlier detection is also performed on each feature type, and an improved Local Outlier Factor (LOF) algorithm is used to identify and address anomalous samples, enhancing the stability of the feature distribution. Furthermore, a feature correlation matrix is ​​calculated to identify redundant information between MFCC and GFCC features, providing a reference for subsequent fusion.

[0134] Based on the dual-objective optimization problem, the NSGA-II algorithm is used to generate the initial population. Each individual contains two parts: feature fusion weight and dimensionality reduction. The fusion weight range is [0, 1] and the sum is 1. The dimensionality reduction range is [8, 20]. The algorithm initializes 100 solutions, each of which contains a weight vector W and a dimensionality reduction number k. The individual encoding adopts real number encoding, where w M ,w G ∈[0,1] and w M +w G = 1, k∈{8,9,...,20}. The algorithm runs for 200 generations, with the crossover probability set to 0.9 and the mutation probability set to 1 / n (n is the number of decision variables).

[0135] The initial population is subjected to non-dominated sorting and crowding calculation. A parent population is generated through binary tournament selection, and the offspring population is generated using simulated binary crossover and polynomial mutation. A local search for solutions on the non-dominated frontier is performed every 10 generations. The search step size is dynamically adjusted with the number of iterations to accelerate algorithm convergence and obtain a Pareto-optimal solution set. This non-dominated sorting and crowding calculation mechanism ensures diversity and convergence in multi-objective optimization, enabling the discovery of a range of high-quality solutions representing different trade-offs.

[0136] During the optimization process, the first objective function evaluates feature discriminability using a modified Fisher discriminant ratio. This metric not only considers the global inter- and intra-class scatter matrices but also incorporates local neighbor structure information. This allows the fused features to maintain global discriminability while preserving the local manifold structure. The second objective function evaluates feature complexity. In addition to considering the dimension k, it also incorporates feature information entropy as an auxiliary metric to comprehensively measure complexity. During the evaluation of each solution, recognition performance metrics are calculated using five-fold cross-validation to ensure the reliability of the evaluation results.

[0137] Based on the Pareto optimal solution set, the optimal solution is selected using a knee point detection method. This method first sorts the solution set in ascending order according to the first objective, then calculates the curvature between adjacent solutions. The point with the maximum curvature is selected as the knee point to generate a personalized feature fusion model. This method first sorts the solutions on the frontier in ascending order according to the first objective (discriminative ability), then calculates the curvature between adjacent solutions, and selects the point with the maximum curvature as the knee point. This point represents the location where marginal benefits change significantly and is the optimal choice for balancing the two objectives.

[0138] Based on the selected optimal solution, the system generates a personalized feature fusion model. This model consists of two parts: a feature weighting module and an adaptive dimensionality reduction module. The feature weighting module weights the MFCC and GFCC features according to the optimized weights; the adaptive dimensionality reduction module uses the optimized dimensionality reduction matrix to project the weighted features into a lower-dimensional space. Unlike traditional PCA, the dimensionality reduction matrix is ​​constructed based on the principle of maximizing the preservation of class information, considering both variance contribution and class differentiation, thereby generating a more discriminative low-dimensional representation.

[0139] Finally, an adaptive adjustment mechanism is implemented to monitor recognition performance. When the cumulative error rate exceeds 5% or after 30 days, an update of the feature fusion model is automatically triggered to ensure stable performance. The system continuously monitors recognition performance during operation and automatically triggers a model update when the cumulative error rate exceeds a preset threshold or after a fixed time interval. During this update, the multi-objective optimization is re-executed based on newly collected data, adjusting feature fusion weights and dimensionality reduction parameters to ensure the system can adapt to long-term changes in the trainees' physiological state and environmental conditions.

[0140] Experimental results show that compared to the original solution, the feature fusion strategy based on multi-objective optimization improved recognition accuracy by an average of 4.2 percentage points, with the most significant improvement in T2 (moderate-intensity exercise) scenarios, from 94.5% to 99.1%. Furthermore, the feature dimension was reduced by an average of 35%, significantly reducing the computational burden of the subsequent i-vector model, validating the effectiveness of this solution in improving recognition performance while optimizing computing resource utilization.

[0141] Example 6 This embodiment provides an i-vector and model parameter optimization method based on co-evolution, which is used to enhance the performance of the i-vector framework and decision model.

[0142] Specifically, the i-vector framework parameter space and the decision model parameter space are first constructed. The i-vector framework parameter space includes the number of Gaussian mixture components, the rank of the T matrix, the eigenvector dimension, and the number of iterations. The decision model parameter space includes the number of LSTM layers, the number of hidden units, the dropout rate, and the learning rate. The first subspace is the i-vector framework parameter space, which includes parameters such as the number of Gaussian mixture components (range 16-128), the rank of the T matrix (range 100-400), the eigenvector dimension (range 8-20), and the number of iterations (range 5-15) of the UBM. The second subspace is the decision model parameter space, which mainly includes parameters such as the number of LSTM layers (range 1-4), the number of hidden units (range 32-256), the dropout rate (range 0.1-0.5), and the learning rate (range 0.0001-0.01).

[0143] The covariance matrix adaptive evolutionary strategy (CMA-ES) is used to optimize the i-vector framework parameter space and the decision model parameter space, respectively. Interactive optimization between subspaces is achieved through a three-layer collaborative mechanism of parameter cross-evaluation, fitness sharing, and search direction coordination. CMA-ES can efficiently handle complex nonlinear and multimodal optimization problems by adaptively adjusting the covariance matrix of the search distribution. For the i-vector subspace, the initial population size is 16 and the initial step size is 0.3; for the decision model subspace, the initial population size is 20 and the initial step size is 0.25. To address the heterogeneity of different parameters, each parameter is normalized so that its distribution is in the interval [0,1] and then mapped back to the original value range in actual application.

[0144] During the co-evolutionary process, a three-layer collaborative mechanism enables effective interaction between subspaces. The first layer is parameter cross-evaluation: elite individuals in each subspace are combined with elite individuals from the other subspace to jointly evaluate system performance and form a cross-fitness matrix. The second layer is fitness sharing: the populations of the two subspaces update their individual fitness based on the cross-fitness matrix, allowing the optimization process to account for the mutual influence of parameters. The third layer is search direction coordination: the CMA-ES algorithms of the two subspaces adjust their covariance matrices based on the evolutionary trends of the other subspace, achieving coordinated adjustment of search directions.

[0145] Based on the progress and performance improvement of the interactive optimization of subspaces, the computational resources allocated to each subspace are dynamically adjusted, and multi-threaded technology is used to simultaneously evaluate different parameter combinations. When the optimization of a subspace reaches a plateau, its computational resources are temporarily reduced to allocate more resources to subspaces with greater improvement potential. When a significantly superior solution is found in a subspace, its resource allocation is increased to accelerate the exploration of the parameter space surrounding that area.

[0146] For different training scenarios, typical scenario configurations are predefined and optimized separately, generating scenario-specific parameter configurations. The system automatically switches parameter configurations based on the training scenario detected in real time. The system predefines various typical scenario configurations (such as high-intensity physical training, fine motor training, and long-term endurance training). The system performs a separate optimization process for each scenario, generating scenario-specific parameter configurations. During deployment, the system can automatically switch parameter configurations based on the training scenario detected in real time, or gradually adjust parameters to adapt to new scenarios through online learning.

[0147] A regularization control strategy is adopted, adding L2 regularization constraints to the i-vector subspace to control the complexity of the T matrix. Batch normalization and early stopping strategies are introduced in the decision model subspace to improve model generalization. In the i-vector subspace, L2 regularization constraints are added to control the complexity of the T matrix to avoid overfitting the training data. In the decision model subspace, in addition to the conventional dropout mechanism, batch normalization and early stopping strategies are also introduced to improve the model's generalization ability.

[0148] Incremental optimization is achieved using a two-tiered strategy: a lightweight update strategy that fine-tunes parameters locally, while a comprehensive update strategy reruns the full co-evolutionary optimization after the accumulated new data exceeds a preset threshold. After initial parameter optimization, the system enters an online learning phase, regularly collecting new PPG data to incrementally update the model. The update process employs a two-tiered strategy: a lightweight update uses the current parameters as a starting point, fine-tuning the parameters within a small range to adapt to the new data; a comprehensive update reruns the full co-evolutionary optimization after sufficient new data has accumulated.

[0149] Experimental results show that compared to the original solution, the co-evolution-based parameter optimization solution improved overall recognition accuracy by 5.3 percentage points, reaching an average of 99.4%. The improvement was particularly significant when processing complex dynamic scenes (such as T3-type high-intensity exercise), increasing recognition accuracy from 92.7% in the original solution to 98.9%. Furthermore, the optimized system reduced training time by 47% and inference latency by 63%, demonstrating the solution's effectiveness in improving system performance while optimizing computing resource utilization.

[0150] Example 7 This embodiment provides an experiment and analysis of a biometric recognition method driven by PPG big data.

[0151] For performance evaluation, we used a dataset from 30 participants to verify the generalization ability of the proposed model for PPG biometric recognition. We employed a sliding window processing strategy based on a 4-second analysis window (500 samples, 125Hz sampling rate), a window step of 25 samples (0.2-second offset), and implemented a modified stratified 5-fold cross-validation scheme. The specific implementation process is as follows: 1. Data preprocessing and window division: For each 20-min PPG signal (150,000 sampling points), a rectangular window was used to strictly intercept 500 consecutive sampling points (4 seconds), and adjacent windows overlapped by 475 sampling points (3.8 seconds), resulting in a total of approximately 2,370 analysis windows.

[0152] 2. Hierarchical grouping strategy: First layer: grouped by age (18-30 / 31-45 / 46-58 years old, 10 people each); The second level: activity intensity (T1 = low intensity, T2 = moderate intensity, T3 = high intensity); Third layer: Specially marked overlapping subjects (No. 9 / 10 / 21 / 22); An improved K-means clustering algorithm was used, with feature space distance as the metric, to ensure that each fold contained: the same proportion of windows for each age group (±2%); the same proportion of windows for each activity type (T1:T2:T3≈1:1:1); and all windows of overlapping subjects were completely retained in a single fold.

[0153] 3. Cross-validation implementation: a) All windows of the 30 subjects (71,100 in total) are divided into 5 folds through stratified sampling, as follows Figure 9 As shown; b) Each fold contains: approximately 14,220 windows (20%) from the complete records of 6 subjects; the number of windows for each activity type: T1≈4,740, T2≈4,740, T3≈4,740; c) Training and validation datasets: training set: 4 folds (56,880 windows); test set: 1 fold (14,220 windows).

[0154] The metrics used for evaluation are: (1) Accuracy: The ratio of correctly classified observations (to be identified) to the total number of observations.

[0155] (2) Precision: The ratio of correctly predicted positive observations to the total predicted positive observations.

[0156] (3) Recall: The ratio of correctly predicted positive observations to all observations in the actual category.

[0157] (4) F1 Score: The weighted average of precision and recall.

[0158] The accuracy, precision, recall, and F1 score of a 5-fold cross validation (CV) evaluation on the dataset with 30 participants are shown in Table 4. Several commonly used machine learning and deep learning methods were used to select effective classifiers for i-vector classification, including support vector machines (SVMs), multilayer perceptrons (MLPs), random forests (RFs), convolutional neural networks (CNNs), and long short-term memory networks (LSTMs).

[0159] Table 4 Evaluation indicators of each model under different decision-making methods

[0160] Table 4 clearly shows that the recognition rates achieved by the MFCC+GFCC fusion feature and LSTM classification models are both optimal, approximately 3.1% higher than the second-place MFCC+GFCC fusion feature and random forest (RF) classification models. They also achieve 6.3% and 3.6% improvements over the other two decision-making models, the multi-layer perceptron (MLP) and convolutional neural network (CNN), respectively, and are approximately 7.7% higher than the traditional SVM model. This indicates that the use of a more rational decision-making mechanism in this study further improves the overall classification capability of the model. Comparing the recognition success rates before and after model optimization for the same feature values ​​clearly shows that the optimized models outperform the pre-optimization models, further enhancing the classification capabilities of the traditional cosine kernel and SVM models.

[0161] To validate the effectiveness of our approach, we conducted qualitative and quantitative comparative experiments on the same dataset using three traditional neural network models that directly process one-dimensional time data: 1D-CNN, LSTM, and Transformer. 1D-CNN extracts local features from raw time series data through one-dimensional convolutions and performs classification through fully connected layers. Its relatively simple model structure makes it suitable for processing time series data with significant local features, but it has limitations in capturing long-term dependencies. LSTM, a recurrent neural network model with a unique gating mechanism (input gate, forget gate, output gate) and memory cells, effectively captures long-term dependencies and dynamic changes in time series, resulting in excellent performance in time-dependent tasks. In contrast, the Transformer model, based on a self-attention mechanism, effectively captures global dependencies in time series data. Unlike traditional recurrent neural networks, the Transformer model utilizes a multi-head self-attention mechanism to simultaneously model dependencies between different time steps, making it suitable for processing long-term time series data. Table 5 shows the performance of these four networks in terms of accuracy, precision, recall, and F1 score.

[0162] Table 5 Performance comparison between the proposed model and traditional neural network model

[0163] Table 5 shows that the proposed method, which utilizes MFCC+GFCC fusion features, outperforms traditional one-dimensional neural network models (1D-CNN, LSTM, and Transformer) that directly input PPG signals in all performance metrics. The 1D-CNN model performs relatively poorly in all metrics, with both accuracy and F1-score around 92%. In contrast, the LSTM model achieves improvements in accuracy (94.53%) and F1-score (94.39%). The Transformer model further improves upon the LSTM model, but remains significantly lower than the proposed method. This result not only validates the superior performance of our model but also demonstrates the effectiveness of our method for extracting i-vectors and making decisions based on MFCC+GFCC fusion feature input. By more effectively extracting and utilizing key features from the data, it improves model performance in identity recognition tasks. This further demonstrates the significant potential and value of transforming data representation to improve model performance in signal processing.

[0164] To validate the performance of the proposed method for extracting i-vectors for decision-making based on MFCC+GFCC fusion feature input, an ablation experiment was conducted to verify whether the various improvements to the model improve the accuracy (Acc), precision (Pre), recall (Recall), and F1 score of identity recognition based on PPG signals.

[0165] Table 6 compares the performance of using MFCC features alone, GFCC features, and MFCC+GFCC fusion features. Using the same SVM model with different feature values, the model using MFCC+GFCC fusion features as input achieved an average 2.1% improvement in accuracy and 6.4% in recall compared to using MFCC and GFCC features alone. This demonstrates that the fused feature values ​​used in this paper retain more voiceprint information and are more conducive to improving biometric recognition accuracy.

[0166] Table 6 Ablation experiment results

[0167] This paper proposes a novel biometric authentication method based on photoplethysmography (PPG), which uses the i-vector framework to achieve robust verification. The i-vector method can compress the feature vector sequence into a compact low-dimensional representation, thereby capturing the overall variability of the PPG signal. After reducing the PPG noise through preprocessing, the Mel-frequency cepstral coefficients (MFCC) and Gammatone frequency cepstral coefficients (GFCC) are used to extract the feature vectors. Although these two techniques were originally designed for audio processing, MFCC can capture the overall spectral envelope of the signal, while GFCC emphasizes the finer spectral details of the low-frequency part through the Gammatone filter bank. These features form the basis of the universal background model (UBM) based on the Gaussian mixture model, from which the low-dimensional i-vector identity authentication vector is derived. Finally, the optimal authentication effect can be achieved by inputting the i-vector vector into the long short-term memory network (LSTM). Results show that the proposed method achieved average accuracy, precision, recall, and F1 score of 98.7%, 98.4%, 98.5%, and 98.4%, respectively, using a five-fold cross-validation dataset collected from 30 participants simulating three different motion scenarios. The proposed model significantly outperforms traditional biometric recognition methods based on pure neural network models. Experimental results demonstrate that this method has excellent biometric recognition performance, strong application value in training practice, and can provide a high-precision identity authentication solution for wearable devices.

[0168] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A biometric recognition method driven by PPG big data, characterized in that: include: Obtaining PPG signal data of the trainee, and performing high-pass filtering and low-pass filtering preprocessing on the PPG signal data to obtain a preprocessed PPG signal; Frame the preprocessed PPG signal, extract Mel-frequency cepstral coefficient features and Gammatone-frequency cepstral coefficient features respectively, fuse the Mel-frequency cepstral coefficient features and Gammatone-frequency cepstral coefficient features, and then perform principal component analysis and dimensionality reduction on the features to obtain fused features; Constructing a universal background model (UBM) of a Gaussian mixture model based on the fused features, calculating zero-order, first-order, and second-order Baum-Welch statistics, estimating a global difference space matrix from the Baum-Welch statistics, and extracting an i-vector identity authentication vector; The i-vector identity authentication vector is input into the long short-term memory network for training and classification to obtain the identity recognition result of the trainee.

2. The method according to claim 1, characterized in that The performing high-pass filtering and low-pass filtering preprocessing on the PPG signal data includes: Filter parameter configuration is performed on the PPG signal data, and the cutoff frequency of the fourth-order Chebyshev type I high-pass filter is set to 0.5 Hz, the passband fluctuation is set to within ±0.5 dB, and the stopband attenuation slope is set to negative 80 dB per decade to obtain the high-pass filter parameters; constructing the fourth-order Chebyshev type I high-pass filter based on the high-pass filter parameters, applying the fourth-order Chebyshev type I high-pass filter to the PPG signal data to eliminate baseline drift, and obtain a first filtered signal; According to the time domain and frequency domain characteristics of the first filtered signal, the filtering effect is evaluated, and the signal-to-noise ratio and the degree of baseline drift suppression are calculated to ensure that the preset signal quality indicators are met.

3. The method according to claim 2, characterized in that The performing high-pass filtering and low-pass filtering preprocessing on the PPG signal data includes: Performing filter parameter configuration on the first filtered signal, setting the cutoff frequency of the fourth-order elliptical low-pass filter to 18 Hz, setting the passband to 0 Hz to 18 Hz with a fluctuation of less than or equal to 0.1 decibel, setting the stopband to start at 22 Hz with an attenuation of greater than or equal to 40 decibels at 22 Hz and greater than or equal to 60 decibels at 50 Hz, and obtaining low-pass filter parameters; constructing the fourth-order elliptical low-pass filter based on the low-pass filter parameters, applying the fourth-order elliptical low-pass filter to the first filtered signal to eliminate motion artifacts, ambient light interference, and muscle noise, thereby obtaining the preprocessed PPG signal; According to the spectral characteristics of the preprocessed PPG signal, the noise suppression effect is evaluated and the degree of high-frequency interference suppression is calculated to ensure that the preset signal quality requirements are met.

4. The method according to claim 1, wherein The pre-processed PPG signal is divided into frames to extract Mel frequency cepstral coefficient features and Gammatone frequency cepstral coefficient features, respectively, including: Performing pre-emphasis processing on the pre-processed PPG signal to enhance the high-frequency component to obtain a pre-emphasized signal; Divide the pre-emphasized signal into frames according to a preset frame length, and add a Hamming window to each frame signal for smooth transition to obtain a first windowed signal; Performing a fast Fourier transform on the first windowed signal to obtain a spectrum, and inputting the spectrum into a Mel filter bank to perform frequency envelope extraction to obtain a Mel spectrum; The Mel-frequency cepstral coefficient feature is obtained by performing discrete cosine transform on the Mel-frequency spectrum to compress the feature dimension.

5. The method according to claim 1, wherein The framing of the pre-processed PPG signal and extracting Mel frequency cepstral coefficient features and Gammatone frequency cepstral coefficient features respectively include: Frame the preprocessed PPG signal and apply a Hamming window to each frame signal to obtain a second windowed signal; Performing a fast Fourier transform on the second windowed signal to obtain a frequency domain signal; Input the frequency domain signal into the Gammatone filter bank, perform frequency band division according to the ERB frequency scale, and obtain the filtered output of each channel; Energy calculation and logarithmic compression are performed on the filtered outputs of each channel, and discrete cosine transform is performed to obtain the Gammatone frequency cepstral coefficient characteristics.

6. The method according to claim 1, characterized in that The method includes: constructing a universal background model (UBM) of a Gaussian mixture model based on the fusion features, calculating zero-order, first-order, and second-order Baum-Welch statistics, estimating a global difference space matrix from the Baum-Welch statistics, and extracting an i-vector identity authentication vector. Based on the fusion features, the number of Gaussian mixture components is set to 32, a universal background model (UBM) is constructed, and its covariance matrix is ​​set to a diagonal matrix to obtain initial model parameters; According to the initial model parameters, the EM algorithm is used to calculate the zero-order, first-order and second-order Baum-Welch statistics, and the global difference space matrix is ​​iteratively estimated until the change in the log-likelihood value is less than a preset threshold, thereby obtaining the global difference space matrix; Based on the global difference space matrix and the Baum-Welch statistic, the i-vector identity authentication vector with a dimension of 200 is extracted through the maximum a posteriori probability estimation method.

7. The method according to claim 3, characterized in that Also includes: Obtaining the high-pass filter parameters, low-pass filter parameters, and extraction parameters of the Mel-frequency cepstral coefficient features and the Gammatone-frequency cepstral coefficient features, constructing a parameter optimization framework based on a genetic algorithm, and encoding the high-pass filter parameters, low-pass filter parameters, and the extraction parameters into chromosomes of an initial population; Based on the chromosomes of the initial population, calculating the fitness function value of each chromosome, wherein the fitness function value comprehensively evaluates three indicators: signal-to-noise ratio gain, signal quality index, and feature discrimination, to obtain a population fitness distribution; According to the population fitness distribution, the 10% individuals with the highest fitness are retained to enter the next generation, and the remaining 90% individuals are subjected to crossover and mutation using tournament selection, where the crossover rate is initially set to 0.8 and dynamically adjusted according to the population diversity, and the mutation rate is initially set to 0.1 and gradually decreases to 0.01 as the number of iterations increases, thereby obtaining a new generation of population; Performing simulated annealing local optimization on the 10% individuals with the highest fitness in the new generation population every 10 generations until the preset number of iterations is reached to obtain the optimal parameter configuration; Based on the optimal parameter configuration, personalized preprocessing and feature extraction are performed on the PPG signal of each trainee to generate optimized PPG signals and features; An incremental optimization mechanism is set up. When the recognition accuracy drops below a preset threshold or reaches a fixed optimization period, parameter re-optimization is triggered, and the optimization process is accelerated through parallel computing.

8. The method according to claim 1, characterized in that Also includes: Obtaining the initial fusion weights and dimensionality reduction matrices of the Mel-frequency cepstral coefficient features and the Gammatone-frequency cepstral coefficient features, constructing a dual-objective optimization problem for feature fusion, setting the first objective to maximize the Fisher discriminant ratio and the second objective to minimize the feature dimension; Based on the dual-objective optimization problem, the NSGA-II algorithm is used to generate the initial population. Each individual contains two parts: feature fusion weight and dimensionality reduction. The fusion weight range is [0, 1] and the sum is 1. The dimensionality reduction range is [8, 20]. The initial population is subjected to non-dominated sorting and crowding calculation, a parent population is generated through binary tournament selection, and a child population is generated using simulated binary crossover and polynomial mutation; a local search is performed on the non-dominated frontier every 10 generations, and the search step size is dynamically adjusted with the number of iterations to accelerate the convergence of the algorithm and obtain the Pareto optimal solution set; Based on the Pareto optimal solution set, a knee point detection method is used to select the optimal solution. The knee point detection method first arranges the solution set in ascending order according to the first objective, then calculates the curvature of adjacent solutions, selects the point with the largest curvature as the knee point, and generates a personalized feature fusion model; An adaptive adjustment mechanism is set to monitor the recognition performance. When the cumulative error rate exceeds 5% or after 30 days, the feature fusion model is automatically triggered to update to ensure stable performance.

9. The method according to claim 1, characterized in that Also includes: Construct the i-vector framework parameter space and the decision model parameter space. The i-vector framework parameter space includes the number of Gaussian mixture components, the rank of the T matrix, the eigenvector dimension, and the number of iterations. The decision model parameter space includes the number of layers, the number of hidden units, the dropout rate, and the learning rate of the LSTM network. The covariance matrix adaptive evolution strategy (CMA-ES) is used to optimize the parameter space of the i-vector framework and the parameter space of the decision model, respectively, and interactive optimization between subspaces is achieved through a three-layer collaborative mechanism of parameter cross-evaluation, fitness sharing, and search direction coordination. Dynamically adjust the computing resources allocated to each subspace based on the progress rate and performance improvement of the interactive optimization of the subspaces, and use multi-threading technology to simultaneously evaluate different parameter combinations; For different scenarios in training, typical scenario configurations are pre-defined and the optimization process is performed separately to generate scenario-specific parameter configuration schemes, and the parameter configurations are automatically switched according to the training scenarios detected in real time; A regularization control strategy is adopted to add L2 regularization constraints in the i-vector subspace to control the complexity of the T matrix. Batch normalization and early stopping strategies are introduced in the decision model subspace to improve the generalization ability of the model. Incremental optimization is achieved by adopting a two-level strategy of lightweight update strategy and comprehensive update strategy, wherein the lightweight update strategy fine-tunes parameters in a local range, and the comprehensive update strategy re-executes the complete co-evolutionary optimization after the accumulated new data exceeds the preset data value.

10. The method according to claim 1, characterized in that The i-vector identity authentication vector is input into the long short-term memory network for training and classification to obtain the identity recognition result of the trainee, including: Normalizing and reordering the i-vector identity authentication vector to generate a training sample sequence and corresponding identity label; Based on the training sample sequence and the corresponding identity label, a network structure is constructed including an input layer, two LSTM layers, and a fully connected layer, wherein the number of LSTM units in the first layer is 128, the number of LSTM units in the second layer is 64, and the dropout rate is set to 0.3 to obtain an initial network model; The cross entropy is used as the loss function, and the Adam optimizer is used to iteratively train the initial network model. The initial value of the learning rate is set to 0.001 and dynamically adjusted using the cosine annealing strategy to obtain a trained LSTM network model; The i-vector identity authentication vector to be identified is input into the trained LSTM network model, the probability distribution of each identity category is calculated through the Softmax classifier, and the category with the highest probability is selected as the identification result.

Citation Information

Patent Citations

  • Voice data processing method and device and voice interaction equipment

    CN107978311A

  • Flour quality detection method based on hybrid simulated annealing and genetic algorithms

    CN109342352A

  • Power equipment sound recognition method and system

    CN115691508A

  • Identity privacy protection method and system for voice data

    CN117831510A

  • Non-contact psychological state assessment method and system based on multi-modal fusion technology

    CN117936032A

Cited By

  • Metabolic equivalent driven endurance training intensity gradient generation system and method thereof

    CN121354803A

  • Metabolic equivalent driven endurance training intensity gradient generation system and method

    CN121354803B

  • Anti-noise PPG biological recognition method based on spectrum-KAN discretization and multi-scale convolution cooperation

    CN121905542A

  • Anti-noise ppg biometric identification method based on spectral-kan discretization and multi-scale convolution cooperation

    CN121905542B

  • PPG identification method based on biomechanical derivative and KAN-xLSTM

    CN122004803A