A method and system for evaluating the authenticity of running data based on biometric identification

By collecting and processing exercise data and biometric data, a fused feature vector is generated, which solves the problem of insufficient data credibility in the campus running APP system, realizes high-precision running data authenticity assessment and adaptive anti-counterfeiting capability, and improves the system's anti-counterfeiting capability and the reliability of assessment results.

CN121350919BActive Publication Date: 2026-06-02XIANGTAN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIANGTAN INST OF TECH
Filing Date
2025-10-22
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing campus running app systems lack dynamic correlation of biometrics and fail to build a strong coupling system between exercise data and individual physiological characteristics, resulting in a strategic shortcoming in data credibility. This leads to cheating methods such as identity theft, using substitutes for other people's transportation, and falsifying GPS parameters, which affects the effectiveness of physical education.

Method used

By collecting motion and biometric data during running, timestamp-aligned data records are generated. Filtering and noise reduction, outlier removal, and missing data completion are performed. Kinematic and biometric features are extracted, and deep learning algorithms are used to generate fused feature vectors. Time consistency detection and physiological and motion data rationality analysis are performed to generate credibility scores and output authenticity verification reports. Model parameters are optimized to feed back into data processing and feature extraction steps.

Benefits of technology

It significantly improves the accuracy of evaluating the authenticity of running data, enhances anti-counterfeiting capabilities, achieves adaptive evolution capabilities, reduces system operation and maintenance and manual intervention costs, and ensures the reliability and interpretability of evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121350919B_ABST
    Figure CN121350919B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent sports monitoring, and in particular to a running data authenticity evaluation method and system based on biometric recognition, which comprises the following steps: collecting motion data and biometric data during running and generating timestamp-aligned collected data records, filtering and denoising the collected data, removing outliers, completing missing data and standardizing the format, obtaining standardized data, extracting kinematic features and biometric features from the standardized data, constructing a multi-modal feature set, generating a fusion feature vector through a deep learning algorithm, performing time consistency detection, data integrity verification and physiological and motion data rationality analysis based on the fusion feature vector, generating a credibility score, and generating an authenticity verification report based on the credibility score. The present application can effectively improve the credibility of running data and provide accurate risk assessment and recommendations for athletes and data users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent sports monitoring technology, and in particular to a method and system for evaluating the authenticity of running data based on biometric recognition. Background Technology

[0002] While current campus running app systems have achieved initial standardized management of exercise data collection, they have exposed strategic shortcomings in data reliability in practice. New cheating methods, such as identity theft, using substitutes for others to run, and falsifying GPS parameters, seriously hinder the effectiveness of physical education reform. These problems not only deviate from the essential requirement of enhancing students' health but also easily lead to secondary problems such as distorted health records. At its root, the existing verification mechanisms lack dynamic correlation with biometrics, failing to construct a strong coupling system between exercise data and individual physiological characteristics, and failing to form a systematic closed loop covering pre-event prevention, in-event monitoring, and post-event evaluation. Summary of the Invention

[0003] To overcome the above shortcomings, this invention provides a method and system for evaluating the authenticity of running data based on biometric recognition. It aims to improve the existing verification mechanisms, which lack dynamic correlation of biometrics, fail to build a strong coupling system between exercise data and individual physiological characteristics, and fail to form a systematic closed loop covering pre-event prevention, in-event monitoring, and post-event evaluation.

[0004] In a first aspect, the present invention provides the following technical solution: a method for evaluating the authenticity of running data based on biometric recognition, comprising the following steps:

[0005] S1. Collect exercise data and biometric data during the running process, and generate timestamp-aligned data records;

[0006] S2. The collected data records are filtered and denoised, outlier removed, missing data filled in, and format standardized to obtain standardized data;

[0007] S3. Extract kinematic and biological features from the standardized data, construct a multimodal feature set, and generate a fused feature vector through a deep learning algorithm;

[0008] S4. Based on the fused feature vector, perform time consistency detection, data integrity verification, and physiological and exercise data rationality analysis to generate a credibility score;

[0009] S5. Based on the credibility score, generate an authenticity verification report and output the credibility assessment results of the running data, risk warnings, and running logs;

[0010] S6. Based on the abnormal patterns identified during the verification process, optimize the model parameters and feed the updated parameters back to the data processing and feature extraction steps.

[0011] Preferably, in step S1, generating timestamp-aligned collected data records specifically includes:

[0012] At the acquisition device end, assign a source timestamp to each data point;

[0013] Based on a unified time reference, source timestamps from different acquisition devices are aligned and corrected.

[0014] The corrected and aligned timestamps are bound to the corresponding motion data and biometric data to form the timestamp-aligned acquisition data record;

[0015] Verify the time series continuity and data synchronization consistency of the collected data records.

[0016] Preferably, in step S2, the filtering and noise reduction, outlier removal, missing data completion, and format standardization of the collected data records specifically include:

[0017] Filtering and noise reduction processing are performed on the biometric signals and motion signals in the collected data records;

[0018] Based on the statistical distribution characteristics of the collected data records, abnormal data points are identified and removed;

[0019] Estimate and fill in missing values ​​based on one or more valid data points surrounding the missing data points;

[0020] The data is formatted and converted to unify its timestamp format, physical units, field naming, and data type, resulting in the standardized data.

[0021] Preferably, in step S3, extracting kinematic and biological features from the standardized data to construct a multimodal feature set specifically includes:

[0022] Extract kinematic features from standardized data, including cadence, stride length, gait cycle, speed, and acceleration;

[0023] Extract biometrics from standardized data, including heart rate, respiratory rate, electromyography signals, and heart rate variability;

[0024] The extracted kinematic and biological features are combined to construct a multimodal feature set.

[0025] Preferably, in step S3, generating the fused feature vector using a deep learning algorithm specifically includes:

[0026] The extracted multimodal feature set is processed by a deep learning algorithm to generate a fusion feature vector representing the runner's behavior and physiological state.

[0027] Preferably, in step S4, based on the fused feature vector, performing time consistency detection, data integrity verification, and physiological and motion data rationality analysis to generate a credibility score specifically includes:

[0028] Verify the continuity of the timestamp sequence and the reasonableness of the time interval contained in the fused feature vector;

[0029] Verify the data integrity of the fused feature vector and perform completion processing on any missing data.

[0030] The coupling relationship between physiological data and motion data in the fused feature vector is analyzed, and its consistency is evaluated by comparing it with a preset physiological-kinematic correlation model.

[0031] The credibility score is generated by calculating the comprehensive verification and evaluation results using a preset model.

[0032] Preferably, in step S5, based on the credibility score, an authenticity verification report is generated, and the running data credibility assessment results, risk warnings, and running logs are output, specifically including:

[0033] The credibility score is compared with a preset credibility threshold to generate a credibility assessment result for the running data.

[0034] Based on the credibility score and evaluation results, an authenticity verification report containing anomaly detection analysis and risk assessment is generated.

[0035] Based on the credibility score, the risk warning is generated, and the risk warning indicates potential areas of data anomaly;

[0036] The operation log is generated, which records key steps, parameter settings, evaluation results, and risk warnings during the verification process.

[0037] The authenticity verification report, credibility assessment results, risk warnings, and operation logs will be visualized and output.

[0038] Preferably, in step S6, optimizing the model parameters based on the abnormal patterns identified during the verification process and feeding the updated parameters back to the data processing and feature extraction steps specifically includes:

[0039] Based on a preset model, abnormal patterns in the fused feature vector are identified;

[0040] Based on the identified abnormal patterns, the model parameters are adjusted, including feature weights, feature selection configuration, and training hyperparameters.

[0041] The adjusted model parameters are fed back to the data processing and feature extraction steps to update the filtering configuration, data completion strategy, data alignment parameters, or feature extraction weights.

[0042] The model is iterated using an online learning mechanism.

[0043] Secondly, this invention provides the following technical solution: a running data authenticity assessment system based on biometric recognition, the system comprising:

[0044] The data acquisition layer is used to collect motion data and biometric data generated during running from the user's end through sensor devices, and generate timestamp-aligned data records.

[0045] The data processing system is used to filter and reduce noise, remove outliers, fill in missing data, and standardize the format of the collected data records to generate standardized data.

[0046] The feature extraction module is used to extract kinematic and biological features from the standardized data and generate a fused feature vector through a deep learning algorithm.

[0047] The verification engine is used to perform time consistency detection, data integrity verification, and physiological and motion data rationality analysis based on the fused feature vector, and generate a credibility score.

[0048] The analysis platform is used to generate an authenticity verification report based on the credibility score, and output the credibility assessment results of the running data, risk warnings, and running logs.

[0049] The feedback application layer is used to optimize model parameters based on the abnormal patterns identified during the verification process, and to feed the updated parameters back to the data processing system and feature extraction module.

[0050] The present invention has the following beneficial effects:

[0051] 1. In this invention, kinematic and biological features are extracted from multimodal data, a fusion feature vector is generated using a deep learning algorithm, and the deep coupling relationship and rationality between physiological data and exercise data are further analyzed based on the vector. This improves the accuracy of evaluating the authenticity of running data and can effectively identify complex cheating patterns that are difficult to detect by traditional methods and where kinematic and physiological features do not match. This significantly enhances the anti-counterfeiting capability and evaluation accuracy of the system.

[0052] 2. In this invention, by setting up a feedback application layer, new abnormal patterns that appear during the verification process can be automatically identified, and the model parameters can be dynamically optimized and iterated using an online learning mechanism. At the same time, the updated parameters are fed back to the data processing system and feature extraction module, enabling the system to have adaptive evolution capabilities. This continuously improves the robustness of detecting unknown or variant cheating methods and reduces the cost of system operation and maintenance and manual intervention.

[0053] 3. In this invention, by performing refined filtering and noise reduction, outlier removal and missing value completion on the collected data during the data processing stage, the data quality of subsequent feature extraction and analysis is ensured. The analysis platform outputs a detailed verification report containing credibility assessment results, risk warnings and operation logs, realizing closed-loop management of the entire process from raw data cleaning to final traceability, which significantly improves the reliability, completeness and interpretability of the assessment results. Attached Figure Description

[0054] Figure 1 This is a flowchart of a method for evaluating the authenticity of running data based on biometric recognition proposed in this invention.

[0055] Figure 2 This is an architecture diagram of a running data authenticity assessment system based on biometric recognition proposed in this invention. Detailed Implementation

[0056] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] Example 1:

[0058] In a first embodiment of the present invention, the present invention provides a method for evaluating the authenticity of running data based on biometric recognition, such as... Figure 1 As shown, it includes the following steps:

[0059] S1. Collect exercise data and biometric data during the running process, and generate timestamp-aligned data records;

[0060] Furthermore, in S1, generating timestamp-aligned collected data records specifically includes:

[0061] At the acquisition device end, assign a source timestamp to each data point;

[0062] Based on a unified time reference, source timestamps from different acquisition devices are aligned and corrected.

[0063] The corrected and aligned timestamps are bound to the corresponding motion data and biometric data to form the timestamp-aligned acquisition data record;

[0064] Verify the time series continuity and data synchronization consistency of the collected data records.

[0065] Specifically, data acquisition can be accomplished through a single integrated device, such as a smartwatch with built-in multiple sensors, or through the collaboration of multiple independent devices, such as a smartphone and a Bluetooth heart rate monitor. These physical devices or their internal sensor modules are collectively referred to as "acquisition device ends" in this invention. The motion data acquisition end is handled by an inertial measurement unit (IMU), which acquires triaxial acceleration and triaxial angular velocity signals at a higher frequency, such as 100Hz. The biometric data acquisition end is handled by a photoplethysmography (PPG) sensor, which acquires heart rate and heart rate variability data at a lower frequency, such as 1Hz. At the start of acquisition, each device independently assigns a source timestamp based on its own internal clock to each data point. Since the clocks of different acquisition device ends have inherent offsets and drifts, a unified time reference must be established, such as the operating system clock of a smartphone, and all source timestamps must be corrected to this reference. In this embodiment, correction is achieved through a linear clock model, and the algorithm is as follows: ;in, This represents the final corrected and aligned timestamp, an output associated with the smartphone's system time. This is an input parameter representing the original source timestamp generated by a certain acquisition device. It is the clock skew correction factor, and These are clock drift correction factors. These two correction factors are calculated at the start of the data acquisition session through the exchange of timestamp information between the master and slave devices. For example, for data points acquired by a heart rate monitor that include heart rate values ​​and source timestamps, their corresponding unified timestamps are obtained through... The calculations show that all data points are mapped onto the same precise time axis.

[0066] After timestamp correction and alignment, the timestamps are bound to the corresponding original data to form structured, timestamp-aligned acquisition data records. This step typically involves resampling, such as unifying all data streams to a frequency of 10Hz, downsampling high-frequency data, and upsampling low-frequency data using methods like linear interpolation, ultimately forming a time-strictly aligned multidimensional data sequence. After generating the records, a final verification is performed, including checking the continuity of the unified timestamp sequence, such as confirming whether the time interval is constant, and the synchronization consistency of all data fields at any given timestamp, ensuring data integrity. Through time synchronization and correction, data accuracy is guaranteed and data quality is improved, ensuring the reliability and accuracy of subsequent data analysis. Especially in complex environments with simultaneous acquisition from multiple devices, it effectively avoids data inconsistencies caused by time errors.

[0067] S2. The collected data records are filtered and denoised, outlier removed, missing data filled in, and format standardized to obtain standardized data;

[0068] Furthermore, in S2, the specific steps of filtering and noise reduction, outlier removal, missing data completion, and format standardization of the collected data records include:

[0069] Filtering and noise reduction processing are performed on the biometric signals and motion signals in the collected data records;

[0070] Based on the statistical distribution characteristics of the collected data records, abnormal data points are identified and removed;

[0071] Estimate and fill in missing values ​​based on one or more valid data points surrounding the missing data points;

[0072] The data is formatted and converted to unify its timestamp format, physical units, field naming, and data type, resulting in the standardized data.

[0073] Specifically, after generating timestamp-aligned acquisition data records in step S1, the raw sensor data stream inevitably contains noise, transient outliers, and data packet loss caused by factors such as device hardware characteristics, environmental electromagnetic interference, or atypical user movements (e.g., impacts from arm swings). The purpose of step S2 is to perform a series of refined preprocessing steps on this high-fidelity but unprocessed data record to generate standardized data that is clean at the signal level, complete at the data level, and uniform at the structural level. This lays a highly reliable data foundation for feature extraction in step S3 and physiological-kinematic correlation analysis in step S4.

[0074] First, targeted filtering and noise reduction processing is performed on the biometric and motion signals in the collected data records. For motion data, such as the triaxial acceleration signal from the inertial measurement unit (IMU), the main noise sources are high-frequency jitter and measurement background noise. In this embodiment, a second-order Butterworth low-pass filter is used to process it, with a cutoff frequency set, for example, 20Hz. This effectively filters out high-frequency noise unrelated to human movement while preserving the main signal components of running gait dynamics (such as ground contact and takeoff impact) without distortion. For biometric data, such as the photoplethysmography (PPG) signal, which is highly susceptible to motion artifacts, a bandpass filter is used to strictly limit its passband frequency to the effective physiological range of 0.5Hz to 3Hz. This range corresponds to the expected heart rate range of 30 to 180 beats per minute, thereby effectively separating the dominant heart rate frequency and suppressing noise interference in other frequency bands.

[0075] After filtering and noise reduction, based on the statistical distribution characteristics of the collected data records, transient outlier data points in the data stream are identified and removed. The Z-score algorithm is used to dynamically identify outliers over a sliding time window. This algorithm determines whether a data point is an outlier by quantifying its deviation from its local mean; the calculation formula is as follows:

[0076] ;

[0077] in, These are input parameters, representing the time points. A specific data point, such as heart rate or raw acceleration. and These represent the data within the current sliding time window. The local mean and local standard deviation within the range. This is the corresponding output result, i.e., the Z-score. Set a preset threshold. ,For example When a certain data point Corresponding The absolute value of the value is greater than the threshold. At that time, the data point Values ​​identified as statistical outliers are removed from the data sequence and their positions are temporarily marked as invalid or left blank.

[0078] Then, based on one or more valid data points surrounding the missing data point, the missing values ​​resulting from outlier removal or loss of the original signal are estimated and filled in. This embodiment uses linear interpolation for data completion to maintain the continuity of the data sequence. This method uses the two nearest valid data points before and after the missing point to linearly estimate the value of the missing point. The algorithm is as follows:

[0079] ;

[0080] in, The timestamp representing the missing data point. This is the missing data value to be calculated; this is one output result. The input parameters are the two nearest valid data points to the missing data point. and ,in and These are timestamps of valid data points. and These are the corresponding valid sensor readings, such as acceleration or heart rate values. This allows for the effective recovery of the continuity and integrity of the data sequence.

[0081] Finally, the data is formatted and transformed to unify its timestamp format, physical units, field naming, and data types, resulting in the standardized data. For example, all timestamps are unified to a UNIX timestamp format with millisecond precision; all acceleration units are unified to the international standard unit meters per second squared. All data fields are named in a standardized manner, such as "heart_rate" and "accel_x"; and all values ​​are standardized to 32-bit floating-point number type. After this process, the output standardized data has a high degree of structural consistency, unit uniformity, and program readability.

[0082] All data undergoes formatting and transformation to standardize timestamp format, physical units, field naming, and data types, ensuring data consistency and usability. These processing steps significantly improve data quality, providing a reliable foundation for subsequent feature extraction and analysis.

[0083] S3. Extract kinematic and biological features from standardized data, construct a multimodal feature set, and generate a fused feature vector through a deep learning algorithm;

[0084] Furthermore, in S3, extracting kinematic and biological features from standardized data to construct a multimodal feature set specifically includes:

[0085] Extract kinematic features from standardized data, including cadence, stride length, gait cycle, speed, and acceleration;

[0086] Extract biometrics from standardized data, including heart rate, respiratory rate, electromyography signals, and heart rate variability;

[0087] The extracted kinematic and biological features are combined to construct a multimodal feature set.

[0088] Furthermore, in S3, the generation of fused feature vectors using deep learning algorithms specifically includes:

[0089] The extracted multimodal feature set is processed by a deep learning algorithm to generate a fusion feature vector representing the runner's behavior and physiological state.

[0090] Specifically, the extraction of kinematic and biological features is performed first. This process is typically carried out over a sliding time window, for example, calculating feature values ​​every 2 seconds, with 50% overlap between windows.

[0091] Extracting kinematic features from standardized data can specifically include acceleration, which refers to a comprehensive index, such as calculating the vector magnitude of a three-axis acceleration signal, i.e., the total magnitude of acceleration. The calculation formula is as follows:

[0092] ;

[0093] in, , , These are input parameters, representing the triaxial acceleration values ​​from the standardized data. The output features include: Step frequency (PF), which is the number of peaks detected per unit time, such as steps per minute; Gait period, which is related to PF and is the time interval between two consecutive ground-touching peaks, measured in milliseconds; Stride length, which is a function of running speed and PF and can be estimated using known speed and calculated PF; and Speed, which can be extracted directly from GPS module data in the standardized data or estimated in specific scenarios by integrating the acceleration signal and combining it with a zero-speed correction algorithm.

[0094] Biometric features extracted from standardized data may include: heart rate, which can be directly obtained from the heart rate signal field in the standardized data, representing the average heart rate value within the current time window, in units of beats per minute; heart rate variability, which is an indicator of changes in heart rhythm, quantified by calculating the standard deviation (SDNN) of the interval between adjacent heartbeats, i.e., the NN interval; respiratory rate, which can be extracted from the baseline drift of the ECG or PPG signal in the standardized data, and this drift is usually modulated by changes in the thoracic cavity caused by respiratory movements; and electromyography (EMG) signals. If the acquisition device includes an EMG sensor, the standardized data will contain the corresponding EMG signals. In this embodiment, this feature is extracted by calculating the root mean square (RMS) value of the EMG signal within the time window to characterize the intensity of muscle activity.

[0095] After all features have been extracted, all kinematic and biological features calculated within the same time window are combined to construct a multimodal feature set. This collection It is a vector, which is at time point Its form can be represented as:

[0096] ;

[0097] Then, the multimodal feature set is processed using a deep learning algorithm to generate a fused feature vector. Because... It is a time-varying sequence, and a recurrent neural network (RNN), specifically a long short-term memory (LSTM) network, is used to process the time series of this feature set. LSTM networks can effectively capture complex nonlinear dependencies between features and between features over time. The abstract process of this deep learning algorithm can be represented as:

[0098] ;

[0099] Among them, input parameters The multimodal feature set is in continuous The sequence at each time step. This represents a pre-trained Long Short-Term Memory (LSTM) network model. Output results It is the network at the last time step The output hidden state vector The output vector This is the fusion feature vector defined in this invention. It is a low-dimensional, dense real-number vector, where each dimension integrates comprehensive information from both kinematic and biological modalities over a period of time, thus efficiently representing the runner's current behavioral and physiological coupling state.

[0100] This allows high-dimensional, heterogeneous, and multimodal raw time-series data to be transformed into a low-dimensional, homogeneous, and information-dense fused feature vector. This vector not only compresses the data, but also, by using deep learning algorithms to process the multimodal feature set, generates a fused feature vector that more accurately represents runners' behavior and physiological state, improving the accuracy of subsequent analysis and validation. This process ensures efficient feature extraction and fusion, helping to improve the overall performance of the running data authenticity assessment system and enhancing its predictive ability and adaptability.

[0101] S4. Based on the fused feature vector, perform time consistency detection, data integrity verification, and rationality analysis of physiological and exercise data to generate a credibility score;

[0102] Furthermore, in S4, based on the fused feature vector, time consistency detection, data integrity verification, and physiological and motion data rationality analysis are performed to generate a credibility score, specifically including:

[0103] Verify the continuity of the timestamp sequence and the reasonableness of the time interval contained in the fused feature vector;

[0104] Verify the data integrity of the fused feature vectors and perform completion processing on any missing data.

[0105] The coupling relationship between physiological and kinematic data in the fused feature vectors was analyzed, and the consistency was evaluated by comparing it with the pre-defined physiological-kinematic association model.

[0106] The credibility score is generated by calculating the comprehensive verification and evaluation results through a preset model.

[0107] Specifically, a time consistency check is first performed. This check verifies the continuity of the timestamp sequence and the reasonableness of the time intervals within the fused feature vector sequence. Since the fused feature vector is generated based on a sliding time window in S3, its timestamps should have a fixed time interval, such as 1 second. This step calculates the consistency between adjacent fused feature vectors. and Time difference between and compare it with the preset standard interval. Compare. If Significantly greater than If so, it is marked as a discontinuous time series or a missing data frame.

[0108] Subsequently, the system performs a data integrity check. This check examines the fused feature vector. To ensure data integrity, this step verifies whether the vector contains invalid values, such as NaN or Inf values. These invalid values ​​may be caused by calculation errors in the deep learning model in S3 or data completion failures in S2. If missing data frames are identified in the temporal consistency check, this step will perform completion processing. For example, linear interpolation is used to estimate in the fused feature vector space:

[0109] ;

[0110] in, It is the output vector to be completed. and These are the two nearest valid fused feature vectors on either side of the missing location; they are the input parameters of this algorithm. and These are their corresponding timestamps. It is the timestamp of the missing vector.

[0111] Next, a rationality analysis of the physiological and exercise data is performed. The coupling relationship between physiological and exercise data in the fused feature vectors is analyzed, and their consistency is evaluated against a pre-defined physiological-kinematic correlation model. This pre-defined model is an anomaly detection model, such as Isolation Forest, pre-trained on a training set containing massive amounts of real, valid running data. This model learns the complex nonlinear correlation patterns between kinematic features (such as cadence and acceleration) and biological features (such as heart rate and heart rate variability) under various exercise intensities. During evaluation, each fused feature vector... This is fed into the isolated forest model as an input parameter:

[0112] ;

[0113] Model output results This is an outlier score. This score quantifies the input... The degree to which a vector deviates from the "normal physiological-motor coupling relationship". For example, a fused feature vector that represents "high cadence, high speed, but extremely low heart rate" will be identified as highly abnormal by the model and given a high abnormality score. .

[0114] Finally, based on the combined results of all the above verifications and evaluations, a final credibility score is calculated using a pre-defined aggregation model. This embodiment uses a weighted summation model to generate this score:

[0115] ;

[0116] Among them, the output results This is the final credibility score, whose value range is typically normalized to between 0 and 1. Input parameters include: This represents the time consistency test result; for example, 1 represents complete continuity and 0 represents severe missing data. This represents the data integrity verification result; for example, 1 represents no invalid values, and 0 represents corrupted data. It is the average anomaly score of all fused feature vectors during the entire running session; , , These are preset weighting coefficients used to adjust the impact of different verification items on the final score.

[0117] This led to the construction of a progressive verification process, from basic time-series verification to deep model analysis. The fused feature vectors generated by S3 were used to assess the rationality of the coupling relationship between physiological and motion data at a higher dimension. This process can accurately identify "fake" data patterns that are difficult to detect using traditional methods, such as genuine motion data but falsified physiological data, thus providing a quantitative basis for credibility assessment.

[0118] S5. Based on the credibility score, generate an authenticity verification report and output the credibility assessment results of the running data, risk warnings, and running logs;

[0119] Furthermore, in S5, based on the credibility score, an authenticity verification report is generated, and the following information is output: running data credibility assessment results, risk warnings, and running logs.

[0120] The credibility score is compared with a preset credibility threshold to generate a credibility assessment result for the running data.

[0121] Based on the credibility score and evaluation results, an authenticity verification report is generated, which includes anomaly detection analysis and risk assessment.

[0122] Based on the credibility score, risk warnings are generated, indicating potential areas of data anomaly.

[0123] Generate an operation log, which records key steps, parameter settings, evaluation results, and risk warnings during the verification process;

[0124] The authenticity verification report, credibility assessment results, risk warnings, and operation logs will be visualized and output.

[0125] Specifically, the purpose of this step is to convert the quantitative credibility score calculated in step S4 into an assessment conclusion, risk identification, and detailed report that provides clear guidance for users or management systems, and to ensure the traceability of the entire verification process.

[0126] First, the credibility score output by S4 is compared with a preset credibility threshold to generate a credibility assessment result for the running data. The algorithm for this process can be expressed as follows:

[0127] ;

[0128] Among them, input parameters It is a comprehensive credibility score from S4 that represents the entire running session. This is a preset threshold value, for example, 0.8. Output result. It is a discrete evaluation label, such as "credible" or "uncredible," which provides a clear binary judgment for the authenticity of this running data.

[0129] Subsequently, based on the credibility score and evaluation results, an authenticity verification report is generated, including anomaly detection analysis and risk assessment. This report is a structured document that aggregates key analytical results from S4, such as the final... value, The average anomaly score of the physiological-kinematic association model output in S4 is denoted by the label. And the data missing rate detected in S2, etc.

[0130] Meanwhile, the system is based on the time-series anomaly score sequence generated in S4. This generates a risk warning. The risk warning indicates potential areas of data anomaly. Specifically, the system will iterate through... A sequence, when at a certain point in time or a period of consecutive time... Above, abnormal scores Exceeding a preset risk threshold The system records the time period. The final risk warning is a list containing all identified abnormal time periods and their possible causes, such as "Time 00:10:15-00:10:45, Risk: Severe mismatch between heart rate and exercise intensity".

[0131] The system also generates a runtime log, which details the key steps, parameter settings, evaluation results, and risk warnings during the verification process. This log is a technical log used for auditing and debugging. It records the time synchronization parameters in S1, the filter type and Z-score threshold in S2, and the version of the LSTM model used in S3. Weighting coefficients and the decision threshold in S5 .

[0132] Finally, the system will provide a visual output of the authenticity verification report, credibility assessment results, risk warnings, and operational logs. = Credibility Assessment Results Present the verification report to the user in the most intuitive way, such as a green "trusted" icon or a red "untrusted" icon. The authenticity verification report can be generated as a PDF file or displayed in a dedicated interface within the application. Risk Warning Then, in a graphical way, such as on a time-series chart of running data (e.g., a heart rate-speed curve), the identified abnormal time periods are highlighted. (Running log) It will then be written to a specified log file or database table for system administrators to view.

[0133] This effectively transforms the abstract evaluation scores output by S4 into multi-layered, actionable, and easily understandable output results. It not only provides a final, clear "credible" or "uncredible" conclusion, but also accurately identifies anomalies in the data through risk alerts, and ensures the transparency, interpretability, and traceability of the entire evaluation process through verification reports and operational logs, enabling the evaluation results to be truly adopted and utilized by users and management systems.

[0134] S6. Based on the abnormal patterns identified during the verification process, optimize the model parameters and feed the updated parameters back to the data processing and feature extraction steps;

[0135] Furthermore, in S6, based on the abnormal patterns identified during the verification process, the model parameters are optimized, and the updated parameters are fed back to the data processing and feature extraction steps. Specifically, this includes:

[0136] Based on a pre-defined model, identify abnormal patterns in the fused feature vectors;

[0137] Based on the identified abnormal patterns, adjust the model parameters, including feature weights, feature selection configuration, and training hyperparameters.

[0138] The adjusted model parameters are fed back to the data processing and feature extraction steps to update the filtering configuration, data completion strategy, data alignment parameters, or feature extraction weights.

[0139] An iterative model using an online learning mechanism is employed.

[0140] Specifically, firstly, based on the pre-defined physiological-kinematic correlation model (e.g., isolated forest) in S4, batch or continuous fused feature vectors are analyzed to identify abnormal patterns. When a certain number or proportion of fused feature vectors are detected... Assigned a high anomaly score At this point, a model optimization process will be triggered. The system will collect vectors with high outlier scores and use an unsupervised clustering algorithm (such as density-based noise spatial clustering DBSCAN) to perform pattern recognition on these outlier vectors. The output of this algorithm is to divide the outlier data into different clusters, each cluster representing a specific outlier pattern, such as "Pattern A: Sudden increase in exercise intensity but delayed heart rate response" or "Pattern B: Stable cadence but drastic fluctuations in heart rate variability."

[0141] After identifying specific abnormal patterns, the relevant model parameters are adjusted according to the characteristics of these patterns. These model parameters may include feature weights, feature selection configurations, and training hyperparameters. For example, if the system detects that "pattern A" occurs frequently, it will automatically increase the weight of the LSTM network in S3 that focuses on the temporal coupling relationship between the features "rate of change of velocity" and "rate of change of heart rate". At the same time, the system can adjust the "contamination" hyperparameter of the isolated forest model in S4 to improve its sensitivity to this specific pattern.

[0142] Subsequently, the adjusted model parameters are fed back to data processing step S2 and feature extraction step S3 to optimize the front-end processing logic. For example, based on an anomaly pattern caused by a certain high-frequency noise, the system can update the filter configuration, specifically the cutoff frequency parameter of the Butterworth filter, and feed it back to S2. Similarly, based on the adjusted feature weights, the multimodal feature set constructed in S3 can be updated. The weighting coefficients of each original feature, or even temporarily masking some features that are judged to have low contribution or noise sources.

[0143] To achieve continuous model iteration, this embodiment employs an online learning mechanism. The system uses a small training batch (mini-batch) of newly identified and confirmed anomalous data samples for incremental training of the LSTM model in S3 or the Isolation Forest model in S4. This process updates the model parameters through one or more iterations of the gradient descent algorithm, with the abstract update rules as follows:

[0144] ;

[0145] in This represents the set of model parameters that need to be optimized, such as the weight matrix of an LSTM network. Input parameters include the old model parameters. And a training batch consisting of anomalous samples. . It is the learning rate hyperparameter. This is the gradient of the preset loss function L with respect to the parameters. This loss function aims to maximize the model's ability to identify anomalous samples in this batch. Output results This refers to the updated model parameters.

[0146] This effectively improves the system's adaptability, ensuring that the model can maintain high performance and reliability when facing new data environments.

[0147] Example 2:

[0148] While current campus running app systems have achieved basic standardized management of exercise data collection, they have exposed strategic shortcomings in data reliability in practice. New cheating methods such as identity theft, using substitutes for others to run, and falsifying GPS parameters seriously hinder the effectiveness of physical education reform. To address these issues, this invention provides a running data authenticity assessment system based on biometric recognition, such as... Figure 2 As shown. The system includes:

[0149] The data acquisition layer is used to collect motion data and biometric data generated during running from the user's end through sensor devices, and generate timestamp-aligned data records.

[0150] The data processing system is used to filter and reduce noise, remove outliers, fill in missing data, and standardize the format of the collected data records to generate standardized data.

[0151] The feature extraction module is used to extract kinematic and biological features from the standardized data and generate a fused feature vector through a deep learning algorithm.

[0152] The verification engine is used to perform time consistency detection, data integrity verification, and physiological and motion data rationality analysis based on the fused feature vector, and generate a credibility score.

[0153] The analysis platform is used to generate an authenticity verification report based on the credibility score, and output the credibility assessment results of the running data, risk warnings, and running logs.

[0154] The feedback application layer is used to optimize model parameters based on the abnormal patterns identified during the verification process, and to feed the updated parameters back to the data processing system and feature extraction module.

[0155] Specifically, the data acquisition layer collects motion and biological data through user-end sensors and aligns the timestamps. The data processing system then uses a Butterworth low-pass filter to reduce noise, employs the Z-score algorithm to remove outliers, and uses linear interpolation to complete missing data. Finally, standardized data is generated in a unified format. The feature extraction module extracts features such as step frequency, acceleration modulus, heart rate, and heart rate variability over a sliding time window, constructs a multimodal feature set, and inputs it into a Long Short-Term Memory (LSTM) network model to generate a fused feature vector representing the coupled physiological and motor states. The verification engine receives this vector, verifies its temporal continuity and data integrity, and inputs it into a pre-trained Isolation Forest model to analyze its physiological-kinematic rationality. Finally, a weighted summation model is used to calculate a comprehensive credibility score. The analysis platform compares this score with a preset threshold, generates a "credible" or "unreliable" assessment result, highlights abnormal time periods as risk warnings, and outputs a complete verification report and operation log. Finally, the feedback application layer collects the feature vectors that are judged to be abnormal, identifies new abnormal patterns through clustering algorithms such as DBSCAN, and uses an online learning mechanism to iteratively optimize the parameters of the LSTM model or the isolated forest model. At the same time, the updated filter configuration and other parameters are fed back to the data processing system to form an adaptive optimization closed loop.

[0156] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for evaluating the authenticity of running data based on biometric recognition, characterized in that, Includes the following steps: S1. Collect exercise data and biometric data during the running process, and generate timestamp-aligned data records; S2. The collected data records are filtered and denoised, outlier removed, missing data filled in, and format standardized to obtain standardized data; S3. Extract kinematic and biological features from the standardized data, construct a multimodal feature set, and generate a fused feature vector through a deep learning algorithm; S4. Based on the fused feature vector, perform time consistency detection, data integrity verification, and physiological and exercise data rationality analysis to generate a credibility score; S5. Based on the credibility score, generate an authenticity verification report and output the credibility assessment results of the running data, risk warnings, and running logs; S6. Based on the abnormal patterns identified during the verification process, optimize the model parameters and feed the updated parameters back to the data processing and feature extraction steps; In step S4, based on the fused feature vector, time consistency detection, data integrity verification, and physiological and exercise data rationality analysis are performed to generate a credibility score, specifically including: Verify the continuity of the timestamp sequence and the reasonableness of the time interval contained in the fused feature vector; Verify the data integrity of the fused feature vector and perform completion processing on any missing data. The coupling relationship between physiological data and motion data in the fused feature vector is analyzed, and its consistency is evaluated by comparing it with a preset physiological-kinematic association model, wherein the preset physiological-kinematic association model is an isolated forest model. Based on the comprehensive verification and evaluation results, the credibility score is calculated and generated through a preset model, which is a weighted summation model. In step S6, optimizing the model parameters based on the abnormal patterns identified during the verification process and feeding the updated parameters back to the data processing and feature extraction steps specifically includes: Based on the preset physiological-kinematic correlation model, abnormal patterns in the fused feature vector are identified; Based on the identified abnormal patterns, the model parameters are adjusted. The model parameters include the parameters of the feature extraction model or the parameters of the physiological-kinematic association model, specifically including feature weights, feature selection configuration, and training hyperparameters. The feature extraction model is a Long Short-Term Memory (LSTM) network model. The adjusted model parameters are fed back to the data processing and feature extraction steps to update the filtering configuration, data completion strategy, data alignment parameters, or feature extraction weights. The feature extraction model or the physiological-kinematic association model is iterated using an online learning mechanism.

2. The method for evaluating the authenticity of running data based on biometric recognition according to claim 1, characterized in that, In step S1, generating timestamp-aligned collected data records specifically includes: At the acquisition device end, assign a source timestamp to each data point; Based on a unified time reference, source timestamps from different acquisition devices are aligned and corrected. The corrected and aligned timestamps are bound to the corresponding motion data and biometric data to form the timestamp-aligned acquisition data record; Verify the time series continuity and data synchronization consistency of the collected data records.

3. The method for evaluating the authenticity of running data based on biometric recognition according to claim 1, characterized in that, In step S2, the filtering and noise reduction, outlier removal, missing data completion, and format standardization of the collected data records specifically include: Filtering and noise reduction processing are performed on the biometric signals and motion signals in the collected data records; Based on the statistical distribution characteristics of the collected data records, abnormal data points are identified and removed; Estimate and fill in missing values ​​based on one or more valid data points surrounding the missing data points; The data is formatted and converted to unify its timestamp format, physical units, field naming, and data type, resulting in the standardized data.

4. The method for evaluating the authenticity of running data based on biometric recognition according to claim 1, characterized in that, In step S3, extracting kinematic and biological features from the standardized data to construct a multimodal feature set specifically includes: Extract kinematic features from standardized data, including cadence, stride length, gait cycle, speed, and acceleration; Extract biometrics from standardized data, including heart rate, respiratory rate, electromyography signals, and heart rate variability; The extracted kinematic and biological features are combined to construct a multimodal feature set.

5. The method for evaluating the authenticity of running data based on biometric recognition according to claim 1, characterized in that, In step S3, generating the fused feature vector using a deep learning algorithm specifically includes: The extracted multimodal feature set is processed by a deep learning algorithm to generate a fusion feature vector representing the runner's behavior and physiological state.

6. The method for evaluating the authenticity of running data based on biometric recognition according to claim 1, characterized in that, In step S5, based on the credibility score, an authenticity verification report is generated, and the running data credibility assessment results, risk warnings, and running logs are output, specifically including: The credibility score is compared with a preset credibility threshold to generate a credibility assessment result for the running data. Based on the credibility score and evaluation results, an authenticity verification report containing anomaly detection analysis and risk assessment is generated. Based on the credibility score, the risk warning is generated, and the risk warning indicates potential areas of data anomaly; The operation log is generated, which records key steps, parameter settings, evaluation results, and risk warnings during the verification process. The authenticity verification report, credibility assessment results, risk warnings, and operation logs will be visualized and output.

7. A running data authenticity evaluation system based on biometric recognition, characterized in that, The system for evaluating the authenticity of running data based on biometric recognition as described in any one of claims 1-6 comprises: The data acquisition layer is used to collect motion data and biometric data generated during running from the user's end through sensor devices, and generate timestamp-aligned data records. The data processing system is used to filter and reduce noise, remove outliers, fill in missing data, and standardize the format of the collected data records to generate standardized data. The feature extraction module is used to extract kinematic and biological features from the standardized data and generate a fused feature vector through a deep learning algorithm. The verification engine is used to perform time consistency detection, data integrity verification, and physiological and motion data rationality analysis based on the fused feature vector, and generate a credibility score. The analysis platform is used to generate an authenticity verification report based on the credibility score, and output the credibility assessment results of the running data, risk warnings, and running logs. The feedback application layer is used to optimize model parameters based on the abnormal patterns identified during the verification process, and to feed the updated parameters back to the data processing system and feature extraction module.

Citation Information

Patent Citations

  • Physical education teaching management system based on big data

    CN119692803A

  • Exercise risk assessment system based on big data

    CN120565082A