An intelligent identity recognition method based on hierarchical heterogeneous feature alignment knowledge distillation

CN122821594APending Publication Date: 2026-09-25HANGZHOU UNIV OF ELECTRONIC SCI & TECH PINGHU DIGITAL TECH INNOVATION RES INST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611086506.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0007]本发明旨在解决现有知识蒸馏方法应用于心电身份识别任务时存在的判别性监督不足、异构特征空间不匹配以及中间层对齐不充分的问题

Benefits of technology

[0018]与现有技术相比,本发明的有益效果主要体现在以下方面:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821594A_ABST
    Figure CN122821594A_ABST
Patent Text Reader

Abstract

The application discloses an intelligent identity recognition method based on layered heterogeneous feature alignment knowledge distillation, which comprises the following steps: firstly, preprocessing the original one-dimensional electrocardio signal to obtain a two-dimensional Gram angle difference field image; secondly, constructing a discriminative enhanced teacher network and a lightweight student network, introducing a triple loss and a cross-entropy loss into the teacher training to jointly optimize, and enhancing the in-class compactness and the inter-class separability of the embedding space; then, fixing the teacher network, mapping the student network's Transform label embedding to the CNN feature map space of the teacher network through a learnable projection layer, and realizing effective knowledge transfer between heterogeneous architectures by adopting a three-level joint supervision strategy; in the reasoning stage, only the distilled student network is reserved, and the subject level identity label is determined through a majority voting mechanism. The application can significantly reduce the model parameter quantity and the reasoning delay while maintaining the high-precision electrocardio identity recognition performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent biometrics and deep learning model compression technology, and specifically relates to a heterogeneous knowledge distillation framework based on hierarchical feature alignment and its application method in individual identity recognition based on electrocardiogram signals. Background Technology

[0002] Biometric authentication technology plays a crucial role in modern security systems. Electrocardiogram (ECG), as an intrinsic bioelectrical signal, possesses inherent uniqueness, persistence, and temporal continuity, providing inherent liveness detection capabilities and exhibiting stronger resistance to spoofing attacks compared to external biometrics such as faces and fingerprints. With the rapid development of wearable and portable data acquisition devices, ECG-based biometric identification (ECGID) has gradually become a research hotspot in the field of biometric recognition. Currently, deep learning-based ECGID methods have made significant progress, mainly including reference point-based and non-reference point methods. Reference point-based methods rely on manual feature engineering, extracting morphological features by detecting and analyzing reference points (such as peak amplitude, interval duration, and slope angle) of characteristic waves in the ECG signal; however, their performance heavily depends on the accuracy of reference point detection. Non-reference point methods automatically learn discriminative representations through deep neural networks, achieving higher recognition performance; however, they typically rely on deep architectures with millions of parameters, resulting in high computational overhead and storage requirements, making them difficult to deploy in resource-constrained embedded or mobile devices.

[0003] To address these challenges, lightweight model techniques have been extensively studied. Network pruning and quantization often lead to unstructured sparsity, potentially causing irreversible performance degradation. While directly adopting lightweight architectures (such as MobileNet and ShuffleNet) improves efficiency, their feature representation capabilities are limited. Knowledge distillation (KD), as a typical model compression paradigm, transfers knowledge learned by a large-scale teacher model to a lightweight student model, significantly reducing model complexity while maintaining competitive performance, and has received increasing attention in recent years.

[0004] However, existing knowledge distillation methods still face the following shortcomings in ECGID tasks: (1) Insufficient discriminative supervision. Traditional teacher models only use cross-entropy loss for training and lack explicit metric learning constraints, resulting in insufficient intra-class compactness and inter-class separability of teacher embeddings, making it difficult to provide high-quality discriminative knowledge sources for subsequent distillation.

[0005] (2) Feature space mismatch. Existing feature alignment methods are mainly designed for homogeneous networks. When dealing with heterogeneous teacher-student architectures (such as CNN teachers and Transformer students), there are problems of dimensional mismatch and semantic misalignment between CNN feature maps and Transformer labeled embeddings.

[0006] (3) Insufficient alignment of intermediate layers. Early KD methods only limited knowledge transfer to the final output distribution (Logit-level distillation), which could not effectively transfer the hierarchical discriminative representation of the intermediate layers of the teacher model, while the intermediate representation is particularly important for fine-grained identity recognition tasks. Summary of the Invention

[0007] This invention aims to address the problems of insufficient discriminative supervision, mismatch between heterogeneous feature spaces, and inadequate alignment of intermediate layers in existing knowledge distillation methods applied to ECG identity recognition tasks. Specifically, traditional teacher models lack metric learning constraints, feature dimensions and semantics are misaligned between heterogeneous architectures, and reliance solely on Logit-level distillation leads to insufficient intermediate representation propagation, all of which contribute to the insufficient generalization performance of lightweight student models in fine-grained identity recognition tasks. This invention provides a high-precision, low-latency ECG identity recognition method for resource-constrained scenarios.

[0008] To achieve the above objectives, this invention proposes an intelligent identity recognition method based on heterogeneous knowledge distillation using hierarchical feature alignment, which mainly includes the following steps: A smart identity recognition method based on hierarchical heterogeneous feature alignment and knowledge distillation includes the following steps: S1: Preprocess the acquired raw one-dimensional electrocardiogram signal to obtain a two-dimensional Gramian Angular Difference Field (GADF) image. S2: Construct a discriminative enhanced heterogeneous teacher-student architecture, which includes a discriminative enhanced teacher network based on a convolutional neural network, a lightweight student network based on a Transformer, and a learnable feature alignment projection layer; S3: Using a two-dimensional Gram difference field image as input, the teacher network is trained by jointly optimizing cross-entropy loss and triplet loss. The parameters of the trained teacher network are fixed, and the student network and the learning feature alignment projection layer are jointly optimized. The knowledge of the teacher network is transferred to the student network through hierarchical heterogeneous feature alignment distillation to complete the distillation. S4: Only retain the distilled student network, preprocess the one-dimensional electrocardiogram signal to be identified and input it into the distilled student network for identity prediction, and determine the subject-level identity label through a majority voting mechanism.

[0009] Preferably, step S1 specifically includes: S11: Bandpass filtering is performed on the original one-dimensional electrocardiogram signal to remove baseline drift, power frequency interference and electromyographic noise; S12: QRS complex detection is performed based on the R wave peak point, dividing the continuous ECG signal into multiple single-cycle ECG signal segments; S13: After normalizing the amplitude of each single-cycle ECG signal segment to the [-1, 1] interval, convert it to polar coordinate angle space representation, calculate the Gram matrix to obtain a two-dimensional Gram angle difference field image; S14: Scale the two-dimensional Gram difference field image to a preset size and use it as input for the subsequent network.

[0010] Preferably, the discriminative enhanced teacher network is constructed based on ResNet34, retaining its four residual stages, removing the original ImageNet classification head, replacing it with a classifier specific to the ECG identity recognition task, and outputting C-dimensional logits, where C is the number of registered identities; a 128-dimensional embedding layer is inserted after the global average pooling layer to calculate the metric constraints of the triplet loss.

[0011] Preferably, the lightweight student network is constructed based on DeiT-tiny, containing 12 Transformer blocks, with an embedding dimension of 192, 3 attention heads, and a feedforward network hidden dimension of 768. The input 224×224 GADF image is divided into 14×14=196 image patches of size 16×16, and a patch embedding sequence is obtained by linear projection. Learnable classification labels are added before the sequence, and learnable position embeddings are added to obtain an input sequence consisting of 197 labels.

[0012] As a preferred option, the total loss function of the teacher network is: in, The cross-entropy loss of the teacher network is used to supervise the classification prediction of identity categories; The triplet loss is used to enhance intra-class compactness and inter-class separability in the embedding space; This is the balance coefficient; Defined as: in, For Euclidean distance, For the interval hyperparameter, , , These are the embedding vectors for the anchor sample, positive sample, and negative sample, respectively.

[0013] Preferably, hierarchical heterogeneous feature alignment distillation includes joint supervision at the following three levels: (a) Task supervision: Calculate the cross-entropy loss of student network outputs using real identity labels. ; (b) Logit-level distillation: The Logit-level distillation loss is calculated by aligning the soft output distributions of the teacher and student networks with the temperature coefficient τ using KL divergence. : in, For temperature coefficient, As the gradient compensation factor, and The logits are for teachers and students, respectively. This represents the Softmax probability distribution after temperature smoothing. (c) Multi-layer feature alignment: Between the outputs of the teacher network at multiple residual stages and the outputs of the corresponding Transformer layers of the student network, the student features are mapped to the teacher feature space through the learnable feature alignment projection layer, and then the L2 distance alignment loss is calculated. : in, For teacher-student level pairs, For the teacher network at the level Output feature map, For student networks in layers Output tag embedding, The student features are mapped by the projection layer. For normalization and reshaping operations; The total loss function of the student network is expressed as: in, and These are the weight coefficients for the Logit-level distillation loss and the feature alignment loss, respectively.

[0014] Preferably, the learnable feature alignment projection layer embeds the Transformer labels of the student network and projects them into the CNN feature map space of the teacher network. The projection direction is a mapping from the student feature space to the teacher feature space, rather than compressing the teacher features into the student space.

[0015] As a preferred approach, in multi-layer feature alignment, the multiple residual stages of the teacher network are selected from Stage 2, Stage 3, and Stage 4, which correspond to the three levels of local morphological patterns, intermediate texture features, and high-level semantic abstraction, respectively.

[0016] Preferably, step S4 specifically includes: The collected one-dimensional electrocardiogram signal is preprocessed according to step S1 to obtain a two-dimensional Gram difference field image. The image is then input into the distilled student network to obtain sample-level identity prediction results. The number of votes obtained for each identity category is counted, and the identity with the highest number of votes is taken as the subject-level predicted identity. Calculate the voting consistency rate, which is the proportion of the highest number of votes to the total number of votes. If the voting consistency rate is greater than or equal to a preset threshold, the identity prediction result is accepted; otherwise, it is rejected.

[0017] Secondly, the present invention provides an intelligent identity recognition system based on hierarchical heterogeneous feature alignment knowledge distillation, comprising: The ECG signal preprocessing module is used to preprocess the acquired raw one-dimensional ECG signal and output a two-dimensional Gram difference field image. The discriminative enhanced teacher network module is built on a convolutional neural network and uses joint optimization training with cross-entropy loss and triplet loss to provide discriminative knowledge sources. A lightweight student network module, built on Transformer, is used to receive the Gram angular difference field image and output the identity prediction result; A learnable feature alignment projection layer module is used to embed and project the Transformer labels of the student network into the CNN feature map space of the teacher network; The hierarchical heterogeneous feature alignment distillation module is used to fix the parameters of the teacher network module and jointly optimize the student network module and the learnable feature alignment projection layer module. It transfers the knowledge of the teacher network to the student network through task supervision, Logit-level distillation and multi-layer feature alignment. The majority voting identity reasoning module is used to aggregate the identity prediction results through majority voting to determine the subject-level identity label.

[0018] Compared with the prior art, the beneficial effects of the present invention are mainly reflected in the following aspects: This invention addresses the problem of insufficient discriminative features in ECG identity recognition. It introduces cross-entropy loss and triplet loss for joint optimization of the teacher network, enabling it to learn identity embeddings with higher intra-class compactness and inter-class separability, providing a discriminative knowledge source for subsequent distillation. Secondly, based on the hierarchical representation of ECG signals from local waveform features and intermediate texture features to high-level identity semantics, this invention establishes multi-layer feature alignment relationships between the teacher network's Stage 2, Stage 3, and Stage 4 and the corresponding layers of the student network, achieving hierarchical transfer of ECG discriminative information. Thirdly, this invention employs a learnable projection method, mapping student Transformer features to the teacher's CNN feature space, rather than compressing teacher features, effectively preserving the ECG discriminative information learned by the teacher network and improving the heterogeneous feature transfer effect. Finally, this invention combines task supervision, Logit distillation, and multi-layer feature supervision to form a joint distillation mechanism for ECG identity recognition, enabling the lightweight student network to achieve high identity recognition performance while maintaining low computational complexity. Therefore, this invention is not a direct application of existing heterogeneous knowledge distillation techniques, but rather a targeted heterogeneous knowledge transfer scheme for ECG identity recognition scenarios. Attached Figure Description

[0019] To further understand the technical solutions and processes of the present invention, the accompanying drawings required in the description of the embodiments or prior art will be briefly introduced below. The accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 The overall flowchart of the intelligent identity recognition method based on hierarchical heterogeneous feature alignment and knowledge distillation provided by the present invention; Figure 2 The present invention provides a teacher-student network model structure diagram based on CNN and Transformer, wherein (a) is a teacher network diagram based on discriminative enhancement; and (b) is a lightweight student network diagram based on Transformer. Detailed Implementation

[0021] The technical solution of the present invention will be further described in detail below through embodiments and in conjunction with the accompanying drawings.

[0022] like Figure 1 As shown, this invention provides an intelligent identity recognition method based on hierarchical heterogeneous feature alignment and knowledge distillation, comprising: S1: Construct a signal preprocessing method based on bandpass filtering noise reduction and GADF transform, and output GADF single-center images. This step converts the original one-dimensional electrocardiogram signal into an image representation suitable for two-dimensional deep networks, and is divided into four sub-steps.

[0023] (1) Bandpass filtering for noise reduction. The original ECG signal usually contains baseline drift (low-frequency noise), power line interference (50 / 60Hz) and electromyographic noise (high-frequency noise). Bandpass filters (usually with cutoff frequencies set at 0.5 Hz and 40 Hz) are used to filter out the above noise components and retain the effective frequency band of the ECG signal.

[0024] (2) Cardiac cycle segmentation. QRS complex detection is performed based on the R wave peak point to locate the start and end points of each cardiac cycle, and the continuous ECG signal is segmented into multiple single-cycle ECG signal segments. Each segment is an independent training / test sample.

[0025] (3) GADF Transform. For each single-cycle segment Perform the following operations in sequence: Amplitude normalized to Interval: Transform to polar coordinate angular space to obtain the time vector. : The GADF image is obtained by calculating the Gram matrix, specifically as follows: in It is the angle mapped to the polar coordinate system after time-series normalization; Using trigonometric identities, we get: Write two vectors: , Inner product form: The core advantage of the GADF transform lies in the fact that the inner product operation preserves the correlation between time points of the signal. Essentially, it's a point in time. and A measure of directional similarity in angular space. Due to the symmetry of the inner product ( The resulting Gram matrix is ​​a symmetric matrix, with its diagonal reflecting the amplitude variation trajectory of the signal and the off-diagonal region encoding the temporal correlation pattern.

[0026] (4) Uniform size. Scale the GADF image to... The pixels match the standard input size of ResNet and DeiT networks, without requiring modifications to existing network structures.

[0027] S2: Constructing a heterogeneous teacher-student architecture based on discriminative enhancement This step constructs a heterogeneous teacher-student architecture that includes a teacher network, a student network, and a feature alignment projection layer, providing a structural foundation for subsequent distillation.

[0028] (1) A discriminative enhanced teacher network based on ResNet34. For example... Figure 2 As shown in (a), the teacher network uses ResNet34 as the convolutional backbone, retains its four residual stages (Stages 1-4), removes the original ImageNet classification head, and replaces it with a classifier specific to the ECG identification task, outputting... dimensional logits ( (Number of registered identities). A 128-dimensional embedding layer is inserted after the global average pooling layer to compute the metric constraint of the triplet loss. The key design of this teacher network is the introduction of joint optimization of triplet loss and traditional cross-entropy loss—cross-entropy loss supervises the class decision boundary, while triplet loss directly optimizes the geometry of the embedding space through metric learning constraints (see step S3 for details).

[0029] (2) Lightweight student networks based on DeiT-tiny. For example... Figure 2 As shown in (b), the student network adopts the DeiT-tiny architecture, containing 12 Transformer blocks, an embedding dimension of 192, 3 attention heads, and a feedforward network with a hidden dimension of 768. Input GADF image, with Size cut into Image patches are linearly projected to obtain a patch embedding sequence. A learnable classification token and a learnable position embedding are added before the sequence to obtain an input sequence consisting of 197 tokens. After processing by 12 Transformer blocks sequentially, the final classification label is output as an identity prediction by a linear classifier. Compared to CNNs, the Transformer's self-attention mechanism can model the global dependencies of electrocardiogram signals.

[0030] (3) Learnable feature alignment projection layer. This is the core component for solving the feature space mismatch between heterogeneous teacher-student architectures. The ResNet34 teacher output is a three-dimensional spatial feature map ( ), while the DeiT-tiny student output is a two-dimensional labeled sequence ( The two pairs are misaligned in both dimensions and semantics. Therefore, an alignment layer is created for each teacher-student pair. Introducing a learnable projection layer (Implemented as a 1×1 convolutional or fully connected layer), the student Transformer labels are reshaped into a spatial grid and then mapped to the teacher feature map space: The key design choice is the projection direction: mapping student features to the teacher feature space, rather than compressing teacher features into the student space. This is because the teacher feature space is more expressive, and back projection would compress discriminative information.

[0031] S3: Online Training for Identifying Enhanced Teachers This step involves discriminative augmentation training of the ResNet34 teacher network, enabling it to learn identity embedding representations with high intra-class compactness and inter-class separability.

[0032] (1) Training data. The GADF image generated in step S1 is used as input, and the batch size is set to 32.

[0033] (2) Cross-entropy loss. Used for supervised identity classification tasks, enabling teachers to learn category-level decision boundaries: in For teachers' logits, This is a label for your real identity.

[0034] (3) Triple Loss. This is one of the core innovations of this invention. Traditional teachers only use cross-entropy for training, which cannot explicitly constrain the geometric structure of the embedding space, resulting in loose intra-class clustering and blurred inter-class boundaries. Triple loss, through the cross-entropy method, applies cross-entropy to each training sample (anchor point) to achieve the desired results. Select positive samples with the same identity. Negative samples with different identities Apply interval constraints: in For Euclidean distance, This is the interval hyperparameter. This loss forces the distance between the anchor point and the positive sample to be less than the distance between the anchor point and the negative sample by at least [percentage missing]. This results in compact intra-class clustering and clear inter-class separation within the embedding space.

[0035] (4) Joint optimization. The total loss function of the teacher network is: in The balancing coefficient is used. The effect of joint optimization is reflected in two aspects: cross-entropy loss ensures classification accuracy, and triplet loss optimizes the geometry of the embedding space. The two complement each other, enabling the teacher network to simultaneously possess correct category decisions and separable identity embeddings, providing a high-quality discriminative knowledge source for subsequent distillation.

[0036] (5) Save intermediate features. After training convergence, save the output feature maps of each residual stage: Stage 2 ( Stage 3 Stage 4 These correspond to discriminative representations at three levels: local morphology, intermediate texture, and high-level semantics, respectively, and are used for subsequent layered distillation.

[0037] S4: Distillation method based on hierarchical heterogeneous feature alignment This step, based on the discriminative augmented teacher network trained in step S3, fixes the teacher network parameters and jointly optimizes the DeiT-tiny student network and the learnable projection layer. The distillation objective includes three levels of joint supervision: (a) Task supervision – using real identity labels to calculate the cross-entropy loss of the student network output, preserving task-specific classification ability; (b) Logit-level distillation – aligning the soft output distributions of the teacher and student networks after smoothing by the temperature coefficient τ using KL divergence, conveying class-level decision knowledge; (c) Multi-layer feature alignment – ​​between the outputs of the teacher network in the 2nd to 4th residual stages and the corresponding Transformer layer outputs of the student network, the L2 distance alignment loss is calculated after mapping the student features to the teacher feature space through the learnable projection layer, conveying a hierarchical discriminative intermediate representation.

[0038] (1) Distillation setup and teacher forward inference. The GADF image generated in step S1 is used as input, and the batch size is set to 32. During the distillation process, the teacher network parameters are frozen and it only acts as a knowledge provider. Each GADF image is simultaneously input into the teacher network and the student network: the teacher network outputs the Softmax probability distribution softened by the temperature coefficient τ. Intermediate feature maps from Stage 2 to Stage 4 Student network output corresponding Intermediate marker embedding with Transformer layer Based on the teacher-student pair set defined in step S2. (in and learnable projection layer Student labels are embedded and projected into the teacher feature space: This provides a dimensionally and semantically consistent representation for subsequent multi-layer feature alignment.

[0039] (2) Task Supervision. The cross-entropy loss of the student network output is calculated using real identity labels to ensure that the student network retains its task-specific classification ability. The cross-entropy loss is defined as: in For student network logits, The true label for category c (one-hot encoded). The predicted student probabilities are calculated using the standard Softmax function (temperature coefficient τ=1).

[0040] (3) Logit-level distillation. Class-level decision knowledge is conveyed by aligning the soft output distributions of teachers and students after temperature coefficient τ with KL divergence. Introducing a temperature coefficient τ>1 softens the output probability distribution, exposing inter-class similarity relationships contained in low-confidence categories, providing richer supervision signals for lightweight student networks than hard labels. The temperature-smoothed Softmax is defined as: Based on temperature smoothing Softmax, the KL divergence loss of Logit-level distillation is defined as: in and Let be the logits for teachers and students, respectively, and τ be the temperature coefficient. In the formula, τ... 2 This is a gradient compensation factor used to counteract the scaling effect of temperature smoothing on the gradient. It is achieved by minimizing... The student network is encouraged to match the temperature-smoothed output distribution of the teacher in order to obtain category-level decision-making knowledge containing inter-class similarity relationships.

[0041] (4) Multi-layer feature alignment. Although Logit-level distillation can convey output layer decision knowledge, it cannot explicitly constrain intermediate layer representations. Therefore, the intermediate representations of teachers and students are further aligned at multiple semantic levels. Based on the teacher-student pair set defined in step S2... and learnable projection layer Projected student characteristics After normalizing and reshaping the corresponding teacher feature map, the L2 distance alignment loss is calculated: in For the teacher network at the level Output feature map, For student networks in layers Output tag embedding, For projection layer The student features mapped to the teacher feature space, Ψ(·) represents the normalization and reshaping operation, converting heterogeneous representations (CNN feature maps and Transformer labeled embeddings) into comparable forms. The total number of teachers and students. Minimize... The student network is encouraged to align with the teacher's hierarchical discriminative representation, thereby acquiring rich intermediate semantic knowledge beyond output-level supervision. The feature alignment layers are selected from the teacher network at Stage 2 (28×28×128), Stage 3 (14×14×256), and Stage 4 (7×7×512), corresponding to the local morphological patterns, intermediate texture features, and high-level semantic abstractions, respectively, ensuring that the student network receives effective teacher guidance at different receptive field scales.

[0042] (5) Joint optimization. The overall loss function for student network training combines the supervision terms at the three levels mentioned above to form an end-to-end heterogeneous knowledge distillation objective: Here, α and β control the contribution weights of the Logit-level distillation loss and feature alignment loss, respectively, and their optimal values ​​are determined through parameter sensitivity analysis. During the distillation process, the teacher network parameters... Keep it fixed, only for student network parameters and projection layer parameters Perform joint optimization. By minimizing... The student network retains task-specific classification capabilities (first aspect) while learning to mimic the teacher's soft output decision-making behavior (second aspect) and acquiring hierarchical intermediate discriminative representations (third aspect), thus achieving efficient ECG identity recognition within a compact DeiT-tiny architecture. After training convergence, the teacher network and all projection layers are discarded, retaining only the distilled student network for inference deployment.

[0043] S5: Subject-level Identity Reasoning Based on Majority Voting This step uses only the distilled DeiT-tiny student network (discarding the teacher and projection layers) during the inference phase, and achieves robust subject-level identity recognition through a multi-sample voting mechanism.

[0044] (1) Sample-level prediction. For the test subject Collected A single-cycle ECG sample Each sample is independently transformed into a GADF image in step S1, and after distillation, the student network obtains sample-level identity predictions. (2) Majority voting aggregation. The identities of each candidate are tallied. Number of votes received: Take the identity with the most votes as the subject-level prediction: (3) Consistency rate threshold rejection. Calculate the voting consistency rate: when At that time, accept The final identification result is determined by the given criteria; otherwise, the identification is rejected. The physical meaning of this rejection mechanism is that if the prediction results of multiple heartbeat samples of a subject are highly dispersed (low consistency rate), it indicates that the subject's electrocardiogram signal quality is poor or the characteristics are atypical. In this case, rejection is safer than forced classification.

[0045] The core value of the majority voting mechanism lies in the fact that while a single heartbeat may produce abnormal predictions due to transient noise, unstable electrode contact, or physiological fluctuations, the statistical aggregation of multiple heartbeats can smooth out these accidental factors. For example, if 4 out of 5 samples are predicted as identity A and 1 is misclassified as B due to noise, the majority vote will still correctly output A; if the consistency rate is 80% and exceeds the threshold... If the percentage is 70%, then accept the result.

[0046] Experimental environmental parameters and experimental results (1) Hardware and software environment: All experiments were conducted under the Python 3.10.19 and PyTorch 2.5.1 framework, accelerated by CUDA 12.1, and run on a workstation equipped with an NVIDIA RTX 4090 GPU (24GB video memory).

[0047] (2) Dataset partitioning and preprocessing: After bandpass filtering to remove noise, the original one-dimensional ECG signal was segmented into single-cycle ECG signal segments based on the R-wave peak point, and then transformed into two-dimensional images using GADF. Five-fold cross-validation was used, with five independent runs. Forty GADF image samples were selected for each subject, with 32 samples from each fold used for training and 8 samples used for testing. The final performance report is the mean ± standard deviation of all folds and runs.

[0048] (3) Batch size and input size: The training batch size is set to 32, and the input GADF image size is uniformly scaled to 224×224 pixels.

[0049] (4) Optimizer and learning rate: The teacher network (ResNet34) uses the AdamW optimizer with an initial learning rate of 1×10⁻⁶. -3 Weight decay 1×10 -4 The student network (DeiT-tiny) was trained using the AdamW optimizer with an initial learning rate of 1×10⁻⁶. -4 Weight decay 1×10 -4 .

[0050] (5) Number of training rounds: 30 epochs for teacher network training and 50 epochs for student network distillation training.

[0051] (6) Distillation hyperparameters: temperature coefficient τ=4, triplet loss interval hyperparameter m=0.5, triplet loss weight (balance coefficient) in teacher training is set to 0.1. Logit-level distillation weight α and feature alignment weight β are set according to the datasets respectively: ECG-ID dataset (α=0.1, β=0.5), PTB dataset (α=1.0, β=1.0), CYBHi dataset (α=0.3, β=0.2).

[0052] (7) Teacher network loss function: joint optimization of cross-entropy loss and triple loss, where triple loss is used to enhance intra-class compactness and inter-class separability of the embedding space, providing a discriminative knowledge source for subsequent distillation.

[0053] (8) Student network distillation loss function: includes three levels of joint supervision -- (a) Task supervision: real label cross-entropy loss; (b) Logit level distillation: KL divergence alignment of the soft output distribution of the teacher network after smoothing by temperature coefficient τ; (c) Multi-layer feature alignment: L2 distance alignment loss between the output of the teacher network Stage 2, Stage 3, Stage 4 and the corresponding Transformer layer output of the student network after mapping by a learnable projection layer.

[0054] (9) Evaluation metrics: Recognition performance is measured by accuracy (ACC), precision (P), recall (R) and F1 score (F1); lightweight metrics are measured by number of parameters (Params, M), number of floating-point operations (FLOPs, G) and inference latency (Latency, ms).

[0055] Table 1 Table 2 The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An intelligent identity recognition method based on hierarchical heterogeneous feature alignment and knowledge distillation, characterized in that, Includes the following steps: S1: Preprocess the acquired raw one-dimensional electrocardiogram signal to obtain a two-dimensional Gram angle difference field image; S2: Construct a discriminative enhanced heterogeneous teacher-student architecture, which includes a discriminative enhanced teacher network based on a convolutional neural network, a lightweight student network based on a Transformer, and a learnable feature alignment projection layer; S3: Using a two-dimensional Gram difference field image as input, the teacher network is trained by jointly optimizing cross-entropy loss and triplet loss. The parameters of the trained teacher network are fixed, and the student network and the learning feature alignment projection layer are jointly optimized. The knowledge of the teacher network is transferred to the student network through hierarchical heterogeneous feature alignment distillation to complete the distillation. S4: Only retain the distilled student network, preprocess the one-dimensional electrocardiogram signal to be identified and input it into the distilled student network for identity prediction, and determine the subject-level identity label through a majority voting mechanism.

2. The intelligent identity recognition method based on hierarchical heterogeneous feature alignment and knowledge distillation according to claim 1, characterized in that, Step S1 specifically includes: S11: Bandpass filtering is performed on the original one-dimensional electrocardiogram signal to remove baseline drift, power frequency interference and electromyographic noise; S12: QRS complex detection is performed based on the R wave peak point, dividing the continuous ECG signal into multiple single-cycle ECG signal segments; S13: After normalizing the amplitude of each single-cycle ECG signal segment to the [-1, 1] interval, convert it to polar coordinate angle space representation, calculate the Gram matrix to obtain a two-dimensional Gram angle difference field image; S14: Scale the two-dimensional Gram difference field image to a preset size and use it as input for the subsequent network.

3. The intelligent identity recognition method based on hierarchical heterogeneous feature alignment and knowledge distillation according to claim 1, characterized in that, The discriminative enhanced teacher network is built on ResNet34, retaining its four residual stages, removing the original ImageNet classifier head and replacing it with a classifier specific to the ECG identity recognition task, outputting C-dimensional logits, where C is the number of registered identities; a 128-dimensional embedding layer is inserted after the global average pooling layer to calculate the metric constraints of the triplet loss.

4. The intelligent identity recognition method based on hierarchical heterogeneous feature alignment and knowledge distillation according to claim 1, characterized in that, The lightweight student network is built on DeiT-tiny and contains 12 Transformer blocks.

5. The intelligent identity recognition method based on hierarchical heterogeneous feature alignment and knowledge distillation according to claim 1, characterized in that, The total loss function for discriminative enhanced teacher networks is: in, To determine the cross-entropy loss of the enhanced teacher network, it is used to supervise the classification prediction of identity categories; The triplet loss is used to enhance intra-class compactness and inter-class separability in the embedding space; This is the balance coefficient; Defined as: in, For Euclidean distance, For the interval hyperparameter, , , These are the embedding vectors for the anchor sample, positive sample, and negative sample, respectively.

6. The intelligent identity recognition method based on hierarchical heterogeneous feature alignment and knowledge distillation according to claim 5, characterized in that, Hierarchical heterogeneous feature alignment distillation includes joint supervision at the following three levels: (a) Task supervision: Calculate the cross-entropy loss of student network outputs using real identity labels. ; (b) Logit-level distillation: The Logit-level distillation loss is calculated by aligning the soft output distributions of the teacher and student networks with the KL divergence and smoothed by the temperature coefficient τ. : in, For temperature coefficient, As the gradient compensation factor, and The logits are for teachers and students, respectively. This represents the Softmax probability distribution after temperature smoothing. (c) Multi-layer feature alignment: Between the outputs of the teacher network at multiple residual stages and the outputs of the corresponding Transformer layers of the student network, the student features are mapped to the teacher feature space through the learnable feature alignment projection layer, and then the L2 distance alignment loss is calculated. : in, For teacher-student level pairs, For the teacher network at the level Output feature map, For student networks in layers Output tag embedding, The student features are mapped by the projection layer. For normalization and reshaping operations; The total loss function of the lightweight student network is expressed as: in, and These are the weight coefficients for the Logit-level distillation loss and the feature alignment loss, respectively.

7. The intelligent identity recognition method based on hierarchical heterogeneous feature alignment and knowledge distillation according to claim 6, characterized in that, The learnable feature alignment projection layer embeds the Transformer labels of the student network and projects them into the CNN feature map space of the teacher network, with the projection direction being a mapping from the student feature space to the teacher feature space.

8. The intelligent identity recognition method based on hierarchical heterogeneous feature alignment and knowledge distillation according to claim 7, characterized in that, In multi-layer feature alignment, the multiple residual stages of the teacher network are selected from Stage 2, Stage 3 and Stage 4, which correspond to the three levels of local morphological patterns, intermediate texture features and high-level semantic abstraction, respectively.

9. The intelligent identity recognition method based on hierarchical heterogeneous feature alignment and knowledge distillation according to claim 1, characterized in that, Step S4 specifically includes: The collected one-dimensional electrocardiogram signal is preprocessed according to step S1 to obtain a two-dimensional Gram difference field image. The image is then input into the distilled student network to obtain sample-level identity prediction results. The number of votes obtained for each identity category is counted, and the identity with the highest number of votes is taken as the subject-level predicted identity. Calculate the voting consistency rate, which is the proportion of the highest number of votes to the total number of votes. If the voting consistency rate is greater than or equal to a preset threshold, the identity prediction result is accepted; otherwise, it is rejected.

10. An intelligent identity recognition system based on hierarchical heterogeneous feature alignment knowledge distillation, implementing the method as described in any one of claims 1-9, characterized in that, include: The ECG signal preprocessing module is used to preprocess the acquired raw one-dimensional ECG signal and output a two-dimensional Gram difference field image. The discriminative enhanced teacher network module is built on a convolutional neural network and uses joint optimization training with cross-entropy loss and triplet loss to provide discriminative knowledge sources. A lightweight student network module, built on Transformer, is used to receive the Gram angular difference field image and output the identity prediction result; A learnable feature alignment projection layer module is used to embed and project the Transformer labels of the student network into the CNN feature map space of the teacher network; The hierarchical heterogeneous feature alignment distillation module is used to fix the parameters of the teacher network module and jointly optimize the student network module and the learnable feature alignment projection layer module. It transfers the knowledge of the teacher network to the student network through task supervision, Logit-level distillation and multi-layer feature alignment. The majority voting identity reasoning module is used to aggregate the identity prediction results through majority voting to determine the subject-level identity label.