Source domain irrelevant cross-domain cardiac beat identification method and system for pseudo label mining
By using pseudo-label mining and data augmentation techniques, and optimizing the target domain model with a local-global semantic perception strategy, the problems of inaccurate pseudo-labels and insufficient model generalization in cross-domain recognition of ECG signals were solved, achieving higher recognition accuracy and generalization ability.
Patent Information
- Application Number
- CN202511003947.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-07
AI Technical Summary
Existing methods for cross-domain recognition of ECG signals suffer from inaccurate pseudo-label mining and insufficient model generalization ability when the target domain label is missing and source domain data cannot be obtained when crossing domains.
A source-domain-independent cross-domain heartbeat recognition method based on pseudo-label mining is proposed. Pseudo-labels are screened by pre-training the source domain model. Pseudo-label revision strategy with local-global semantic awareness and data augmentation technology are combined to optimize the pseudo-labels of the target domain model. Data manipulation between sample representations and between categories is used to improve the reliability of pseudo-labels and the generalization ability of the model.
In cases where target domain data labels are missing and source domain data is unavailable, the accuracy of cross-domain heartbeat recognition of ECG signals and the generalization ability of the model are significantly improved, the problem of false label mining errors is solved, and the recognition performance of target domain data is enhanced.
Smart Images

Figure CN120910646A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of pseudo-label learning, in particular to a source domain independent cross-domain heartbeat recognition method and system based on pseudo-label mining. BACKGROUND
[0002] Cardiovascular disease (CVD) is a serious threat to global human life and health. The changes in the waveform of the electrocardiogram (ECG) signal can reflect the pathology of CVD and is a commonly used diagnostic tool in clinical practice. ECG is usually operated by medical personnel and is used to identify cardiac electrical activity and detect arrhythmia. With the development of machine learning, numerous heartbeat recognition methods based on deep learning have been widely used in arrhythmia detection and have made significant progress to an advanced level, but they often face the challenge of domain shift caused by changes in ECG waveforms and features. Unsupervised domain adaptation (UDA) provides a possible solution to this challenge, which transfers knowledge from a labeled source domain to a target domain without labels. Given the privacy of electrocardiogram data, it is of great significance to explore the use of pre-trained source models for knowledge transfer instead of ECG data in the source domain under the condition that the target domain lacks labels and the source domain data cannot be obtained for the clinical application of intelligent diagnosis of CVD.
[0003] Under the condition that the target domain lacks labels and the source domain data cannot be obtained when crossing domains, existing researchers have proposed a source domain independent unsupervised domain adaptation (SFUDA) method, which has made some progress in image classification tasks. However, in the task of intelligent recognition of electrocardiogram signals, existing SFUDA methods can be divided into two categories: methods based on the idea of reconstruction and methods based on pseudo-label learning. The former usually designs a generative adversarial model to reconstruct virtual source domain data, and then uses these additional samples together with the target data to reduce the distribution difference between the cross-domain data by using traditional domain adaptation strategies. The second method usually uses the class discrimination ability provided by the knowledge of the trained source domain model to obtain target pseudo-labels, and then optimizes the model through self-training. Although existing methods have achieved encouraging performance, they also face their own unique challenges. For example, the former reconstruction-based method involves the complexity of source data generation and introduces additional network parameters. The latter pseudo-label-based method ignores the structural representation of the target domain, which is crucial for mining reliable pseudo-labels. SUMMARY
[0004] To solve the problems mentioned in the background, the purpose of the present application is to provide a source domain independent cross-domain heartbeat recognition method and system based on pseudo-label mining.
[0005] In a first aspect, the object of the present application can be achieved by the following technical solution: a source domain independent cross-domain heartbeat recognition method for pseudo-label mining, comprising the following steps:
[0006] Obtaining electrocardio data with heartbeat labels in a source domain, inputting the electrocardio data with heartbeat labels in the source domain into a pre-established source domain model for pre-training, and obtaining a pre-trained source domain model;
[0007] Obtaining electrocardio data without labels in a target domain, filtering the target domain data based on the pre-trained source domain model and a preset category threshold, obtaining data with high pseudo-label confidence and data with low pseudo-label confidence, initializing parameters of a target domain model with the pre-trained source domain model, updating the pseudo-labels of the data with low pseudo-label confidence based on a local-global semantic perception pseudo-label revision strategy, merging the data with high pseudo-label confidence, and obtaining updated pseudo-labeled target domain data;
[0008] Performing data augmentation on the updated pseudo-labeled target domain data using a sample representation mixup operation and an intra-class sample data mixup operation, obtaining augmented target domain data, inputting the augmented target domain data into a target domain feature extractor and a target domain classifier, outputting a target domain model optimization total loss function, and realizing source domain independent cross-domain heartbeat intelligent recognition.
[0009] In combination with the first aspect, in some implementations of the first aspect, the method further comprises: the inputting of the electrocardio data with heartbeat labels in the source domain into the pre-established source domain model for pre-training comprises:
[0010] The pre-training of the source domain model comprises pre-training the source domain model M s using the electrocardio data D s with heartbeat labels in the source domain according to the following formula (1) using a label-smoothed cross-entropy loss:
[0011]
[0012] In formula (1), represents the kth category heartbeat label processed by a smoothing operation, a represents a smoothing hyperparameter, K represents the number of all heartbeat categories, and δ k represents the probability prediction result of the kth category output by a softmax function.
[0013] In combination with the first aspect, in some implementations of the first aspect, the method further comprises: the filtering of the target domain data based on the pre-trained source domain model and the preset category threshold comprises:
[0014] Given a target domain data set X tFor any unlabeled data x in the source domain, if the pre-trained source domain model M... s If two labeling networks P1 and P2 predict the same heartbeat category, and the predicted classification probability value of the corresponding category is higher than a pre-set threshold, then the data x is classified into the pseudo-label sample set with high confidence according to equation (2). Otherwise, according to equation (3), they are assigned to the sample set with low confidence. middle:
[0015]
[0016] In equation (2), τ c This represents the classification probability threshold set for category c, where c∈{1,2,...,K}.
[0017] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: initializing the target domain model parameters with a pre-trained source domain model, updating the pseudo-labels of low-confidence data based on a local-global semantic awareness pseudo-label revision strategy, and updating the pseudo-labels of data with low pseudo-label confidence, including:
[0018] Target domain model parameter initialization and representation acquisition: The target domain model is initialized using pre-trained source domain model parameters. A target domain feature extractor is then used to acquire deep representations of samples from datasets with low pseudo-label confidence and datasets with high pseudo-label confidence, respectively. For datasets with low pseudo-label confidence in the target domain... The data in the dataset is used to obtain corresponding deep representations using a target domain feature extractor, based on target domain datasets with high confidence levels from pseudo-labels. From the data in the dataset, obtain the corresponding depth representation;
[0019] Calculate the nearest neighbor representation: for any data Corresponding depth representation Calculate its nearest neighbor representation in the high-confidence sample space.
[0020] Update the pseudo-labels for low-confidence data in the target domain: First, based on the high-confidence dataset... Obtain the representation prototype μ for each category k ,right any data Corresponding depth representation Calculate the similarity between each class and its representation center to obtain the normalized class prediction probability q. i ,calculate Nearest neighbor high confidence representation The corresponding normalized class prediction probability value qi' , q i , and q i' Weighted calculation, update target domain low confidence sample Corresponding pseudo label
[0021] In combination with the first aspect, in some implementations of the first aspect, the method further comprises: the pseudo label revision strategy based on local-global semantic perception comprises:
[0022] For any low confidence data Deep representation According to the nearest neighbor distance evaluation distance function Nearest according to formula (4), the nearest neighbor
[0023]
[0024] Based on the target domain high confidence dataset According to formula (5), the class prototype μ of class c is calculated c :
[0025]
[0026] Wherein, Indicates the sample Corresponding pseudo label, I represents an indicator function;
[0027] For any one data in the low confidence data set According to formula (6), the feature representation of the data is calculated The distance Dis between each class prototype μ c Obtain the normalized similarity score
[0028]
[0029] For Nearest neighbor representation According to formula (7), the distance Dis between each class prototype μ c Obtain the normalized similarity score
[0030]
[0031] In combination with the two normalized similarity scores And According to formula (8), the pseudo label of the target domain low confidence sample Is updated:
[0032]
[0033] where w reflects the adaptive weight of each sample confidence, which is calculated by equation (9):
[0034]
[0035] where, represents the entropy of the probability distribution generated by the target domain model prediction of the sample, p i = [p i1 , p i2 ,..., p iK ] is the probability vector output by the softmax function, representing the probability of the sample belonging to each of the K different categories.
[0036] In combination with the first aspect, in some implementations of the first aspect, the method further comprises: using the updated pseudo-labeled target domain data for data augmentation by sample representation intermixup operation and same category sample intra-data mixup operation to obtain augmented target domain data, including:
[0037] Augmenting the pseudo-labeled target domain data: given a mini-batch of samples (x1, y1),..., (x B , y B ) in the pseudo-labeled target domain data, where y i represents the pseudo-label corresponding to the sample x i , for each sample in each category c, construct the intra-class augmented feature representation by equation (10)
[0038]
[0039] In equation (10), the weight parameter β i is independently sampled from a uniform distribution U(0,1);
[0040] In addition, randomly shuffle the given mini-batch data to obtain reordered sequence samples (x'1, y'1),..., (x' B , y' B ), and obtain the inter-class augmented samples (x' mix , y' mix ) by equation (11) and equation (12):
[0041] x′ mix = λ·x i + (1-λ)·x′ i (11)
[0042] y′ mix = λy · y i + (1 - λ y ) · y' i (12)
[0043] In formula (11), λ is independently sampled from Beta(1, 1), λ y is obtained according to formula (13);
[0044]
[0045] In formula (13), N i and N j respectively represent the number of samples y i and y' i corresponding categories in the target domain data set.
[0046] In combination with the first aspect, in some implementations of the first aspect, the method further comprises: the target domain model optimizing the calculation process of the total loss function, as follows:
[0047] The knowledge contained in the source domain model is transferred to the target domain by the information maximization loss of formula (14);
[0048]
[0049] In formula (14), M t (x) = P t (G t (x)) represents the K-dimensional category prediction of the target domain model M t to the sample x, represents the average output of the target domain data set X t after being processed by δ(M t (x));
[0050] Based on the pseudo-label of the target domain data, the label prediction result of the target domain model M t is supervised classification by the cross-entropy loss in formula (15);
[0051]
[0052] In formula (15), l is a standard cross-entropy loss function, and respectively represent the high-confidence pseudo-labeled samples and the updated low-confidence pseudo-labeled samples in the target domain data, and respectively represent the number of high-confidence samples and the number of low-confidence samples in the target domain data;
[0053] Based on the augmented target domain data, the in-class inter-class semantic structure regularization loss in formula (16) is used to make the target domain model M t Learning more robust feature representation;
[0054]
[0055] In formula (16), l fc represents the focal loss function, y c is a single-hot encoding label vector corresponding to the class c, and alpha is a weight parameter;
[0056] In summary, the optimization total loss function of the target domain model can be represented as: L total = L im + L ce + L cls .
[0057] In a second aspect, to achieve the above object, the application discloses a source domain independent cross-domain heartbeat recognition system based on pseudo-label mining, comprising:
[0058] A data processing module is configured to obtain electrocardiogram data with heartbeat annotations in a source domain, input the electrocardiogram data with heartbeat annotations in the source domain into a pre-established source domain model for pre-training, and obtain a pre-trained source domain model.
[0059] A data updating module is configured to filter and classify target domain data based on the pre-trained source domain model and a preset class threshold, obtain data with high pseudo-label confidence and data with low pseudo-label confidence, initialize parameters of a target domain model using the pre-trained source domain model, update the pseudo-labels of the data with low pseudo-label confidence based on a local-global semantic perception pseudo-label revision strategy, and merge the data with high pseudo-label confidence to obtain updated pseudo-labeled target domain data.
[0060] An intelligent recognition module is configured to perform data augmentation on the updated pseudo-labeled target domain data using a sample representation mixup operation and an intra-class sample data mixup operation, obtain augmented target domain data, input the augmented target domain data into a target domain feature extractor and a target domain classifier, output an optimization total loss function of the target domain model, and realize source domain independent cross-domain heartbeat intelligent recognition.
[0061] In another aspect of the application, to achieve the above object, a terminal device is disclosed, which comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor, the memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, a source domain independent cross-domain heartbeat recognition method based on pseudo-label mining is adopted.
[0062] In still another aspect of the present application, in order to achieve the above object, a computer readable storage medium is disclosed, wherein the computer readable storage medium stores a computer program, and the computer program is loaded and executed by a processor, and a source domain independent cross-domain heartbeat recognition method based on pseudo label mining is adopted.
[0063] The present application has the following beneficial effects:
[0064] The present application realizes intelligent recognition of cross-domain heartbeat types in the case of missing target domain data labels and unavailability of source domain electrocardiogram data in the cross-domain process. By utilizing the local-global semantic perception knowledge of source domain knowledge and target domain data structure, the present application significantly enhances the reliability of pseudo label prediction of unlabelled data in the target domain. Even under the condition of missing target domain data labels and limited access to source domain data during target domain model training, the recognition accuracy of the central heartbeat type of the target domain electrocardiogram data can still be improved. The present application can well solve the problem of pseudo label mining error in the existing source domain independent cross-domain unsupervised intelligent recognition method. By effectively utilizing the distribution information of target domain samples in their global semantic space and the local semantic information of the nearest neighbor samples of the samples, more reliable pseudo labels of target domain samples are mined. At the same time, in order to reduce the unreliability problem of the source domain model in predicting the target domain samples as much as possible, a pre-trained source domain model and a class threshold filtering pseudo label strategy are adopted, and a data augmentation is used to construct a local-global semantic structure consistency loss for unlabelled target domain data, so that the source domain independent cross-domain heartbeat recognition method has stronger generalization. BRIEF DESCRIPTION OF DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings;
[0066] Figure 1 is a method flowchart of the present application;
[0067] Figure 2 is a work flowchart of the present application;
[0068] Figure 3 is an electrocardiogram raw data processing method flowchart of the present application;
[0069] Figure 4 is a source domain model framework diagram of the present application;
[0070] Figure 5 is a performance comparison diagram of the method of the present application and common pseudo label learning methods;
[0071] Figure 6 is a schematic diagram of the system structure of the present application. DETAILED DESCRIPTION
[0072] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0073] Embodiment one:
[0074] As shown in the figure, a pseudo-label mining source domain independent cross-domain heartbeat recognition method, the method comprises the following steps: Figure 1
[0075] S101: acquiring electrocardio data with heartbeat annotation in a source domain, inputting the electrocardio data with heartbeat annotation in the source domain into a pre-established source domain model for pre-training, and obtaining a pre-trained source domain model;
[0076] The inputting of the electrocardio data with heartbeat annotation in the source domain into the pre-established source domain model for pre-training comprises:
[0077] The pre-training of the source domain model comprises using the electrocardio data D s with heartbeat annotation in the source domain to pre-train the source domain model M s according to the following formula (1) using label-smoothed cross-entropy loss:
[0078]
[0079] In formula (1), represents the kth class heartbeat label processed by the smoothing operation, a represents a smoothing hyperparameter, which is set to 0.1 according to experience, K represents the number of all heartbeat classes, and δ k represents the probability prediction result of the kth class output by the softmax function.
[0080] S102: acquiring electrocardio data without annotation in a target domain, filtering and classifying the target domain data based on the pre-trained source domain model and a preset class threshold, obtaining data with high pseudo-label confidence and data with low pseudo-label confidence, initializing target domain model parameters using the pre-trained source domain model, updating the pseudo-label of the data with low pseudo-label confidence based on a local-global semantic perception pseudo-label revision strategy, merging the data with high pseudo-label confidence, and obtaining updated pseudo-annotation target domain data;
[0081] The filtering and classifying of the target domain data based on the pre-trained source domain model and the preset class threshold comprises:
[0082] Given a target domain dataset X t For any unlabeled data x in the source domain, if the pre-trained source domain model M... s If two labeling networks P1 and P2 predict the same heartbeat category, and the predicted classification probability value of the corresponding category is higher than a pre-set threshold, then the data x is classified into the pseudo-label sample set with high confidence according to equation (2). Otherwise, according to equation (3), they are assigned to the sample set with low confidence. middle:
[0083]
[0084] In equation (2), τ c This represents the classification probability threshold set for category c, where c∈{1,2,...,K}.
[0085] The process of initializing the target domain model parameters using the source domain model and updating the pseudo-labels of data with low pseudo-label confidence based on a local-global semantic awareness pseudo-label revision strategy includes:
[0086] Target domain model parameter initialization and representation acquisition: The target domain model is initialized using pre-trained source domain model parameters. A target domain feature extractor is then used to acquire deep representations of samples from datasets with low pseudo-label confidence and datasets with high pseudo-label confidence, respectively. For datasets with low pseudo-label confidence in the target domain... The data in the dataset is used to obtain corresponding deep representations using a target domain feature extractor, based on target domain datasets with high confidence levels from pseudo-labels. From the data in the dataset, obtain the corresponding depth representation;
[0087] Calculate the nearest neighbor representation: for any data Corresponding depth representation Calculate its nearest neighbor representation in the high-confidence sample space.
[0088] Update the pseudo-labels for low-confidence data in the target domain: First, based on the high-confidence dataset... Obtain the representation prototype μ for each category k ,right any data Corresponding depth representation Calculate the similarity between each class and its representation center to obtain the normalized class prediction probability q. i ,calculate Nearest neighbor high confidence representation The corresponding normalized class prediction probability value qi' , for q i and q i' Weighted calculation to update low-confidence samples in the target domain Corresponding pseudo tags
[0089] The pseudo-tag revision strategy based on local-global semantic awareness includes:
[0090] For any low confidence data depth representation Based on the nearest neighbor distance evaluation function Nearest (4) in equation (4), the nearest neighbors of the features with high confidence are obtained.
[0091]
[0092] To update the pseudo-labels for low-confidence target domain data, a high-confidence target domain dataset is used. The category prototype μ of category c is calculated according to equation (5). c :
[0093]
[0094] in, Indicates sample The corresponding pseudo-label, I, represents the characteristic function;
[0095] For any data point in the low-confidence dataset Its characteristic representation is calculated according to equation (6). With each category of prototype μ c The distance Dis between them is used to obtain the normalized similarity score.
[0096]
[0097] for Nearest neighbor representation Calculate its relationship with the prototypes of each category according to equation (7). c The distance Dis between them is used to obtain the normalized similarity score.
[0098]
[0099] Combining two normalized similarity scores and Update the low-confidence samples in the target domain according to equation (8). pseudo-tags:
[0100]
[0101] where w reflects the adaptive weight of each sample confidence, which is calculated by equation (9):
[0102]
[0103] where, denotes the entropy of the probability distribution generated by the target domain model prediction of the sample, p i = [p i1 , p i2 ,..., p iK ] is the probability vector output by the softmax function, representing the probability of the sample belonging to each of the K different categories.
[0104] S103: The updated pseudo-labeled target domain data is subjected to sample representation inter-mixup operation and same category sample intra-mixup operation for data augmentation, obtaining augmented target domain data, and the augmented target domain data is input into the target domain feature extractor and the target domain classifier, outputting the target domain model optimization total loss function, realizing source domain independent cross-domain heartbeat intelligent recognition.
[0105] The augmented target domain data obtained by using the sample representation inter-mixup operation and the same category sample intra-mixup operation on the updated pseudo-labeled target domain data includes:
[0106] Augment the pseudo-labeled target domain data: Given a mini-batch sample (x1, y1),..., (x B , y B ) in the pseudo-labeled target domain data, where y i represents the pseudo-label corresponding to the sample x i , and for each category c, the intra-class augmented feature representation is constructed by equation (10)
[0107]
[0108] where the weight parameter β i is independently sampled from a uniform distribution U(0, 1);
[0109] In addition, the given mini-batch data is randomly shuffled to obtain reordered sequence samples (x'1, y'1),..., (x' B , y' B ), and the two groups of sequence data are obtained by equation (11) and equation (12) to obtain inter-class augmented samples (x' mix , y' mix ):
[0110] x′ mix = λ · x i + (1 - λ) · x′ i (11)
[0111] y′ mix = λ y · y i + (1 - λ y ) · y′ i (12)
[0112] where λ is independently sampled from Beta(1, 1), λ y is obtained according to formula (13);
[0113]
[0114] In formula (13), N i and N j respectively represent the number of samples y i and y' i corresponding categories in the target domain dataset. Parameters κ and τ are empirically set to 3 and 0.5 respectively.
[0115] The target domain model optimizes the total loss function, as follows:
[0116] The information maximization loss of formula (14) transfers the knowledge contained in the source domain model to the target domain;
[0117]
[0118] In formula (14), M t (x) = P t (G t (x)) represents the K-dimensional category prediction of the target domain model M t to the sample x, represents the average output of the target domain dataset X t sample processed by δ(M t (x));
[0119] Based on the pseudo label of the target domain data, the cross entropy loss in formula (15) is used to supervise the classification of the label prediction result of the target domain model M t ;
[0120]
[0121] In formula (15), l is the standard cross entropy loss function, and respectively represent the high confidence pseudo labeled samples and the updated low confidence pseudo labeled samples in the target domain data, and respectively represent the number of high-confidence samples and the number of low-confidence samples in the target domain data;
[0122] Based on the augmented target domain data, the target domain model M is regularized by the within-class-between-class semantic structure loss in formula (16) t Learn more robust feature representation to improve its generalization ability;
[0123]
[0124] In formula (16), l fc represents the focal loss function, y c is a one-hot encoding label vector corresponding to the class c, and a is a weight parameter for balancing the relative importance of the within-class and between-class semantic structure regularization loss;
[0125] Specifically, the scheme of the present application is further described below through embodiments:
[0126] A method for preprocessing original electrocardiogram data into three-dimensional input. The processing flow of the method is shown in Figure 3 , including the beat segmentation of the original electrocardiogram data and the three-dimensional input of the model.
[0127] In this embodiment, the original electrocardiogram data is segmented into single heartbeats. In order to reduce the interference of baseline drift noise on the electrocardiogram signal, the baseline offset noise is first removed by subtracting the average value of the segment from each sample point, and the normalized electrocardiogram segment with an average value of 0 is output. After obtaining the denoised electrocardiogram signal, in order to obtain the heartbeat segment, the electrocardiogram signal is intercepted according to the labeled R wave peak position in the electrocardiogram database to obtain the heartbeat segment. Let the position of the i-th R wave peak in the electrocardiogram signal be R i , the position of the i-1-th R wave peak be R i-1 , and the position of the i+1-th R wave peak be R i+1 . The i-1-th R wave peak is intercepted from the position of 0.14 seconds, that is, the starting position is R i-1 +0.14. The i-th R wave peak is intercepted from the position of 0.28 seconds, that is, the ending position is R i +0.28. Therefore, the interception interval of the i-th heartbeat segment can be expressed as:
[0128] [R i +0.14,R i +0.28](17)
[0129] In order to meet the requirement of the neural network for fixed length input, the intercepted heartbeat segment is resampled to 128 sample points. Let the resampled heartbeat segment be x i , then:
[0130] x i = resample([R i +0.14, R i +0.28], 128) (18)
[0131] where resample denotes a resampling operation to resample the truncated segment to 128 sample points.
[0132] The three-dimensional input features of the model include heartbeat sample features, pre-RR interval ratio features, and near-pre-RR interval ratio features. The heartbeat sample features are the heartbeat segment x i resampled to 128 sample points as described above; the pre-RR interval ratio features refer to the ratio of the current pre-RR interval to all average pre-RR intervals before the current heartbeat to be measured. The pre-RR interval is obtained by subtracting the current R-wave peak position from the previous heartbeat R-wave peak position; the near-pre-RR interval ratio features refer to the ratio of the current heartbeat pre-RR interval to the average pre-RR interval of the previous 10 heartbeats. Since the heartbeat sample features are vector features and the other two features are scalar features, and the input of the model network requires regular height, width, and channel number, in order to meet the input requirements of the model network, the pre-RR interval ratio features and the near-pre-RR interval ratio features are first expanded to vectors with the same width as the heartbeat sample features, and then the obtained features are concatenated in the channel dimension as the three-dimensional input of the neural network.
[0133] In order to verify the effectiveness of the source domain independent cross-domain heartbeat recognition based on the local-global semantic perception based pseudo-label mining strategy of the present application, the present application method and the existing pseudo-label mining method are compared in the cross-domain heartbeat recognition task. Figure 5
[0134] In order to verify the effectiveness of the present application, the inventors implemented the present application on the MIT-BIH arrhythmia dataset (MITDB) public electrocardiogram dataset, and compared it with the unsupervised cross-domain heartbeat recognition models proposed by Niu et al., Chen et al., and He et al., and compared it with the source domain independent unsupervised cross-domain heartbeat recognition models proposed by Liang et al., Yang et al., Prabhu et al., and Yuan et al., to test the model recognition performance ability of the present application from each recognized heartbeat type. Table 1 shows the experimental results of the comparative methods and the model proposed by the present application on the MITDB dataset (blackened as the optimal result).
[0135] Table 1
[0136]
[0137] [1] Unsupervised cross-domain method
[0138] [2] Unsupervised cross-domain method independent of source domain
[0139] From the results presented in Table 1, it can be clearly seen that the method of the present application exhibits a significant advantage in distinguishing the four heartbeat types, which strongly confirms the effectiveness of the method of the present application in the source domain-independent cross-domain heartbeat recognition task. In addition, the method of the present application compares the model performance of the pseudo-label samples generated by two different pseudo-label learning strategies to verify the pseudo-label learning ability of the present application. From the results of Table 2, it can be seen that the method of the present application performs more outstandingly on class F and significantly better than the methods using other pseudo-label learning strategies in terms of class sensitivity. In the visualization results, the classes with good performance are presented with lighter color intensity in the corresponding performance area. These results fully verify the ability of the method of the present application to capture reliable semantic representations in the target domain electrocardiogram data, and ultimately provide strong support for achieving better adaptation performance. Figure 5
[0140] Embodiment Two: In order to achieve the above purpose, based on the basis of Embodiment One, as shown in Figure 6 The present application discloses a source domain-independent cross-domain heartbeat recognition system based on pseudo-label mining, comprising:
[0141] The data processing module 11 is used for obtaining electrocardiogram data with heartbeat annotation in the source domain, inputting the electrocardiogram data with heartbeat annotation in the source domain into a pre-established source domain model for pre-training, and obtaining a pre-trained source domain model.
[0142] The data updating module 12 is used for filtering and classifying the target domain data based on the pre-trained source domain model and a preset class threshold, obtaining data with high pseudo-label confidence and data with low pseudo-label confidence, initializing the target domain model parameters with the pre-trained source domain model, updating the pseudo-label of the data with low pseudo-label confidence based on a local-global semantic perception pseudo-label revision strategy, and merging the data with high pseudo-label confidence to obtain updated pseudo-labeled target domain data.
[0143] The intelligent recognition module 13 is used for performing data augmentation on the updated pseudo-labeled target domain data by using a sample representation mixup operation and an intra-class sample data mixup operation, obtaining augmented target domain data, inputting the augmented target domain data into a target domain feature extractor and a target domain classifier, outputting a target domain model optimization total loss function, and realizing source domain-independent cross-domain heartbeat intelligent recognition.
[0144] Based on the same inventive concept, the present application further provides a computer device, comprising: one or more processors, and a memory for storing one or more computer programs; the program comprises program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, and are configured to implement one or more instructions, and are specifically configured to load and execute one or more instructions in the computer storage medium to implement the above method.
[0145] It needs to be further explained that, based on the same inventive concept, the present application further provides a computer storage medium, which stores a computer program, and the computer program is executed by the processor to perform the above method. The storage medium can adopt any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may, for example, be but is not limited to an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples (non-exhaustive list) of the computer readable storage medium include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component.
[0146] In the description of the present application, the description of the terms "one embodiment", "an example", "a specific example" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In the present specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any suitable manner in one or more embodiments or examples.
[0147] The foregoing presents and describes the basic principles, main features and advantages of the present disclosure. It should be understood by those skilled in the art that the present disclosure is not limited to the above-mentioned embodiments, and the above-mentioned embodiments and descriptions in the specification are only to illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, various changes and improvements can be made to the present disclosure, and all these changes and improvements fall within the scope of the present disclosure.
Claims
1. A source domain independent cross-domain heartbeat recognition method of pseudo-label mining, characterized in that, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:
2. The pseudo-label mining source domain independent cross-domain heartbeat recognition method according to claim 1, characterized in that, The method comprises the following steps: The pre-training source domain model comprises using the electrocardiogram data D with heart beat annotation in the source domain s The pre-training source domain model M is pre-trained using the cross-entropy loss with label smoothing according to the following formula (1) s : In formula (1), represents the kth category heartbeat label processed by the smoothing operation, a represents a smoothing hyperparameter, K represents the number of all heartbeat categories, and δ k represents the probability prediction result of the kth category output by the softmax function. 3.The pseudo-label mining source domain independent cross-domain heartbeat recognition method according to claim 1, characterized in that, The method comprises the following steps: Any one of the unlabeled data x in the target domain data set X t , if the pre-trained source domain model M s , the corresponding class prediction classification probability value is higher than the pre-set threshold τ c , then the data x is divided into the pseudo-label sample set with high confidence according to formula (2) , otherwise, according to formula (3) is divided into the sample set with low confidence : In formula (2), τ c denotes a classification probability threshold set for the class c, c e {1, 2,..., K}.
4. The source domain independent cross-domain heartbeat recognition method of pseudo label mining according to claim 1, characterized in that, The method comprises the following steps: Target domain model parameter initialization and representation acquisition: the pre-trained source domain model parameters are used to initialize the parameters of the target domain model, the deep representations of the samples in the data set with low pseudo label confidence and the data set with high pseudo label confidence are obtained through the target domain feature extractor, for the data in the data set with low pseudo label confidence in the target domain data , the corresponding deep representation is obtained by using the target domain feature extractor, and the corresponding deep representation is obtained according to the data in the data set with high pseudo label confidence in the target domain . Compute nearest neighbor representation: for X t Any one of the data Corresponding deep representation Compute its nearest neighbor representation in the high-confidence sample representation space updating the pseudo label of low confidence data in the target domain: first, according to the high confidence data set X t obtain the representation prototype μ of each category k , the pseudo label of the low confidence data in the target domain is updated as follows: any one of the data the corresponding deep representation calculate the similarity of each category representation center, get the normalized category prediction probability q i , calculate nearest neighbor high confidence representation the corresponding normalized category prediction probability value q i' , q i and q i' weighted calculation, update the target domain low confidence sample the corresponding pseudo label 5. The pseudo-label mining source domain independent cross-domain heartbeat recognition method according to claim 4, characterized in that, The method comprises the following steps: Deep characterization of any low confidence data The distance function Nearest is evaluated according to the nearest neighbor distance of formula (4) to obtain its nearest neighbor in the high confidence feature representation Based on target domain high-confidence dataset Calculate the class prototype μ of the class c according to formula (5) c : wherein, representing a sample a corresponding pseudo-label, I denotes an indicator function; For any one data in the low confidence data set The characteristic representation is calculated according to formula (6) The distance Dis between each category prototype μ c and the distance Dis between each category prototype μ For nearest neighbor representation The distance Dis between each category prototype μ c and the query vector q is calculated according to equation (7) Combining two normalized similarity scores and Update the low-confidence samples in the target domain according to equation (8). pseudo-tags: The method comprises the following steps: wherein, represents the entropy of the probability distribution generated by the target domain model prediction of the sample, p i = [p i1 , p i2 ,..., p iK ] is a probability vector output by the softmax function, representing the probability value of the sample belonging to each of the K different categories.
6. The pseudo-label mining source domain independent cross-domain heartbeat recognition method according to claim 1, characterized in that, The method comprises the following steps: Augment the pseudo-labeled target domain data: Given a mini-batch of samples (x1, y1),..., (x B , y B ) in the pseudo-labeled target domain data, where y i denotes the pseudo-label corresponding to sample x i , for each sample in class c, construct the intra-class augmented feature representation y In formula (10), the weight parameter β i is obtained from independent sampling from a uniform distribution U(0, 1); In addition, the given mini-batch data is randomly shuffled to obtain reordered sequence samples (x'1, y'1),..., (x'N, y'N), and the two groups of sequence data are obtained by formula (11) and formula (12) to obtain inter-class augmented samples (x'1, y'1),..., (x'N, y'N): B B mix mix ) x ′ mix = λ · x i + (1 - λ) · x ′ i (11) y ′ mix = λ y · y i + (1 - λ y )· y ′ i (12) In formula (11), λ is obtained independently from a Beta(1,1) distribution, λ y is obtained according to formula (13) In formula (13), N i and N j respectively represent the number of samples y i and y' i corresponding to the category in the target domain dataset.
7. The pseudo-label mining source domain independent cross-domain heartbeat recognition method according to claim 1, characterized in that, The method comprises the following steps: The method comprises the following steps: In formula (14), M t (x) = P t (G t (x)) represents the target domain model M t K-dimensional class prediction for a sample x, denotes the target domain dataset X t The average output after the sample is processed by δ(M t (x)). Based on the target domain data pseudo label, the cross entropy loss in formula (15) is used to supervise the classification of the target domain model M t label prediction result In formula (15), l is a standard cross-entropy loss function, and respectively represent the high-confidence pseudo-labeled samples and the updated low-confidence pseudo-labeled samples in the target domain data, and respectively represent the number of high-confidence samples and the number of low-confidence samples in the target domain data; Based on the augmented target domain data, the in-class-inter-class semantic structure regularization loss in formula (16) is used to make the target domain model M t learning more robust feature representation; In formula (16), l fc represents a focal loss function, y c is a one-hot encoding label vector corresponding to the class c, and a is a weight parameter; In summary, the target domain model optimization total loss function can be represented as: L total = L im + L ce + L cls .
8. A source domain independent cross-domain heartbeat recognition system for pseudo-label mining, characterized in that, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the The intelligent identification module is used for data augmentation of the updated pseudo-labeled target domain data by using sample representation intermixup operation and same category sample intra-data mixup operation, obtaining augmented target domain data, inputting the augmented target domain data into a target domain feature extractor and a target domain classifier, outputting a target domain model optimization total loss function, and realizing source domain independent cross-domain heartbeat intelligent identification.
9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, a source domain independent cross-domain heartbeat identification method of pseudo-label mining is adopted.
10. A computer-readable storage medium having stored therein a computer program, characterized in that, The computer program is loaded and executed by the processor, and a source domain independent cross-domain heartbeat identification method of pseudo-label mining is adopted.
Citation Information
Cited By
Data set automatic annotation construction method and system
CN121597976A
Double contrast learning Transform cross-dataset electroencephalogram emotion recognition method
CN121705849A