Data processing methods, devices, equipment, media and products based on deep learning

By introducing domain feature guidance in self-supervised comparison learning, the problem of poor model training effect in traditional methods is solved, and more accurate feature extraction and effective application of downstream tasks in medical and health data is achieved.

CN119833160BActive Publication Date: 2025-08-22OXFORD UNIV (SUZHOU) SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510324153.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-08-22
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The traditional self-supervised comparative learning method ignores key nuances in medical health data, resulting in poor model training and inability to effectively learn semantic relationships between samples, especially in healthcare applications, where model promotion capabilities are limited.

Method used

In combination with domain features, self-supervised comparison learning is guided, and positive and negative pairs are determined by obtaining the domain features of medical health data samples, and the model is trained using contrast loss to learn local and global semantic relationships to improve the accuracy of the model in feature extraction.

Benefits of technology

The semantic representation ability of the model when extracting medical health data features is improved, and high-quality features can be better extracted, suitable for downstream health care tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119833160B_ABST
    Figure CN119833160B_ABST
Patent Text Reader

Abstract

The present application relates to a data processing method, apparatus, device, medium and product based on deep learning. The method comprises: obtaining medical and health data samples, and obtaining domain features corresponding to each medical and health data sample for reflecting semantic information; extracting sample features of each medical and health data sample through a model to be trained; determining positive and negative pairs based on the domain features corresponding to the medical and health data samples, and determining contrast loss according to the features of each sample in the positive and negative pairs; wherein the positive and negative pairs include at least one of a first positive and negative pair for learning local semantic relationships or a second positive and negative pair for learning global semantic relationships, and the features of at least some of the samples in the positive and negative pairs are obtained based on the sample features; training the model to be trained through contrast loss, and the model obtained after the training is completed is used to extract features of medical and health data. The use of this method can improve the model training effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a data processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product based on deep learning. Background Art

[0002] With the development of artificial intelligence (AI), deep learning algorithms have become increasingly widely used. For example, they can automatically extract semantic information from raw data, enabling more accurate classification and prediction. Currently, methods using deep learning algorithms and large-scale annotated datasets have achieved performance levels comparable to or even exceeding that of domain experts.

[0003] In the healthcare sector, converting healthcare data collected by testing devices into actionable information suitable for back-end applications is a major challenge. Manually annotating healthcare data collected by testing devices (e.g., physiological signals collected by wearable devices, such as ECG (electrocardiogram) or EEG (electroencephalogram) signals) consumes significant labor and time, and often requires extensive clinical expertise. This significantly hinders the expansion of deep learning applications in healthcare. Consequently, traditional techniques often employ SSCL (Self-Supervised Contrastive Learning) to avoid the inconvenience of manual annotation.

[0004] However, traditional SSCL methods often ignore key nuances in medical and health data and treat all samples equally, which may miss important relationships between data points that share similar medical concepts and lead to poor model training results. Summary of the Invention

[0005] Based on this, it is necessary to provide a data processing method, device, computer equipment, computer-readable storage medium and computer program product based on deep learning that can improve the model training effect in response to the above technical problems.

[0006] In a first aspect, the present application provides a data processing method based on deep learning, comprising:

[0007] Acquire medical and health data samples, and obtain domain features corresponding to each medical and health data sample for reflecting semantic information;

[0008] Extracting sample features of each of the medical and health data samples through the model to be trained;

[0009] Determining positive and negative pairs based on domain features corresponding to medical and health data samples, and determining contrastive loss based on features of each sample in the positive and negative pairs; wherein the positive and negative pairs include at least one of a first positive and negative pair for learning local semantic relationships or a second positive and negative pair for learning global semantic relationships, and features of at least some of the samples in the positive and negative pairs are obtained based on the sample features;

[0010] The model to be trained is trained using the contrast loss; the model obtained after training is used to extract features of medical health data.

[0011] In a second aspect, the present application further provides a data processing device based on deep learning, comprising:

[0012] An acquisition module is used to acquire medical and health data samples and obtain domain features corresponding to each medical and health data sample for reflecting semantic information;

[0013] An extraction module, configured to extract sample features of each of the medical and health data samples through a model to be trained;

[0014] a determination module, configured to determine positive and negative pairs based on domain features corresponding to the medical and health data samples, and determine a contrastive loss based on features of each sample in the positive and negative pairs; wherein the positive and negative pairs include at least one of a first positive and negative pair for learning local semantic relationships or a second positive and negative pair for learning global semantic relationships, and features of at least some of the samples in the positive and negative pairs are obtained based on the sample features;

[0015] A training model is used to train the model to be trained through the contrast loss; the model obtained after the training is used to extract features of medical health data.

[0016] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0017] Acquire medical and health data samples, and obtain domain features corresponding to each medical and health data sample for reflecting semantic information;

[0018] Extracting sample features of each of the medical and health data samples through the model to be trained;

[0019] Determining positive and negative pairs based on domain features corresponding to medical and health data samples, and determining contrastive loss based on features of each sample in the positive and negative pairs; wherein the positive and negative pairs include at least one of a first positive and negative pair for learning local semantic relationships or a second positive and negative pair for learning global semantic relationships, and features of at least some of the samples in the positive and negative pairs are obtained based on the sample features;

[0020] The model to be trained is trained using the contrast loss; the model obtained after training is used to extract features of medical health data.

[0021] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0022] Acquire medical and health data samples, and obtain domain features corresponding to each medical and health data sample for reflecting semantic information;

[0023] Extracting sample features of each of the medical and health data samples through the model to be trained;

[0024] Determining positive and negative pairs based on domain features corresponding to medical and health data samples, and determining contrastive loss based on features of each sample in the positive and negative pairs; wherein the positive and negative pairs include at least one of a first positive and negative pair for learning local semantic relationships or a second positive and negative pair for learning global semantic relationships, and features of at least some of the samples in the positive and negative pairs are obtained based on the sample features;

[0025] The model to be trained is trained using the contrast loss; the model obtained after training is used to extract features of medical health data.

[0026] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:

[0027] Acquire medical and health data samples, and obtain domain features corresponding to each medical and health data sample for reflecting semantic information;

[0028] Extracting sample features of each of the medical and health data samples through the model to be trained;

[0029] Determining positive and negative pairs based on domain features corresponding to medical and health data samples, and determining contrastive loss based on features of each sample in the positive and negative pairs; wherein the positive and negative pairs include at least one of a first positive and negative pair for learning local semantic relationships or a second positive and negative pair for learning global semantic relationships, and features of at least some of the samples in the positive and negative pairs are obtained based on the sample features;

[0030] The model to be trained is trained using the contrast loss; the model obtained after training is used to extract features of medical health data.

[0031] The above-mentioned data processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product based on deep learning obtain domain features corresponding to each medical and health data sample for reflecting semantic information, wherein the domain features are based on the domain knowledge of clinical insights that have long guided data interpretation and feature engineering. Using domain features to guide the selection of positive and negative pairs in the comparative learning process can enable the model to learn local semantic relationships and / or global semantic relationships between samples. The local semantic relationship between samples can be understood as the semantic relationship between individuals (i.e., between samples), and the global semantic relationship between samples can be understood as the semantic relationship between individuals and clusters (i.e., between samples and sample clusters). This improves the training effect of the model, so that when the model extracts features from medical and health data, it can well represent the semantic information therein in the feature space, thereby extracting high-quality features. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 This is a diagram of an application environment of a data processing method based on deep learning in one embodiment;

[0034] Figure 2 1 is a flow chart of a data processing method based on deep learning in one embodiment;

[0035] Figure 3 A flowchart illustrating the steps of determining positive and negative pairs based on domain features corresponding to medical and health data samples and determining contrast loss based on features of each sample in the positive and negative pairs in one embodiment;

[0036] Figure 4 A flowchart illustrating the steps of determining positive and negative pairs based on domain features corresponding to medical and health data samples and determining contrast loss according to features of each sample in the positive and negative pairs in another embodiment;

[0037] Figure 5 A schematic diagram of a data processing method based on deep learning in one embodiment;

[0038] Figure 6 FIG1 is a schematic diagram showing a performance comparison of different methods in terms of class-averaged F1, accuracy, AUC, recall, and precision on different medical and health data in one embodiment;

[0039] Figure 7 is a schematic diagram of visualization results of waveform significance maps of two representative samples in one embodiment;

[0040] Figure 8 t-SNE plot showing feature representations derived from a self-supervised contrastive learning backbone for ECG, IMU, and EEG data in one embodiment;

[0041] Figure 9 is a schematic diagram of comparison results of semi-supervised learning based on an ECG dataset in one embodiment;

[0042] Figure 10 A schematic diagram showing the importance of different domain features to direct logistic regression and the framework solution of the present application in one embodiment;

[0043] Figure 11 is a structural block diagram of a data processing device based on deep learning in one embodiment;

[0044] Figure 12 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0046] In the field of artificial intelligence, a promising solution to the scarcity of labeled data when it comes to model training is self-supervised learning (SSL). It reduces reliance on annotations by learning useful representations from raw, unlabeled data. In healthcare applications, SSL can simplify the analysis of medical and health data collected by testing equipment, enabling faster and more accurate analysis and diagnosis without the need for labor-intensive data labeling. Among traditional SSL techniques, contrastive learning has been particularly successful. It is popular within SSL for its ability to learn representations by aligning similar pairs and distancing dissimilar pairs.

[0047] Generally speaking, self-supervised contrastive learning approaches require identifying positive and negative pairs without direct access to downstream labels. Therefore, inappropriate or suboptimal pairing when applying self-supervised contrastive learning approaches may lead to conflicts with the optimization objective, which may limit the model's ability to effectively generalize to real-world healthcare tasks. Various strategies have been proposed to alleviate this problem, including data augmentation techniques (such as adding noise or time flipping) and sampling strategies that group positive pairs based on shared medical concepts. Typical examples of the latter include treating samples from the same patient or from adjacent time frames as positive pairs. However, these methods often introduce inductive biases that may conflict with the underlying medical semantics of the data. For example, two electrocardiogram recordings of the same patient may differ significantly due to a sudden cardiac event, while recordings from different patients may have similar features if both patients were healthy. Therefore, traditional SSCL methods may not be consistent with the semantic knowledge required for practical downstream healthcare applications.

[0048] Unlike the success of deep learning, traditional supervised learning methods rely heavily on hand-crafted features based on specific domain knowledge. Data analysts (often involving domain experts such as clinicians) design and extract these features by applying domain-specific expertise, examining the data, or interpreting the signals using "textbook" empirical knowledge. These domain features, carefully crafted through extensive reliance on domain expertise, encode meaningful and clinically relevant concepts unique to the healthcare domain. While deep learning methods have far surpassed these traditional techniques in many aspects, domain features based on domain knowledge still contain important semantic information and have great potential for guiding pair selection in self-supervised contrastive learning. We argue that while advanced self-supervised learning solutions, particularly "new-school" contrastive methods for efficient label learning, have become increasingly popular, "old-school" domain knowledge-based feature engineering remains crucial for extracting valuable information from healthcare data.

[0049] Based on this, this application develops a new framework that combines and utilizes domain features that reflect semantic information to guide self-supervised contrastive learning. Specifically, when forming positive and negative pairs for self-supervised contrastive learning, domain knowledge of clinical insights that have long guided data interpretation and feature engineering is incorporated into self-supervised contrastive learning to help train the model. We tested our approach in various medical and health testing devices (including cardiac, neurological, and activity monitors) and found that it consistently outperformed traditional SSCL techniques (see below for detailed experimental details). These findings point to a novel and critical direction for developing efficient and scalable learning models in the field of digital health, highlighting the enduring value of domain-specific expertise in modern machine learning frameworks.

[0050] The data processing method based on deep learning provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the computer device 102 communicates with the electronic device 104 and / or the medical and health detection device 106 through the network. The data storage system can store the data that the computer device 102 needs to process. The data storage system can be integrated on the computer device 102, or it can be placed on the cloud or other network servers. For example, the computer device 102 can pre-deploy a model to be trained, and the computer device 102 can obtain medical and health data samples from at least one electronic device 104 and / or detection device 106 through network communication, and then extract the sample features of each medical and health data sample through the model to be trained; determine the positive and negative pairs based on the domain features corresponding to the medical and health data samples, and determine the contrast loss according to the features of each sample in the positive and negative pairs; train the model to be trained by the contrast loss, and the model obtained after the training is completed is used to extract the features of the medical and health data. Among them, the medical and health data can be collected by the detection device 106 or collected by the electronic device 104 and transmitted to the computer device 102.

[0051] Computer device 102 may be a terminal and / or a server. Terminals may include, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, smart car devices, projectors, and the like. Portable wearable devices may include smart watches, smart bracelets, head-mounted devices, and the like. Head-mounted devices may include virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, and the like. A server may be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services. Electronic device 104 may be a terminal and / or a server. Detection device 106 may be a medical imaging device or a signal measurement device, such as a device capable of detecting biological medical and health data. Medical imaging devices may include X-ray equipment, ultrasound imaging equipment, magnetic resonance imaging equipment, and the like; signal measurement devices may include desktop measurement instruments or wearable devices.

[0052] In an exemplary embodiment, Figure 2 As shown, a data processing method based on deep learning is provided, which is applied to Figure 1 Taking the computer device 102 in the example as an example, the following steps 202 to 208 are included. Among them:

[0053] Step 202: Acquire medical and health data samples, and obtain domain features corresponding to each medical and health data sample for reflecting semantic information.

[0054] Medical and health data samples are sample data used to train the model to be trained. Specifically, these sample data can be medical and health data. Medical and health data is data reflecting the medical and / or health status of an individual, and can include at least one of a person's physiological signals, activity signals, medical images, or medical text. Physiological signals are signals reflecting the physiological state of an individual, such as electrocardiogram (ECG), electroencephalogram (EEG), electromyography (EMG), and pulse signals. Activity signals are collected by sensors such as accelerometers and gyroscopes and reflect the external movement or behavior of an individual, such as steps, walking, running, sitting, etc. Medical images are images of internal tissues of a particular part of an organism, obtained non-invasively. Specifically, they can include anatomical images (such as X-ray images, computed tomography (CT) images, and magnetic resonance imaging (MRI) images), functional imaging data (such as positron emission tomography (PET) images and single-photon emission computed tomography (SPECT) images), and ultrasound imaging data. Medical text is text data that describes biomedical conditions, which may include medical records, test reports, etc. Its format can be structured or unstructured.

[0055] For example, if the medical health data sample Including physiological signal samples, the physiological signal samples of this application can specifically be time series data collected from different medical care modes of signal measurement equipment. Consider an unlabeled dataset, which is a collection of physiological signal samples, and express it as , from time series data Composition, of which Indicates the number of signal channels, represents the time step, Indicates the number of samples.

[0056] The model to be trained can be a machine learning model to be trained, such as a deep model to be trained. The model to be trained in this application at least includes a neural network for feature extraction. The neural network for feature extraction can be a transformer-based neural network or a neural network based on other structures. The embodiment of this application does not limit its structure. In some applications, the goal of this application is to develop a feature encoder. (That is, the "model" trained in this article) (where, represents the real space dimension of feature z), and after training, it can extract semantically meaningful representations for downstream healthcare-related tasks.

[0057] In some embodiments, the medical and health data samples include at least original samples. In some cases, the medical and health data samples include original samples and enhanced samples. Original samples are medical and health data samples directly collected by a detection device, while enhanced samples are samples obtained by enhancing the original samples. Enhancement processing can specifically include cropping, rotation, masking, noise addition, or time flipping, etc., which are not limited in the embodiments of the present application.

[0058] Domain features are features extracted based on specific domain knowledge. In this application, they can be features extracted based on medical domain knowledge. It should be noted that in order to use domain knowledge to improve the contrastive learning framework, for each medical and health data sample , computer equipment can extract domain features based on domain knowledge These domain features are selected based on their relevance to the specific domain knowledge of the corresponding modality and are identified as important features based on specific domain expertise.

[0059] In some embodiments, these domain features (covering at least one of morphological, temporal, spectral, and signal quality dimensions, among others) have been developed and standardized and can be easily accessed through readily available tools and software packages.

[0060] It should be noted that the model training process typically involves multiple training epochs, with multiple iterations performed within each epoch. Each iteration extracts a batch of samples from the training dataset and processes them to update the model's parameters (such as weights and / or biases). The entire training process is a continuous, iterative cycle. It is understood that the computer device can pre-acquire the full training dataset and extract domain features for each medical and health data sample in the training dataset. Consequently, during each iteration of each epoch of the model training process, the medical and health data samples and corresponding domain features required for the current iteration are obtained, and subsequent steps are performed to determine the contrastive loss to train the model using the contrastive loss. After the current iteration completes, the next iteration will proceed until the current epoch of training is complete. After the current epoch of training completes, the next epoch will proceed, and this cycle will continue until the training end condition is met, resulting in a fully trained model. The following details each iteration within each epoch of the training process.

[0061] Step 204: extracting sample features of each medical and health data sample through the model to be trained.

[0062] Specifically, each medical and health data sample (of the current batch) can be input into the model to be trained. The model to be trained then extracts features from each medical and health data sample. These extracted features can be referred to as sample features. In some embodiments, the computer device can further transform the sample features extracted by the model for subsequent loss determination. Exemplarily, this further transformation can include projection and normalization.

[0063] Step 206: determine positive and negative pairs based on the domain features corresponding to the medical and health data samples, and determine the contrast loss according to the features of each sample in the positive and negative pairs; wherein the positive and negative pairs include at least one of a first positive and negative pair for learning local semantic relationships or a second positive and negative pair for learning global semantic relationships, and the features of at least some of the samples in the positive and negative pairs are obtained based on the sample features.

[0064] Among them, the positive and negative pairs include positive pairs and negative pairs. In this application, positive pairs are samples that have semantic similarity to each other, and negative pairs are samples that do not have semantic similarity to each other. In some embodiments, the positive and negative pairs include at least one of a first positive and negative pair for learning local semantic relationships or a second positive and negative pair for learning global semantic relationships. Among them, the first positive and negative pairs include a first positive pair and a first negative pair. The first positive pair may include medical and health data samples and samples similar to it; the first negative pair may include medical and health data samples and samples different from it. The second positive and negative pairs include a second positive pair and a second negative pair. The second positive pair may include medical and health data samples and the prototype of the cluster to which the sample belongs, and the second negative pair may include medical and health data samples and prototypes of other clusters. Among them, the cluster can be obtained by clustering the full amount of medical and health data samples based on the domain features corresponding to each medical and health data sample in the full amount of medical and health data samples. The prototype of the cluster can refer to the centroid of the cluster or the centroid after iterative update.

[0065] Specifically, the computer device can determine positive and negative pairs based on the domain features corresponding to the medical and health data samples, and then construct a contrastive loss based on the features of each example in the positive and negative pairs. The contrastive loss is used to minimize the distance between each example in the positive pair and maximize the distance between each example in the negative pair.

[0066] In some embodiments, if the positive-negative pairs include only the first positive-negative pair, or only the second positive-negative pair, the contrast loss can be constructed based on the characteristics of each sample in the first positive-negative pair or the characteristics of each sample in the second positive-negative pair. If the positive-negative pairs include the first positive-negative pair and the second positive-negative pair, the contrast loss can be obtained by combining the loss determined by the characteristics of each sample in the first positive-negative pair and the loss determined by the characteristics of each sample in the second positive-negative pair.

[0067] Step 208: The model to be trained is trained using contrast loss; the model obtained after training is used to extract features of medical and health data.

[0068] Specifically, the computer device may iteratively train the model to be trained using contrastive loss to continuously adjust the model's network parameters (such as weights and / or biases). Training is terminated when a training termination condition is met, resulting in a trained model. The trained model can then be applied to extract features from the medical and health data to be processed. The training termination condition refers to the condition for terminating model training, which may specifically include reaching a preset number of training rounds or achieving a preset model performance.

[0069] As mentioned above, the model training process involves multiple training rounds (epochs), and multiple iterations (iterations) are performed in each training round. The above steps are the steps required to perform one iteration. After one iteration is completed, the next iteration will be performed until the current round of training is completed. After the current round of training is completed, the next round will be entered. This cycle will continue until the training end condition is met, and the trained model will be obtained.

[0070] In some embodiments, the model trained in this application can be applied to downstream tasks. For example, the model trained in this application can extract semantically meaningful features for downstream healthcare-related tasks. For example, these downstream tasks may include at least one of electrocardiogram-based cardiovascular disease classification, electroencephalogram-based sleep staging, or IMU (Inertial Measurement Unit)-based activity recognition.

[0071] The above-mentioned deep learning-based data processing method obtains domain features corresponding to each medical and health data sample for reflecting semantic information. The domain features are based on domain knowledge that has long guided clinical insights in data interpretation and feature engineering. Using domain features to guide the selection of positive and negative pairs in the comparative learning process can enable the model to learn local semantic relationships and / or global semantic relationships between samples. The local semantic relationship between samples can be understood as the semantic relationship between individuals (i.e., between samples), and the global semantic relationship between samples can be understood as the semantic relationship between individuals and clusters (i.e., between samples and sample clusters). This improves the training effect of the model, allowing the model to well represent the semantic information in the feature space when extracting features from medical and health data, thereby extracting high-quality features.

[0072] It is understood that in actual implementation, domain features of each medical and health data sample in the training dataset can be pre-extracted for use in subsequent model training. If prototype loss is also involved in the contrastive loss, multiple clusters can be pre-clustered based on the domain features of each medical and health data sample to be used for subsequent prototype-level contrastive learning.

[0073] In some embodiments, domain features corresponding to each medical and health data sample for reflecting semantic information are obtained, including: for any medical and health data sample, extracting features of the medical and health data sample or at least one of the cross-modal data samples associated with the medical and health data sample on a preset dimension to obtain the domain features corresponding to the medical and health data sample; wherein the preset dimension is determined based on domain knowledge.

[0074] Among them, the cross-modal data sample is a data sample whose modality is different from that of the medical health data sample. For example, when the medical health data sample is a medical imaging sample, its associated cross-modal data sample can specifically be at least one of a physiological signal sample, an activity signal sample or a medical text sample; when the medical health data sample is a physiological signal sample, its associated cross-modal data sample can specifically be at least one of an activity signal sample, a medical imaging sample or a medical text sample; when the medical health data sample is an activity signal sample, its associated cross-modal data sample can specifically be at least one of a physiological signal sample, a medical imaging sample or a medical text sample; when the medical health data sample is a medical text sample, its associated cross-modal data sample can specifically be at least one of a medical imaging sample, an activity signal sample or a physiological signal sample.

[0075] In some embodiments, for any biological object, multiple modal data from the biological object can be pre-collected to serve as training samples, such as medical images, physiological signals, activity signals, and medical text from the same biological object, and these multimodal data can be stored in association. For example, for cardiovascular diseases, multiple modal data from the same biological object (such as patients with cardiovascular disease) can be pre-collected, such as medical images (cardiac ultrasound), physiological signals (electrocardiogram), and medical text (corresponding reports of cardiac ultrasound and electrocardiogram). Therefore, when needed, cross-modal domain features can be applied and modeled based on the multimodal data corresponding to a large number of biological objects stored in association.

[0076] The following will provide detailed descriptions of single-modal application and modeling, as well as multi-modal application and modeling:

[0077] For single-modal applications and modeling: In some embodiments, for each medical and health data sample, the computer device can extract the features of the medical and health data sample on the preset dimensions of the corresponding modality and obtain the domain features of the medical and health data sample. When there are multiple preset dimensions, the features of multiple preset dimensions can be combined to obtain the domain features, such as splicing the features of each preset dimension in a preset order to obtain the domain features of the medical and health data sample. In this way, professional domain knowledge can be combined to extract domain features in the medical and health data sample that can reflect its inherent semantic information, which is convenient for guiding the selection of positive and negative pairs in the subsequent training process.

[0078] In other embodiments, for each medical and health data sample, the computer device may search a storage medium (such as a database or memory) for any cross-modal data sample associated with the medical and health data sample, and then extract the features of the cross-modal data sample along the preset dimensions of the corresponding modality to obtain the domain features corresponding to the medical and health data sample. When there are multiple preset dimensions, the features of the multiple preset dimensions may be combined to obtain the domain features, such as by concatenating the features of each preset dimension in a preset order to obtain the domain features of the medical and health data sample. In this way, professional domain knowledge can be combined to extract domain features that can reflect the inherent semantic information of any cross-modal data sample of the medical and health data sample, thereby facilitating the selection of positive and negative pairs during subsequent training. It can be understood that in situations where domain features cannot be extracted based on the medical and health data sample itself, or where the domain features extracted based on the medical and health data sample itself are insufficiently accurate, extracting the features of the cross-modal data sample associated with the medical and health data sample as domain features can effectively compensate for the shortcomings of this situation and greatly improve the adaptability of the solution.

[0079] For multimodal applications and modeling: In some embodiments, for any medical and health data sample, the computer device can extract domain features on at least two modalities, and combine the domain features on at least two modalities to obtain the domain features corresponding to the medical and health data sample. It can be understood that the at least two modalities may include the modality of the medical and health data sample itself, or at least two modalities of the modality of the corresponding cross-modal data sample; the combination method can be splicing or fusion, etc., which is not limited in the embodiments of the present application. In this way, the domain features obtained by combining the professional domain knowledge of multiple modalities can more comprehensively reflect the intrinsic semantic information of the medical and health data sample, which is convenient for guiding the selection of positive and negative pairs in the subsequent training process.

[0080] It should be noted that the preset dimensions of different modalities may be different, and the preset dimensions of the corresponding modality are determined based on the domain knowledge of the corresponding modality.

[0081] In some embodiments, the medical health data samples and / or cross-modal data samples may be any one or more of physiological signal samples, activity signal samples, medical imaging samples, and medical text samples. The following exemplifies how to extract features from physiological signal samples, activity signal samples, medical imaging samples, and medical text samples to obtain domain features. It is understood that the domain features of this application may be extracted based on at least one modality sample among physiological signal samples, activity signal samples, medical imaging samples, or medical text samples.

[0082] In some embodiments, the physiological signal samples include electrocardiogram (ECG) samples, electroencephalogram (EEG) samples, etc. The domain features extracted from the ECG samples may include at least one of AF (atrial fibrillation)-related features, morphological features, RR interval-related features, beat similarity-related features, etc.; the domain features extracted from the EEG samples may include at least one of time domain statistical features, time domain feature entropy, frequency domain statistical features, frequency domain feature entropy, specific frequency band spectral energy, etc.

[0083] In some embodiments, the activity signal samples include data collected by sensors such as accelerometers and gyroscopes, and the domain features extracted from the activity signal samples may include at least one of time statistical features, angle features, frequency domain features, peak features, etc.

[0084] In some embodiments, the medical image samples include ultrasound imaging data, magnetic resonance imaging images, etc. For example, when the medical image samples are cardiac ultrasound imaging data, the domain features extracted from the cardiac ultrasound imaging data may include the size of each atrium and ventricle, the maximum thickness of the ventricular septum, the ascending aorta, the E peak, the maximum regurgitant pressure gradient of the tricuspid valve, aortic valve degeneration, etc.; when the medical image samples are brain magnetic resonance imaging images, the domain features extracted from the brain magnetic resonance imaging images may include the size of brain gray matter and white matter, cerebral cortex thickness, cerebrospinal fluid distribution, perfusion flow rate, functional connectivity, intrinsic coherence volume fraction, mean diffusivity, isotropic volume fraction, etc.; when the medical image samples are lung CT, the domain features extracted from the lung CT may include lung density (HU (Hounsfield Unit) value distribution), lung volume, pulmonary artery diameter, lung nodule size, etc.

[0085] In some embodiments, domain features extracted from medical text samples may include diagnostic conclusions, local and / or global word / term frequencies, etc. These features may be obtained through processing by an NLP (Natural Language Processing) model. Local word / term frequencies may be extracted through a word2vec model (a related model for generating word vectors), and global word / term frequencies may be extracted through a Gram model (a statistical language model). Furthermore, in embodiments in which domain features are obtained based on medical text samples, we may use the similarity of medical text samples (such as test reports) to define the distance between patients, thereby determining positive and negative pairs, and performing clustering. For example, similarity calculation methods may include BLEU (Bilingual Evaluation Understudy) scores, cosine similarity, and Jaccard distance, etc.

[0086] In the above embodiment, combined with professional domain knowledge, the features of at least one modality sample can be extracted to obtain the domain features corresponding to the medical and health data samples, so that the domain features can accurately and comprehensively reflect their intrinsic semantic information, which is convenient for guiding the selection of positive and negative pairs in subsequent training processes.

[0087] It should be noted that traditional self-supervised contrastive learning methods often face challenges in healthcare applications due to sampling bias. Specifically, traditional instance-based contrastive learning usually assumes that all non-positive samples (i.e., negative samples) are dissimilar, which may lead to the incorrect separation of samples that actually have similar clinical or physiological characteristics. This problem is particularly serious in healthcare data. For example, different individuals with the same symptoms may have similar patterns in their healthcare data, such as patients with the same type of arrhythmia in ECG data. If based on traditional methods, incorrectly classifying such semantically similar data as dissimilar may hinder the model's ability to learn clinically meaningful representations.

[0088] To overcome this limitation, this application performs a nearest neighbor search (NNS) in the domain feature space. By leveraging the domain knowledge encoded in the domain features, samples with higher semantic similarity are identified as additional positive pairs, guiding instance-level contrastive learning so that the model can learn local semantic relationships between samples. Furthermore, this application also performs unsupervised clustering of domain features to ensure that semantically similar samples remain clustered. This introduces domain knowledge-guided prototype contrastive learning, allowing the model to learn global semantic relationships between samples and sample clusters.

[0089] It can be understood that in some embodiments, instance-level contrastive learning and / or prototype-level contrastive learning can be guided by domain knowledge.

[0090] refer to Figure 3 In some embodiments, when the positive and negative pairs include a first positive and negative pair, determining the positive and negative pairs based on the domain features corresponding to the medical and health data samples, and determining the contrast loss based on the features of each sample in the positive and negative pairs includes steps 2062 and 2064, wherein:

[0091] Step 2062: Determine semantically similar samples of each medical and health data sample based on the domain features corresponding to the medical and health data sample.

[0092] It is understandable that during the model training process, the training is usually done in batches. For each anchor sample in the current batch (that is, the medical and health data samples currently being focused on), the computer device can identify the nearest neighbors in the batch as semantically similar samples of the anchor sample based on the distance in the domain feature space (such as the Euclidean distance).

[0093] For example, the computer device may perform the nearest neighbor search using the following formula:

[0094] (Formula 1)

[0095] in, Indicates the number of anchor samples in the batch A collection of different medical and health data samples, also known as a negative sample set, Represents anchor samples By selecting the samples that are most similar to the anchor samples based on the domain features, we can ensure that semantically similar positive pairs are identified.

[0096] In some embodiments, it can be understood that the medical health data sample and its enhanced sample can be considered to be semantically consistent, so the semantically similar samples that are semantically similar to the anchor sample can include A sample identified from a collection of different medical and health data samples, and an enhanced sample of the identified sample.

[0097] In some embodiments, after performing the nearest neighbor search, the computer device may search for the sample with the smallest distance from the anchor sample as the semantically similar sample, and may also select the first M (M is an integer greater than 1) samples with the smallest distance as the semantically similar samples.

[0098] Step 2064: determine a first positive-negative pair based on the medical and health data sample and the corresponding semantically similar sample, and determine an instance-level contrast loss based on the sample features of each sample in the first positive-negative pair; the instance-level contrast loss is used to determine the contrast loss.

[0099] Specifically, the computer device may construct a first positive pair by combining a medical data sample with a corresponding semantically similar sample, and a first negative pair by combining the medical data sample with another medical data sample. An instance-level contrastive loss is then determined based on the sample features of each example in the first positive and negative pairs. It will be appreciated that the instance-level contrastive loss is used to minimize the distance between the first positive pairs and maximize the distance between the first negative pairs.

[0100] It should be noted that the samples in the first positive-negative pair are typically medical data samples. Therefore, when determining the instance-level contrastive loss, it can be determined based on the sample features of each sample in the first positive-negative pair. The sample features are features extracted by the model to be trained. In some embodiments, after the model to be trained extracts features from each medical data sample to obtain the corresponding sample features, the computer device can further perform nonlinear projection and normalization (e.g., l2 normalization) on the sample features for use in calculating various contrastive losses (including instance-level contrastive loss and / or prototype-level contrastive loss).

[0101] In some embodiments, a first positive-negative pair is determined based on medical and health data samples and corresponding semantically similar samples, including: determining a first positive pair based on an anchor sample and a semantically similar sample of at least the anchor sample, and determining a first negative pair based on the anchor sample and a medical and health data sample different from the anchor sample; wherein the anchor sample is any medical and health data sample in the training process, and the first positive pair and the first negative pair constitute the first positive-negative pair.

[0102] For example, for any anchor sample in the current training batch, the computer device may use the anchor sample and a semantically similar sample to the anchor sample as a first positive pair; alternatively, use the anchor sample and a semantically similar sample to the anchor sample as a first positive pair, and use the anchor sample and an enhanced sample to the anchor sample as a first positive pair; alternatively, use the anchor sample and a semantically similar sample to the anchor sample as a first positive pair, use the anchor sample and an enhanced sample to the anchor sample as a first positive pair, and use the anchor sample and an enhanced sample to the semantically similar sample as a first positive pair. If multiple first positive pairs are constructed, the model can more fully and accurately learn the local semantic relationships between samples during training, achieving better results.

[0103] Furthermore, the computer device may use the other medical and health data samples in the current training batch, except those used to construct the first positive pair, together with the anchor sample, to construct at least one first negative pair. It is understood that the computer device may also use the other medical and health data samples in the current training batch, except those used to construct the first positive pair, and the enhanced samples of the other medical and health data samples, together with the anchor sample, to construct multiple first negative pairs.

[0104] In this way, at least one first positive pair and at least one first negative pair are constructed by using an anchor sample and at least a semantically similar sample to the anchor sample, so that a model trained based on the first positive pair and the first negative pair can learn local semantic relationships between different samples, such as local semantic information and local semantic structure.

[0105] Furthermore, after constructing the first positive-negative pair, the computer device may determine an instance-level contrastive loss based on the features of each example in the first positive-negative pair. In some embodiments, the computer device may use a variant of the InfoNCE (Information Noise Contrastive Estimation) loss to define the instance-level contrastive loss, aiming to bring the features of the first positive pair closer together while simultaneously separating the features of the first negative pair. Of course, this application may also employ other forms of contrastive loss, such as triplet loss, which is not limited in this application.

[0106] For example, positive pairs can be generated based on the enhanced SSCL method, and considering that the medical health data samples and their enhanced samples share the same domain features , so we can select the two most similar samples in the domain feature space to form (i.e., a medical and health data sample whose domain characteristics are most similar to the domain characteristics of the anchor sample and the enhanced sample of this sample).

[0107] On this basis, the instance-level contrast loss can be determined by the following formula:

[0108] (Formula 2)

[0109] in, , Represent the indexes of anchor samples and positive samples respectively; The set of samples that form the first positive pair with the anchor sample may include the original positive pair and the sample that is most similar in terms of domain characteristics , that is, a set consisting of the enhanced samples of the anchor sample, the semantically similar samples of the anchor sample, and the enhanced samples of the semantically similar samples; and Represents the set of samples that form the first negative pair with the anchor sample, that is, the set of original negative samples excluding The resulting collection. Represents the features of the medical data samples obtained from the model being trained. This can be the sample features, or the sample features after nonlinear projection and normalization. By bringing the features of positive samples and anchor samples closer together, the model is guided to better align samples with meaningful semantic relationships.

[0110] It should be noted that the first positive pair constructed by the enhanced sample can be omitted in the instance-level contrast loss, or multiple semantically similar samples similar to the anchor sample can be selected to construct multiple first positive pairs, etc. The above examples are only used to illustrate the instance-level contrast loss of this application and are not intended to limit this application.

[0111] After constructing the instance-level contrastive loss, it can be directly used as the contrastive loss of the model, that is, the contrastive loss can be determined by the above formula 2. In other words, only domain features are used to guide the instance-level contrastive loss.

[0112] In some embodiments, the computer device may further determine a prototype-level contrast loss, and obtain a comprehensive contrast loss by combining the instance-level contrast loss and the prototype-level contrast loss. In some cases, the prototype-level contrast loss can be obtained by guiding domain features (see the description below for details), and in other cases, it can be obtained by guiding features learned by the model. A specific method for the latter is to cluster the features of each medical and health data sample extracted using the model during the training process to obtain multiple clusters, and then form positive and negative pairs with the medical and health data samples according to the prototypes of each cluster during the training process to determine the prototype-level contrast loss.

[0113] It can be understood that the above-mentioned various combinations are used to illustrate the technical concept of this application, rather than to limit this application.

[0114] In the above embodiment, when the positive-negative pairs include the first positive-negative pair, enforcing the similarity between samples with similar domain features can ensure that the local semantic relationship is well represented in the learned feature space.

[0115] refer to Figure 4 In some embodiments, when the positive-negative pair includes a second positive-negative pair, the positive-negative pair is determined based on the domain features corresponding to the medical and health data sample, and the contrast loss is determined according to the features of each sample in the positive-negative pair, including steps 2066 and 2068, wherein:

[0116] Step 2066: Determine the prototype of each cluster in the current iteration process; the cluster is obtained by clustering the domain features corresponding to each medical and health data sample in the full amount of medical and health data samples.

[0117] Here, a cluster refers to a cluster, or sample cluster. Specifically, it can be clustered based on the domain characteristics of all medical and health data samples. A prototype can be considered the representative of the cluster. In practice, it can be the centroid of the cluster or an updated centroid.

[0118] In some embodiments, during the offline phase, domain features of all medical and health data samples can be clustered, with each medical and health data sample assigned to the nearest centroid, resulting in multiple clusters. Domain features corresponding to medical and health data samples within the same cluster are relatively close, while domain features corresponding to medical and health data samples in different clusters are far apart.

[0119] For example, in this embodiment, an unsupervised k-means clustering method can be applied to derive K (according to experience, K can be set to 128) clusters based on the similarity between medical and health data samples. Assigned to the nearest cluster centroid , which can reflect the proximity of similar domain features. This ensures that samples in the same cluster share domain-informed semantics and vice versa. Subsequently, for each cluster, an initial prototype is derived , for example, by using k-means clustering to obtain the characteristics of the initial prototype.

[0120] In some embodiments, the prototype of each cluster is updated during the training phase. That is, in the current iteration, the prototype of each cluster is updated based on the prototype obtained in the previous iteration. Specifically, the prototype obtained in the previous iteration is updated based on the sample features generated during the learning process. It is understood that in the first iteration of the first round, the initial prototype can be the centroid of the cluster.

[0121] In some embodiments, during the training process, the prototypes of each cluster can be updated online based on the sample clusters in each iteration. For example, determining the prototype of each cluster in the current iteration includes: obtaining the prototypes of each cluster in the previous iteration; wherein the prototype of the cluster in the first iteration is the centroid of the cluster; and for the current iteration, updating the prototypes of each cluster in the previous iteration based on the sample characteristics of the medical and health data samples belonging to each cluster in the current iteration to obtain the prototypes of each cluster in the current iteration.

[0122] It should be noted that the initial prototype of a cluster is the centroid of each of the K clusters obtained at the end of clustering. The features corresponding to the cluster centroids can be the average domain features of all medical and health data samples belonging to that cluster. It can be understood that during the offline phase, K clusters are obtained by clustering based on the full set of medical and health data samples, and the prototype of each cluster is determined based on all the medical and health data samples in that cluster. The training phase, on the other hand, is iterative training based on batches. The sample size in each batch is smaller than the full set of samples, so the clustering of the batch actually varies compared to the clustering of the full set of samples. Therefore, during application, the prototype can be updated based on the distribution of samples in the current batch. That is, for any cluster in the current iteration, the sample features of the medical and health data samples belonging to that cluster in the medical and health data samples during that iteration can be used to update the prototype of that cluster to obtain an updated prototype. This updated prototype is then used to construct the second positive-negative pair in the current iteration.

[0123] In this way, by updating the prototypes of each cluster based on the features learned by the model in the current iteration, the prototypes used in the training process can be closer to the samples in the current batch. This process ensures that the prototypes continue to evolve when generating new representations while maintaining consistency with past observations, allowing the model to learn the global semantic relationship between samples more flexibly.

[0124] In some embodiments, the prototype can be updated online using an exponential moving average (EMA) representing the average of the clusters within each batch. For the current iteration, the prototypes of each cluster in the previous iteration are updated based on the sample features of the medical and health data samples belonging to each cluster in the current iteration to obtain the prototypes of each cluster in the current iteration, including: for the current iteration, for any cluster, the sample features of the medical and health data samples belonging to the cluster in the current iteration are integrated to obtain adjusted features, and the prototype of the cluster is updated based on the adjusted features.

[0125] For example, for any cluster, the average of the features of the medical and health data samples belonging to the cluster in the current iteration can be calculated to obtain an adjusted feature, and the prototype of the cluster can be updated based on the adjusted feature. The features of the updated prototype of the cluster can be obtained by taking a weighted sum of the features of the prototype of the cluster generated in the previous iteration and the adjusted feature.

[0126] For example, the prototype can be updated using the following formula:

[0127] (Formula 3)

[0128] in, represents the prototype of cluster j in the previous iteration, Represents the prototype of cluster j during the iteration; is the indicator function, Refers to the sample The cluster to which it belongs, when Returns 1 if yes, otherwise returns 0; Indicates the number of samples belonging to cluster j in each batch; is the momentum parameter (can be set to 0.5).

[0129] Step 2068: Determine a second positive-negative pair based on the medical and health data sample and the prototype of each cluster, and determine a prototype-level contrast loss based on the characteristics of each sample in the second positive-negative pair; wherein the characteristics of the prototype are determined based on the centroid of the corresponding cluster; the prototype-level contrast loss is used to determine the contrast loss.

[0130] In some embodiments, a second positive and negative pair is determined based on the medical and health data samples and the prototypes of each cluster, including: determining a second positive pair based on the anchor sample and the prototype of the cluster to which the anchor sample belongs, and determining a second negative pair based on the anchor sample and the prototypes of other clusters; wherein the anchor sample is any medical and health data sample in the training process, and the second positive pair and the second negative pair constitute the second positive and negative pair.

[0131] Furthermore, the computer device can determine a prototype-level contrastive loss based on the features of each sample in the second positive and negative pairs. It will be appreciated that the prototype-level contrastive loss is used to minimize the distance between the second positive pairs and maximize the distance between the second negative pairs. The features of the anchor sample in the second positive pair can be transformed from the sample features extracted from the anchor sample by the model to be trained; (for non-first iterations) the features of the prototype of the cluster to which the anchor sample belongs in the second positive pair are the features of the prototype updated in the current iteration. (For non-first iterations) the features of the prototype of the other cluster in the second negative pair are also the features of the prototype updated in the current iteration.

[0132] In some embodiments, the features of the anchor samples may specifically be features obtained by extracting features from the anchor samples by the model to be trained, and then performing nonlinear projection and normalization (such as l2 normalization).

[0133] In this way, applying prototype-level contrastive learning to each sample can ensure that their features are closely consistent with the assigned prototype while maintaining an appropriate distance from other clusters, which is conducive to the model learning the global semantic relationship between samples.

[0134] In some embodiments, the prototype-level contrast loss is determined based on the characteristics of each sample in the second positive-negative pair, including: obtaining the temperature corresponding to each cluster; wherein the temperature corresponding to the cluster is determined by the number of samples in the cluster and the prototype obtained at the end of the previous round; the prototype-level contrast loss is determined based on the characteristics of each sample in the second positive-negative pair and the temperature corresponding to each cluster.

[0135] To optimize the learning process, the temperature can be dynamically adjusted based on the number of samples within each cluster. This adjustment is based on the observation that lower temperatures lead to higher concentrations of corresponding prototypes. This dynamic temperature strategy ensures that the latent feature distribution around each prototype has a similar concentration.

[0136] It should be noted that the temperature update can be performed once in an epoch, and the updated temperature of each cluster is applied to the next epoch.

[0137] Specifically, for any cluster, the temperature corresponding to the cluster can be determined based on the number of medical and health data samples in the cluster and the distance between each medical and health data sample and the prototype of the cluster. For example, the temperature is defined as follows:

[0138] (Formula 4)

[0139] in, is the number of samples belonging to cluster k in the total amount of medical and health data samples, and m is a smoothing parameter used to avoid Too big. That is, in any round, the characteristics of each medical and health data sample obtained after multiple iterations in the round, Features of the latest prototype in the current round.

[0140] It should be noted that at the beginning of the first epoch, since no updated prototypes were generated, all You can follow the constant To set, then in the second epoch and later, The value of each cluster can be calculated according to the above formula by continuously iterating and updating the prototype value. Will be normalized so that all The average value is equal to .

[0141] It can be understood that the current round uses the prototype-level contrast loss Obtained through the previous round of updates; and the current round of updates Used for calculation of prototype contrast loss in the next round.

[0142] For example, the prototype-level contrast loss can be determined by the following formula:

[0143] (Formula 5)

[0144] in, is the anchor sample obtained from the model to be trained Features, Anchor sample The characteristics of the prototype of cluster i, is the characteristic of the prototype of cluster j, K is the total number of clusters, is the temperature corresponding to cluster i, is the temperature corresponding to cluster j.

[0145] In some embodiments, the prototype-level contrast loss can be directly used as the contrast loss, that is, the contrast loss can be determined by the above formula 5. In other words, only domain features are used to guide the prototype-level contrast loss.

[0146] In some embodiments, the computer device may further determine an instance-level contrastive loss, combining the instance-level contrastive loss with the prototype-level contrastive loss to obtain a comprehensive contrastive loss. In some cases, this instance-level contrastive loss can be derived through guidance from domain features (see the above description for details), while in other cases, it can be derived through augmented samples. The latter method can specifically use augmented views of the same instance as the positive pair. For example, the computer device may obtain the instance-level contrastive loss through the following methods:

[0147] (Formula 6)

[0148] in, , Represent the indexes of anchor samples and positive samples respectively; Represents samples in a small batch (i.e., in a batch) The negative set of is the temperature that controls the sharpness of the similarity. For simplicity and without loss of generality, before similarity calculation, another nonlinear projection is used by default to process the latent representation , and conduct Normalization. Usually, in the absence of actual ground truth, and is defined as augmented views of the same instance. Therefore, the loss function encourages the model to learn representations that make augmented views similar in the latent space while separating them from the representations of other samples in the mini-batch.

[0149] It can be understood that the above-mentioned various combinations are used to illustrate the technical concept of this application, rather than to limit this application.

[0150] In the above embodiment, when the positive-negative pairs include a second positive-negative pair, domain features guide the determination of prototypes, enabling the capture of global semantic relationships. By applying unsupervised clustering to these domain features and then generating cluster-specific prototypes online, the features are aligned with global semantic clusters, enabling the model to effectively learn global semantic relationships between samples.

[0151] In some embodiments, the positive-negative pairs include a first positive-negative pair and a second positive-negative pair, and the contrast loss is determined according to the features of each sample in the positive-negative pairs, including: determining the instance-level contrast loss according to the features of each sample in the first positive-negative pair, and determining the prototype-level contrast loss according to the features of each sample in the second positive-negative pair; combining the instance-level contrast loss and the prototype-level contrast loss to obtain the contrast loss.

[0152] In some embodiments, a computer device may determine a first positive-negative pair based on a medical and health data sample whose domain features meet similarity conditions, and determine an instance-level contrast loss based on the features of each sample in the first positive-negative pair, wherein the domain features meeting the similarity conditions may specifically be that the similarity between the domain features reaches a preset threshold or is ranked less than a preset rank; determine a second positive-negative pair based on the medical and health data sample and the prototype of each cluster, and determine a prototype-level contrast loss based on the features of each sample in the second positive-negative pair, wherein the cluster is obtained based on clustering of domain features, and the prototype of the cluster is determined based on the centroid of the cluster; combine the instance-level contrast loss and the prototype-level contrast loss to obtain a comprehensive contrast loss. For example, the instance-level contrast loss and the prototype-level contrast loss may be weightedly summed to obtain a comprehensive contrast loss. The coefficient of the weighted summation can be set as needed.

[0153] In some embodiments, determining a first positive-negative pair based on medical and health data samples whose domain features meet similarity conditions includes: determining a semantically similar sample for each medical and health data sample based on the domain features corresponding to the medical and health data samples; and determining a first positive-negative pair based on the medical and health data samples and the corresponding semantically similar samples.

[0154] The process of determining the instance-level contrastive loss and the prototype-level contrastive loss can be referred to in the above description. After determining the instance-level contrastive loss and the prototype-level contrastive loss, the computer device can combine the instance-level contrastive loss and the prototype-level contrastive loss to obtain a comprehensive contrastive loss.

[0155] In this way, by jointly training the model through instance-level contrastive loss and prototype-level contrastive loss, the model can ensure that the features are not only locally aligned but also aligned with the global semantics of the cluster when encoding features.

[0156] In some embodiments, the instance-level contrast loss and the prototype-level contrast loss are combined to obtain the contrast loss, including: determining a dynamic weight based on the current round in the training process; the dynamic weight increases as the current round increases; and applying the dynamic weight to the prototype-level contrast loss to perform weighted processing on the instance-level contrast loss and the prototype-level contrast loss to obtain the contrast loss.

[0157] Given that prototype representations are typically of low quality during the initial training phase, prototype contrastive learning can be applied gradually. The dynamic weight is greater than or equal to 0 and less than or equal to 1, and can be determined by the ratio of the difference between the current round and the starting round of applying prototype-level contrastive loss to the total number of training rounds. For example, the contrastive loss function can be constructed using the following formula:

[0158] (Formula 7)

[0159] in, Refers to the current round (current epoch), Refers to the starting round (starting epoch) of applying the prototype-level contrastive loss, is the total number of training rounds (total number of epochs).

[0160] In the above embodiment, as the training rounds progress, the application of prototype-level contrast loss is gradually increased, and dynamic adjustment can be made based on the quality of each of the instance-level contrast loss and the prototype-level contrast loss, so that the comprehensive contrast loss is more accurate, thereby improving the training quality.

[0161] In one embodiment, the data processing method based on deep learning may include the following steps: obtaining a training data set, which includes multiple medical and health data samples. For any medical and health data sample, extract the features of the medical and health data sample or at least one of the cross-modal data samples associated with the medical and health data sample in a preset dimension to obtain the domain features corresponding to the medical and health data sample, and cluster the full amount of medical and health data samples based on the domain features to obtain multiple clusters. The model to be trained is trained for multiple epochs, wherein multiple iterations are performed within each epoch. During each iteration, the current batch of samples is extracted from the training data set for processing, and the sample features of each sample are extracted based on the model to be trained, which are recorded as ,based on Calculated Then, the first positive and negative pairs are determined based on the domain characteristics, and the instance-level contrast loss is determined based on the sample characteristics of each example in the first positive and negative pairs. During each iteration, the EMA updates the prototypes of each cluster, calculates the distance between the anchor sample and the prototype of each cluster, and calculates the temperature of each cluster to determine the prototype-level contrastive loss. A comprehensive contrastive loss is derived from the instance-level and prototype-level contrastive losses, and the model is trained using this comprehensive contrastive loss. After the last iteration of the epoch, the temperature of each cluster is updated, completing the epoch. The next epoch then proceeds, and this cycle continues until the final epoch completes training, resulting in a fully trained model.

[0162] In one embodiment, the deep learning-based data processing method includes: obtaining a training dataset, the training dataset including multiple medical and health data samples. For any medical and health data sample, extracting features of the medical and health data sample or at least one of the cross-modal data samples associated with the medical and health data sample on a preset dimension to obtain domain features corresponding to the medical and health data sample, and clustering the full set of medical and health data samples based on the domain features to obtain multiple clusters. For any iteration in any round, extracting sample features of the medical and health data samples used by the model to be trained, determining semantically similar samples for each medical and health data sample based on the domain features of the medical and health data samples, determining first positive and negative pairs based on the medical and health data samples and the corresponding semantically similar samples, and determining instance-level contrastive loss based on the sample features of each sample in the first positive and negative pairs. For any iteration in any round, determining a prototype for each cluster, clustering the full set of medical and health data samples based on the domain features of the full set of medical and health data samples, determining second positive and negative pairs based on the medical and health data samples and the prototypes of each cluster, and determining prototype-level contrastive loss based on the features of each sample in the second positive and negative pairs. Combining instance-level contrastive loss and prototype-level contrastive loss yields a comprehensive contrastive loss, which is then used to train the model to be trained. After all iterations in all rounds are completed, a model for extracting features from medical and health data is obtained.

[0163] In a specific embodiment, the processing method of the model includes the following steps:

[0164] Obtain a medical and health data sample, and for any medical and health data sample, extract features of the medical and health data sample or at least one of the cross-modal data samples associated with the medical and health data sample on a preset dimension to obtain domain features corresponding to the medical and health data sample; wherein the preset dimension is determined based on domain knowledge.

[0165] The sample features of each medical and health data sample are extracted through the model to be trained.

[0166] Based on the domain features corresponding to the medical and health data samples, semantically similar samples of each medical and health data sample are determined.

[0167] The anchor sample and its semantically similar sample are considered as the first positive pair, the anchor sample and its enhanced sample are considered as the first positive pair, and the anchor sample and its semantically similar enhanced sample are considered as the first positive pair. A first negative pair is determined based on the anchor sample and a medical data sample different from the anchor sample. The anchor sample is any medical data sample used in the training process, and the first positive pair and the first negative pair constitute the first positive-negative pair. An instance-level contrastive loss is determined based on the sample features of each sample in the first positive-negative pair.

[0168] Obtain the prototype of each cluster in the previous iteration; wherein, the prototype in the first iteration is the centroid of the cluster; for the current iteration, based on the sample characteristics of the medical and health data samples belonging to each cluster in the current iteration, update the prototype of each cluster in the previous iteration to obtain the prototype of each cluster in the current iteration. Clusters are obtained by clustering the entire amount of medical and health data samples based on the domain characteristics corresponding to each medical and health data sample in the entire amount of medical and health data samples. Determine the second positive pair based on the prototype of the anchor sample and the cluster to which the anchor sample belongs, and determine the second negative pair based on the prototype of the anchor sample and other clusters; wherein, the anchor sample is any medical and health data sample in the training process, and the second positive pair and the second negative pair constitute the second positive-negative pair. Obtain the temperature corresponding to each cluster; wherein, the temperature corresponding to the cluster is determined by the number of samples in the cluster and the prototype obtained at the end of the previous round; determine the prototype-level contrast loss based on the characteristics of each sample in the second positive-negative pair and the temperature corresponding to each cluster.

[0169] A dynamic weight is determined based on the current round in the training process; the dynamic weight increases as the current round increases. The dynamic weight is applied to the prototype-level contrastive loss to weight the instance-level contrastive loss and the prototype-level contrastive loss to obtain the contrastive loss. The model to be trained is trained using the contrastive loss. Multiple iterations are performed within a round, and multiple rounds are performed throughout the training period. At the end of training, a model for extracting features from healthcare data is obtained.

[0170] This application uses the domain characteristics of medical and health data samples to guide the nearest neighbor search when forming self-supervised comparison pairs. By enforcing similarity between samples with similar domain characteristics, it ensures that local semantic relationships are well represented in the learned feature space. In addition, to capture global semantic relationships, this application applies unsupervised clustering to these domain characteristics and then generates cluster-specific prototypes online. This prototype-based contrastive learning ensures that representations are not only locally aligned, but also aligned with global semantic clusters.

[0171] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0172] Next, combine Figure 5 One of the concepts of this application is described in detail: Figure 5 As shown in the figure, this application develops a domain knowledge-guided SSL framework that combines "old-school" domain feature engineering with "new-school" SSCL. We leverage domain features extracted through mature related techniques (such as biomedical signal processing methods) and integrate these domain features into the SSL process to guide model learning and improve downstream validation performance. Building on the enhanced SSCL method, we use the Euclidean distance between these domain features to guide the selection of positive pairs for instance-level contrastive learning. In addition, to further strengthen the learning of global semantic relationships, we introduce prototype-level contrastive learning. This forces each example to be more closely aligned with its assigned prototype, which is generated by clustering based on domain features.

[0173] exist Figure 5 In the algorithmic framework diagram, a) "Old-fashioned" domain feature engineering is used to interpret healthcare data (such as physiological signals). For example, electrocardiograms (ECGs) are commonly used to detect cardiovascular anomalies. A series of features have been developed to indicate abnormal heart rate variability (HRV) and the presence / amplitude of P waves. b) Domain features based on domain knowledge are integrated into the supervised contrastive learning paradigm. Compared to traditional approaches, guided by domain features related to cardiologist expertise, we aim to preserve semantic relationships during the contrastive learning process. For example, by considering domain features indicative of a normal ECG recording (such as heart rate variability and the presence of P waves), the model should position the anchor (typically a normal ECG) closer to another normal sample than to an abnormal sample exhibiting irregular RR intervals and missing P waves. c) An illustration of the algorithmic framework guided by domain knowledge. In this application, conventional approaches are followed to generate two enhanced views of the same sample and derive their representations (i.e., sample features). At the same time, following standard healthcare data processing practices (such as biomedical signal processing), domain features reflecting the underlying semantic information are acquired offline from all samples. In order to effectively utilize domain knowledge in self-supervised contrastive learning, this application first identifies the nearest samples in the domain feature space to establish instance-level positive pairs, which together with the enhanced counterparts contribute to instance-level contrastive learning. On the other hand, in order to preserve global and local semantic structures, this application can also adopt prototype-level contrast. Specifically, it can be to perform unsupervised clustering based on domain features to obtain multiple clusters and prototypes of each cluster, and then use the exponential moving average (EMA) to derive the features of the prototypes online. The introduction of prototype-level contrast loss can encourage each sample to be closely aligned with its corresponding prototype while maintaining a distance from other samples. The importance of domain knowledge can be seen from this figure, not only in traditional machine learning methods, but also in advanced deep learning techniques for interpreting medical time series data.

[0174] In addition, the researchers also conducted experiments on one of the implementation plans of this application, and the experimental results are as follows:

[0175] Experimental Settings:

[0176] This model can be run on any GPU (Graphics Processing Unit) model, such as the NVIDIA V100 GPU, using a deep learning framework such as PyTorch. For self-supervised training experiments, the model was trained for 1000 epochs with a batch size of 1024, using the Adam optimizer, an initial learning rate of 0.001, and a warm-up strategy for the first five epochs. Logistic regression was applied to the validation set every 10 epochs to monitor model selection. For transfer learning and semi-supervised learning, the learning rate was reduced to 0.0005 and training was limited to 40 epochs. Given the imbalanced nature of some datasets in downstream tasks, balanced softmax (normalized exponential function) cross entropy was used to address class imbalance. Performance was evaluated using the class-averaged F1 score. All data were obtained from publicly available sources, including MIMIC-III WDB, PhysioNet17, CPSC, SleepEDF, and Capture24.

[0177] Implementation results:

[0178] 1. Datasets: Please refer to Table 1. To evaluate the feasibility and generalizability of the proposed framework, experiments were conducted on three types of wearable sensors: electrocardiogram (ECG), electroencephalogram (EEG), and inertial measurement unit (IMU). The datasets in Table 1 are described below: The CinC17 dataset consists of short ECG recordings collected from a portable chest patch, annotated into four categories, including normal sinus rhythm and atrial fibrillation. The CPSC dataset provides 12-lead ECG recordings annotated by cardiologists for nine different cardiac abnormalities, although we only use Lead II to simulate wearable scenarios. The MIMIC-III-WDB dataset consists of ICU (Intensive Care Unit) ECG recordings from bedside monitors. Due to the lack of segment-level annotations, it is primarily used for self-supervised training. For EEG, we used the SleepEDF dataset (two EEG channels, Fpz-Cz and Pz-Oz), which includes long-term recordings segmented into 30-second windows and annotated into one of the five sleep stages. Finally, for the IMU, we used the Capture24 dataset (xyz accelerometer), which contains wrist-worn accelerometer recordings segmented into 10-second windows for human activity recognition. Standard preprocessing was applied across all datasets, including resampling, bandpass filtering, segmentation, and z-normalization for consistency. All datasets were randomly split into training, validation, and test subsets in a 6:2:2 ratio, ensuring no individual overlap. These subsets were used for model training, hyperparameter tuning, and performance evaluation, respectively.

[0179] Table 1 Datasets used in the experiment

[0180]

[0181] To benchmark the performance of our framework across various healthcare scenarios, we evaluated the discriminability of features derived from SSL using the class-averaged F1 score on the test set (see Table 2 for details). The discriminative power of the learned features was directly evaluated using simple classifiers, including a linear classifier and a KNN (K-Nearest Neighbors) classifier (n = 10), as these classifiers do not involve further complex feature abstraction. This allows us to assess the quality of the features themselves, rather than introducing additional layers of abstraction that could obscure their inherent effectiveness. Our approach was then compared to a range of SSCL methods, a domain-feature-based model (DomainFeat.), and a fully supervised learning model with a randomly initialized backbone (FullySup.), all of which share the same architecture as our approach. Our approach demonstrated superior performance compared to other SSL methods, even outperforming the fully supervised model in some cases, particularly on IMU data, where it achieved a class-averaged F1 score of 0.526. This improvement may be attributed to the domain knowledge-guided approach in this application, which helps to mitigate the impact of inter-subject variability, especially in physical activity data where individuals may have very different movement patterns. These results suggest that incorporating domain knowledge can help reduce bias and enhance the overall robustness of the model even in a self-supervised learning setting. In addition to the class-averaged F1, we also evaluated the F1 scores of the best-performing models (e.g., Figure 6 ). Our approach consistently outperforms other SSL methods on almost all of these metrics, further demonstrating its versatility and effectiveness in learning robust features from wearable device measurement data. Compared to state-of-the-art SSCL methods, our approach leverages domain knowledge to drive features, which enhances the selection of comparison pairs and ensures that the model captures clinically meaningful patterns. This finding points to a promising direction for advancing SSL in time series data, particularly in healthcare applications where domain knowledge plays a key role in improving model performance.

[0182] Table 2 Comparison results based on self-supervised learning

[0183]

[0184] To test the effectiveness of the downstream knowledge acquired during self-supervised learning, we froze the backbone into a feature extractor (the model mentioned above) and applied supervised training using either a kNN or linear classifier. Each method was run with five different random seeds to evaluate performance. We compared the class-averaged F1 scores of the different models, and the results are shown in Table 2 above. Bold indicates the best result, and underlined indicates the second-best result.

[0185] In addition, please refer to Figure 6 , Figure 6 This section compares the performance of different methods on various healthcare data sets, such as physiological and activity signals, in terms of class-averaged F1, accuracy, AUC, recall, and precision. We compare the performance of the model with the best F1 score on the validation set on the test set. These values ​​are normalized by the deviation of each corresponding metric from the best performance.

[0186] Please refer to the following Figure 7 , Figure 7 Figure 2 shows a visualization of waveform saliency maps for two representative samples from one embodiment. Using a fully trained classifier attached to each self-supervised learning backbone (model) to classify electrocardiograms, the present invention employs the method to assign higher salience to clinically relevant regions, such as ST-segment elevation, while avoiding noise and irrelevant waveform components, thereby improving model interpretability and predictive accuracy.

[0187] refer to Figure 8 , Figure 8 In one embodiment, a t-SNE (t-distributed stochastic neighbor embedding) plot of the feature representations derived from the self-supervised contrastive learning backbone for ECG, IMU, and EEG data is shown. A. For the first three columns, the top row shows the results using SimCLR, while the bottom row shows the results using the method proposed in this application. Each dot represents a sample, and different colors correspond to different categories. Compared with SimCLR, the clusters in the t-SNE plot of the method of this application are better separated and compact, indicating that the feature representations are more discriminative and coherent. This enhanced discriminability indicates that the model of this application is better at capturing meaningful patterns from the data, which is critical for downstream classification tasks of medical wearable devices. B. The last column shows points colored according to the subject / individual. It can be observed that compared with the method proposed in this application, the traditional SSCL method may tend to cluster samples from the same individual more prominently. This suggests that traditional methods may be more likely to learn individual-specific patterns, while our method can better generalize across individuals. For Figure 7In the atrial fibrillation case in

[15] , our method effectively captures clinically relevant morphological abnormalities, such as the absence of P waves—a key indicator of atrial fibrillation—demonstrating its ability to focus on key features. Here, our method consistently assigns higher significance to waveform subsegments that directly reflect ST-segment elevation, while effectively ignoring irrelevant or noisy portions. This selective attention to clinically meaningful subsegments not only enhances the interpretability of the model’s decisions, but also improves its overall prediction accuracy. Furthermore, we visualize the t-SNE plots of the learned representations derived from our SSL framework on the test dataset, as shown in Figure 2. Figure 8 Even without access to the ground truth labels used for training, the features learned by our approach can show better separation and discrimination between different categories compared to the popular SimCLR method. This demonstrates that our proposed framework is effective in capturing meaningful features for downstream tasks.

[0188] refer to Figure 9 , Figure 9 Figure 2 shows the comparison results of semi-supervised learning based on ECG dataset in one embodiment. Figure 9 As can be seen in the results, our proposed method consistently outperforms other methods, especially when labeled data is scarce. For example, with only 10% labeled data, our method shows significant improvement over other methods. Furthermore, compared to other models, our method is less sensitive to the proportion of labeled data and maintains high performance even with small amounts of labeled data, reducing the reliance on large labeled datasets.

[0189] Furthermore, to evaluate the effectiveness of each module in the proposed method, we conducted an ablation study, with the results shown in Table 3. This study clearly demonstrates the effectiveness of each module and confirms their importance during training. Comparisons based on NNS (Nearest Neighbor Search) and Prototype help understand semantic information at the local and global levels, respectively.

[0190] Table 3 Ablation study during model training

[0191]

[0192] Impact of Domain Knowledge Components on the Model

[0193] We further compared the method of this application with direct regression as a pre-training proxy task. The results are shown in Table 3, and direct regression performs poorly. Although domain features encode semantically meaningful information, they often lack sufficient separability. Relying solely on these features will reduce the overall performance, which further demonstrates the importance of integrating them into the SSL framework to enhance semantic understanding and discrimination capabilities. On the other hand, in order to understand the key role of domain features in the integration framework of this application, we conducted a comprehensive ablation study on these features in the ECG domain. Specifically, we divided the features into different groups (AF-related features, morphological features, RR interval-related features, beat similarity-related features), and systematically ablated each group separately, as follows Figure 10 As shown in the figure, removing individual feature groups has a significant negative impact on direct logistic regression, demonstrating the sensitivity of traditional models to the loss of specific clinical features. In contrast, our SSCL framework demonstrates resilience to such feature removal, maintaining stable performance even when certain clinically important features are removed (that is, using a model trained using our training method, even when extracting features from input data with these missing features, it can still extract features with good performance). This robustness is particularly important from a clinical perspective. It demonstrates that our framework is more capable of handling realistic variations in clinical data, such as those caused by noise, artifacts, or incomplete data collection.

[0194] refer to Figure 10 , which shows the importance of different domain features for direct logistic regression and our framework. The heatmap on the left illustrates which feature combinations (AF features, morphology, RR interval, and similarity) were used, with purple cubes indicating which features were used. The bar chart on the right shows the class-averaged F1 score for each feature combination, with direct logistic regression represented in blue (hatched) and our approach in green. Higher values ​​indicate better performance, highlighting the advantage of integrating domain features into the self-supervised learning framework for downstream classification tasks.

[0195] Based on the same inventive concept, the present application also provides a deep learning-based data processing device for implementing the deep learning-based data processing method described above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations in the one or more embodiments of the deep learning-based data processing device provided below can be found in the above-mentioned limitations on deep learning-based data processing and will not be repeated here.

[0196] In an exemplary embodiment, Figure 11As shown, a data processing device based on deep learning is provided, including: an acquisition module 1101, an extraction module 1102, a determination module 1103 and a training module 1104, wherein:

[0197] The acquisition module 1101 is used to acquire medical and health data samples and obtain domain features corresponding to each medical and health data sample for reflecting semantic information.

[0198] The extraction module 1102 is used to extract sample features of each medical and health data sample through the model to be trained.

[0199] Determination module 1103 is used to determine positive and negative pairs based on the domain features corresponding to the medical and health data samples, and determine the contrast loss according to the features of each sample in the positive and negative pairs; wherein the positive and negative pairs include at least one of the first positive and negative pairs for learning local semantic relationships or the second positive and negative pairs for learning global semantic relationships, and the features of at least some of the samples in the positive and negative pairs are obtained based on the sample features.

[0200] The training module 1104 is used to train the model to be trained by contrast loss; the model obtained after training is used to extract features of medical and health data.

[0201] In some embodiments, the determination module is also used to determine the semantically similar samples of each medical and health data sample based on the domain characteristics corresponding to the medical and health data sample; determine a first positive and negative pair based on the medical and health data sample and the corresponding semantically similar sample, and determine the instance-level contrast loss based on the sample characteristics of each sample in the first positive and negative pair; the instance-level contrast loss is used to determine the contrast loss.

[0202] In some embodiments, the determination module is further used to determine a first positive pair based on the anchor sample and at least a semantically similar sample of the anchor sample, and to determine a first negative pair based on the anchor sample and a medical health data sample different from the anchor sample; wherein the anchor sample is any medical health data sample in the training process, and the first positive pair and the first negative pair constitute a first positive-negative pair.

[0203] In some embodiments, the determination module is further configured to take the anchor sample and the semantically similar sample of the anchor sample as the first positive pair; take the anchor sample and the enhanced sample of the anchor sample as the first positive pair; take the anchor sample and the enhanced sample of the semantically similar sample as the first positive pair.

[0204] In some embodiments, the determination module is also used to determine the prototype of each cluster during the current iteration; the clusters are clustered based on the domain features corresponding to each medical and health data sample in the full amount of medical and health data samples; a second positive and negative pair is determined based on the medical and health data samples and the prototypes of each cluster, and a prototype-level contrast loss is determined based on the features of each sample in the second positive and negative pairs; wherein the features of the prototype are determined based on the centroid of the corresponding cluster; the prototype-level contrast loss is used to determine the contrast loss.

[0205] In some embodiments, the determination module is also used to obtain the prototype of each cluster in the previous iteration process; wherein, the prototype of the cluster in the first iteration process is the centroid of the cluster; for the current iteration, according to the sample characteristics of the medical and health data samples belonging to each cluster in the current iteration, the prototype of each cluster in the previous iteration process is updated to obtain the prototype of each cluster in the current iteration process.

[0206] In some embodiments, the determination module is further used to determine a second positive pair based on the anchor sample and the prototype of the cluster to which the anchor sample belongs, and to determine a second negative pair based on the prototype of the anchor sample and other clusters; wherein the anchor sample is any medical health data sample in the training process, and the second positive pair and the second negative pair constitute a second positive-negative pair.

[0207] In some embodiments, the determination module is further used to obtain the temperatures corresponding to each cluster; wherein the temperature corresponding to the cluster is determined by the number of samples in the cluster and the prototype obtained at the end of the previous round; and the prototype-level contrast loss is determined based on the characteristics of each sample in the second positive-negative pair and the temperature corresponding to each cluster.

[0208] In some embodiments, the determination module is further used to determine the instance-level contrast loss based on the features of each sample in the first positive-negative pair, and determine the prototype-level contrast loss based on the features of each sample in the second positive-negative pair; and combine the instance-level contrast loss and the prototype-level contrast loss to obtain the contrast loss.

[0209] In some embodiments, the determination module is further used to determine a dynamic weight based on the current round in the training process; the dynamic weight increases as the current round increases; and the dynamic weight is applied to the prototype-level contrast loss to perform weighted processing on the instance-level contrast loss and the prototype-level contrast loss to obtain the contrast loss.

[0210] In some embodiments, the acquisition module is also used to extract, for any medical and health data sample, features of the medical and health data sample or at least one of the cross-modal data samples associated with the medical and health data sample in a preset dimension to obtain domain features corresponding to the medical and health data sample; wherein the preset dimension is determined based on domain knowledge.

[0211] Each module in the above-mentioned deep learning-based data processing device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0212] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 12 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store training data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a data processing method based on deep learning is implemented.

[0213] Those skilled in the art will understand that Figure 12 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0214] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0215] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0216] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0217] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0218] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0219] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0220] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A data processing method based on deep learning, characterized in that: The method comprises: Obtaining medical and health data samples and obtaining domain features corresponding to each medical and health data sample for reflecting semantic information; the domain features are extracted based on medical domain knowledge; Extracting sample features of each of the medical and health data samples through the model to be trained; Determine positive and negative pairs based on domain features corresponding to medical and health data samples, and determine contrast loss based on features of each sample in the positive and negative pairs; wherein the positive and negative pairs include at least a second positive and negative pair for learning global semantic relationships, and features of at least some samples in the positive and negative pairs are obtained based on the sample features; The model to be trained is trained by using the contrast loss; the model obtained after the training is used to extract features of medical health data; Wherein, in the case where the positive-negative pairs include a second positive-negative pair, determining the positive-negative pairs based on the domain features corresponding to the medical and health data samples, and determining the contrast loss according to the features of each sample in the positive-negative pairs include: Obtaining the prototype of each cluster in the previous iteration; wherein, in the first iteration, the prototype of the cluster is the centroid of the cluster, and the cluster is obtained by clustering based on the domain features corresponding to each medical and health data sample in the full amount of medical and health data samples; For any cluster in the current iteration, the sample features of the medical and health data samples belonging to the cluster in the current iteration are integrated to obtain an adjusted feature, and the prototype of the cluster in the previous iteration is updated based on the adjusted feature; A second positive-negative pair is determined based on the medical and health data sample and the updated prototype of each cluster, and a prototype-level contrast loss is determined based on the characteristics of each sample in the second positive-negative pair; the prototype-level contrast loss is used to determine the contrast loss.

2. The method according to claim 1, characterized in that In a case where the positive-negative pairs further include a first positive-negative pair for learning a local semantic relationship, determining the positive-negative pairs based on the domain features corresponding to the medical and health data samples, and determining the contrast loss according to the features of each sample in the positive-negative pairs further includes: Determining semantically similar samples of each medical and health data sample based on the domain features corresponding to the medical and health data sample; A first positive-negative pair is determined based on the medical health data sample and the corresponding semantically similar sample, and an instance-level contrast loss is determined based on the sample features of each sample in the first positive-negative pair; the instance-level contrast loss is used to determine the contrast loss.

3. The method according to claim 2, characterized in that The determining of a first positive-negative pair according to the medical health data sample and the corresponding semantically similar sample includes: A first positive pair is determined based on an anchor sample and a semantically similar sample of at least the anchor sample, and a first negative pair is determined based on the anchor sample and a medical health data sample different from the anchor sample; wherein the anchor sample is any medical health data sample in the training process, and the first positive pair and the first negative pair constitute the first positive-negative pair.

4. The method according to claim 3, characterized in that The determining a first positive pair based on an anchor sample and at least a semantically similar sample to the anchor sample comprises: Taking the anchor sample and a semantically similar sample of the anchor sample as the first positive pair; Taking the anchor sample and the enhanced sample of the anchor sample as the first positive pair; The anchor sample and the enhanced sample of the semantically similar sample are used as the first positive pair.

5. The method according to claim 1, wherein The determining of a second positive-negative pair based on the medical and health data sample and the updated prototype of each cluster includes: A second positive pair is determined based on the anchor sample and the updated prototype of the cluster to which the anchor sample belongs, and a second negative pair is determined based on the updated prototypes of the anchor sample and other clusters; wherein the anchor sample is any medical health data sample in the training process, and the second positive pair and the second negative pair constitute the second positive-negative pair.

6. The method according to claim 1, characterized in that Determining the prototype-level contrast loss according to the features of each sample in the second positive-negative pair includes: Obtain the temperature corresponding to each cluster; the temperature corresponding to the cluster is determined by the number of samples in the cluster and the prototype obtained at the end of the previous round; A prototype-level contrast loss is determined according to the characteristics of each sample in the second positive-negative pair and the temperature corresponding to each cluster.

7. The method according to claim 2, characterized in that The determining of positive and negative pairs based on the domain features corresponding to the medical and health data samples, and determining the contrast loss according to the features of each sample in the positive and negative pairs further includes: The instance-level contrast loss and the prototype-level contrast loss are combined to obtain the contrast loss.

8. The method according to claim 7, characterized in that The combining of the instance-level contrast loss and the prototype-level contrast loss to obtain the contrast loss includes: Determining a dynamic weight based on a current round in the training process; the dynamic weight increases as the current round increases; The dynamic weight is applied to the prototype-level contrast loss to perform weighted processing on the instance-level contrast loss and the prototype-level contrast loss to obtain a contrast loss.

9. The method according to any one of claims 1 to 8, characterized in that The obtaining of domain features corresponding to each medical and health data sample for reflecting semantic information includes: For any medical and health data sample, extract the features of the medical and health data sample or at least one of the cross-modal data samples associated with the medical and health data sample on a preset dimension to obtain the domain features corresponding to the medical and health data sample; wherein the preset dimension is determined based on domain knowledge.

10. A data processing device based on deep learning, characterized in that: The device comprises: An acquisition module is used to acquire medical and health data samples and obtain domain features corresponding to each medical and health data sample for reflecting semantic information; the domain features are extracted based on medical domain knowledge; An extraction module, configured to extract sample features of each of the medical and health data samples through a model to be trained; a determination module for determining positive and negative pairs based on domain features corresponding to the medical and health data samples, and determining a contrast loss based on features of each sample in the positive and negative pairs; wherein the positive and negative pairs include at least a second positive and negative pair for learning global semantic relationships, and features of at least some samples in the positive and negative pairs are obtained based on the sample features; A training model is used to train the model to be trained using the contrast loss; the model obtained after the training is used to extract features of the medical health data; Among them, when the positive and negative pairs include a second positive and negative pair, the determination module is specifically used to: obtain the prototype of each cluster in the previous iteration process; wherein, the prototype of the cluster in the first iteration process is the centroid of the cluster, and the cluster is clustered based on the domain features corresponding to each medical and health data sample in the full amount of medical and health data samples; for the current iteration, for any cluster, the sample features of the medical and health data samples belonging to the cluster in the current iteration are integrated to obtain adjustment features, and the prototype of the cluster in the previous iteration process is updated based on the adjustment features; the second positive and negative pairs are determined according to the medical and health data samples and the updated prototypes of each cluster, and the prototype-level contrast loss is determined according to the features of each sample in the second positive and negative pairs; the prototype-level contrast loss is used to determine the contrast loss.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Supervised learning method and device for image features, equipment and storage medium

    CN113822325A

  • Target classification model training method, target identification method and electronic equipment

    CN117523340A