Training Method, Device and Electronic Device Based on Fine-Grained Contrastive Learning

Fine-grained contrastive learning enhances heart rate data analysis by aligning waveform features with textual data, improving the model's ability to capture detailed features and align heart rate data with textual reports.

CN119940578BActive Publication Date: 2025-07-15GUANGDONG TRANSTEK MEDICAL ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510429369.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-15
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing ECG waveform alignment learning model is difficult to capture specific waveform features, affecting the analysis effect, resulting in insufficient alignment between the ECG waveform and text information.

Method used

Through a training method based on fine-grained comparison learning, the waveform feature information of text data is obtained using the preset large language model and added it to the text data. The pre-trained model is trained in combination with the correction loss function and the comparison loss function to optimize the model's alignment ability.

Benefits of technology

The model's ability to interpret fine-grained waveform features is improved, the model's alignment is optimized, and the degree of matching between the electrocardiogram waveform and text information is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940578B_ABST
    Figure CN119940578B_ABST
Patent Text Reader

Abstract

The present invention provides a training method, device, and electronic device based on fine-grained contrastive learning. By obtaining multiple groups of data sets, each data set includes waveform data and text data. For the text data in each data set, the waveform feature information corresponding to the text data is obtained by using a preset large language model, and the waveform feature information is added to the text data to update the text data. The pre-trained model is trained using the updated multiple groups of data sets to obtain a fine-grained model. This solution reversely obtains the waveform feature information related to the text by using the preset large language model, thereby enhancing the fine-grained degree of the text data, improving the interpretation ability of the final model for fine-grained waveform features, and optimizing the alignment ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular, to a training method, device, and electronic device based on fine-grained contrast learning. Background Art

[0002] In medical scenarios, it is usually involved in using relevant waveform data for auxiliary diagnosis. For example, electrocardiogram, as an important non-invasive diagnostic tool, is widely used in the detection and management of cardiovascular diseases such as arrhythmia. In recent years, with the development of artificial intelligence technology, using machine learning algorithms to automatically analyze relevant waveform data, text data, etc. has become a research hotspot. However, there are still many deficiencies in this research field.

[0003] For example, taking the analysis of waveform data in electrocardiogram as an example, currently existing analysis models mostly learn based on the alignment of waveforms in electrocardiogram and their text reports. However, in actual scenarios, text reports mostly focus on the global features of electrocardiogram, such as diagnosis and rhythm, etc. As a result, when the analysis model performs alignment learning between waveforms and text reports, due to the difficulty in capturing specific waveform features, the learning effect is affected, and further the alignment ability between the electrocardiogram waveform and text information of the analysis model is affected. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a training method, device, and electronic device based on fine-grained contrast learning, so as to improve the interpretation ability of the final model for fine-grained waveform features and optimize the alignment ability of the model.

[0005] In a first aspect, the present invention provides a training method based on fine-grained contrast learning, and the method includes:

[0006] Obtain multiple groups of data groups, each group of the data groups including waveform data and text data;

[0007] For the text data in each of the data groups, use a preset large language model to obtain waveform feature information corresponding to the text data;

[0008] Add the waveform feature information to the text data to update the text data;

[0009] Use the updated multiple groups of data groups to train a pre-trained model to obtain a fine-grained model.

[0010] In an optional implementation manner, the method further includes the step of pre-training the pre-trained model, and this step includes:

[0011] Use the obtained multiple groups of data groups to train an initial model constructed to obtain a pre-trained model;

[0012] The step of adding the waveform feature information to the text data to update the text data includes:

[0013] Based on the pre-trained model, screening out the waveform feature information corresponding to the waveform data in the data group from the waveform feature information;

[0014] Adding the screened waveform feature information to the text data to update the text data.

[0015] In an alternative embodiment, the step of obtaining the waveform feature information corresponding to the text data by using a preset large language model includes:

[0016] Extracting the diagnosis result included in the text data;

[0017] Generating an inquiry statement based on the diagnosis result;

[0018] Importing the inquiry statement into the preset large language model and outputting the waveform feature information corresponding to the diagnosis result.

[0019] In an alternative embodiment, the step of importing the inquiry statement into the preset large language model and outputting the waveform feature information corresponding to the diagnosis result includes:

[0020] Generating a format command statement;

[0021] Importing the inquiry statement and the format command statement into the preset large language model and outputting the waveform feature information corresponding to the diagnosis result under the format command statement.

[0022] In an alternative embodiment, the waveform feature information includes at least one waveform feature item;

[0023] The step of screening out the waveform feature information corresponding to the waveform data in the data group from the waveform feature information based on the pre-trained model includes:

[0024] Importing the waveform data and at least one waveform feature item in the data group into the pre-trained model;

[0025] Obtaining the similarity between the waveform data and the at least one waveform feature item through the pre-trained model;

[0026] Taking the waveform feature items with a similarity greater than or equal to a preset threshold as the waveform feature information corresponding to the waveform data.

[0027] In an alternative embodiment, the step of training the constructed initial model with the obtained multiple data groups to obtain a pre-trained model includes:

[0028] Import the obtained multiple data groups into the constructed initial model to obtain the waveform features of the waveform data and the text features of the text data in each data group;

[0029] Calculate the similarity based on the waveform features and text features in each data group;

[0030] Construct a loss function based on the similarity;

[0031] Perform iterative training of the initial model under the guidance of the loss function until a pre-trained model is obtained when the preset requirements are met.

[0032] In an alternative embodiment, the similarity includes inter-group similarity and text similarity. The inter-group similarity is the similarity between the waveform features and text features of different data groups, and the text similarity is the similarity between the text features of different data groups. The loss function includes a correction loss function;

[0033] The step of constructing a loss function based on the similarity includes:

[0034] For each two of the multiple data groups, calculate the difference value between the inter-group similarity and the text similarity of the two data groups;

[0035] Construct a correction loss function based on the difference values of each two data groups. The correction loss function is a minimization function.

[0036] In an alternative embodiment, the similarity includes intra-group similarity and inter-group similarity. The intra-group similarity is the similarity between the waveform features and text features in the same data group, and the inter-group similarity is the similarity between the waveform features and text features of different data groups. The loss function includes a contrastive loss function;

[0037] The multiple data groups include multiple positive sample data groups and multiple negative sample data groups. Each positive sample data group and each negative sample data group have corresponding sample label identifiers. The text data includes at least one diagnosis result;

[0038] The step of constructing a loss function based on the similarity includes:

[0039] Construct a basic function based on the intra-group similarity, inter-group similarity, and sample label identifier corresponding to each data group. The basic function is used to evaluate the matching probability of the waveform data in the data group with each diagnosis result in at least one diagnosis result;

[0040] The contrast loss function is obtained by adding the set bias term and temperature parameter to the basic function, and the bias term and temperature parameter are used to adjust the ratio of the positive sample data group and the negative sample data group.

[0041] In a second aspect, the present invention provides a training device based on fine-grained contrast learning, and the device includes:

[0042] An acquisition module, configured to acquire multiple groups of data groups, and each group of the data groups includes waveform data and text data;

[0043] A processing module, configured to, for the text data in each of the data groups, obtain waveform feature information corresponding to the text data by using a preset large language model;

[0044] The processing module is further configured to add the waveform feature information to the text data to update the text data;

[0045] A training module, configured to train a pre-trained model by using the updated multiple groups of data groups to obtain a fine-grained model.

[0046] In a third aspect, the present invention provides an electronic device, including a memory and a processor, where a computer program that can run on the processor is stored in the memory, and when the processor executes the computer program, the steps of the method according to any one of the foregoing embodiments are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 It is a flowchart of the training method based on fine-grained contrast learning provided by the embodiment of the present invention;

[0049] Figure 2 It is another flowchart of the training method based on fine-grained contrast learning provided by the embodiment of the present invention;

[0050] Figure 3 For Figure 2 It is a flowchart of the sub-steps included in S01;

[0051] Figure 4 For Figure 3 It is one of the flowcharts of the sub-steps included in S013;

[0052] Figure 5 ForFigure 3 The second flowchart of the sub-steps included in S013;

[0053] Figure 6 For Figure 1 The flowchart of the sub-steps included in S12;

[0054] Figure 7 For Figure 6 The flowchart of the sub-steps included in S123;

[0055] Figure 8 For Figure 1 The flowchart of the sub-steps included in S13;

[0056] Figure 9 For Figure 8 The flowchart of the sub-steps included in S131;

[0057] Figure 10 The overall logic schematic diagram of the training method based on fine-grained contrastive learning provided by the embodiments of the present invention;

[0058] Figure 11 The functional module block diagram of the training device based on fine-grained contrastive learning provided by the embodiments of the present invention;

[0059] Figure 12 The structural block diagram of the electronic device provided by the embodiments of the present invention. Detailed implementation manners

[0060] Next, the technical solutions in the embodiments of the present invention will be described in combination with the accompanying drawings in the embodiments of the present invention.

[0061] Please refer to Figure 1 , which is the flowchart of the training method based on fine-grained contrastive learning provided by the embodiments of the present invention. The training method based on fine-grained contrastive learning can be executed by a training device based on fine-grained contrastive learning. The training device based on fine-grained contrastive learning can be implemented by software and / or hardware and can be configured in an electronic device. The electronic device can be a computer device, a server, etc. The detailed steps of the training method based on fine-grained contrastive learning are introduced as follows.

[0062] S11, Obtain multiple groups of data groups, and each group of the data groups includes waveform data and text data.

[0063] S12, For the text data in each of the data groups, obtain the waveform feature information corresponding to the text data by using a preset large language model.

[0064] S13, Add the waveform feature information to the text data to update the text data.

[0065] S14. Use the updated multiple data groups to train the pre-trained model to obtain a fine-grained model.

[0066] In this embodiment, the obtained data group can be composed of data of multiple patient-related examinations obtained in historical detections. Each data group includes waveform data and text data. Hereinafter, taking the waveform data as the waveform data in the electrocardiogram and the text data as the text report related to the electrocardiogram as an example for illustration, but it is not limited thereto. This solution can also be applied to the processing of waveform data and text data related to other examination items.

[0067] Among them, the waveform data and text data in the data group can be matched, that is, the situation described in the text data is consistent with the situation represented by the waveform data. For example, they belong to the same patient. Or, the waveform data and text data in the data group can be unmatched, that is, the situation described in the text data is inconsistent with the situation represented by the waveform data.

[0068] Among them, the matching relationship between the waveform data and the text data can be specifically characterized by the level of matching degree, and the level of matching degree between the waveform data and the text data can be achieved through model learning. That is, train a model that can obtain the level of matching degree between the input waveform data and text data.

[0069] Therefore, based on the data groups in which the waveform data and text data in the multiple data groups are unmatched, as well as the data groups in which the waveform data and text data are matched, the model can learn the characteristics of the unmatched waveform data and text data, and the characteristics of the matched waveform data and text data. Thus, ultimately, it can be achieved that based on the input waveform data and text data, the matching degree between the waveform data and the text data can be obtained by analyzing the characteristics of the two.

[0070] The text data is the data in the text report generated during the doctor's diagnosis process. In the actual application scenario, after obtaining an electrocardiogram signal, a doctor usually first observes the waveform characteristics of the electrocardiogram signal, identifies the key waveform characteristics therein (such as the RSR' waveform pattern), to infer possible diagnostic results, such as right bundle branch block (RBBB). When writing the report, the specific waveform characteristics are usually not recorded in the report, only the final diagnostic result is recorded. Therefore, the granularity of the information in the obtained text data is very large, lacking some detailed and specific waveform feature-related information. In this way, it will lead to less feature information that the model can learn when learning the waveform data and text data, and less information materials that can be relied on, resulting in affecting the final evaluation performance of the model.

[0071] Based on the above considerations, in this embodiment, for the text data in each group of data obtained, the waveform feature information corresponding to the text data is obtained by using a preset large language model. Among them, the preset large language model can be any commonly used large language model at present, such as GPT, BERT, etc.

[0072] Through the preset large language model, based on the diagnostic information described in the text data, the waveform feature information that may cause the diagnosis can be obtained in reverse. The obtained waveform feature information is added to the text data to obtain text data with finer granularity.

[0073] After performing the above processing on the text data in each group of data, multiple groups of data can be used to train the pre-trained model until a trained fine-grained model is obtained when the preset conditions are met. Among them, the pre-trained model can be a preliminary model pre-trained using data group samples, and this pre-trained model has a certain ability to align and evaluate waveform data and text data. Continuing to train on the basis of the pre-trained model can improve the performance of the finally trained model.

[0074] The preset conditions can be, for example, that the number of training iterations reaches a preset number, the training duration reaches a preset duration, or the training reaches convergence, etc.

[0075] The training method based on fine-grained contrast learning provided in this embodiment uses a preset large language model to obtain waveform feature information related to the text in reverse, which can enhance the fine-grained degree of the text data, enabling the model to learn more specific and detailed feature information when learning text data and waveform data, thereby improving the final model's ability to interpret fine-grained waveform features and optimizing the model's alignment ability.

[0076] As can be seen from the above, in this embodiment, a pre-trained model can be pre-trained in advance. First, the training process of the pre-trained model will be described below. Please refer to Figure 2 , in this embodiment, after obtaining multiple groups of data, the pre-trained model is obtained through the following method.

[0077] S01, Use the multiple groups of data obtained to train the constructed initial model to obtain a pre-trained model.

[0078] In this embodiment, the initial model can include, but is not limited to, an autoencoder, a convolutional neural network model, a graph neural network model, etc. The pre-trained model is obtained by training on the basis of the initial model using the contrast learning method.

[0079] Specifically, please refer to Figure 3 , Training the initial model with multiple groups of data can be achieved through the following method:

[0080] S011. Import the obtained multiple groups of data into the constructed initial model to obtain the waveform features of the waveform data and the text features of the text data in each group of data.

[0081] S012. Calculate the similarity based on the waveform features and text features in each group of data.

[0082] S013. Construct a loss function based on the similarity.

[0083] S014. Perform iterative training of the initial model guided by the loss function until a pre-trained model is obtained when the preset requirements are met.

[0084] In this embodiment, the initial model may include a waveform signal encoder (such as an electrocardiogram signal encoder). Through the waveform signal encoder, the waveform data in each group of data (such as an electrocardiogram signal) can be feature-encoded to obtain encoded features . Then, through the mapping layer in the initial model map the encoded features into the feature space to obtain waveform features . Its implementation process can be characterized as follows:

[0085] ( )

[0086]

[0087] In addition, the initial model also includes a text encoder , and the text encoder can perform feature encoding on the input text data to obtain encoded features , and then through the mapping layer in the initial model map the features into the feature space to obtain text features . Its implementation process can be characterized as follows:

[0088] ( )

[0089]

[0090] The waveform features and text features obtained in the above manner are both in the form of feature vectors. The similarity can be calculated based on the waveform features and text features in each group of data. Among them, the similarity can be the similarity for measuring the similarity degree between different waveform features, the similarity for measuring the similarity degree between different text features, and the similarity for measuring the similarity degree between waveform features and text features.

[0091] Based on the calculated similarity, a loss function can be constructed. The loss function can be a minimization function, that is, the training of the initial model is guided by this loss function, and the goal of training is to minimize the loss function. When the preset requirements are met, a pre-trained model that has completed training can be obtained. Among them, the preset requirements can be that the number of training iterations reaches a preset number, or the training duration reaches a preset duration, or the loss function converges, etc.

[0092] In this embodiment, the multiple groups of data sets obtained include multiple groups of positive sample data sets and multiple groups of negative sample data sets. The positive sample data set is such that the waveform data and text data it includes are matched, that is, the situation described by the text data is consistent with the situation characterized by the waveform data. The negative sample data set is such that the waveform data and text data it includes are not matched, that is, the situation described by the text data is inconsistent with the situation characterized by the waveform data.

[0093] Electrocardiogram data has a long-tailed distribution, that is, normal electrocardiograms account for the majority, and abnormal electrocardiograms also mostly originate from common diseases. Furthermore, many electrocardiograms may exhibit similar symptoms, resulting in a high semantic similarity in their corresponding diagnostic reports. This makes the traditional multi-modal contrast learning assumption of "different samples are negative samples" not completely valid, thus triggering the problem of pseudo-negative samples, which in turn affects the model training effect.

[0094] That is, in the traditional method, electrocardiograms and reports from different patients are regarded as unmatched and thus classified as negative samples. However, in actual situations, reports from different patients may have similar or consistent situations. Therefore, after swapping the reports of two patients, the newly formed data set may also be a positive sample rather than a negative sample.

[0095] Therefore, in the traditional method, only pairing the electrocardiogram with its corresponding report and embedding them in alignment while distinguishing them from unpaired reports will lead to the problem of false negatives.

[0096] Based on the above considerations, a modified loss function is designed in this embodiment to help identify and correct the problem of false negatives.

[0097] Specifically, the similarity calculated in this embodiment includes the inter-group similarity and the text similarity. The inter-group similarity is the similarity between the waveform features and text features belonging to different data sets, and the text similarity is the similarity between the text features belonging to different data sets.

[0098] In this embodiment, the multiple groups of data sets obtained constitute a data set, which can be expressed as , where n represents the number of data sets, represents paired waveform data and text data, where the waveform features obtained after encoding the waveform data are represented as , the text feature obtained after encoding the text data is represented as .

[0099] The similarity matrix can be calculated in the following way , and this similarity matrix is composed of the similarities between every two text features, which is characterized as follows:

[0100]

[0101]

[0102] Among them, represents the text feature and the text feature the similarity between them.

[0103] In addition, the inter-group similarity of the data groups, that is, the calculation method of the similarity between the waveform features and the text features belonging to different data groups is the same as the above, and can be characterized as:

[0104]

[0105] Among them, represents the waveform feature and the text feature the similarity between them.

[0106] On this basis, please refer to Figure 4 , as a possible implementation method, the steps of constructing a loss function based on the similarity can be implemented in the following way:

[0107] S0131A, for each two of the multiple data groups, calculate the difference value between the inter-group similarity and the text similarity of the two data groups.

[0108] S0132A, construct a modified loss function based on the difference values of each two data groups, and the modified loss function is a minimization function.

[0109] In this embodiment, the inter-group similarity of two data groups can be characterized as , and the text similarity of two data groups can be characterized as . The difference value between the inter-group similarity and the text similarity can be the L1 distance value. Among them, the L1 distance is also called the Manhattan distance, which is a measurement method in the loss function and is used to measure the difference between two vectors.

[0110] On this basis, accumulate the difference values between all pairs of data groups to form a modified loss function, and the obtained modified loss function is as follows:

[0111]

[0112] Among them, B represents the number of data groups, represents the L1 distance.

[0113] The meaning represented by the above-mentioned corrected loss function is that, on the basis of obtaining the similarity between the text features in two data groups, the similarity between the waveform features and the text features of different data groups is made as close as possible to the similarity between the text features. Theoretically, the data pairs composed of the waveform features and the text features of different data groups do not match and are classified as negative samples. However, through the above-mentioned corrected loss function, if the similarity of the text features in two data groups is very high, the similarity between the waveform features and the text features between these two data groups will increase, and the data pairs originally classified as negative samples will be corrected, so as to alleviate the false negative problem.

[0114] In addition, for an electrocardiogram signal, its corresponding diagnosis results may include multiple items. For example, it may include sinus bradycardia, RBBB and LAFB, left ventricular hypertrophy, etc. Therefore, the downstream tasks of electrocardiogram usually have multi-label characteristics. The loss functions of traditional contrastive learning are generally designed based on Softmax and cannot adapt to multi-label scenarios, thus limiting the model performance.

[0115] Based on the above considerations, in this embodiment, a loss function that can adapt to the multi-label scenario of electrocardiogram is designed to enable the model to have multi-label evaluation capabilities.

[0116] In this embodiment, the calculated similarity may include the intra-group similarity, which is the similarity between the waveform features and the text features within the same data group, and the constructed loss function may include a contrastive loss function.

[0117] The multiple data groups include multiple positive sample data groups and multiple negative sample data groups, and each positive sample data group and each negative sample data group have corresponding sample label identifiers. For example, the sample label identifier of the positive sample data group is 1, and the sample label identifier of the negative sample data group is -1.

[0118] Among them, for the positive sample data group, the waveform data and the text data it includes are matched, that is, the situation described by the text data is consistent with the situation represented by the waveform data. For the negative sample data group, the waveform data and the text data it includes are not matched, that is, the situation described by the text data is inconsistent with the situation represented by the waveform data.

[0119] The text data in each data group includes at least one diagnosis result (usually including multiple items).

[0120] As a possible implementation, please refer to Figure 5 , the steps of constructing a loss function based on similarity can be implemented in the following way:

[0121] S0131B constructs a basic function based on the within-group similarity, between-group similarity, and sample label identification corresponding to each data group. The basic function is used to evaluate the matching probability between the waveform data in the data group and each of the at least one diagnostic result.

[0122] S0132B adds a set bias term and temperature parameter to the basic function to obtain a contrastive loss function. The bias term and temperature parameter are used to adjust the ratio of the positive sample data group and the negative sample data group.

[0123] In this embodiment, the calculation method of the within-group similarity corresponding to each data group is the same as the above calculation method of text similarity and can be characterized as . The basic function can be constructed based on the sigmoid function. On this basis, a sample label identification is added to the basic function to distinguish between the positive sample data group and the negative sample data group. The form of the constructed basic function is as follows:

[0124]

[0125] Among them, represents the sample label identification. In , when the values of i and j are the same, it represents the within-group similarity, and when the values of i and j are different, it represents the between-group similarity.

[0126] A basic function is constructed based on the sigmoid function. Under the guidance of the basic function, the model can be trained to solve the multi-label problem and obtain the matching probability between the waveform features and each diagnostic result in the text features. The matching probability between the waveform features and each diagnostic result can be between 0 and 1.

[0127] In the actual scenario, the number of negative samples is much larger than the number of positive samples. During initialization, due to the existence of a large number of negative samples, significant imbalance occurs, and this imbalance dominates in the loss function, causing the model to learn more features of negative samples and insufficiently learn the features of positive samples. To correct this bias, in this embodiment, a bias term and a temperature parameter are added to the basic function, so that the training process can start from a situation where the ratio between positive samples and negative samples is relatively close. The finally constructed contrastive loss function is as follows:

[0128]

[0129] Among them, b represents the bias term, represents the temperature parameter.

[0130] In this embodiment, when training the pre-trained model, any one of the above-mentioned correction loss function and comparison loss function, or a combination of the two can be used for training.

[0131] When training is performed using a combination of the two functions, the final loss function can be shown as follows:

[0132]

[0133] Among them, is a hyperparameter used to control the proportion of the correction loss function and the contrast loss function in the final loss function.

[0134] The above is the process of pre-training the pre-trained model. During the training process of the pre-trained model, considering the multi-label problem in application scenarios such as electrocardiogram, a contrast loss function is designed to guide the model training, so that the model can be used for the evaluation of multi-label problems. In addition, considering the false negative problem in application scenarios such as electrocardiogram, a correction loss function is designed to guide the model training, so that the similarity between the waveform features and text features between different data groups can be tilted towards the similarity between the text features between different data groups, thereby alleviating the false negative problem and improving the learning effect of the model on the semantic representation of electrocardiogram signals.

[0135] The pre-trained model obtained by training in the above manner has a certain evaluation ability. When it is applied to the inference stage, for the input electrocardiogram and text, the pre-trained model can calculate their similarity in the above manner and use the sigmoid function to determine the probability of their matching. Specifically, the matching probability score P can be obtained through the following manner:

[0136]

[0137] Among them, represents function. This probability score is used to indicate the possibility that the electrocardiogram is associated with the text description. The higher the probability score, the stronger the confidence of the model in the matching of the two.

[0138] On the basis of the pre-trained model, after the text data in each data group is updated by using the preset large language model in the above manner, the pre-trained model is continuously trained by using the updated data group to obtain the final fine-grained model.

[0139] Please refer to Figure 6 In this embodiment, the above step of obtaining the waveform feature information corresponding to the text data by using the preset large language model can be implemented in the following manner:

[0140] S121, extract the diagnosis result included in the text data.

[0141] S122, generate inquiry statements based on the diagnosis result.

[0142] S123, import the inquiry statements into a preset large language model, and output waveform feature information corresponding to the diagnosis result.

[0143] In an actual application scenario, the text data is data in a report generated by a doctor based on obtained waveforms, such as electrocardiograms, after diagnosis. As can be seen from the above, the report includes the final diagnosis result, and the diagnosis result may be one or more, but there is a lack of records of specific waveform features. Therefore, it is necessary to restore and supplement this waveform feature information.

[0144] In one possible implementation, the diagnosis result in the text data can be extracted, and inquiry statements can be generated based on the diagnosis result. For example, the generated inquiry statement can be "In an electrocardiogram with symptoms of right bundle branch block, which waveform features are most likely to appear?". Input the generated inquiry statement into the preset large language model, and after being processed by the preset large language model, the corresponding waveform feature information can be output.

[0145] In another possible implementation, the report containing the text data can also be directly input into the preset large language model in the form of a picture. At this time, the inquiry statement can be like "In an electrocardiogram with the symptoms in the report, which waveform features are most likely to appear?". The preset large language model can recognize the report in the form of a picture, identify the text data related to the symptoms therein, and then perform analysis and processing to output the relevant waveform feature information.

[0146] To format the result, when using the preset large language model, specific instructions on the output form can also be given to make the output result more standardized. Specifically, please refer to Figure 7 , which can be achieved in the following way:

[0147] S1231, generate format command statements.

[0148] S1232, import the inquiry statements and the format command statements into the preset large language model, and output the waveform feature information corresponding to the diagnosis result under the format command statements.

[0149] For example, the format command statement can be like "Organize the output waveform feature information into a Python list, where each item represents an independent waveform feature information". In this way, import the inquiry statements and the format command statements into the preset large language model for processing. Through this clear chain-of-thought instruction, a list of potential waveform feature information can be obtained.

[0150] Finally, the obtained waveform feature information can be added to the text data to update the text data.

[0151] Considering that the relationship between waveform features and diagnostic results is not a simple one-to-one correspondence, but involves complex logical reasoning, therefore, each waveform feature item included in the waveform feature information predicted by the preset large language model does not necessarily appear in the electrocardiogram.

[0152] Based on this consideration, in order to verify these waveform feature items, in this embodiment, a pre-trained model obtained in advance is used to screen out the waveform feature information that meets the requirements. Specifically, please refer to Figure 8 In this embodiment, it is implemented in the following manner:

[0153] S131, based on the pre-trained model, screen out the waveform feature information in the waveform feature information that corresponds to the waveform data in the data group.

[0154] S132, add the screened waveform feature information to the text data to update the text data.

[0155] After pre-training, the pre-trained model has a certain alignment ability, and the obtained waveform feature information is also in text form. The pre-trained model can analyze and process to obtain the waveform feature information corresponding to the waveform data. Specifically, please refer to Figure 9 It can be implemented in the following manner:

[0156] S1311, import the waveform data and at least one waveform feature item in the data group into the pre-trained model.

[0157] S1312, obtain the similarity between the waveform data and the at least one waveform feature item through the pre-trained model.

[0158] S1313, use the waveform feature items with similarity greater than or equal to the preset threshold as the waveform feature items corresponding to the waveform data.

[0159] In this embodiment, the waveform data and each waveform feature item are imported into the pre-trained model. The pre-trained model can process the waveform data to obtain the corresponding waveform features, and can also process each waveform feature item to obtain the corresponding text features.

[0160] Based on the above similarity calculation method, calculate the similarity between the waveform features and each text feature, and screen out the waveform feature items with similarity greater than or equal to the preset threshold. Integrate the screened waveform feature items into the original report to obtain updated waveform data.

[0161] Use the updated multiple data groups to continue training the pre-trained model. The process of continued training is similar to the training process of the above pre-trained model, and will not be elaborated here.

[0162] Please refer to Figure 10 the overall logical implementation process of the training method provided in this embodiment shown in the figure. The training method of this embodiment includes two stages. The first stage is the training process of the pre-trained model, and the second stage is the training process of the fine-grained model.

[0163] Among them, in the first stage, based on multiple sample data groups composed of waveform data of electrocardiograms and text data in reports, the initial model is preliminarily trained to obtain a pre-trained model.

[0164] In the second stage, after obtaining the fine-grained report, the pre-trained model is continuously trained based on the data group composed of the fine-grained report and waveform data to obtain the final fine-grained model.

[0165] Among them, doctors generally observe and analyze the obtained electrocardiograms, and infer the diagnosis results according to the waveform features therein (such as the RSR' pattern), such as sinus bradycardia, RBBB and LAFB, and may be left ventricular hypertrophy. The diagnosis results are recorded in the report. However, the specific waveform features are ignored and not recorded in the report.

[0166] Therefore, the generated report is imported into the preset large language model, and inquiry statements are used to guide the preset large language model to output waveform feature items related to the diagnosis results in the report, such as the RSR' pattern, prolonged QRS duration, and high-voltage QRS complex.

[0167] The similarity between each waveform feature item and the waveform data is obtained by using the pre-trained model obtained in the first stage, so as to screen out the waveform feature items with similarity greater than or equal to the preset threshold. The screened waveform feature items are added to the report to obtain a fine-grained report, realizing report enhancement.

[0168] To verify the performance of the fine-grained model under the training method based on fine-grained contrastive learning provided in this embodiment, the MIMIC-ECG dataset is used for pre-training and tested on the PTB-XL, CPSC2018, and CSN datasets. All electrocardiograms in the datasets are recorded in 12 leads. The MIMIC-ECG dataset contains nearly 800,000 pairs of electrocardiogram and report samples. To improve the data quality, samples containing empty reports or reports with less than three words are excluded, reports without useful information are deleted, and electrocardiogram samples with abnormal conditions are discarded.

[0169] The PTB-XL dataset can be further divided into four subsets (including PTBXL-Super, PTBXL-Sub, PTBXL-Form, and PTBXL-Rhythm in Tables 1 and 2), following the official division of the training set, validation set, and test set. For the CPSC2018 and CSN datasets, they are divided into a training set, a validation set, and a test set in the ratio of 70%:10%:20%. The model performance and the comparison results with the current state-of-the-art methods are shown in Tables 1 and 2.

[0170] Table 1 Comparison Table of Validation Results 1

[0171]

[0172] Table 2 Comparison Table of Validation Results 2

[0173]

[0174] Among them, in Tables 1 and 2, the first column represents the methods adopted, including Random Init, SimCLR, BYOL, BarlowTwins, MoCo-v3, SimSiam, TS-TCC, CLOCS, ASTCL, CRT, ST-MEM, MERL, CLEP, FG-CLEP. Among them, CLEP is a conventional contrastive learning training method, and FG-CLEP is the fine-grained contrastive learning method provided in this embodiment.

[0175] In Tables 1 and 2, 1%, 10%, and 100% respectively represent performing validation using the corresponding proportions of data in the corresponding datasets. The validation results of various methods in Tables 1 and 2 in each dataset are evaluated using macro AUC as an indicator. As can be seen from Tables 1 and 2, the macro AUC of the model under the training method provided in this embodiment is relatively higher, indicating better model performance.

[0176] Based on the same inventive concept, please refer to Figure 11 , the embodiment of the present invention also provides a schematic diagram of the functional modules of a training device based on fine-grained contrastive learning. In this embodiment, the functional modules of the training device based on fine-grained contrastive learning can be divided according to the above method embodiment. For example, corresponding functional modules can be divided for each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present invention is schematic, only a logical functional division, and there may be other division methods in actual implementation.

[0177] For example, in the case of dividing each functional module according to each function, Figure 11 The training device based on fine-grained contrast learning shown is only a schematic diagram of a device. The training device based on fine-grained contrast learning may include an acquisition module, a processing module, and a training module. The functions of each functional module of the training device based on fine-grained contrast learning will be elaborated in detail below.

[0178] The acquisition module is used to acquire multiple groups of data sets, and each group of the data sets includes waveform data and text data;

[0179] The processing module is used to obtain waveform feature information corresponding to the text data by using a preset large language model for the text data in each of the data sets;

[0180] The processing module is further used to add the waveform feature information to the text data to update the text data;

[0181] The training module is used to train a pre-trained model by using the updated multiple groups of data sets to obtain a fine-grained model.

[0182] It can be understood that the above acquisition module, processing module, and training module can be used to execute the above S11 to S14. The detailed implementation manners of the acquisition module, processing module, and training module can refer to the content related to the above S11 to S14.

[0183] In a possible implementation manner, the training module can also be used to train the pre-trained model. Specifically, the training module can be used to:

[0184] Train an initial model constructed by using the acquired multiple groups of data sets to obtain a pre-trained model;

[0185] Specifically, the above processing module can be used to:

[0186] Screen out the waveform feature information corresponding to the waveform data in the data set from the waveform feature information based on the pre-trained model;

[0187] Add the screened waveform feature information to the text data to update the text data.

[0188] In a possible implementation manner, the above processing module is used to obtain waveform feature information in the following manner:

[0189] Extract the diagnosis result included in the text data;

[0190] Generate an inquiry statement based on the diagnosis result;

[0191] Import the inquiry statement into a preset large language model to output waveform feature information corresponding to the diagnostic result.

[0192] In a possible implementation manner, the above processing module is specifically configured to obtain waveform feature information in the following manner:

[0193] Generate a format command statement;

[0194] Import the inquiry statement and the format command statement into a preset large language model to output waveform feature information corresponding to the diagnostic result under the format command statement.

[0195] In a possible implementation manner, the waveform feature information includes at least one waveform feature item; the above processing module is used to screen out waveform feature information in the following manner:

[0196] Import the waveform data and at least one waveform feature item in the data group into the pre-trained model;

[0197] Obtain the similarity between the waveform data and the at least one waveform feature item through the pre-trained model;

[0198] Use the waveform feature items with a similarity greater than or equal to a preset threshold as the waveform feature information corresponding to the waveform data.

[0199] In a possible implementation manner, the above training module is used to train and obtain a pre-trained model in the following manner:

[0200] Import multiple groups of data groups obtained into a constructed initial model to obtain the waveform features of the waveform data and the text features of the text data in each group of data groups;

[0201] Calculate the similarity based on the waveform features and text features in each group of data groups;

[0202] Construct a loss function based on the similarity;

[0203] Perform iterative training of the initial model under the guidance of the loss function until a pre-trained model is obtained when the preset requirements are met.

[0204] In a possible implementation manner, the similarity includes inter-group similarity and text similarity. The inter-group similarity is the similarity between the waveform features and text features of different data groups, and the text similarity is the similarity between the text features of different data groups. The loss function includes a correction loss function. The above training module is used to construct the loss function in the following manner:

[0205] For every two of the multiple data groups, calculate the difference value between the inter-group similarity and the text similarity of the two data groups;

[0206] Construct a corrected loss function based on the difference values of every two data groups, and the corrected loss function is a minimization function.

[0207] In a possible implementation manner, the similarity includes intra-group similarity and inter-group similarity. The intra-group similarity is the similarity between the waveform feature and the text feature in the same data group, and the inter-group similarity is the similarity between the waveform features and text features of different data groups. The loss function includes a contrastive loss function.

[0208] The multiple data groups include multiple positive sample data groups and multiple negative sample data groups. Each positive sample data group and each negative sample data group have corresponding sample label identifiers. The text data includes at least one diagnosis result. The above training module is used to construct a loss function in the following manner:

[0209] Construct a basic function based on the intra-group similarity, inter-group similarity, and sample label identifier corresponding to each data group. The basic function is used to evaluate the matching probability between the waveform data in the data group and each diagnosis result in the at least one diagnosis result;

[0210] Add a set bias term and temperature parameter to the basic function to obtain a contrastive loss function. The bias term and temperature parameter are used to adjust the ratio of the positive sample data group and the negative sample data group.

[0211] Please refer to Figure 12 , which is the structural block diagram of the electronic device provided by the embodiment of the present invention. The electronic device can be a computer device, a server, etc. The electronic device includes a memory, a processor, and a communication module. The memory, the processor, and the communication module are directly or indirectly electrically connected to each other to realize data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines.

[0212] Among them, the memory is used to store computer programs or data. The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electric Erasable Programmable Read-Only Memory (EEPROM), etc.

[0213] The processor is used to read / write the data or programs stored in the memory and execute the training method based on fine-grained contrastive learning provided in any embodiment of the present invention.

[0214] The communication module is used to establish a communication connection between the electronic device and other communication terminals through the network and is used to send and receive data through the network.

[0215] It should be understood that Figure 12 The structure shown is only a schematic diagram of the structure of the electronic device, and the electronic device may also include more or fewer components than those shown Figure 12 in it, or have a different configuration from that shown Figure 12 in it.

[0216] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium, and the computer-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed, the training method based on fine-grained contrastive learning provided in the above embodiment is implemented.

[0217] Specifically, the computer-readable storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the computer-readable storage medium is run, it can execute the above training method based on fine-grained contrastive learning. Regarding the process involved when the computer-readable storage medium and its executable instructions are run, reference can be made to the relevant descriptions in the above method embodiments, which will not be elaborated here.

[0218] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0219] In addition, the units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0220] Furthermore, in each embodiment of the present invention, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0221] It should be noted that if the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the existing technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs and other various media that can store program codes.

[0222] The above are only the embodiments of the present invention and are not used to limit the protection scope of the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A training method based on fine-grained contrastive learning, characterized in that, The method includes: Obtaining multiple groups of data sets, each data set including waveform data and text data. The data sets are composed of data of examination items of multiple patients obtained in historical detections. The waveform data is the waveform data of the examination items, and the text data is the text report of the examination items; For the text data in each data set, using a preset large language model to obtain the waveform feature information corresponding to the text data; Adding the waveform feature information to the text data to update the text data; Training a pre-trained model with the updated multiple groups of data sets to obtain a fine-grained model. The pre-trained model is obtained by training an initial model constructed with the obtained multiple groups of data sets. Specifically, it includes: importing the obtained multiple groups of data sets into the constructed initial model, obtaining the waveform features of the waveform data and the text features of the text data in each data set, calculating the similarity based on the waveform features and text features in each data set, and using the contrast learning method to perform iterative training on the basis of the initial model under the guidance of a loss function until the preset requirements are met; the loss function is constructed from the similarity between the waveform features of the waveform data and the text features of the text data in each data set; The similarity includes inter-group similarity and text similarity. The inter-group similarity is the similarity between the waveform features and text features of different data sets, and the text similarity is the similarity between the text features of different data sets. The loss function includes a correction loss function, and the correction loss function is constructed in the following way: For every two data sets in the multiple groups of data sets, calculating the difference value between the inter-group similarity and the text similarity of the two data sets; constructing a correction loss function based on the difference values of every two data sets, and the correction loss function is a minimization function.

2. The training method based on fine-grained contrastive learning according to claim 1, wherein The step of adding the waveform feature information to the text data to update the text data includes: Based on the pre-trained model, screening out the waveform feature information corresponding to the waveform data in the data set from the waveform feature information; Adding the screened waveform feature information to the text data to update the text data.

3. The training method based on fine-grained contrastive learning according to claim 1, wherein The step of using a preset large language model to obtain the waveform feature information corresponding to the text data includes: Extracting the diagnosis result included in the text data; Generating an inquiry statement based on the diagnosis result; Importing the inquiry statement into the preset large language model and outputting the waveform feature information corresponding to the diagnosis result.

4. The training method based on fine-grained contrastive learning according to claim 3, wherein The step of importing the inquiry statement into the preset large language model and outputting the waveform feature information corresponding to the diagnosis result includes: Generating a format command statement; Importing the inquiry statement and the format command statement into the preset large language model and outputting the waveform feature information corresponding to the diagnosis result under the format command statement.

5. The training method based on fine-grained contrastive learning according to claim 2, wherein The waveform feature information includes at least one waveform feature item; The step of screening out the waveform feature information corresponding to the waveform data in the data set from the waveform feature information based on the pre-trained model includes: Import the waveform data and at least one waveform feature item in the data group into the pre-trained model; Obtain the similarity between the waveform data and the at least one waveform feature item through the pre-trained model; Use the waveform feature item with a similarity greater than or equal to the preset threshold as the waveform feature information corresponding to the waveform data.

6. The training method based on fine-grained contrastive learning according to claim 2, wherein The similarity also includes intra-group similarity and inter-group similarity. The intra-group similarity is the similarity between the waveform feature and the text feature in the same data group, and the inter-group similarity is the similarity between the waveform features and text features of different data groups. The loss function also includes a contrastive loss function; The multiple data groups include multiple positive sample data groups and multiple negative sample data groups. Each positive sample data group and each negative sample data group have corresponding sample label identifiers, and the text data includes at least one diagnosis result; The step of constructing the loss function based on the similarity includes: Construct a basic function based on the intra-group similarity, inter-group similarity, and sample label identifier corresponding to each data group. The basic function is used to evaluate the matching probability between the waveform data in the data group and each diagnosis result in at least one diagnosis result; Add a set bias term and temperature parameter to the basic function to obtain a contrastive loss function. The bias term and temperature parameter are used to adjust the ratio of the positive sample data group and the negative sample data group.

7. A training device based on fine-grained contrastive learning, characterized in that A device for implementing the training method based on fine-grained contrastive learning according to any one of claims 1-6, the device includes: An acquisition module for acquiring multiple data groups, each data group including waveform data and text data; A processing module for using a preset large language model to obtain the waveform feature information corresponding to the text data for the text data in each data group; The processing module is further configured to add the waveform feature information to the text data to update the text data; A training module for training the pre-trained model using the updated multiple data groups to obtain a fine-grained model.

8. An electronic device, comprising a memory and a processor, wherein a computer program that can run on the processor is stored in the memory, and is characterized in that The processor implements the steps of the method according to any one of claims 1 to 6 when executing the computer program.

Citation Information

Patent Citations

  • Image-text semantic alignment model training method and device

    CN115187839A

  • Method and apparatus for acquiring pre-trained model

    US20220292269A1