Method, apparatus, electronic device and program product for enhancing medical data set
By sorting the importance of feature items in the medical dataset according to the importance of feature items and generating enhanced feature values, the data scarcity and logical consistency problems in the medical dataset are solved, and the accuracy of data analysis and model training is improved.
Patent Information
- Application Number
- CN202510285922.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
In medical cross-table generation tasks, the problems of scarcity of data, poor logical consistency and insufficient correlation of multi-dimensional features lead to limited accuracy of data analysis and model training.
By obtaining medical data sets, sorting them according to the importance of feature items, the enhanced feature values are generated based on feature values, ensuring the logical consistency of the data, and dynamically enhancing features through the big model, the enhanced medical data set is generated.
It solves the problems of data scarcity and logical consistency, improves the accuracy of data analysis and model training, and ensures the integrity of multi-dimensional feature associations.
Smart Images

Figure CN120216872A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this specification generally relate to the field of computer technology, and more particularly to a method, apparatus, electronic device, and computer program product for enhancing medical data sets. Background Art
[0002] As the degree of dependence on data in various fields deepens, data augmentation technology has become a hot topic in the field of information technology today. The core of data augmentation technology is to automatically generate and expand various types of data through algorithms and models, so as to provide training resources for various deep learning models.
[0003] With the continuous progress of data augmentation technology, the application scenarios of data augmentation are also constantly expanding, gradually extending from the initial text generation to multiple fields such as images, audio, and video. In addition, data augmentation technology is also closely combined with technologies such as big data and cloud computing, jointly promoting the innovation and development of information technology. Summary of the Invention
[0004] Embodiments of this specification provide a method, apparatus, electronic device, and computer program product for enhancing medical data sets.
[0005] In the first aspect of this specification, a method for enhancing a medical data set is provided. The method includes obtaining a medical data set, which includes a plurality of feature items sorted by importance and corresponding multiple feature values. The method further includes generating, according to the sorting of the plurality of feature items, multiple enhanced feature values for the plurality of feature items based on the multiple feature values. In addition, the method further includes determining an enhanced medical data set based on the multiple feature values corresponding to the plurality of feature items and the enhanced multiple feature values.
[0006] In the second aspect of this specification, an apparatus for enhancing a medical data set is provided. The apparatus includes a data set acquisition module configured to obtain a medical data set, which includes a plurality of feature items sorted by importance and corresponding multiple feature values. The apparatus further includes a feature value generation module configured to generate, according to the sorting of the plurality of feature items, multiple enhanced feature values for the plurality of feature items based on the multiple feature values. In addition, the apparatus further includes a data set determination module configured to determine an enhanced medical data set based on the multiple feature values corresponding to the plurality of feature items and the enhanced multiple feature values.
[0007] In the third aspect of this specification, an electronic device is provided. The electronic device includes at least one processor. The electronic device further includes a memory coupled to the at least one processor and having instructions stored thereon that, when executed by the at least one processor, cause the device to execute the method according to the first aspect of this specification.
[0008] In a fourth aspect of the present specification, a computer program product is provided. The computer program product includes machine-executable instructions that, when executed, cause the method according to the first aspect of the present specification to be implemented.
[0009] In a fifth aspect of the present specification, a computer storage medium is provided. The computer-readable storage medium stores computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method provided according to the first aspect of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present specification will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0011] Figure 1 A schematic diagram of an example environment in which some embodiments of the present specification can be implemented is shown;
[0012] Figure 2 A flowchart of a method for enhancing a medical data set according to some embodiments of the present specification is shown;
[0013] Figure 3 A schematic diagram of a causal diagram according to some embodiments of the present specification is shown;
[0014] Figure 4 A flowchart of a method for determining the total deviation of a plurality of feature items according to some embodiments of the present specification is shown;
[0015] Figure 5 A schematic diagram of the application effect 500 of medical data synthesis according to some embodiments of the present specification is shown; and
[0016] Figure 6 A schematic block diagram of a computing device according to some embodiments of the present specification is shown.
[0017] In all the drawings, the same or similar reference numerals represent the same or similar elements. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The embodiments of the present specification will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present specification are shown in the drawings, it should be understood that the present specification can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present specification. It should be understood that the drawings and embodiments of the present specification are only for exemplary purposes and are not used to limit the protection scope of the present specification.
[0019] It is understandable that before using the technical solutions disclosed in the embodiments of this specification, the types, scope of use, usage scenarios, etc. of the personal information involved in this specification should be informed to users and the authorization of users should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0020] In the description of the embodiments of this specification, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions below.
[0021] In medical research and clinical trials, especially in the fields of rare disease research, medical auxiliary diagnosis, etc., constructing a cross-table data set with logical consistency, covering multi-dimensional features and capable of supporting complex correlation analysis is of great significance for promoting the application of medical artificial intelligence. The quality and quantity of data are crucial for the accuracy of research results and medical decisions. In related technologies, it is usually table data generation, mainly generating records in a single table, aiming to replicate the underlying distribution of the original data. The generated records are similar to the original table in terms of data distribution to simulate real data.
[0022] In the field of synthetic tabular data generation, there is Graph-based Relational Data Generation for Tabular Data (GReaT) for generating realistic synthetic data. There are Fourier-Transformers (FT-Transformer) which show unique advantages in tabular data processing with their powerful architectures. Variational Information Maximizing Exploration (VIME), as a self- and semi-supervised framework, has been deeply studied specifically for tabular data processing. There is also research on the contrastive learning method based on feature corruption, Self-supervised Contrastive Learning via Feature Reconstruction and Feature Fusion (SCARF). Another research focuses on Self-Attentive Instance Neighborhood Transformer (SAINT) which focuses on attention mechanisms for handling rows and columns. In addition, there is work applying additive attention mechanisms to tabular tasks, focusing on global context. There is research proposing Tabular Network (TabNet) which uses sequential attention to enhance interpretability. The research on Tabular Transformer (TabTransformer) and Tabular Predictive Flow Networks (TabPFN) respectively emphasizes the importance of fast data processing.
[0023] In terms of transfer learning for cross-tabular tasks, there is work proposing Cross-Tabular Bidirectional Encoder Representations from Transformers (CT-BERT) which integrates contrastive learning and is applicable to supervised and self-supervised settings. There is also work developing eXtensible Tabular (XTab), a framework for pre-training tabular transformers applicable to various tasks including regression and classification. In addition, there is work emphasizing the importance of transferring knowledge across different tables to a specific target. However, these methods for synthetic tables are mainly applied in general domains rather than the medical domain which requires more meticulous data synthesis methods.
[0024] To this end, an embodiment of this specification provides a method for enhancing a medical data set. First, a medical data set is obtained. Then, based on the sorting order of the importance of feature items, the feature values of multiple feature items are enhanced based on multiple feature values to generate multiple enhanced feature values. Further, an enhanced medical data set is determined based on the feature values corresponding to the feature items and the enhanced feature values.
[0025] In this way, an enhanced data set can be generated based on the feature items and feature values in the original data set. The enhanced data set is different from the original data set, ensuring the sufficiency of the data and making up for the scarcity in the real data distribution. Moreover, the enhanced data set is generated based on the importance order of the feature items, maintaining the logical consistency of the data. Therefore, the embodiments of this specification can solve the problems of data scarcity, poor logical consistency, and insufficient multi-dimensional feature association in medical cross-table generation tasks, improving the accuracy of data analysis and model training in the medical field.
[0026] Figure 1 FIG. shows a schematic diagram of an example environment 100 in which some embodiments of this specification can be implemented. Refer to Figure 1 , the example environment 100 includes a computing device 102, and the computing device 102 includes but is not limited to a communication device, an arithmetic chip, a computing system, a single server, a distributed server, or a cloud-based server, etc. A feature table 104, a medical document 106, a large model 110, and a medical data synthesis table 112 are deployed in the computing device 102. It can be understood that the feature table 104 includes feature items 114, and the feature items 114 can be medical feature items, including "disease name", "white blood cell value", "tumor size", etc. They can also be personal information feature items, such as "name", "occupation", etc. The medical document 106 can include the medical records of multiple patients, and the medical record of each patient can include information describing the feature items 114, such as detailed information on drug-taking history, imaging examinations, biochemical indicators, etc.
[0027] In some embodiments, the large model can be a generative large model or a multi-modal large model. The large model can identify the required features from the unstructured medical document 106 and determine their corresponding feature values. The large model can also be used for dynamic enhancement and feature expansion, enriching the feature dimensions of the data set, enhancing the diversity and coverage of the data set, and providing more comprehensive data support for medical research and analysis. The medical data synthesis table 112 is based on the unique identifier of the patient, the features in the structured medical document 106, and the corresponding feature values.
[0028] In some embodiments, the large model can dynamically enhance features for the dynamic enhancement and feature expansion processes. It can generate disease-related feature items based on a given disease label and a set of relevant context features. When given the disease label "diabetes" and context features such as the patient's blood sugar level and family medical history, the large model can generate feature items such as "diabetes complication risk assessment" and "recommended blood sugar control target value".
[0029] Figure 2 FIG. 200 is a flowchart of a method for enhancing a medical data set according to the present specification. In some embodiments, in the exemplary environment 100 shown in Figure 1 FIG. 1, the method 200 may be executed by the computing device 102. It should be understood that although the following description is based on the computing device 102 as the execution entity, the method 200 may also be executed by other devices. The method 200 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present specification is not limited in this regard.
[0030] At step 202, a medical data set is obtained, which includes a plurality of feature items sorted by importance and corresponding feature values. It can be understood that the medical data set contains numerous patient summaries, which cover rich medical information, from basic patient personal information such as name, to detailed disease symptoms, diagnosis processes, treatment plans, and records of disease progression. For example, some patient summaries detail the situation from the initial appearance of symptoms to the diagnosis of the disease and then to the recovery after various treatment measures, providing a comprehensive and real data basis for subsequent medical research and data analysis.
[0031] In some embodiments, the feature items and corresponding feature values in the medical data set can be basic personal information such as "Name: Xiaoming", or disease-related information such as "Medical history: 3 years", "Blood pressure: high", etc. The feature items and feature values describing different types of information can be in different subsets of the medical data set. It can be understood that the medical data set can be the above-mentioned unstructured documents or a structured table constructed based on unstructured documents.
[0032] The medical data set includes feature items and feature values sorted by importance. It can be understood that in actual medical research and applications, not all feature items have the same importance. Therefore, these feature items are evaluated for importance and sorted according to specific research purposes or clinical needs. When studying the pathogenesis of a rare disease, feature items such as specific gene markers and particular symptoms directly related to the disease may be considered the most important and ranked at the forefront; while some general demographic features, although of certain value, have relatively lower importance and are ranked behind.
[0033] At step 204, according to the sorting of multiple feature items, multiple enhanced feature values are generated based on multiple feature values. In some embodiments, the large model enhances the feature values of the feature item with the highest importance level to obtain multiple enhanced feature values of the feature item with the highest importance level. Then, the large model enhances the feature values of the feature item with the second highest importance level to obtain multiple enhanced feature values of the feature item with the second highest importance level. And so on, the large model enhances the feature values of each feature item according to the importance order of the feature items to obtain multiple enhanced feature values of each feature item.
[0034] At step 206, based on the multiple feature values corresponding to multiple feature items and the multiple enhanced feature values, an enhanced medical dataset is determined. In some embodiments, the computing device 102 integrates the feature values corresponding to each feature item and the multiple enhanced feature values to obtain a set of feature values. Based on certain conditions, the computing device 102 selects a target feature value from the set of feature values. It can be understood that the computing device 102 selects a target feature value based on the set of feature values of each feature item, and then the computing device 102 integrates the target feature values corresponding to each feature item to form an enhanced medical dataset. It can be understood that the enhanced medical dataset includes the feature items sorted by importance level and the corresponding target feature values, and the target feature values may be the same as or different from the feature values in the medical dataset.
[0035] In the embodiments of this specification, an enhanced dataset can be generated based on the feature items and feature values in the original dataset. The enhanced dataset is different from the original dataset, which ensures the sufficiency of the data and makes up for the scarcity in the real data distribution. Moreover, the enhanced dataset is generated based on the importance order of the feature items, maintaining the logical consistency of the cross-table data and meeting the requirements of complex logical associations between multiple tables. Therefore, the embodiments of this specification can solve the problems of data scarcity, poor logical consistency, and insufficient multi-dimensional feature associations in the medical cross-table generation task, and improve the efficiency of data-driven analysis and artificial intelligence modeling in the medical field.
[0036] In some embodiments, the computing device 102 obtains medical record files from the medical file library. The medical file library is a resource library that stores a large amount of medical-related materials and contains rich case files. These case files have a wide range of sources and may come from clinical records of different hospitals, experimental reports of medical research institutions, case analyses published in academic journals, etc. These medical record files may contain detailed information such as patient basic information, symptom descriptions, diagnosis results, treatment processes, etc. To focus on the research of a specific disease, it is necessary to screen out the documents directly related to the disease from the obtained medical record files.
[0037] In some embodiments, the computing device 102 selects a set of disease-related documents from the case files through keywords related to the disease. These keywords are determined based on the characteristics, symptoms, diagnostic criteria, etc. of the disease. Through these keywords, documents related to the target disease can be accurately identified. For example, for the study of diabetes, keywords such as "diabetes", "blood glucose", "insulin", "glycated hemoglobin" can be used. During the screening process, the computing device 102 scans the text content in the case files, searches for documents containing these keywords, and collects them to form a document set D. D = {d i |d i ∈Dataset, match(d i , Keyword) = 1}, where d i represents a single document, and match(d i , Keyword) indicates whether the document d i contains keywords related to the target disease. This can effectively exclude documents unrelated to the target disease and improve the efficiency and accuracy of subsequent data processing.
[0038] In some embodiments, after obtaining the set of disease-related documents, for more convenient data extraction and analysis, the computing device 102 divides the document set according to the types of paragraphs in the documents. The types of paragraphs can be classified according to their content and use, such as patient basic information paragraphs, biochemical index paragraphs, imaging examination paragraphs, treatment plan paragraphs, etc. Taking a diabetes case document as an example, the document may contain basic information paragraphs such as the patient's name and occupation, symptom description paragraphs such as polyuria, polydipsia, and polyphagia, and diagnostic result paragraphs such as whether the patient has diabetes and the type of diabetes. By dividing these paragraphs according to their types, the document set D can be divided into multiple document subsets S(d i ), S(d i ) = {s j |s j ∈d i , type(s j )∈T}, where s j represents a paragraph in the document d i , and T represents a predefined set of paragraph types. Each subset contains paragraphs of the same type, facilitating subsequent targeted processing of different types of information.
[0039] In some embodiments, the computing device 102 determines disease-related feature items based on a medical term library, which are used to describe various attributes or metrics of disease-related information. For example, for the study of diabetes, feature items such as "blood glucose level", "glycated hemoglobin value", and "whether there is a family history" may be defined in the medical term library. When determining feature values from multiple document subsets, the computing device 102 analyzes and extracts the content of each document subset based on a large model to find the specific values or descriptions corresponding to the feature items. E(d i ) = {(e k , v k ) | e k ∈ Ontology, v k ∈ extract(s j )}, where E(d i ) represents feature information, including feature items and corresponding feature values, e k represents the feature item, v k represents the feature value, Ontology represents the medical term library, and extract(s j ) represents the s j excerpt.
[0040] In some embodiments, if the computing device 102 finds the information "the patient's blood glucose is 4.5 mmol / L" in the document subset, then the feature value corresponding to the feature item "blood glucose" is "4.5 mmol / L". The computing device 102 collects the feature values corresponding to all feature items to form a medical data subset. By performing the same processing on different types of document subsets, the computing device 102 can obtain multiple medical data subsets and finally form a complete medical data set.
[0041] In some embodiments, the computing device 102 cannot find the feature values corresponding to some feature items in the document subset. This may be due to incomplete case file records, missing information, etc. For example, in some case documents, the family history information of the patient may not be recorded, so there is no corresponding feature value for the feature item "whether there is a family history". The computing device 102 utilizes the powerful context understanding and learning ability of the neural network model to automatically analyze based on the context information related to the feature item. It can be understood that the context information may include other relevant symptom descriptions in the same document, the overall disease progression record of the patient, etc. Taking the processing of research data on a certain rare disease as an example, if the feature value corresponding to the feature item "specific gene mutation" is missing in the document subset, but the document mentions the special clinical manifestations related to the disease and the situation of other associated genes, the neural network model will conduct in-depth learning and reasoning based on this context information. v′ k = f context (E(di ),C(s j )), where v′ k is the possible value inferred for the missing feature, f context It is an inference model based on deep learning, C(s j ) is j The context of the paragraph. The set of inference features F = {(e k ,v′ k )|e k ∈Ontology}.
[0042] The neural network model uses deep learning algorithms to train a large amount of similar contextual data to grasp the potential relationship between feature items and context. In practical applications, it will perform feature extraction, pattern recognition and other processing on the input context information to generate reasonable feature values. For example, in the above rare disease cases, the model may generate possible feature values of "specific gene mutations" based on existing clinical manifestations and related gene information, combined with its learned knowledge, such as specific mutation types or mutation probability ranges.
[0043] The method of generating missing feature values based on the neural network model improves the integrity and availability of the data, making subsequent data processing and analysis more accurate and reliable, providing strong support for building a cross-table data set with consistent logic and covering multi-dimensional features, and can better meet the needs of various medical applications such as medical research, clinical trial simulation, and medical model development.
[0044] It is understandable that the computing device 102 enhances the feature value in a cyclic iteration manner. In this way, it is possible to fully utilize the known information in the medical data set and explore the potential relationship between the features. In some embodiments, the computing device 102 obtains the current feature item and the corresponding feature value from the medical data set, and obtains the feature item and the corresponding feature value and the enhanced feature value before the current feature item according to the order of the feature items. According to the order of multiple feature items, the feature value corresponding to the feature item before the current feature item and the enhanced feature value are obtained.
[0045] When constructing a medical data set, the logical relationship between the features in each table corresponding to each medical data subset is determined, and the order of feature generation is sorted according to the logical weight. This sorting reflects the causal relationship and importance between the features. For example, when constructing a data set related to patient disease diagnosis, the "symptom manifestation" feature item is often generated before the "disease diagnosis result" feature item, because symptoms are an important basis for disease diagnosis. Obtaining the feature value before the current feature item and the enhanced feature value helps to comprehensively consider the relationship between the data and provide comprehensive information for generating more reasonable enhanced feature values.
[0046] The computing device 102 utilizes a large model to enhance the eigenvalue of the current feature item based on the current feature item and its corresponding eigenvalue, as well as the previous feature item and its corresponding eigenvalue and the enhanced eigenvalue. The enhanced eigenvalue not only contains the information of the original data but also learns and fuses other relevant features through the model. The enhanced eigenvalue of the current feature item is used for the next iteration. It can be understood that the computing device 102 enhances the eigenvalue of the next feature item based on the eigenvalue and the enhanced eigenvalue corresponding to the current feature item, as well as the eigenvalue corresponding to the next feature item. The entire enhancement process forms a dynamic loop, and the enhanced eigenvalue generated in each iteration will be used as the input information for the next iteration, continuously optimizing and enriching the data.
[0047] In some embodiments, after generating multiple enhanced eigenvalues, the computing device 102 selects multiple target eigenvalues corresponding to multiple feature items from the eigenvalues and the enhanced eigenvalues. The computing device 102 performs screening according to the occurrence probability of the target disease. The occurrence probability of the target disease is a key measurement index. By statistically analyzing a large amount of medical data and combining medical expertise and research experience, a suitable probability condition is determined. For example, when studying cardiovascular diseases, considering multiple feature items such as blood pressure, blood lipids, and family medical history, through the study of a large number of cases, it is found that when these feature items are within a specific range, the occurrence probability of cardiovascular diseases reaches a certain value, and this value can be set as the condition for screening target eigenvalues.
[0048] In some embodiments, the computing device 102 evaluates the occurrence probability of the target disease for each combination of eigenvalues of the feature item (including the eigenvalue and the enhanced eigenvalue) one by one. For each possible combination of eigenvalues, the computing device 102 uses the large model for calculation. For example, using a logistic regression model or a probability prediction model based on deep learning, inputting the current combination of eigenvalues, and the model outputs the occurrence probability of the target disease. Only when this occurrence probability meets the pre-set condition will this combination of eigenvalues be selected as the target eigenvalue v′ i , v′ i = argmax v P(v|y,context), where context is a set of context features related to the target disease. The data quality of the target eigenvalue is evaluated by the following formula: where match and consistency represent the matching degree and consistency of the target eigenvalue with the real data.
[0049] In some embodiments, after the computing device 102 determines multiple target feature values corresponding to multiple feature items, it determines an enhanced medical data set related to the target disease based on these feature items and the corresponding target feature values. Taking cardiovascular disease as an example, feature items such as blood pressure and blood lipid are respectively corresponded to the target feature values that meet the conditions of the occurrence probability of cardiovascular disease, forming data records containing multiple key information, which constitute the enhanced medical data set.
[0050] Figure 3 FIG. 300 is a schematic diagram of a causal graph according to some embodiments of the present specification. The computing device 102 establishes a causal graph based on a medical data subset. Nodes in the causal graph represent feature items in the medical data subset, and edges in the causal graph represent causal relationships between two feature items connected by the edge. As Figure 3 shown, the causal graph 300 includes feature A 302, feature B 304, and feature C 306. The edge 306 between feature A 302 and feature B 304 represents the causal relationship between feature A and feature B. The edge 310 between feature A 302 and feature C 306 represents the causal relationship between feature A and feature C. The edge 310 between feature B 304 and feature C 306 represents the causal relationship between feature B and feature C.
[0051] In some embodiments, the feature item "working years" (from the basic information subset), the feature item "blood pressure value" (from the biochemical index subset), the feature item "degree of stenosis of the heart blood vessels" (from the imaging examination subset), etc. can all be used as nodes in the causal graph. It can be understood that an increase in working years may lead to a decrease in blood vessel elasticity, which in turn affects the increase in blood pressure value. This causal connection is represented by the edge connecting the "working years" node and the "blood pressure value" node; if the blood pressure remains at a relatively high level for a long time, it may also cause the degree of stenosis of the heart blood vessels to increase. This causal connection is represented by the edge connecting the "blood pressure value" node and the "degree of stenosis of the heart blood vessels" node. In this way, the causal graph comprehensively and intuitively shows the causal connections between feature items in the medical data subset.
[0052] In the above example of cardiovascular disease, the weight of the edge can be calculated by statistically counting the frequency of co-occurrence of two feature items in the data of cardiovascular disease patients. Suppose the two feature items "hypertension" and "coronary heart disease", in the statistically counted patient data, the frequency of co-occurrence of "hypertension" and "coronary heart disease" is relatively high, and the frequency of occurrence of "hypertension" is relatively stable, then the weight of the edge between "hypertension" and "coronary heart disease" will be relatively large, indicating that hypertension has a strong causal impact on the occurrence of coronary heart disease.
[0053] In some embodiments, in order to more accurately quantify the causal relationship between feature items, the computing device 102 needs to determine the weights of multiple edges in the causal graph. The weight w of the edgeij Indicates the degree of causal influence between features. Where the weight w of the edge ij represents the conditional probability of features v i and v j , and freq(v i, v j ) represents the co-occurrence frequency of v i and v j , and freq(v i ) is the occurrence frequency of v i . The computing device 102 determines the weight of the edge between two feature items based on the ratio of the frequency of simultaneous occurrence of the feature values corresponding to the two feature items and the frequency of occurrence of the feature values corresponding to a single feature item.
[0054] Taking cardiovascular disease research as an example, assume that the first feature item is "hypertension", and its corresponding feature value can be systolic blood pressure greater than 140 mmHg and / or diastolic blood pressure greater than 90 mmHg; the second feature item is "coronary heart disease", and its corresponding feature value is being diagnosed with coronary heart disease through medical examinations. Count the number of patients who simultaneously meet the "hypertension" feature value and the "coronary heart disease" feature value, and then divide by the total number of patient data records. The resulting ratio is the first frequency. This frequency reflects the frequency of simultaneous occurrence of hypertension and coronary heart disease in the medical dataset.
[0055] Next, determine the frequency of occurrence of the feature value corresponding to the first feature item (the second frequency). Count the number of patients in the dataset with systolic blood pressure greater than 140 mmHg and / or diastolic blood pressure greater than 90 mmHg, that is, the number of patients with the "hypertension" feature value, and then divide by the total number of patient data records. The resulting value is the second frequency. It represents the frequency of occurrence of the feature of hypertension alone in the dataset. Based on the ratio of the first frequency to the second frequency, determine the original weight (the first weight) of the edge between the nodes corresponding to the first feature item and the second feature item. The larger the ratio of the first frequency to the second frequency, the higher the probability of simultaneous occurrence of coronary heart disease in the case of hypertension, which means that the causal influence of hypertension on coronary heart disease is stronger.
[0056] In some embodiments, the computing device 102 sorts multiple feature items based on the weight of the edge and obtains the order of feature generation by optimizing the deviation . By adjusting the order of the features, the deviation L orderMinimization. The purpose is to clarify the importance and order of each feature item in the entire causal relationship network. The feature items connected by edges with larger weights are more closely related causally, meaning they have a greater impact on subsequent results. In cardiovascular disease research, if the weight of the edge between "hypertension" and "coronary heart disease" is high, it indicates that hypertension has a significant impact on the occurrence and development of coronary heart disease, so the feature item "hypertension" will be relatively ranked higher in the sorting.
[0057] In some embodiments, after the initial enhancement of the medical data set is completed, the computing device 102 further optimizes the enhanced data set. The computing device 102 builds an enhanced causal graph based on the enhanced medical data set, analyzes the weight difference of the edges to adjust the sorting of the feature items, so that the sorting of the feature items in the enhanced medical data set is more reasonable and more in line with the sorting of the feature items in the medical data set.
[0058] The computing device 102 builds an enhanced causal graph based on the enhanced medical data set. The principle is similar to building a causal graph based on the medical data set, but the enhanced causal graph reflects the causal relationship between the feature items in the enhanced medical data set. The computing device 102 calculates the enhanced weight (second weight) of the edge connecting the corresponding nodes of two feature items based on the ratio of the frequency of occurrence of the feature values in the enhanced medical data set and the frequency of co-occurrence of the feature values of two causally related feature items. The computing device 102 determines the change in the causal relationship before and after data enhancement based on the difference between the original weight (first weight) and the enhanced weight (second weight) of the edge between the nodes corresponding to two causally related feature items. If the difference is large, it means that the causal relationship between the two feature items has changed significantly before and after data enhancement. If the difference is positive, it may indicate an enhanced causal relationship, and vice versa, it may indicate a weakened causal relationship.
[0059] The computing device 102 sums up multiple differences between the multiple original weights and the enhanced weights, and the obtained sum is used to determine the sorting deviation D of multiple feature items in the enhanced medical data set. logic , where w i ′ j is the weight of the enhanced data (second weight). As Figure 3As shown, the computing device 102 sums the differences between the original weights and the enhanced weights of edge 308, the differences between the original weights and the enhanced weights of edge 310, and the differences between the original weights and the enhanced weights of edge 312 to obtain the deviation in the sorting of multiple feature items in the enhanced medical dataset. If the deviation is large, it indicates that there is a significant difference in the sorting of feature items between the enhanced medical dataset and the original medical dataset, and further adjustment is required; if the deviation is small, it indicates that the enhanced dataset is relatively stable in terms of causal relationships and the sorting of feature items. It can be understood that the computing device 102 adjusts the sorting of feature items in the enhanced medical dataset to minimize the sorting deviation, making the sorting of multiple feature items in the enhanced medical dataset more reasonable.
[0060] In some embodiments, after completing the construction of the enhanced medical dataset and the preliminary adjustment of the sorting of feature items, the computing device 102 determines the consistency deviation D between the multiple target feature values and the multiple original feature values of the multiple feature items. consistency . The target feature value is the key data selected for constructing the enhanced medical dataset, while the original feature value comes from the initially obtained medical dataset. Taking the research on cardiovascular diseases as an example, a certain feature value of the "blood pressure value" feature item in the original medical dataset is "systolic blood pressure 130 mmHg, diastolic blood pressure 85 mmHg". After data enhancement and screening, the target feature value of this feature item is "systolic blood pressure 135 mmHg, diastolic blood pressure 90 mmHg". The computing device 102 compares these two sets of feature values and determines the consistency deviation of this feature item through a set calculation method, such as calculating the difference between the two or using a more complex similarity measurement algorithm. Such operations are performed on all feature items in the medical dataset to obtain a set of consistency deviations for multiple feature items.
[0061] Figure 4 The flowchart of method 400 for determining the total deviation of multiple feature items in some embodiments of this specification is shown. After determining the consistency deviation, the computing device 102 combines it with the sorting deviation of the multiple feature items to calculate the total deviation L of the multiple feature items. opt . L opt = λ1D logic + λ2D consistency , where λ1 and λ2 are weight parameters. The sorting deviation reflects the difference between the sorting of feature items in the enhanced dataset and the ideal causal relationship sorting, and the consistency deviation reflects the degree of deviation between the target feature value and the original feature value. Together, they constitute the comprehensive index total deviation for measuring the quality of the enhanced medical dataset.
[0062] Suppose in the cardiovascular disease dataset, the change in the weight of the edge between feature items such as "hypertension" and "coronary heart disease" before and after enhancement results in a sorting deviation of 0.3. At the same time, the comprehensive calculation of the consistency deviation between the target feature values and the original feature values of multiple feature items such as "blood pressure value" and "blood lipid level" is 0.2. Through specific weighted calculation (the weight can be set according to actual needs and data characteristics), the total deviation is 0.25 (for example, assume the sorting deviation weight is 0.6 and the consistency deviation weight is 0.4, total deviation = 0.3×0.6 + 0.2×0.4 = 0.25).
[0063] The computing device 102 adjusts the sorting of multiple feature items and multiple target feature values so that the total deviation meets certain conditions. During the adjustment process, in order to reduce the sorting deviation, a certain feature item needs to be advanced, but this may increase the consistency deviation between its target feature value and the original feature value. The computing device 102 will comprehensively weigh the impact of the changes of both on the total deviation. For example, in the cardiovascular disease dataset, if the "family medical history" feature item is advanced, it can better reflect its causal relationship with other disease-related feature items and reduce the sorting deviation, but it may slightly increase the consistency deviation between the "family medical history" target feature value and the original feature value. The computing device will decide whether to accept this adjustment according to the change of the total deviation.
[0064] By continuously adjusting the feature item sorting and target feature values to make the total deviation meet the conditions, the enhanced medical dataset obtained has been greatly improved in terms of logical relationship and data accuracy. Such a dataset can provide more reliable data support for medical research and help researchers analyze diseases more accurately. In the training of medical artificial intelligence models, it can also improve the accuracy and stability of the models.
[0065] Figure 5 The schematic diagram of the medical data synthesis application effect 500 of some embodiments of this specification is shown. The computing device 102 uses the publicly available medical literature library as the data source, screens and de-duplicates the documents related to the target disease through keyword matching, extracts medical features using the large model entity recognition tool, complements rare features and missing feature values through the context inference model, constructs a multi-table structure dataset with the patient unique identifier as the primary key, models the logical relationship between the features of each table with the help of the causal graph, generates rare features using the generative adversarial network, expands features with the help of the large model, and ensures the data quality through logical consistency verification and automatic optimization, so as to generate synthetic patient data containing rich features.
[0066] In some embodiments of this specification, synthetic patient data can support clinical trials. The synthetic data covers rich features such as occupation type, disease, examinations, etc. These data can simulate real patient situations, helping researchers conduct pre-experiments before the trial to evaluate the feasibility and potential effects of different treatment plans. The synthetic medical data can provide researchers with a large number of samples, contributing to the in-depth exploration of the pathogenesis, risk factors, and development laws of diseases.
[0067] In the field of rare disease research, due to the insufficient number of real cases, it is difficult to obtain enough data for comprehensive research. The medical data generated by this solution can supplement the deficiency of real data, and researchers can use these data to mine potential features and causal relationships related to rare diseases. In terms of medical auxiliary diagnosis, these synthetic data can be used to train medical models. Through the training of a large amount of synthetic data, the model can learn more disease features and diagnostic patterns, improving the accuracy and reliability of diagnosis, providing more valuable diagnostic suggestions for doctors, and assisting doctors in making more accurate medical decisions.
[0068] Figure 6 A block diagram of a computing device 600 suitable for implementing the embodiments of the present invention is schematically shown. The computing device 600 can be used to implement the computing device 102. The computing device 600 can be a device for implementing the Figure 2 method 200 shown. As Figure 6 shown, the computing device 600 includes a processing unit (CPU) 601, which can execute various appropriate actions and processes according to computer program instructions stored in the read-only memory (ROM) 602 or computer program instructions loaded from the storage unit 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computing device 600 can also be stored. The CPU 601, ROM 602, and RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0069] Multiple components in the computing device 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608. The processing unit 601 executes the various methods and processes described above, such as executing the method 200. For example, in some embodiments, the various processes or operations described above may be implemented as a computer software program, which is stored in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed onto the computing device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the CPU 601, the various methods and processes described above may be executed, such as one or more operations of the method 200. Alternatively, in other embodiments, the CPU 601 may be configured to execute the various methods and processes described above, such as one or more actions of the method 200, by any other suitable means (e.g., by means of firmware).
[0070] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0071] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.
[0072] These computer - readable program instructions can be provided to the processing unit of a processor, a general - purpose computer, a special - purpose computer, or other programmable data - processing devices in a voice interaction device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data - processing devices, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, and these instructions cause the computer, programmable data - processing devices, and / or other devices to work in a specific manner.
[0073] The above - described specific embodiments of the present specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0074] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
[0075] The above are only alternative embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for enhancing a medical dataset, comprising: Acquire the medical data set, wherein the medical data set includes a plurality of feature items and corresponding plurality of feature values sorted by importance; Generating a plurality of enhanced feature values of the plurality of feature items based on the plurality of feature values according to the order of the plurality of feature items; as well as An enhanced medical data set is determined based on the multiple feature values corresponding to the multiple feature items and the enhanced multiple feature values.
2. The method according to claim 1, wherein generating the enhanced plurality of feature values comprises: In a current iteration, obtaining a current feature item and a corresponding feature value from the medical data set; According to the order of the multiple feature items, obtain the feature value and the enhanced feature value corresponding to the feature item before the current feature item; as well as Based on the feature value corresponding to the current feature item, and the feature value and enhanced feature value corresponding to the feature item before the current feature item, the enhanced feature value of the current feature item is generated, and the enhanced feature value of the current feature item is used for the next iteration.
3. The method according to claim 1 or 2, wherein determining the enhanced medical data set comprises: Selecting a plurality of target feature values corresponding to the plurality of feature items from the plurality of feature values and the plurality of enhanced feature values, wherein the plurality of target feature values make the occurrence probability of the target disease meet a condition; as well as Based on the multiple feature items and the corresponding multiple target feature values, an enhanced medical data set related to the target disease is determined.
4. The method according to claim 3, wherein the medical data set comprises a plurality of medical data subsets of different types, and the method further comprises: Based on the multiple medical data subsets, establishing a causal graph, wherein multiple nodes in the causal graph indicate the multiple feature items in the multiple medical data subsets, and multiple edges in the causal graph indicate causal relationships between the multiple feature items; as well as The plurality of feature items are sorted based on the weights of the plurality of edges in the causal graph.
5. The method of claim 4, wherein establishing the cause-effect graph comprises: Determine a first frequency at which characteristic values corresponding to a first characteristic item and a second characteristic item of the plurality of characteristic items appear simultaneously; Determine a second frequency of occurrence of a characteristic value corresponding to the first characteristic item; and Based on the ratio of the first frequency to the second frequency, a first weight of an edge between nodes corresponding to the first feature item and the second feature item is determined, wherein the edge indicates a causal relationship in which a feature value corresponding to the first feature item causes a feature value corresponding to the second feature item to appear.
6. The method according to claim 5, wherein sorting the plurality of feature items comprises: In the causal graph, determining multiple pairs of nodes corresponding to multiple pairs of first characteristic items and second characteristic items having a causal relationship; Determining a deviation of the order of the plurality of feature items based on a sum of a plurality of first weights corresponding to a plurality of edges between the plurality of pairs of nodes; as well as The order of the plurality of feature items is adjusted so that a deviation of the order of the plurality of feature items satisfies a condition.
7. The method of claim 4, wherein determining the enhanced medical dataset comprises: Based on the enhanced medical data set, establishing an enhanced causal graph, wherein a plurality of nodes in the enhanced causal graph indicate the plurality of feature items in the enhanced medical data set, and a plurality of edges in the enhanced causal graph indicate causal relationships between the plurality of feature items; Based on the enhanced causal graph, determine multiple second weights of multiple edges between multiple pairs of nodes corresponding to multiple pairs of first feature items and second feature items having causal relationships; Determine a difference between a first weight and a second weight of an edge between nodes corresponding to a first characteristic item and a second characteristic item having a causal relationship; Determining a deviation of the ranking of the plurality of feature items in the enhanced medical dataset based on a sum of a plurality of differences between the plurality of first weights and the plurality of second weights; as well as The order of the plurality of feature items in the enhanced medical data set is adjusted so that a deviation of the order of the plurality of feature items satisfies a condition.
8. The method according to claim 7, further comprising: Determining consistency deviations between a plurality of target feature values of the plurality of feature items and the plurality of feature values; And adjusting the order of the plurality of feature items includes: Determining a total deviation of the plurality of feature items based on the consistency deviation and the deviation of the order of the plurality of feature items; as well as The order of the plurality of feature items and the plurality of target feature values are adjusted so that a total deviation of the plurality of feature items satisfies a condition.
9. The method according to claim 1, wherein acquiring the medical data set comprises: Obtain case files from the medical document repository; Filtering a document collection related to the disease from the case files by using keywords related to the disease; Dividing the document collection into a plurality of document subsets according to the types of paragraphs in the document collection; as well as The plurality of feature values corresponding to the plurality of feature items are determined from the plurality of document subsets to form the medical data set including a plurality of medical data subsets, wherein the plurality of feature items are determined based on a medical term library related to the disease.
10. The method of claim 9, wherein determining the plurality of feature values comprises: In response to a feature item among the multiple feature items not having a corresponding feature value in the multiple document subsets, a neural network model generates a feature value corresponding to the feature item based on a context related to the feature item.
11. An apparatus for enhancing a medical dataset, comprising: A data set acquisition module is configured to acquire the medical data set, wherein the medical data set includes a plurality of feature items and corresponding plurality of feature values sorted by importance; A feature value generating module is configured to generate a plurality of feature values after the plurality of feature items are enhanced based on the plurality of feature values according to the order of the plurality of feature items; as well as The data set determination module is configured to determine an enhanced medical data set based on the multiple feature values corresponding to the multiple feature items and the enhanced multiple feature values.
12. An electronic device comprising: at least one processor; as well as A memory coupled to the at least one processor and having instructions stored thereon, the instructions, when executed by the at least one processor, causing the apparatus to perform the method according to any one of claims 1 to 10.
13. A computer program product comprising machine executable instructions which, when executed, cause the method according to any one of claims 1 to 10 to be implemented.