Medical image classification method based on knowledge graph

By combining medical imaging data and knowledge graphs, using image feature vectors, medication data and living habit information, the impact of drugs on diseases is simulated, and the problem of insufficient factors in traditional medical imaging classification methods is solved, and more accurate disease classification and diagnostic support is achieved.

CN120452746AInactive Publication Date: 2025-08-08SHANGHAI EAST HOSPITAL EAST HOSPITAL TONGJI UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510546110.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional medical imaging classification methods rely on a single data source and do not fully consider factors such as patients' medication, genetic background and living habits, resulting in inaccurate and comprehensive classification results.

Method used

By combining medical imaging data with knowledge graphs, using image feature vectors to map them to the knowledge graphs, combining drug use data, genetic background and lifestyle information, the impact of drugs on diseases is simulated, the disease collection is screened and verified, and the final disease classification results are output.

Benefits of technology

It improves the accuracy of disease classification and personalized evaluation ability, provides doctors with a more comprehensive diagnostic basis, dynamically adjusts candidate disease collections, and improves diagnostic efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452746A_ABST
    Figure CN120452746A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image classification method based on a knowledge graph, and relates to the technical field of medical image processing, and the method comprises the steps: mapping a medical image feature vector into a medical knowledge graph, and screening a candidate disease set; mapping the medical knowledge graph by utilizing a synergistic and antagonistic relationship between drugs, simulating the influence of drug use of a patient on diseases, and updating a candidate disease set; determining the disease induction probability of the patient based on the genetic background and the living habits, screening the updated candidate disease set, and marking a first disease set; according to the micro tissue features of the patient and the associated disease data, the diseases in the first disease set are verified, a second disease set is output, and the second disease set is a disease classification result of the patient. Through multi-source data fusion and knowledge graph application, the accuracy and personalized evaluation ability of disease classification are significantly improved, and a more comprehensive and customized disease diagnosis basis is provided for doctors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing technology, and specifically to a medical image classification method based on knowledge graphs. Background Art

[0002] With the continuous development of medical imaging technology, medical imaging data has exploded. How to effectively classify and interpret these massive imaging data to assist doctors in disease diagnosis has become an important issue facing the medical field. At the same time, knowledge graphs, as a powerful knowledge representation and reasoning tool, have gradually attracted attention in the medical field.

[0003] Traditional medical image classification methods rely solely on the image data itself, ignoring other important information such as the patient's medication status, genetic background, and lifestyle habits. They do not fully consider the impact of the patient's medication on the disease and the dynamic development of the disease in the individual, resulting in inaccurate and incomplete classification results. In addition, some existing methods do not effectively utilize rich medical knowledge to assist in image classification.

[0004] Therefore, to address the above problems, a medical image classification method based on knowledge graph is urgently needed. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention provides a medical image classification method based on knowledge graphs, which solves the problem that traditional medical image classification methods rely on a single data source and lack consideration of comprehensive factors such as patient medication, genetic background and lifestyle habits, resulting in insufficient classification accuracy.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a medical image classification method based on a knowledge graph, comprising the following steps: step S1, obtaining the patient's medical image data, extracting the medical image feature vector, mapping the medical image feature vector to the medical knowledge graph, and screening the candidate disease set based on the similarity of the corresponding nodes of the medical knowledge graph; step S2, obtaining the patient's medication data, identifying the drug efficacy coefficient, side effect coefficient and interaction coefficient based on the medication data, and then mapping the medical knowledge graph using the synergistic and antagonistic relationships between the drugs, simulating the impact of the patient's medication on the disease, and updating the candidate disease set; step S3, identifying the patient's genetic background and lifestyle habits, and then determining the probability of the patient having the disease based on the genetic background and lifestyle habits, screening the updated candidate disease set based on the probability of the patient having the disease, and marking the first disease set; step S4, using the medical knowledge graph to identify the micro-tissue effects and related diseases of the disease in the first disease set, and then verifying the diseases in the first disease set based on the patient's micro-tissue characteristics and related disease data, and outputting a second disease set, which is the patient's disease classification result.

[0007] Furthermore, step S1 is specifically analyzed as follows: using medical imaging equipment to obtain the patient's medical imaging data, and then inputting the medical imaging data into a convolutional neural network to output a medical imaging feature vector; associating the medical imaging feature vector with the node attributes in the medical knowledge graph, and then identifying the similarity between the medical imaging feature vector and each disease node in the medical knowledge graph, wherein the node attributes include a typical imaging feature description related to the disease; comparing the similarity between the medical imaging feature vector and each disease node in the medical knowledge graph with a similarity threshold, and marking the diseases represented by the disease nodes whose similarity is higher than the similarity threshold to constitute a candidate disease set.

[0008] Furthermore, step S2 is specifically analyzed as follows: medication data includes the patient's medication records and the corresponding treatment effect and side effect occurrence data, and the medication data is used to obtain the drug's efficacy coefficient, side effect coefficient and interaction coefficient. If the interaction coefficient between the drugs is positive, the drugs have a synergistic effect, and if the interaction coefficient is negative, there is an antagonistic effect; the corresponding drug nodes and disease nodes in the medical knowledge graph are identified, and for drug nodes with synergistic or antagonistic relationships, corresponding edges are established in the medical knowledge graph to represent them. At the same time, the efficacy coefficient and side effect coefficient are added as attributes of the drug nodes to the medical knowledge graph, and then based on the relationship between drugs and diseases and the interaction relationship between drugs in the knowledge graph, combined with the patient's medication records, the impact path of the drug on the disease is simulated. If it is identified that the effect of a certain drug combination on a certain disease is consistent with the patient's actual symptoms, the disease is added to the candidate disease set. If the side effect of a certain drug is related to the abnormal condition of the patient, and the disease corresponding to the side effect is not in the candidate disease set, the disease is added to the candidate disease set, and at the same time, the diseases in the candidate disease set whose patient's medication situation is inconsistent with the actual performance are screened out.

[0009] Furthermore, the specific analysis of the drug efficacy coefficient, side effect coefficient and interaction coefficient obtained by using medication data is as follows: for the efficacy coefficient, a drug efficacy evaluation model is constructed with the therapeutic effect as the dependent variable and the medication data as the independent variable, and the drug efficacy coefficient is obtained by using the drug efficacy evaluation model; for the side effect coefficient, a side effect evaluation model is constructed with the occurrence of side effects as the dependent variable and the medication data as the independent variable, and the side effect coefficient is obtained by using the side effect evaluation model; for the interaction coefficient, the combination of drugs used simultaneously by the patient is used as a new feature, and an interaction evaluation model is constructed with the therapeutic effect or the occurrence of side effects as the dependent variable, and the drug interaction coefficient is obtained by using the interaction evaluation model.

[0010] Furthermore, step S3 is specifically analyzed as follows: obtaining the patient's genetic adaptability coefficient and comprehensive score of lifestyle habits, combining the patient's genetic adaptability coefficient and comprehensive score of lifestyle habits to determine the patient's susceptibility values for various diseases, comparing the susceptibility values for various diseases with the disease risk threshold, marking the diseases with susceptibility values higher than the disease risk threshold as the susceptibility disease set, and then summarizing the susceptibility disease set and the updated candidate disease set to obtain the first disease set.

[0011] Furthermore, the patient's genetic adaptability coefficient and comprehensive lifestyle score are specifically obtained in the following manner: obtaining disease data of the patient's relatives, including the type of disease suffered by the relatives, the age of disease onset, and the relationship with the patient, and determining the patient's genetic adaptability coefficient based on the disease data of the patient's relatives; obtaining the patient's lifestyle information, including dietary structure, amount of exercise, smoking and drinking habits, and work and rest patterns, and determining the patient's comprehensive lifestyle score based on the lifestyle information.

[0012] Furthermore, step S4 is specifically analyzed as follows: the microtissue impact information includes the tissue type affected by the disease and the pathological changes of the tissue, and the related diseases include complications and other diseases caused by related causes; the patient's microtissue feature data and related disease data are obtained, and the microtissue impact of each disease in the first disease set is compared and analyzed with the patient's microtissue feature to determine whether the patient's microtissue feature matches the microtissue impact of the disease, and at the same time, based on the patient's related disease data, determine whether the patient has an associated disease related to the disease; if the patient's microtissue feature matches the microtissue impact of the disease and there is a corresponding associated disease, the disease passes the verification; the diseases that pass the verification in the first set are screened out, and merged to obtain the second disease set.

[0013] The present invention has the following beneficial effects:

[0014] This medical image classification method based on knowledge graph can more comprehensively evaluate the patient's disease status by combining multi-source information such as medical imaging data, patient medication data, genetic background and lifestyle habits, thereby improving the accuracy of disease classification; by simulating the impact of patient medication on the disease, it can predict the therapeutic effect or potential side effects of the drug on the disease, and provide doctors with more comprehensive treatment recommendations; based on the patient's genetic background and lifestyle habits, the probability of disease induction is determined, which can more accurately predict the disease the patient may have, providing a basis for early intervention and treatment; using the medical knowledge graph to identify the subtle tissue effects and related diseases of the disease, it can more quickly verify the diseases in the first disease set and improve the efficiency of disease verification; the final output of the second disease set as the patient's disease classification result can provide doctors with more accurate and comprehensive disease information, which is helpful for formulating more reasonable treatment plans; a large amount of medical data is accumulated in the classification process, which provides rich data resources for medical research, contributes to medical education and training, and improves the professional quality and clinical ability of medical students. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flow chart of the medical image classification method based on knowledge graph in the present invention. DETAILED DESCRIPTION

[0016] The embodiment of the present application uses a medical image classification method based on a knowledge graph to achieve multi-source data fusion and knowledge graph application, significantly improving the accuracy of disease classification and personalized assessment capabilities, and providing doctors with a more comprehensive and customized basis for disease diagnosis.

[0017] The overall idea of the embodiment of the present application is to use the knowledge graph to integrate and associate multi-source heterogeneous medical data, and to achieve accurate classification of medical images through step-by-step screening and verification.

[0018] See also Figure 1An embodiment of the present invention provides a technical solution: a medical image classification method based on a knowledge graph, comprising the following steps: Step S1, obtaining the patient's medical image data, extracting medical image feature vectors, mapping the medical image feature vectors to the medical knowledge graph, and screening a candidate disease set based on the similarity of corresponding nodes in the medical knowledge graph; Step S2, obtaining the patient's medication data, identifying the drug efficacy coefficient, side effect coefficient, and interaction coefficient based on the medication data, and then mapping the medical knowledge graph using the synergistic and antagonistic relationships between drugs, simulating the impact of the patient's medication on the disease, and updating the candidate disease set; Step S3, identifying the patient's genetic background and lifestyle habits, and then determining the probability of the patient having the disease based on the genetic background and lifestyle habits, screening the updated candidate disease set based on the probability of the patient having the disease, and marking a first disease set; Step S4, using the medical knowledge graph to identify the micro-tissue effects and associated diseases of the diseases in the first disease set, and then verifying the diseases in the first disease set based on the patient's micro-tissue characteristics and associated disease data, and outputting a second disease set, which is the patient's disease classification result.

[0019] Specifically, the medical knowledge graph is a structured knowledge representation used to integrate and organize various knowledge and information in the medical field. It represents knowledge in the form of triples, namely "entity, relationship, entity" or "entity, attribute, value". For example, "heart disease, symptoms, chest pain" means that there is a symptom association between heart disease and chest pain; "drug A, treatment, disease B" means that drug A has a therapeutic effect on disease B; "patient, age, 35 years old" means that the patient's age attribute is 35 years old. Entities include various medical-related concepts such as diseases, symptoms, drugs, genes, and anatomical structures. Relationships describe the semantic connections between these entities. Semantic connections include treatment, triggering, association, belonging, etc. Attributes are used to describe the characteristics and properties of entities.

[0020] Medical knowledge graphs provide a comprehensive medical knowledge base for medical image classification, covering a wide range of information such as disease symptoms, diagnostic criteria, treatment methods, drug effects, genetic factors, etc., which helps to explore the potential relationship between image features and diseases; they can integrate and associate medical data from different sources and types, including medical imaging data, medication data, genetic data, lifestyle data, etc., so that these data can be integrated and analyzed under a unified knowledge framework; through the relationship network of the knowledge graph, reasoning and inference can be performed, for example, based on the patient's imaging characteristics, medication status, genetic background and other information, possible diseases can be inferred, providing auxiliary decision support for doctors' diagnosis and improving the accuracy and efficiency of diagnosis; it helps to discover potential knowledge and laws in the medical field, such as rare associations between diseases, new indications for drugs, etc., providing new ideas and directions for medical research and clinical practice.

[0021] The steps for constructing a medical knowledge graph are as follows: collect various sources of medical knowledge, such as medical literature, clinical guidelines, medical records, medical databases, professional books, etc. These sources contain rich medical knowledge, but the data format and quality may vary and need further processing and integration; extract relevant entity, relationship and attribute information from the knowledge sources through natural language processing technology; integrate the extracted knowledge to resolve conflicts and inconsistencies between knowledge from different sources. For example, the same disease may have different names or descriptions in different documents, which need to be unified and standardized. At the same time, map the knowledge into a unified knowledge graph framework and establish associations between entities; select a suitable knowledge graph model and storage method, and represent and store the integrated knowledge in a graph structure. Commonly used knowledge graph models include RDF (Resource Description Framework) and OWL (Web Ontology Language), and storage methods include graph databases such as Neo4j to store and manage knowledge graphs for efficient query and reasoning operations; regularly update and maintain the knowledge graph, including adding new knowledge, correcting error information, and updating relationships and attributes.

[0022] Specifically, step S1 is specifically analyzed as follows: using medical imaging equipment to obtain the patient's medical imaging data, and then inputting the medical imaging data into a convolutional neural network to output a medical imaging feature vector; associating the medical imaging feature vector with the node attributes in the medical knowledge graph, and then identifying the similarity between the medical imaging feature vector and each disease node in the medical knowledge graph, where the node attributes include a typical imaging feature description related to the disease; comparing the similarity between the medical imaging feature vector and each disease node in the medical knowledge graph with a similarity threshold, and marking the diseases represented by the disease nodes whose similarity is higher than the similarity threshold to constitute a candidate disease set.

[0023] In this embodiment, medical imaging data includes X-rays, CT (computed tomography) images, MRI (magnetic resonance imaging) images, ultrasound images, PET (positron emission tomography) images, etc. Different imaging technologies are suitable for different body parts and disease diagnosis. For example, CT is often used to examine structures such as bones and lungs, MRI has better imaging effects on soft tissues, ultrasound is often used to examine the abdomen, cardiovascular system, etc., and PET is mainly used for tumor detection and staging. They are obtained through various medical imaging devices. For example, X-rays penetrate the human body and make photosensitive film sensitive to light to form images; CT uses X-rays to perform tomographic scans of the human body and then generates three-dimensional images through computer reconstruction technology; MRI uses the magnetic resonance phenomenon of hydrogen protons in the human body excited by radio frequency pulses in a magnetic field to generate signals, which are then reconstructed into images through computer processing; ultrasound uses the principles of reflection and refraction of ultrasound in human tissue to form real-time two-dimensional or three-dimensional images; PET generates images by injecting radioactive nuclide-labeled drugs and then detecting their distribution in the body.

[0024] The specific steps of inputting medical image data into the convolutional neural network and outputting the medical image feature vector are as follows: preprocessing the acquired medical image data, including adjusting the image size, grayscale normalization, data enhancement and other operations; inputting the preprocessed medical image data into the convolution layer of the convolutional neural network. The convolution layer performs convolution operations by sliding the convolution kernel on the image to extract local features of the image, such as edges, textures, etc. Each convolution kernel can learn different feature patterns. Through the combination of multiple convolution kernels, rich image features can be extracted. The convolution operation will generate a series of feature maps, which retain the spatial information and feature information of the image; after passing the convolution layer, the size of the feature map is usually large. In order to reduce the amount of data and calculation while retaining important feature information, the pooling layer will be used to downsample the feature map. Common pooling methods include maximum pooling and average pooling. Maximum pooling takes the maximum value in the local area as the pooling result. , average pooling takes the average value in the local area. The pooling operation can effectively compress the size of the feature map, reduce the complexity of the model, and prevent overfitting to a certain extent; after multiple convolution and pooling operations, the obtained feature map is flattened into a one-dimensional vector and then input into the fully connected layer. The neurons in the fully connected layer are connected to all the neurons in the previous layer. The extracted features can be integrated and nonlinearly transformed, and the features can be mapped to a higher-dimensional space to further explore the relationship between the features. The fully connected layer contains multiple neurons, and nonlinearity can be introduced through activation functions (such as ReLU, Sigmoid, etc.) to enhance the expressive power of the model; after processing by the fully connected layer, a low-dimensional vector is finally output, namely the medical image feature vector. The feature vector contains the most representative feature information in the medical image, which can reflect the disease-related features contained in the image, and is used for subsequent similarity matching with disease nodes in the medical knowledge graph.

[0025] The specific steps for identifying the similarity between medical image feature vectors and disease nodes in the medical knowledge graph are as follows: convert the attributes of each disease node in the medical knowledge graph (such as typical image feature descriptions, etc.) into vector representations, perform normalization and standardization on the features to ensure that the medical image feature vectors and disease node feature vectors are in the same feature space, and then select a suitable similarity measurement method to calculate the similarity between the medical image feature vectors and the disease node feature vectors. Similarity measurement methods include cosine similarity and Euclidean distance.

[0026] The similarity threshold is set as follows: based on professional knowledge in the medical field and previous clinical experience, medical experts or researchers directly set an initial similarity threshold. For example, in the diagnosis of certain diseases, based on the analysis and research of a large number of cases, it is found that when the similarity reaches above 0.7, the correlation between the image features and the disease has a high degree of credibility, so the threshold is set to 0.7; the training data set can also be used for cross-validation to determine the optimal similarity threshold. The data set is divided into multiple subsets, and training and verification are performed on different subsets. By adjusting the similarity threshold, the performance indicators of the model on the verification set (such as accuracy, recall rate, F1 value, etc.) are observed, and the threshold that makes the performance indicators reach the optimal level is selected as the final similarity threshold.

[0027] Convolutional neural networks can automatically extract highly abstract feature vectors from medical imaging data. These feature vectors can effectively capture key information in medical images, reducing the subjectivity and inaccuracy of manual feature extraction. By associating the extracted image feature vectors with disease nodes in the medical knowledge graph and utilizing the rich medical knowledge in the knowledge graph, the relationship between disease and imaging features can be considered more comprehensively, improving the accuracy and reliability of disease diagnosis. By setting a similarity threshold to screen the candidate disease set, diseases that are irrelevant to the current medical imaging features can be quickly excluded, narrowing the scope of diagnosis, providing more targeted information for subsequent diagnostic work, and improving diagnostic efficiency.

[0028] Specifically, step S2 is specifically analyzed as follows: medication data includes the patient's medication records and the corresponding treatment effect and side effect occurrence data, and the medication data is used to obtain the drug's efficacy coefficient, side effect coefficient and interaction coefficient. If the interaction coefficient between the drugs is positive, the drugs have a synergistic effect, and if the interaction coefficient is negative, there is an antagonistic effect; the corresponding drug nodes and disease nodes in the medical knowledge graph are identified, and for drug nodes with synergistic or antagonistic relationships, corresponding edges are established in the medical knowledge graph to represent them. At the same time, the efficacy coefficient and side effect coefficient are added as attributes of the drug nodes to the medical knowledge graph, and then based on the relationship between drugs and diseases and the interaction relationship between drugs in the knowledge graph, combined with the patient's medication records, the impact path of the drug on the disease is simulated. If it is identified that the effect of a certain drug combination on a certain disease is consistent with the patient's actual symptoms, the disease is added to the candidate disease set. If the side effect of a certain drug is related to the abnormal condition of the patient, and the disease corresponding to the side effect is not in the candidate disease set, the disease is added to the candidate disease set, and at the same time, the diseases in the candidate disease set whose patient's medication situation is inconsistent with the actual performance are screened out.

[0029] The specific analysis of the drug efficacy coefficient, side effect coefficient and interaction coefficient obtained by using medication data is as follows: for the efficacy coefficient, the therapeutic effect is taken as the dependent variable and the medication data as the independent variable, a drug efficacy evaluation model is constructed, and the drug efficacy coefficient is obtained by using the drug efficacy evaluation model; for the side effect coefficient, the side effect occurrence is taken as the dependent variable and the medication data as the independent variable to construct a side effect evaluation model, and the side effect coefficient is obtained by using the side effect evaluation model; for the interaction coefficient, the combination of drugs used simultaneously by the patient is used as a new feature, and the therapeutic effect or side effect occurrence is taken as the dependent variable to construct an interaction evaluation model, and the drug interaction coefficient is obtained by using the interaction evaluation model.

[0030] In this implementation scheme, the medication record specifically records the name, dosage, medication time, medication frequency and other information of the drug used by the patient. It is obtained from the hospital's electronic medical record system, pharmacy dispensing records or patient's medication diary based on user authorization. The drug name is encoded, and the dosage, medication time and frequency are represented by specific numerical values. For example, drug A is coded as 1, the dosage is in grams or milligrams, the medication time can be represented by the number of days from a specific time point, and the medication frequency can be represented by the number of times the drug is taken per day.

[0031] The therapeutic effect reflects the improvement of the patient's condition after using the drug, such as symptom relief, improvement of indicators, etc. It is obtained through regular doctor's examinations of patients, test reports and patient self-reports, and is quantified according to the assessment indicators of specific diseases. For example, for patients with hypertension, the decrease in blood pressure values can be used as a quantitative indicator of the treatment effect. For cancer patients, changes in tumor size and reduction in the number of cancer cells can be used as quantitative basis. A scoring system can also be used, such as 0-10 points, where 0 means no improvement and 10 means complete cure.

[0032] The occurrence of side effects records the adverse reactions that patients experience during medication, such as nausea, vomiting, rash, etc. Through patient self-reports, doctor's observations and related examinations and tests, different side effects are classified and coded, and then the frequency or severity of each side effect is recorded. For example, mild side effects are recorded as 1, moderate side effects are recorded as 2, and severe side effects are recorded as 3.

[0033] The specific steps for simulating the path of drug impact on disease are as follows: find the corresponding drug node in the medical knowledge graph based on the patient's medication record; determine the impact of the drug on its directly related disease nodes based on the drug efficacy coefficient, side effect coefficient and interaction relationship between drugs; gradually propagate the impact of the drug on the disease along the association relationship between disease nodes in the knowledge graph to simulate the impact path; comprehensively consider the various effects of the drug and the patient's actual symptoms and manifestations to determine whether the disease represented by the disease node is consistent with the actual situation.

[0034] The consistency between the effect of a certain drug combination on a certain disease and the patient's actual symptoms means that the predicted disease symptoms and development trends based on the effects of the drugs on the disease simulated in the knowledge graph are consistent with the actual symptoms and changes in examination indicators exhibited by the patient. For example, a certain drug combination is predicted in the knowledge graph to relieve cough and lower body temperature. After the patient uses this drug combination, the cough symptoms are alleviated and the body temperature returns to normal, which indicates that the effect of the drug combination on the disease is consistent with the patient's actual symptoms. Specifically, this is determined by comparing the symptoms, treatment responses and other information described by the disease nodes in the knowledge graph with the patient's actual medical records, examination reports and other content, using similarity calculation methods, such as calculating the text similarity of symptom descriptions, or comparing whether the values of key symptom indicators are within a reasonable range.

[0035] The side effects of a certain drug are related to the abnormal conditions experienced by the patient when certain abnormal symptoms or signs experienced by the patient match the known side effects of the drug. For example, if a patient develops a rash after taking a certain drug, and the adverse reaction of rash is clearly recorded in the side effects of the drug, it can be considered that the side effects of the drug are related to the abnormal conditions experienced by the patient. The abnormal conditions experienced by the patient are compared with the detailed list of drug side effects, and factors such as the relationship between the time when the side effects occur and the time of medication are considered, such as calculating the probability of the side effects and the abnormal conditions occurring at the same time, to determine the correlation between the two.

[0036] The steps for building a drug efficacy evaluation model are as follows: collect a large amount of patient medication data (including drug type, dosage, medication time, etc.) as independent variables, quantify the treatment effect (such as symptom improvement, indicator changes, etc.) as the dependent variable, and preprocess the data, including data cleaning and missing value processing; divide the data into training and test sets, use the training set data to fit the logistic regression model, adjust the model parameters to make the model fit the data best, and finally verify the accuracy and generalization ability of the model on the test set to obtain the drug efficacy evaluation model. The specific drug efficacy evaluation model expression example is as follows: Where X=(x1,x2,...,x n ) are independent variables, representing different medication-related characteristics, n represents the total number of medication data characteristics, Y is the dependent variable, representing the treatment effect (Y = 1 means effective, Y = 0 means ineffective), β0, β1, β2, ..., β n It represents the efficacy coefficient, which is obtained by methods such as maximum likelihood estimation, and indicates the degree of influence of each medication feature on the treatment effect.

[0037] The steps for building a side effect evaluation model are as follows: Similar to the efficacy evaluation model, the side effect occurrence is used as the dependent variable and the medication data is used as the independent variable. After collecting the data, preprocess it, divide it into training and test sets, fit the logistic regression model with the training set, and evaluate it on the test set to obtain the side effect evaluation model. The specific side effect evaluation model expression example is:

[0038] Where Z represents the occurrence of side effects (Z = 1 means side effects occur, Z = 0 means no side effects occur), γ0,γ1,γ2,...,γ n It represents the side effect coefficient, which is obtained by methods such as maximum likelihood estimation and reflects the relationship between medication characteristics and the occurrence of side effects.

[0039] The steps for constructing the interaction evaluation model are as follows: add the drug combination used by the patient as a new feature to the independent variable, and use the treatment effect or side effect occurrence as the dependent variable. Similarly, data collection, preprocessing, and division into training and test sets are performed. The logistic regression model is trained using the training set, and the model performance is verified on the test set to obtain the interaction evaluation model. The specific expression example of the interaction evaluation model is as follows (with the treatment effect as the dependent variable):

[0040] Where X=(x1,x2,...,x m ,x m+1 ,...,x m+k ) are independent variables, where x1, x2, ..., x m is the characteristic of conventional medication, x m+1 ,...,x m+k is the drug combination feature, Y is the dependent variable, δ0, δ1, ..., δ m+k is the model parameter, and the interaction coefficient is: the parameter δ corresponding to the drug combination characteristics m+i , i=1,2,...,k, used to measure the impact of drug interactions on therapeutic effects.

[0041] By considering the efficacy, side effects, and drug interactions of drugs, the impact of drugs on diseases is comprehensively evaluated, making disease diagnosis more accurate. Based on the patient's actual medication use and performance, the set of candidate diseases is dynamically adjusted, which improves the flexibility and accuracy of diagnosis, helps to timely detect potential diseases and eliminate unreasonable disease hypotheses.

[0042] Specifically, step S3 is specifically analyzed as follows: obtaining the patient's genetic adaptability coefficient and comprehensive score of lifestyle habits, combining the patient's genetic adaptability coefficient and comprehensive score of lifestyle habits to determine the patient's susceptibility values for various diseases, comparing the susceptibility values for various diseases with the disease risk threshold, marking the diseases with susceptibility values higher than the disease risk threshold as the susceptibility disease set, and then summarizing the susceptibility disease set and the updated candidate disease set to obtain the first disease set.

[0043] The specific method of obtaining the patient's genetic adaptability coefficient and comprehensive lifestyle score is as follows: obtain the disease data of the patient's relatives, which includes the type of disease the relatives suffer from, the age of disease onset, and the relationship with the patient, and determine the patient's genetic adaptability coefficient based on the disease data of the patient's relatives; obtain the patient's lifestyle information, which includes diet structure, exercise volume, smoking and drinking habits, and work and rest patterns, and determine the patient's comprehensive lifestyle score based on the lifestyle information.

[0044] In this implementation plan, the method for quantifying disease types is: coding different diseases, such as coding common hypertension as 001, coding diabetes as 002, etc., so as to distinguish the impact of different diseases on genetic adaptability in subsequent calculations; the method for quantifying the age of disease onset is: standardizing the actual age, the younger the age, the greater the influence of genetic factors, and giving a higher weight when calculating the genetic adaptability coefficient; the method for quantifying kinship is: using numerical values to represent the closeness of kinship, such as parents as 1, grandparents / grandparents as 0.5, brothers and sisters as 1, cousins / siblings as 0.25, etc.

[0045] The way to quantify the dietary structure is to establish a dietary scoring system, for example, divide food into healthy food (such as vegetables, fruits, whole grains), general food (such as lean meat, fish) and unhealthy food (such as fried food, high-sugar beverages), and count the proportion of each type of food in the daily diet. The higher the proportion of healthy food, the higher the score. Set a scoring standard with a total score of 10 points, and score according to the proportion. For example, if the proportion of healthy food exceeds 60%, 8-10 points will be given, 40%-60% will be given 5-7 points, and less than 40% will be given 0-4 points; the way to quantify the amount of exercise is: quantify according to the number of hours of exercise per week and the intensity of exercise; the way to quantify smoking and drinking conditions is: smoking conditions are quantified according to the number of cigarettes smoked per day and the number of years of smoking Quantification: similar to drinking, the score is calculated based on the daily alcohol consumption and the number of years of drinking, and finally the smoking and drinking scores are combined, such as adding the two scores and then mapping them to a 0-10 score; the quantification method for work and rest regularity is: by investigating the patient's bedtime, wake-up time, and number of nighttime awakenings and other information, for example, falling asleep between 22:00-23:00 gets 3 points, between 23:00-0:00 gets 2 points, and after 0:00 gets 1 point, getting up between 6:00-7:00 gets 3 points, between 7:00-8:00 gets 2 points, and after 8:00 gets 1 point. The scores are added to get the total score of work and rest regularity, ranging from 3 to 6 points, which is then mapped to a 0-10 score range.

[0046] An example of the steps for determining the genetic compatibility coefficient is as follows: disease data for all relatives of the patient, including disease type, age of disease onset, and kinship information, are collected. For each disease suffered by a relative, a weighted calculation is performed based on the heritability of the disease (heritability data for common diseases can be obtained from medical research literature), the degree of kinship, and the age of disease onset. The calculated results for all relatives' corresponding diseases are accumulated to obtain the patient's genetic compatibility coefficient for all diseases. An example expression for the genetic compatibility coefficient is: Where G j represents the genetic fitness coefficient of the patient for the jth disease, D represents the total number of relatives of the patient, the bth relative suffers from the jth disease, h j The heritability of a disease refers to the extent to which genetic factors play a role in the occurrence of a disease, reflecting the degree of contribution of genetic factors to the occurrence of the disease. It is estimated through large-scale epidemiological studies, twin studies, family studies, etc. b is the kinship coefficient, which is used to measure the degree of blood relationship between the patient and relatives. bj is the age of the bth relative when he or she develops the jth disease.

[0047] An example of the steps for determining a comprehensive lifestyle score is: obtain quantitative scores for the patient's diet, exercise, smoking and drinking habits, and daily routine, assign weights to each factor based on its importance to health, multiply each factor score by the corresponding weight, and then add them up to obtain a comprehensive lifestyle score. An example expression for a comprehensive lifestyle score is:

[0048] L=ω1*S1+ω2*S2+ω3*S3+ω4*S4, where L is the comprehensive score of lifestyle habits, S1, S2, S3, and S4 are the quantitative scores of the patient's diet structure, exercise volume, smoking and drinking habits, and work and rest habits, respectively, and ω1, ω2, ω3, and ω4 represent the weight values of each factor, respectively.

[0049] The steps for obtaining the susceptibility values of various diseases are as follows: for each type of disease, obtain its corresponding genetic adaptability coefficient and comprehensive score of lifestyle habits. According to the relative importance of the influence of genetic factors and lifestyle factors on the susceptibility of this type of disease, determine the weight of genetic factors and the weight of lifestyle factors. The susceptibility value of this type of disease is obtained through weighted calculation. The expression of the susceptibility value is as follows: E j =η j1 *G j +η j2 *L, where E j represents the induced value of the jth disease, η j1 ,η j2 They respectively represent the weight values of the genetic adaptability coefficient of the jth disease and the comprehensive score of lifestyle habits.

[0050] Disease risk thresholds can be set by compiling a large amount of population-based disease research data to analyze the incidence of various diseases in the general population and the distribution of risk factors influenced by genetic and lifestyle factors. Based on this data, a threshold that can distinguish between high-risk and low-risk groups can be determined. For example, if research finds that the probability of developing a common disease increases significantly when the risk threshold exceeds 0.6, the disease risk threshold can be set at 0.6. Medical experts can also be convened to discuss and determine risk thresholds for various diseases based on their clinical experience and understanding of disease pathogenesis. These experts will then comprehensively assess and determine a reasonable threshold range, taking into account factors such as the characteristics of different regions and populations. Alternatively, a large amount of existing patient data, including genetic data, lifestyle data, and disease diagnosis results, can be used to construct disease prediction models using machine learning algorithms (such as logistic regression and decision trees). Through model training and validation, an optimal disease risk threshold can be found, ensuring that the model achieves the best performance metrics (such as precision, recall, and F1 value) in distinguishing between patients with and without the disease.

[0051] By combining the disease data of the patient's relatives and their own lifestyle information, and comprehensively considering the impact of genetic factors and acquired lifestyle on disease induction, compared with single factor assessment, it can more accurately identify the patient's potential disease risk and improve the accuracy of disease prediction; by determining the set of inducing diseases and summarizing it with the candidate disease set to obtain the first disease set, we can focus on high-risk diseases, provide a more targeted scope for subsequent precise diagnosis and intervention, save medical resources, and improve the efficiency of medical services.

[0052] Specifically, step S4 is specifically analyzed as follows: the microtissue impact information includes the tissue type affected by the disease and the pathological changes of the tissue, and the related diseases include complications and other diseases caused by related causes; the patient's microtissue feature data and related disease data are obtained, and the microtissue impact of each disease in the first disease set is compared and analyzed with the patient's microtissue feature to determine whether the patient's microtissue feature matches the microtissue impact of the disease. At the same time, based on the patient's related disease data, it is determined whether the patient has an associated disease related to the disease. If the patient's microtissue feature matches the microtissue impact of the disease and there is a corresponding associated disease, the disease passes the verification; the diseases that pass the verification in the first set are screened out and merged to obtain the second disease set.

[0053] In this implementation, microtissue characteristic data is obtained through biopsies and pathological section examinations of patients. Microtissue characteristic information, such as the morphology, structure, and pathological characteristics of tissue cells, is then extracted through microscopic observation, image analysis, and other technical means. Related disease data is obtained from medical records, past diagnostic reports, examination results, and other medical information. Alternatively, data can be obtained through interviews with patients and their families to determine whether the patient has other symptoms, medical histories, or diagnosed diseases related to the diseases in the first disease set.

[0054] The steps for determining whether the microtissue feature matches are as follows: preprocessing the patient's microtissue feature data, including data cleaning, normalization and other operations, to ensure data consistency and comparability, and extracting key features of the disease's microtissue effect information, such as specific identifiers of tissue types, characteristic parameters of pathological changes, etc.; using similarity calculation algorithms, such as cosine similarity and Euclidean distance, to compare and calculate the patient's microtissue features with the key features of the disease's microtissue effect to obtain a similarity score, and setting a matching threshold. If the similarity score is higher than the threshold, it is considered that the patient's microtissue feature matches the disease's microtissue effect.

[0055] The steps for determining the existence of associated diseases are: using the medical knowledge graph to identify information on various diseases and their associated diseases, extracting diagnosed disease information and current symptom information from the patient's medical data, matching the diseases in the first disease set with the associated disease knowledge base, searching for the associated disease list corresponding to each disease, and checking the patient's disease information and symptom information one by one to see if there are any diseases in the associated disease list. If so, it is determined that the patient has an associated disease related to the disease.

[0056] An example of this implementation scheme is as follows: assuming that there are diseases A and B in the first disease set, the microtissue impact information of disease A shows that it will cause nodular lesions in the lung tissue, and the related diseases are chronic obstructive pulmonary disease (COPD); disease B will cause fibrosis in the liver tissue, and the related diseases are cirrhosis; for patient A, a tissue biopsy was found to have nodular lesions in his lung tissue, and the patient has a long history of smoking and has been diagnosed with COPD. After comparing microtissue features, the lesion features of patient A's lung tissue match the microtissue impact of disease A, and there is also COPD, an associated disease of disease A, so disease A passes verification; for patient B, his liver tissue biopsy shows signs of fibrosis, but there are no symptoms and diagnostic records related to cirrhosis. When judging the microtissue features, it matches the microtissue impact of disease B, but since there is no associated disease of disease B, cirrhosis, disease B cannot pass verification. Finally, patient A's second disease set contains disease A, and patient B's second disease set does not contain disease B.

[0057] By comparing microtissue characteristics and related disease data to verify the disease, the patient's disease can be determined more comprehensively and accurately, reducing the probability of misdiagnosis and missed diagnosis; judging the disease from multiple dimensions, comprehensively considering the disease's microtissue impact and related diseases, makes the final disease classification results more reliable and scientific.

[0058] It should be noted that any relevant data in the embodiments of this application are obtained based on the authorization of the patient and the patient's relatives.

[0059] In summary, this application has at least the following effects:

[0060] By integrating multi-dimensional data, such as medical imaging feature vectors, medication data, genetic background, and lifestyle habits, and comprehensively considering the impact of multiple factors on disease classification, we can more comprehensively grasp the patient's condition, thereby improving the accuracy of disease classification. By using the impact of drugs on diseases to update the candidate disease set and combining the patient's individual characteristics to screen the disease set, we can dynamically reflect the development and changes of the disease in the patient, making the classification results more consistent with the actual condition. The medical knowledge graph provides a rich medical knowledge system, which helps to explore the potential relationship between diseases and the association between diseases and various feature data, providing comprehensive knowledge support for disease classification and reducing the subjectivity and limitations of human judgment.

[0061] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0062] The present invention is described with reference to flowcharts of methods according to embodiments of the present invention. It should be understood that each combination of processes in the flowcharts can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts. Figure 1 A device that specifies functions in a process or multiple processes.

[0063] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A function specified in a process or multiple processes.

[0064] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 The steps of a specified function in a process or multiple processes.

[0065] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0066] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A medical image classification method based on knowledge graph, characterized in that: The following steps are involved: Step S1: Obtain the patient's medical imaging data, extract the medical imaging feature vector, map the medical imaging feature vector to the medical knowledge graph, and screen the candidate disease set based on the similarity of the corresponding nodes in the medical knowledge graph; Step S2: Obtain patient medication data, identify the drug efficacy coefficient, side effect coefficient, and interaction coefficient based on the medication data, and then use the synergistic and antagonistic relationships between drugs to map the medical knowledge graph, simulate the impact of patient medication on the disease, and update the candidate disease set; Step S3, identifying the patient's genetic background and lifestyle habits, and then determining the probability of the patient having a disease based on the genetic background and lifestyle habits, screening the updated set of candidate diseases based on the probability of the patient having a disease, and marking a first set of diseases; Step S4, using the medical knowledge graph to identify the micro-tissue effects and related diseases of the diseases in the first disease set, and then verifying the diseases in the first disease set based on the patient's micro-tissue characteristics and related disease data, and outputting a second disease set, which is the patient's disease classification result.

2. A medical image classification method based on knowledge graph according to claim 1, characterized in that: Step S1 is specifically analyzed as follows: using medical imaging equipment to obtain medical imaging data of the patient, and then inputting the medical imaging data into a convolutional neural network to output a medical imaging feature vector; Associating the medical image feature vector with the node attributes in the medical knowledge graph, and then identifying the similarity between the medical image feature vector and each disease node in the medical knowledge graph, wherein the node attributes include typical image feature descriptions related to the disease; The similarity between the medical image feature vector and each disease node in the medical knowledge graph is compared with the similarity threshold, and the diseases represented by the disease nodes with similarity higher than the similarity threshold are marked to form the candidate disease set.

3. The medical image classification method based on knowledge graph according to claim 1, characterized in that: Step S2 is specifically analyzed as follows: medication data includes the patient's medication records and the corresponding treatment effect and side effect data. The medication data is used to obtain the drug efficacy coefficient, side effect coefficient, and interaction coefficient. If the interaction coefficient between the drugs is positive, the drugs have a synergistic effect. If the interaction coefficient is negative, there is an antagonistic effect. Identify the corresponding drug nodes and disease nodes in the medical knowledge graph. For drug nodes with synergistic or antagonistic relationships, establish corresponding edges in the medical knowledge graph to represent them. At the same time, add the efficacy coefficient and side effect coefficient as attributes of the drug node to the medical knowledge graph. Then, based on the relationship between drugs and diseases and the interaction between drugs in the knowledge graph, combined with the patient's medication records, simulate the impact path of drugs on diseases. If the effect of a certain drug combination on a certain disease is consistent with the patient's actual symptoms, then add the disease to the candidate disease set. If the side effect of a certain drug is related to the abnormal condition of the patient, and the disease corresponding to the side effect is not in the candidate disease set, then add the disease to the candidate disease set, and at the same time screen out diseases in the candidate disease set whose patient's medication situation is inconsistent with the actual performance.

4. The medical image classification method based on knowledge graph according to claim 3, characterized in that: The specific analysis of obtaining the drug efficacy coefficient, side effect coefficient and interaction coefficient of the drug using the medication data is as follows: for the efficacy coefficient, taking the treatment effect as the dependent variable and the medication data as the independent variable, constructing a drug efficacy evaluation model, and using the drug efficacy evaluation model to solve and obtain the drug efficacy coefficient; For the side effect coefficient, a side effect evaluation model was constructed with the occurrence of side effects as the dependent variable and the medication data as the independent variable, and the side effect coefficient was obtained by solving the side effect evaluation model; For the interaction coefficient, the combination of drugs used simultaneously by the patient is taken as a new feature, and the treatment effect or side effect occurrence is used as the dependent variable to construct an interaction evaluation model, and the drug interaction coefficient is obtained by using the interaction evaluation model.

5. The medical image classification method based on knowledge graph according to claim 1, characterized in that: The specific analysis of step S3 is as follows: obtaining the patient's genetic adaptability coefficient and comprehensive score of lifestyle habits, combining the patient's genetic adaptability coefficient and comprehensive score of lifestyle habits to determine the patient's susceptibility values for various diseases, comparing the susceptibility values for various diseases with the disease risk threshold, marking the diseases with susceptibility values higher than the disease risk threshold as the susceptibility disease set, and then summarizing the susceptibility disease set and the updated candidate disease set to obtain the first disease set.

6. The medical image classification method based on knowledge graph according to claim 5, characterized in that: The genetic compatibility coefficient and comprehensive score of lifestyle habits of the patient are specifically obtained by obtaining disease data of the patient's relatives, including the type of disease suffered by the relatives, the age of disease onset, and the relationship with the patient, and determining the genetic compatibility coefficient of the patient based on the disease data of the relatives; Obtain the patient's living habit information, which includes diet structure, amount of exercise, smoking and drinking habits, and work and rest patterns, and determine the patient's comprehensive living habit score based on the living habit information.

7. The medical image classification method based on knowledge graph according to claim 1, characterized in that: Step S4 is specifically analyzed as follows: the micro-tissue impact information includes the tissue type affected by the disease and the pathological changes of the tissue; the associated diseases include complications and other diseases caused by related causes; Obtaining the patient's microtissue feature data and associated disease data, comparing and analyzing the microtissue effect of each disease in the first disease set with the patient's microtissue feature, determining whether the patient's microtissue feature matches the microtissue effect of the disease, and determining whether the patient has an associated disease related to the disease based on the patient's associated disease data. If the patient's microtissue feature matches the microtissue effect of the disease and a corresponding associated disease exists, the disease passes verification; The diseases that have passed the verification in the first set are screened out and merged to obtain the second disease set.