Medical multi-modal knowledge graph construction method and system, electronic equipment and storage medium
By integrating text and visual information into a multimodal medical knowledge graph construction method, the problem that existing medical knowledge graphs cannot effectively integrate visual images is solved, and more comprehensive knowledge expression and intelligent diagnosis support are achieved.
Patent Information
- Application Number
- CN202510791514.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-26
AI Technical Summary
Existing medical knowledge graphs mainly rely on text descriptions, which makes it difficult to effectively integrate visual image information, unable to meet the needs of multimodal medical application scenarios, and lack intuitive display of disease visual characteristics and comprehensive diagnostic support.
By constructing a multimodal medical knowledge graph, integrating text information and medical images, using the BERT-BiLSTM-CRF model for medical entity recognition, combining the large language model (LLM) to extract triplets, and using the ResNet-50 model to screen semantically related medical images, a multimodal knowledge graph was constructed.
It significantly improves the interpretability and clinical reference value of knowledge graphs, provides more comprehensive knowledge support, and provides efficient and accurate multimodal reasoning capabilities for intelligent diagnosis.
Smart Images

Figure CN120706515A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of knowledge graphs, and specifically relates to a method, system, electronic device and storage medium for constructing a medical multimodal knowledge graph. Background Art
[0002] A knowledge graph is a semantic network structure used to express entities and their relationships. It models real-world knowledge in the form of triples and is widely used in intelligent question-answering, recommendation systems, information extraction, and medical decision-making support.
[0003] In the medical field, knowledge graphs can integrate different types of medical data, construct structured knowledge systems, and assist doctors in making efficient and accurate inferences and diagnoses in complex data. Existing medical knowledge graphs usually focus on text sources such as electronic medical records, medical literature, and disease encyclopedias. They only contain medical entities such as symptoms, causes, treatment plans, and sites of disease, as well as the semantic relationships between them, while ignoring the information expression at the visual level. In actual medical scenarios, medical images (such as skin surface images, CT images, pathological images, etc.) play a vital role in disease identification, etiology analysis, and auxiliary diagnosis. Especially for superficial diseases, visual features are often the first basis for diagnosis. The multimodal information carried by medical images can effectively make up for the shortcomings of text information and improve the interpretability and practicality of knowledge graphs.
[0004] Therefore, the existing medical knowledge graph based only on text can no longer meet the needs of multimodal medical application scenarios. It is urgent to propose a multimodal medical knowledge graph construction method that integrates text information and visual image information to achieve more comprehensive, realistic, and semantically rich medical knowledge modeling capabilities, thereby better supporting downstream medical tasks such as multimodal reasoning, disease identification, and intelligent diagnosis. Summary of the Invention
[0005] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology and provide a method, system, electronic device and storage medium for constructing a medical multimodal knowledge graph. The present invention integrates text information and image information, thereby improving the expressive ability and multimodal application effect of the medical knowledge graph.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a method for constructing a medical multimodal knowledge graph, comprising the following steps:
[0008] Build a disease seed library based on the International Classification of Diseases standards and screen disease terms with standardized expressions;
[0009] Using the terms in the disease seed library as keywords, disease-related text data is obtained from the general knowledge platform and stored in a structured manner; the text data includes disease definitions, clinical manifestations, complications, pathogenesis, and treatment methods;
[0010] Use medical named entity recognition tools to perform medical entity recognition on text data, extract symptoms, causes, diagnosis and treatment departments, onset sites and related disease entities, and retain entity context information;
[0011] A command fine-tuning method based on a large language model extracts medical knowledge triplets from identified medical entities. The triplet structure is (head entity, relationship, tail entity), and the relationship types include symptoms, causes, department, site of onset, and related diseases.
[0012] Obtain medical images using disease terms as keywords and screen images that meet semantic relevance and quality standards;
[0013] Medical entities, triple relationships and related medical images are integrated and stored as a graph database to construct a multimodal medical knowledge graph.
[0014] As a preferred technical solution, the disease seed library focuses on surface diseases, including skin diseases, subcutaneous tissue diseases, surface trauma and superficial tumors.
[0015] As a preferred technical solution, the medical named entity recognition tool is a Chinese medical entity recognition model based on the BERT-BiLSTM-CRF architecture. The process of medical entity recognition is as follows:
[0016] The pre-trained language model BERT is used to extract context-related word vectors, which are then input into a bidirectional LSTM network to capture context sequence information. Finally, the optimal label sequence is obtained through the CRF layer, which is expressed as follows:
[0017] BERT(w1,w2,...,w n )→H=[h1,h2,...,h n ]
[0018] BiLSTM(H)→H′=[h′1,h′2,...,h′ n ]
[0019] CRF(H′)→[y1,y2,...,y n ]
[0020] where w i represents the i-th word in the input text, h i 、h i ′ are word vector representations of BERT output and BiLSTM output, respectively, iis the final entity label identified; the identified medical entity is retained together with the original sentence to enhance the semantic consistency in subsequent triple construction.
[0021] As a preferred technical solution, the instruction fine-tuning method based on the large language model is specifically as follows:
[0022] Medical knowledge triples are constructed from the identified medical entity pairs. The triple structure is in the form of "head entity-relationship-tail entity", where the relationship set is:
[0023]
[0024] In the specific implementation process, the instruction input template is input="Given text: x, please extract all medical triples of the form (entity 1, relationship, entity 2)"
[0025] Use the large language model LLM to encode the input text x and generate a set of triples:
[0026] T = LLM(x)
[0027] in represents all medical triples extracted from the text, h i and t i are the head and tail entities respectively, For its corresponding medical relationship;
[0028] After the model output, redundant and erroneous relationships are removed through a confidence filtering mechanism based on semantic similarity and medical dictionary rules.
[0029] As a preferred technical solution, the screening of images that meet the semantic relevance and quality standards is specifically as follows:
[0030] The resolution of the screened images must be no less than 512×512 pixels;
[0031] Image source: Trusted Medical Information Website;
[0032] An image classifier based on the ResNet-50 model is used to determine the semantic relevance of downloaded images.
[0033] 7. As a preferred technical solution, the image classifier based on the ResNet-50 model is used to determine the semantic relevance of the downloaded images, specifically:
[0034] First, extract the features of the medical image: h = f resnet (I), output image feature vector
[0035] Relevance score: s = w Th+b, where w is the learnable weight and b is the bias value;
[0036] Threshold determination: If s ≥ τ, where τ is the set threshold, it is determined to be "highly correlated", otherwise it is considered unrelated or requires manual review;
[0037] The classifier inputs image I and outputs the predicted label in Indicates that the image is highly correlated with the target disease, and the image discriminant function is defined as:
[0038]
[0039] where f resnet (·) represents the image classifier of the ResNet-50 model.
[0040] As a preferred technical solution, the graph database is specifically:
[0041] Entities are stored as nodes with text attributes and image URLs.
[0042] Relationships are stored in the form of edges, and the relationship type is marked;
[0043] Image features are extracted as vectors through a pre-trained model and embedded into corresponding nodes.
[0044] In a second aspect, the present invention provides a medical multimodal knowledge graph construction system, which is applied to the medical multimodal knowledge graph construction method, including a disease seed library construction module, a text data extraction module, a medical entity recognition module, a triple extraction module, a medical image processing module, and a multimodal knowledge graph generation module;
[0045] The disease seed library construction module is used to construct a disease seed library based on the International Classification of Diseases standards and screen disease terms with standardized expressions;
[0046] The text data extraction module is used to obtain disease-related text data from the general knowledge platform using terms in the disease seed library as keywords and store them in a structured manner; the text data includes disease definitions, clinical manifestations, complications, pathogenesis, and treatment methods;
[0047] The medical entity recognition module is used to perform medical entity recognition on text data using a medical named entity recognition tool, extract symptoms, causes, diagnosis and treatment departments, onset sites and related disease entities, and retain entity context information;
[0048] The triple extraction module extracts medical knowledge triples from identified medical entities based on the instruction fine-tuning method of the large language model. The triple structure is (head entity, relationship, tail entity), and the relationship types include symptoms, causes, department, site of onset, and related diseases.
[0049] The medical image processing module is used to acquire medical images using disease terms as keywords and screen images that meet semantic relevance and quality standards;
[0050] The multimodal knowledge graph generation module is used to integrate and store medical entities, triple relationships and associated medical images into a graph database to construct a multimodal medical knowledge graph.
[0051] In a third aspect, the present invention provides an electronic device, comprising:
[0052] at least one processor; and,
[0053] a memory communicatively connected to the at least one processor; wherein,
[0054] The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the medical multimodal knowledge graph construction method.
[0055] In a fourth aspect, the present invention provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the method for constructing a medical multimodal knowledge graph.
[0056] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0057] 1. The present invention constructs a structured multimodal medical knowledge graph by integrating text data and visual image information. Traditional medical knowledge graphs mainly rely on text descriptions, which makes it difficult to intuitively display the visual characteristics of the disease. The present invention associates and stores text information such as disease definitions, symptoms, and treatment plans with medical images (such as skin lesion images, CT images, etc.), making the knowledge expression more three-dimensional. The multimodal fusion method of the present invention makes up for the limitations of a single text modality, significantly improves the interpretability and clinical reference value of the knowledge graph, and provides more comprehensive knowledge support for intelligent diagnosis.
[0058] 2. This paper uses the BERT-BiLSTM-CRF model for high-precision medical entity recognition, combined with triple extraction technology from the Large Language Model (LLM), to achieve the automatic conversion from unstructured text to structured knowledge. Compared with traditional manual annotation methods, this method is much more efficient and ensures the accuracy of triple relationships through a dual screening mechanism of semantic similarity calculation and medical rule verification.
[0059] 3. This paper uses the ResNet-50 model to determine the semantic relevance of downloaded medical images, and then performs dual filtering based on source credibility (limited to professional medical websites such as WebMD and Mayo Clinic) and image quality (resolution ≥ 512×512). This intelligent filtering mechanism effectively addresses the uneven quality of online images and provides professional and reliable visual information for the knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0061] Figure 1 This is a flowchart of a method for constructing a medical multimodal knowledge graph according to an embodiment of the present invention;
[0062] Figure 2 This is a block diagram of a medical multimodal knowledge graph construction system according to an embodiment of the present invention.
[0063] Figure 3 2 is a structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0064] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0065] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0066] ICD-10 (International Classification of Diseases, 10th Revision) is a globally accepted standardized classification system for diseases, symptoms, abnormal signs, and other health problems developed by the World Health Organization (WHO). It is used in disease statistics, clinical diagnosis, medical management, and public health research. Its classification covers diseases, injuries, symptoms, external causes (such as accidental injuries), social factors (such as social environments that affect health), etc., and contains approximately 68,000 codes. The code format is a combination of letters and numbers (such as A00.0). The first letter indicates the major category of the disease, followed by numbers to refine the classification, and the decimal point further distinguishes subtypes.
[0067] Figure 1 As shown, this embodiment provides a method for constructing a medical multimodal knowledge graph, which mainly includes the following steps:
[0068] S1. Build a disease seed library based on the International Classification of Diseases standards and screen disease terms with standardized expressions.
[0069] Furthermore, this example uses the ICD-10 International Classification of Diseases to screen for standardized disease terms and form a disease seed library. This disease seed library is based on the ICD-10 chapter structure related to superficial diseases, focusing on categories such as skin diseases, subcutaneous tissue diseases, superficial trauma, and superficial tumors. Ultimately, a total of 3,817 disease names were collected to form the seed library.
[0070] It is understandable that the disease seed library of this embodiment is dynamically updated. New disease terms not covered by ICD-10 are identified using NLP technology through medical literature (such as PubMed) or electronic medical record data, and the seed library is dynamically updated.
[0071] S2. Using the terms in the disease seed library as keywords, obtain disease-related text data from the general knowledge platform and store them in a structured manner.
[0072] Furthermore, in step S2, each professional disease name in the disease seed library is used as a keyword to automatically obtain text information including disease definitions, clinical manifestations, complications, pathogenesis, treatment methods, etc. on general knowledge platforms such as Baidu Encyclopedia, Wikipedia, and Wikidata, and store them in a structured manner.
[0073] The disease definition is an authoritative description of the nature of the disease, including its pathological characteristics, taxonomic status, and core diagnostic elements. Clinical manifestations are the objective manifestations and subjective symptoms of the disease that can be observed or perceived in patients. Complications are secondary pathological conditions or complications that occur during the development of the disease, including direct and indirect complications. Pathogenesis is the biological process and mechanism of the occurrence and development of the disease, including genetic or environmental factors. Treatment methods include traditional Chinese medicine and Western medicine.
[0074] S3. Use medical named entity recognition tools to perform medical entity recognition on text data, extract symptoms, causes, diagnosis and treatment departments, onset sites and related disease entities, and retain entity context information.
[0075] Furthermore, in step S3, the recognition tool used is a Chinese medical entity recognition model based on the BERT-BiLSTM-CRF architecture. First, the pre-trained language model BERT is used to extract context-related word vectors, which are then input into a bidirectional LSTM network to capture context sequence information. Finally, the optimal label sequence is obtained through the CRF layer. The process can be expressed as follows:
[0076] BERT(w1,w2,...,w n )→H=[h1,h2,...,h n ]
[0077] BiLSTM(H)→H′=[h′1,h′2,...,h′ n ]
[0078] CRF(H′)→[y1,y2,...,y n ]
[0079] where w i represents the i-th word in the input text, h i 、h i ′ are word vector representations of BERT output and BiLSTM output, respectively, i is the final entity label identified. The identified medical entity is retained together with the original sentence to enhance the semantic consistency in the subsequent triple construction.
[0080] Furthermore, to address the specificity of medical text, the BERT model requires domain-adaptive pre-training. Based on the general BERT model, a second pre-training is performed using medical corpora (such as Chinese medical papers, electronic medical records, and clinical guidelines) to learn the semantics of specialized terms such as "colorectal cancer" and "EGFR inhibitors." Furthermore, to address ambiguous Chinese word segmentation (for example, "diabetic nephropathy" should be considered as a single entity), character-level input is used to avoid segmentation errors.
[0081] S4. An instruction fine-tuning method based on a large language model extracts medical knowledge triplets from the identified medical entities. The triplet structure is (head entity, relationship, tail entity), and the relationship types include symptoms, causes, departments, sites of onset, and related diseases.
[0082] Furthermore, in step S4, the triple structure is in the form of "head entity-relationship-tail entity". The relationship set is:
[0083]
[0084] In the specific implementation process, the instruction input template is input="Given text: x, please extract all medical triples of the form (entity 1, relationship, entity 2)"
[0085] Use the large language model LLM to encode the input text x and generate a set of triples:
[0086] T = LLM(x)
[0087] in represents all medical triples extracted from the text, h i and t i are the head and tail entities respectively, To ensure the quality of the triples, the model output is filtered through a confidence filter based on semantic similarity and medical dictionary rules to remove redundant and erroneous relationships.
[0088] S5. Obtain medical images using disease terms as keywords and screen images that meet semantic relevance and quality standards.
[0089] Furthermore, in step S5, the acquired images must meet the following filtering criteria: the image size must be no less than 512×512, the image source must include credible medical information websites such as WebMD, Mayo Clinic, NIH, etc., and the downloaded images must be semantically relevant using an image classifier based on the ResNet-50 model. The determination process is as follows:
[0090] Step 1: First extract the features of the medical image: h = f resnet (I), output image feature vector (Before the last layer of pooling).
[0091] Step 2: Relevance score: s = w T h+b, where w is the learnable weight and b is the bias value.
[0092] Step 3: Threshold determination: If s ≥ τ (τ = 0.8) (such as τ = 0.8), it is determined to be "highly correlated", otherwise it is considered unrelated or requires manual review.
[0093] The classifier inputs image I and outputs the predicted label in Indicates that the image is highly correlated with the target disease, and the image discriminant function is defined as:
[0094]
[0095] where f resnet (·) represents the image classifier of the ResNet-50 model.
[0096] Furthermore, the image classifier based on the ResNet-50 model is used to determine the semantic relevance of the downloaded images, specifically:
[0097] S6. Integrate and store medical entities, triple relationships, and related medical images into a graph database to construct a multimodal medical knowledge graph.
[0098] Furthermore, in step S6, the graph database is specifically:
[0099] Entities are stored as nodes with text attributes and image URLs.
[0100] Relationships are stored in the form of edges, and the relationship type is marked;
[0101] Image features are extracted as vectors through a pre-trained model and embedded into corresponding nodes.
[0102] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.
[0103] Based on the same concept as the medical multimodal knowledge graph construction method in the above embodiment, the present invention also provides a medical multimodal knowledge graph construction system, which can be used to execute the above medical multimodal knowledge graph construction method. For ease of explanation, the structural diagram of the embodiment of the medical multimodal knowledge graph construction system only shows the parts related to the embodiment of the present invention. Those skilled in the art will understand that the illustrated structure does not constitute a limitation of the device, and it can include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0104] See also Figure 2 In another embodiment of the present application, a medical multimodal knowledge graph construction system 100 is provided, which includes a disease seed library construction module 101, a text data extraction module 102, a medical entity recognition module 103, a triple extraction module 104, a medical image processing module 105, and a multimodal knowledge graph generation module 106;
[0105] The disease seed library construction module 101 is used to construct a disease seed library based on the International Classification of Diseases standard and screen disease terms with standardized expression;
[0106] The text data extraction module 102 is used to obtain disease-related text data from the general knowledge platform using terms in the disease seed library as keywords and store them in a structured manner; the text data includes disease definitions, clinical manifestations, complications, pathogenesis, and treatment methods;
[0107] The medical entity recognition module 103 is used to perform medical entity recognition on text data using a medical named entity recognition tool, extract symptoms, causes, diagnosis and treatment departments, onset sites and related disease entities, and retain entity context information;
[0108] The triple extraction module 104 extracts medical knowledge triples from the identified medical entities based on the instruction fine-tuning method of the large language model. The triple structure is (head entity, relationship, tail entity), and the relationship types include symptoms, causes, department, site of onset, and related diseases.
[0109] The medical image processing module 105 is used to acquire medical images using disease terms as keywords and screen images that meet semantic relevance and quality standards;
[0110] The multimodal knowledge graph generation module 106 is used to integrate and store medical entities, triple relationships and associated medical images into a graph database to construct a multimodal medical knowledge graph.
[0111] It should be noted that the medical multimodal knowledge graph construction system of the present invention corresponds one-to-one to the medical multimodal knowledge graph construction method of the present invention. The technical features and beneficial effects described in the above-mentioned embodiments of the medical multimodal knowledge graph construction method are all applicable to the embodiments of the medical multimodal knowledge graph construction. For specific contents, please refer to the description in the embodiments of the method of the present invention. No further details will be given here. This is hereby declared.
[0112] In addition, in the implementation of the medical multimodal knowledge graph construction system in the above embodiment, the logical division of each program module is only an example. In actual application, the above functions can be assigned to different program modules as needed, for example, for the configuration requirements of the corresponding hardware or the convenience of software implementation. That is, the internal structure of the medical multimodal knowledge graph construction system is divided into different program modules to complete all or part of the functions described above.
[0113] See also Figure 3 In one embodiment, an electronic device for implementing a method for constructing a medical multimodal knowledge graph is provided. The electronic device 200 may include a first processor 201, a first memory 202 and a bus, and may also include a computer program stored in the first memory 202 and executable on the first processor 201, such as a medical multimodal knowledge graph construction program 203.
[0114] The first memory 202 includes at least one type of readable storage medium, including flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the first memory 202 can be an internal storage unit of the electronic device 200, such as a mobile hard disk of the electronic device 200. In other embodiments, the first memory 202 can also be an external storage device of the electronic device 200, such as a plug-in mobile hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 200. Furthermore, the first memory 202 can also include both an internal storage unit of the electronic device 200 and an external storage device. The first memory 202 can not only be used to store application software and various types of data installed on the electronic device 200, such as the code of the medical multimodal knowledge graph construction program 203, but can also be used to temporarily store data that has been output or is about to be output.
[0115] In some embodiments, the first processor 201 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The first processor 201 is the control core (Control Unit) of the electronic device, connecting the various components of the entire electronic device using various interfaces and lines, and executing or executing programs or modules stored in the first memory 202, as well as calling data stored in the first memory 202, to perform various functions of the electronic device 200 and process data.
[0116] Figure 3 Only the electronic device with components is shown, and it can be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the electronic device 200 , and the electronic device 200 may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0117] The medical multimodal knowledge graph construction program 203 stored in the first memory 202 of the electronic device 200 is a combination of multiple instructions. When executed in the first processor 201, it can achieve the following:
[0118] Build a disease seed library based on the International Classification of Diseases standards and screen disease terms with standardized expressions;
[0119] Using the terms in the disease seed library as keywords, disease-related text data is obtained from the general knowledge platform and stored in a structured manner; the text data includes disease definitions, clinical manifestations, complications, pathogenesis, and treatment methods;
[0120] Use medical named entity recognition tools to perform medical entity recognition on text data, extract symptoms, causes, diagnosis and treatment departments, onset sites and related disease entities, and retain entity context information;
[0121] A command fine-tuning method based on a large language model extracts medical knowledge triplets from identified medical entities. The triplet structure is (head entity, relationship, tail entity), and the relationship types include symptoms, causes, department, site of onset, and related diseases.
[0122] Obtain medical images using disease terms as keywords and screen images that meet semantic relevance and quality standards;
[0123] Medical entities, triple relationships and related medical images are integrated and stored as a graph database to construct a multimodal medical knowledge graph.
[0124] Furthermore, if the modules / units integrated in the electronic device 200 are implemented as software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0125] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0126] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0127] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A method for constructing a medical multimodal knowledge graph, characterized in that: The steps include: Build a disease seed library based on the International Classification of Diseases standards and screen disease terms with standardized expressions; Using the terms in the disease seed library as keywords, disease-related text data is obtained from the general knowledge platform and stored in a structured manner; the text data includes disease definitions, clinical manifestations, complications, pathogenesis, and treatment methods; Use medical named entity recognition tools to perform medical entity recognition on text data, extract symptoms, causes, diagnosis and treatment departments, onset sites and related disease entities, and retain entity context information; A command fine-tuning method based on a large language model extracts medical knowledge triplets from identified medical entities. The triplet structure is (head entity, relationship, tail entity), and the relationship types include symptoms, causes, department, site of onset, and related diseases. Obtain medical images using disease terms as keywords and screen images that meet semantic relevance and quality standards; Medical entities, triple relationships and related medical images are integrated and stored as a graph database to construct a multimodal medical knowledge graph.
2. The method for constructing a medical multimodal knowledge graph according to claim 1, wherein: The disease seed library focuses on surface diseases, including skin diseases, subcutaneous tissue diseases, surface wounds and superficial tumors.
3. The method for constructing a medical multimodal knowledge graph according to claim 1, wherein: The medical named entity recognition tool is a Chinese medical entity recognition model based on the BERT-BiLSTM-CRF architecture. The process of medical entity recognition is as follows: The pre-trained language model BERT is used to extract context-related word vectors, which are then input into a bidirectional LSTM network to capture context sequence information. Finally, the optimal label sequence is obtained through the CRF layer, which is expressed as follows: BERT(w1,w2,...,w n )→H=[h1,h2,...,h n ] BiLSTM(H)→H′=[h′1,h′2,...,h′ n ] CRF(H′)→[y1,y2,...,y n ] where w i represents the i-th word in the input text, h i 、h i ′ are word vector representations of BERT output and BiLSTM output, respectively, i is the final entity label identified; the identified medical entity is retained together with the original sentence to enhance the semantic consistency in subsequent triple construction.
4. The method for constructing a medical multimodal knowledge graph according to claim 1, characterized in that: The instruction fine-tuning method based on the large language model is specifically as follows: Medical knowledge triples are constructed from the identified medical entity pairs. The triple structure is in the form of "head entity-relationship-tail entity", where the relationship set is: In the specific implementation process, the instruction input template is constructed as input = "given text: x, please extract all medical triples in the form of (entity 1, relationship, entity 2)"; Use the large language model LLM to encode the input text x and generate a set of triples: T = LLM(x) in represents all medical triples extracted from the text, h i and t i are the head and tail entities respectively, For its corresponding medical relationship; After the model output, redundant and erroneous relationships are removed through a confidence filtering mechanism based on semantic similarity and medical dictionary rules.
5. The method for constructing a medical multimodal knowledge graph according to claim 1, characterized in that: The screening of images that meet the semantic relevance and quality standards is specifically as follows: The resolution of the screened images must be no less than 512×512 pixels; Image source: Trusted Medical Information Website; An image classifier based on the ResNet-50 model is used to determine the semantic relevance of downloaded images.
6. The method for constructing a medical multimodal knowledge graph according to claim 1, characterized in that: The image classifier based on the ResNet-50 model is used to determine the semantic relevance of the downloaded images, specifically: First, extract the features of the medical image: h = f resnet (I), output image feature vector Relevance score: s = w T h+b, where w is the learnable weight and b is the bias value; Threshold determination: If s ≥ τ, where τ is the set threshold, it is determined to be "highly correlated", otherwise it is considered unrelated or requires manual review; The classifier inputs image I and outputs the predicted label in Indicates that the image is highly correlated with the target disease, and the image discriminant function is defined as: where f resnet (·) represents the image classifier of the ResNet-50 model.
7. The method for constructing a medical multimodal knowledge graph according to claim 1, characterized in that: The graph database is specifically: Entities are stored as nodes with text attributes and image URLs. Relationships are stored in the form of edges, and the relationship type is marked; Image features are extracted as vectors through a pre-trained model and embedded into corresponding nodes.
8. A medical multimodal knowledge graph construction system, characterized by: A method for constructing a medical multimodal knowledge graph as described in any one of claims 1 to 7, comprising a disease seed library construction module, a text data extraction module, a medical entity recognition module, a triple extraction module, a medical image processing module, and a multimodal knowledge graph generation module; The disease seed library construction module is used to construct a disease seed library based on the International Classification of Diseases standards and screen disease terms with standardized expressions; The text data extraction module is used to obtain disease-related text data from the general knowledge platform using terms in the disease seed library as keywords and store them in a structured manner; the text data includes disease definitions, clinical manifestations, complications, pathogenesis, and treatment methods; The medical entity recognition module is used to perform medical entity recognition on text data using a medical named entity recognition tool, extract symptoms, causes, diagnosis and treatment departments, onset sites and related disease entities, and retain entity context information; The triple extraction module extracts medical knowledge triples from identified medical entities based on the instruction fine-tuning method of the large language model. The triple structure is (head entity, relationship, tail entity), and the relationship types include symptoms, causes, department, site of onset, and related diseases. The medical image processing module is used to acquire medical images using disease terms as keywords and screen images that meet semantic relevance and quality standards; The multimodal knowledge graph generation module is used to integrate and store medical entities, triple relationships and associated medical images into a graph database to construct a multimodal medical knowledge graph.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the medical multimodal knowledge graph construction method as described in any one of claims 1-7.
10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the medical multimodal knowledge graph construction method described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Method and device for constructing medical knowledge graph
CN120930759A
A medical knowledge graph construction method and device
CN120930759B
Multi-modal domain knowledge graph construction method and system
CN121599077A