Rare disease knowledge graph construction method based on modal injection and multi-modal fusion

By employing modal injection and multimodal fusion methods, the problems of data scarcity and heterogeneity in rare disease knowledge graphs were solved, achieving efficient fusion and intelligent diagnosis of rare diseases, and constructing a rare disease knowledge graph with semantic consistency and traceability.

CN120806104AActive Publication Date: 2025-10-17湖南工商大学

Patent Information

Application Number
CN202511303571.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-17
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively construct multimodal knowledge graphs for rare diseases, especially in cases of data scarcity, heterogeneity, and modality loss. They fail to achieve semantic consistency, missing data completion, and structural representation, resulting in insufficient intelligent diagnostic capabilities.

Method used

We employ modality injection and multimodal fusion methods, using RareGAN-VAE to generate models to complete missing modality data, combining Bio_ClinicalBERT, 3D ResNet+U-Net and MLP to extract unified semantic representations, using modality injection mechanisms to enhance cross-modal semantic perception, and using cross-modal alignment modules to achieve shared semantic space alignment, ultimately constructing a two-layer rare disease knowledge graph.

Benefits of technology

It achieves semantic consistency enhancement, modality missing completion, and structural traceability of rare disease knowledge graphs, improving the accuracy and adaptability of rare disease diagnosis and supporting efficient fusion and intelligent diagnosis of multimodal medical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806104A_ABST
    Figure CN120806104A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical artificial intelligence and knowledge graph construction, in particular to a rare disease knowledge graph construction method based on modal injection and multi-modal fusion. Comprising the following steps: S1, collecting multi-modal medical information including texts, images and genes; s2, standardization processing is carried out, and a three-layer metadata structure is constructed; s3, complementing missing modal data, and performing feature extraction and unified dimension conversion on the modal data to realize representation alignment in a shared semantic space; s4, performing multi-level semantic fusion to obtain a unified fusion semantic vector; and S5, constructing a double-layer structure system rare disease knowledge graph comprising an ontology layer and an instance layer. According to the method, multi-modal medical information of texts, images and genes is selected to construct the knowledge graph of the rare disease, the application range, coverage and accuracy of the knowledge graph are improved, correspondence adaptation of rare cases during clinical diagnosis and treatment of the rare disease can be achieved, and the method has high recognition capacity.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical artificial intelligence and knowledge graph construction, and particularly relates to a rare disease knowledge graph construction method based on modal injection and multi-modal fusion. BACKGROUND

[0002] With the development of medical artificial intelligence and knowledge engineering, medical knowledge graphs have gradually become an important support means for clinical auxiliary diagnosis, scientific research knowledge discovery and medical resource integration. However, in the field of rare diseases, the construction and application of knowledge graphs still face many technical challenges, mainly in the aspects of data scarcity, high heterogeneity and semantic fragmentation, which seriously restrict the promotion of the graph in the intelligent diagnosis and treatment scene.

[0003] Rare diseases usually have the characteristics of extremely small number of clinical samples, long diagnosis cycle and significant phenotypic heterogeneity, which makes it difficult for traditional graph construction methods relying on large-scale structured corpus and single modal information to be applicable. Especially in real clinical scenarios, patients' medical data usually exist in the form of multi-modal, covering structured clinical text (such as diagnosis records, symptom descriptions), medical images (such as CT, MRI) and genomic sequencing data (such as mutation sites in VCF format). These modalities have significant differences in structure, dimension and semantic system, showing the complex characteristics of "multi-source heterogeneity".

[0004] Current research attempts to model text and image or text and gene in a dual-modal manner, but has not yet formed a knowledge graph construction scheme that can simultaneously fuse text, image and genomic three modalities. This is mainly due to the fact that the following technical difficulties have not been effectively broken through: Firstly, the three types of modalities are highly heterogeneous in terms of data organization method, semantic representation form and structural granularity. Text is a time-series semantic structure, image is spatial distribution information, and gene data is high-sparse, strongly structured encoded data, which are difficult to establish equivalent relationships or share semantic representation spaces.

[0005] Secondly, existing modal fusion methods mostly use simple concatenation or linear weighting methods, which cannot fully model the deep semantic complementarity and linkage mechanism between modalities. The lack of a unified modal alignment mechanism makes it difficult to maintain semantic consistency in entity relationships in the graph, especially in applications such as diagnosis and reasoning that require high semantic requirements, which shows obvious shortcomings.

[0006] Thirdly, the problem of missing modal data is particularly serious in the field of rare diseases. Most rare disease patients only have partial modal information, such as missing image examination or incomplete gene sequencing. Existing multi-modal graph construction schemes generally rely on complete modal data input, cannot handle incomplete modal samples, and lack an automatic evaluation mechanism for the credibility of synthetic completed data, thereby affecting the rigor and explainability of the graph.

[0007] In addition, in the aspect of knowledge graph structure construction, the traditional scheme relies on static rules and structured ontology extraction, and it is difficult to automatically generate supplementary semantic edges or enhanced semantic connections according to the semantic linkage relationship between different modalities, so that the adaptability of the graph in reasoning, path retrieval and clinical application is weak.

[0008] In summary, there is an urgent need for a graph construction method that can enhance the semantic consistency between modalities, complete the missing data, and have a structure traceable mechanism for the characteristics of rare disease multi-modal medical data, to break through the applicability bottleneck of existing technology in the environment of high sparsity, multi-heterogeneous and weak annotation, and to improve the intelligent diagnosis ability and knowledge expression efficiency in the rare disease scenario. SUMMARY

[0009] The present application provides a rare disease knowledge graph construction method based on modality injection and multi-modal fusion, which aims to process multi-modal medical data and fuse and enhance modality features at the same time, and adapt to the rare disease knowledge graph, to solve the problems of weak fusion ability, unclear structure and poor clinical adaptability of existing methods in multi-modal medical data processing, modality missing completion, semantic consistency alignment and graph structure expression.

[0010] In order to achieve the above purpose, the present application provides a rare disease knowledge graph construction method based on modality injection and multi-modal fusion, comprising the following steps: S1, collecting multi-modal medical information including text, image and gene; S2, standardizing the multi-modal medical information and constructing a three-layer metadata structure; S3, completing the missing modality data, extracting and converting the modality data to a unified dimension respectively, obtaining a global semantic modality representation vector in a unified format, processing the modality representation vector using a modality injection mechanism, enhancing cross-modality semantic perception, obtaining an enhanced representation vector, and realizing cross-modality representation alignment in a shared semantic space through a cross-modality alignment module; S4, multi-level semantic fusion of the modality features after cross-modality representation alignment to obtain a unified fusion semantic vector; S5, based on the unified fusion semantic vector, constructing a double-layer structure system rare disease knowledge graph including ontology layer and instance layer, and establishing edge relationships such as modality reference, semantic enhancement and traceability path in the graph.

[0011] Preferably, the text, image and gene include structured clinical text, medical image and genomic data derived from rare disease patients, and the genomic data includes low-frequency mutation sites and pathogenicity annotation information.

[0012] Rare diseases have some special requirements for the construction of knowledge graph due to the small sample size, special conditions of onset and other reasons: on the one hand, the data of rare diseases itself is very scarce, and often there is a modal missing; on the other hand, its diagnosis largely depends on genetic mutation information, unlike common diseases that can be determined by relying on images or symptoms. Therefore, mechanisms such as modal completion, modal injection, and fusion representation are introduced in our invention to better solve the problems of data structure and semantic fusion of rare diseases, improve the accuracy of the entire process, and have a core advantage in the field of rare disease diagnosis.

[0013] More preferably, the multi-modal medical information is derived from a hospital clinical information system (HIS), a picture archiving and communication system (PACS), and a gene sequencing center. Structured clinical text data includes patient basic information, diagnosis records, disease course descriptions, diagnosis descriptions, treatment descriptions, treatment plans, and follow-up records; medical image data includes typical examination results such as chest CT and brain MRI, stored in DICOM format; genomic data includes variant site information in the target region or whole genome, usually represented in VCF file format. After the above multi-modal data is collected, irrelevant information is cleaned, named entity is desensitized, and unified file format conversion is performed, and the physical and semantic mapping relationship is established through path indexing and semantic labeling, and stored in the multi-modal metadata framework, facilitating cross-modal analysis and calling in subsequent tasks.

[0014] Preferably, the three-layer metadata structure in S2 includes a disease type layer, a case layer, and a modal layer, the disease type layer is used to record standard disease codes, core symptom terms, and related pathogenic gene identifiers; the case layer is used to record patient basic information, treatment timelines, and family medical history; the modal layer is used to record modal data paths, semantic labels, feature vectors, and quality scores.

[0015] All standardized data will be uniformly organized and stored in the following "three-layer metadata structure": the disease type layer (Schema) defines disease ontology, symptom terms, pathogenic genes, treatment plans, and other standard medical entities, achieving standardized expression and unified reference of upper-layer medical knowledge; the case layer (Instance) carries specific individual cases, including patient information, timelines, family history, etc., building a bridge from standard knowledge to actual cases; the modal layer (Evidence) records the multi-modal raw data and semantic labels, quality scores corresponding to each case, providing a basis for graph reasoning and traceability. This structure effectively avoids the problem of "flat structure and flat entities" in traditional graphs, making the graph have the three-layer semantic capabilities of ontology reuse, instance expansion, and modal traceability.

[0016] The modality layer explicitly models each data path, supporting path mapping between raw text, images, and genetic files and semantic entities. Combined with the modality alignment module, it generates semantically enhanced edges (e.g., semantically_related_to, is_similar_to, and textual_equivalent_to), automatically generating semantic links between heterogeneous data. This allows doctors or systems to trace back to the original data source and quality indicators when querying graph nodes.

[0017] Preferably, the completing missing modal data in S3 specifically includes: for the missing modality, completing the missing modal data based on the RareGAN-VAE conditional generation model, and marking its credibility using a three-level verification mechanism including automatic evaluation, semantic rule verification, and expert scoring; The model takes segment features and noise vectors from existing modalities as input and synthesized data from missing modalities as output. The automatic assessment uses Frechet Inception Distance to evaluate image similarity and uses the ClinVar database to verify the matching of gene mutations and UMLS definitions of genes and disease phenotypes. The semantic rule verification includes gene-disease association rules and image-pathology feature matching rules. The gene-disease association rule stipulates that the synthesized gene mutation should be explicitly associated with the target disease in the ClinVar / HGNC database. The image-pathology feature matching rule stipulates that if a disease clearly indicates a certain pathological sign in Radiopaedia (an encyclopedia and teaching platform for medical imaging) or UMLS (Unified Medical Language System), the synthesized image should show the corresponding feature in the target area. The expert scoring is performed double-blind by experts with clinical experience in rare diseases, and the average score is no less than a preset threshold. The RareGAN-VAE model addresses the severe problem of missing modalities in rare disease cases by employing a conditional generation strategy to complete missing modalities. Its input consists of a random noise vector and existing modality embeddings (such as gene sequences and text descriptions), and its output is synthetic data of the corresponding modality. The RareGAN-VAE model employed in this invention is based on a well-known architecture combining a variational autoencoder (VAE) and a generative adversarial network (GAN), incorporating a conditional generation mechanism to address the severe problem of missing modalities in rare disease cases. The model structure itself is not the novelty of this invention; rather, it serves as a supporting completion module in the multimodal graph construction process, adapting to the practical problem of incomplete imaging or genetic data in rare disease scenarios. The model input consists of a random noise vector and existing modality embeddings (such as gene sequences and text descriptions), and its output is synthetic data of the target modality (such as images or missing text), which is used to improve cross-modal alignment and the integrity of graph node information.

[0018] To ensure the quality and clinical reliability of the generated data, a three-level verification mechanism is introduced: (1) For image modalities, the FID (Frechet Inception Distance) is used to calculate the similarity between the synthetic images and the real case distribution; Frechet Inception Distance (FID) is an index used to evaluate the quality of images generated by a generative model (such as GANs, VAEs). It measures the similarity between the generated images and the real images in the feature space by comparing the distribution difference. The lower the FID, the closer the quality and diversity of the generated images to the real data. (2) For the synthetic mutation sites of the gene modality, match the known pathogenic mutations in authoritative databases such as ClinVar to verify their clinical rationality; (3) For the synthetic case text and structural description, semantic consistency detection is performed, and three specialists with clinical experience in rare diseases conduct blind review and scoring. According to the five-point scoring system, only when the average score is greater than 4 points is the data determined to be usable. All verification results and generation parameters are stored in the metadata to form a traceable synthetic data meta record.

[0019] In the field of rare diseases, not only is the sample extremely small, but often the modalities are incomplete, such as the lack of medical images, missing genetic sequencing, incomplete clinical text description, etc. In this case, it is necessary to use a generative model (such as RareGAN-VAE) to complete the missing modalities, but it also carries a higher risk: once the synthetic data quality is not high, it may mislead the subsequent knowledge graph construction, relationship extraction, and even diagnosis reasoning results.

[0020] Therefore, this application puts forward higher reliability requirements for the completion results, not only requiring structural "likeness", but also semantic "reasonableness" and clinical "usability". Under this high requirement background, a single verification method (such as only using automatic indicators) is far from enough. Therefore, the invention designs a three-level verification mechanism to ensure the quality of the generated data from three angles: 1. Statistical evaluation (such as FID) - to ensure that the generated results are similar in structure to the real distribution; 2. Knowledge rule comparison (such as ClinVar matching) - to ensure that the generated content conforms to medical semantics; 3. Expert manual blind review - finally judged by clinical doctors whether it has real usability.

[0021] The advantages of this mechanism are: Comprehensive coverage of the structure, semantics, and actual application layer of the generated data; Three barriers superimposed, greatly reducing the risk of false information entering the graph; Verification results are quantifiable, recordable, and traceable, which is beneficial to data governance and continuous optimization.

[0022] The feature extraction and unified dimension conversion of the modal data respectively specifically include: Feature extraction is performed on text, image, and gene data respectively, and the original representations of the three types of modalities are mapped to a unified dimension, and a modal representation head is designed to extract a global semantic representation vector in a unified format; Text modality (Text-MRH): Bio_ClinicalBERT encoder is used to obtain clinical text semantic representation, and a 256-dimensional vector is output after aggregating the context vector; Bio_ClinicalBERT is a BERT model specially pre-trained for biomedical and clinical text, trained by Stanford University and Google Research team based on clinical notes (such as MIMIC-III) and PubMed literature. It optimizes medical entity recognition, relationship extraction and question answering tasks, and is an important tool for clinical natural language processing (NLP).

[0023] Image modality (Img-MRH): Based on 3D ResNet, global anatomical features are extracted, combined with U-Net to segment the lesion area, and finally output a 256-dimensional vector after dimension reduction processing by a lightweight convolutional layer; 3D ResNet is a three-dimensional extension of traditional ResNet (Residual Network), which is specifically used to process 3D input such as video, medical images (such as CT / MRI), and spatio-temporal sequence data. Its core idea is to solve the gradient vanishing problem of deep network through residual connection (Skip Connection), and capture spatio-temporal features through 3D convolution. U-Net is a classic convolutional neural network (CNN) for medical image segmentation.

[0024] Gene modality (Gene-MRH): The ClinVar annotated pathogenic mutation information is fused with the HGNC encoded gene site embedding vector, and a multi-layer perceptron (MLP) is used for encoding to output a 256-dimensional representation vector; HGNC (International Human Genome Nomenclature Committee) is an authoritative gene naming organization that provides unique and standardized names for genes, gene families, and non-coding RNAs in the human genome.

[0025] Let them be respectively: Text modality representation vector : ; Image modality representation vector : ; Gene modality representation vector : ; Wherein, , 、 text, image and gene respectively.

[0026] Preferably, the modality injection mechanism in S3 specifically comprises: introducing global semantic representation of other modalities to each modality representation vector to enhance cross-modality semantic perception ability; for the original representation vector of modality , the enhanced representation vector is defined as: ; wherein, is a modality injection adjustment coefficient, represents a global vector of other modalities.

[0027] Preferably, the cross-modality alignment module in S3 specifically comprises: using representation mapping based on bidirectional attention mechanism to map enhanced representation vectors of text, image and gene three types of modalities to shared semantic space to realize feature alignment between modalities; the input of the cross-modality alignment module in text, image and gene three types of modality features is respectively: text modality enhanced representation vector : ; image modality enhanced representation vector : ; gene modality enhanced representation vector : ; wherein, , , are the number of text , the number of image , the number of gene mutation sites, is the unified embedding dimension; for any modality pair , the attention weight mapping from modality to modality is defined as: ; ; ; wherein, are learnable projection parameters, represents a real matrix space; is the modality query matrix with dimensions ; are projection parameters; is a key matrix of modalities with dimensions ; are projection parameters; T denotes the transpose of a matrix, in the attention mechanism, is transpose of for computing the dot product between and ; represents the unified embedding dimension; and represent the enhanced representation vectors of modalities and softmax denotes the normalization function, is the attention weight matrix; further computes the cross-modal aligned representation under this mapping: ; ; where are learnable projection parameters; is the weighted sum of value vectors from modalities to modalities ; is the attention weight matrix; is the value vector of modality ; represents the enhanced representation vector of modality ; is the projection matrix of value vectors; denotes a real matrix space; for the consistency and reversibility of enhanced alignment, introduce path-independent consistency loss : ; where is the th mapping of modality to modality ; corresponds to the aligned result after the corresponding reverse mapping back to modality ; measures the alignment between modalities and alignment consistency loss, and The superscript represents the generalization variable of the modal, which is a general term for three modalities 、 、 , and are selected from 、 、 ; Finally, sum all modal combinations to get the overall alignment consistency loss : ; wherein, is the overall alignment consistency loss; is the alignment consistency loss of modal and modal ; is the alignment consistency loss of modal and modal ; is the alignment consistency loss of modal and modal .

[0028] Preferably, the multi-level semantic fusion in step S4 specifically includes: Using a multi-level fusion module FusionBLK to perform deep semantic fusion on the aligned modal features; the module is composed of fusion blocks (Fusion Block), denotes the total number of fusion modules or the number of fusion blocks, and each layer contains four steps of modal quality weighting, sequence alignment, encoder fusion and residual connection.

[0029] The aligned representation of each modality is dynamically adjusted according to its metadata quality weight: ; ; ; wherein, is the aligned modal feature sequence, is the modal quality score; is the aligned modal feature sequence, is the modal quality score; Is the aligned modal Feature sequence, is modal Quality rating; is the enhanced text modality representation vector; is the enhanced image modality representation vector; is the enhanced gene modality representation vector; Adjust modal sequences of different lengths to a consistent size and unify them to the same length : ; is the enhanced text modality representation vector; is the enhanced image modality representation vector; is the enhanced gene modality representation vector; Indicates a The real matrix space of The meaning is to unify the sequence length, represents the unified embedding dimension; The features of the three modalities are concatenated and fed into the fusion encoder: ; in," " means press Dimensional splicing; It is Joint modality representation vector of layer fusion; is the index of the fusion layer; Indicates a The real matrix space of represents the uniform sequence length, represents the unified embedding dimension; is the enhanced text modality representation vector; is the enhanced image modality representation vector; is the enhanced gene modality representation vector; The concatenated multimodal representation is fed into a fusion encoder module (which can be a Transformer, Restormer, or Gated MLP): ; ; where, is the output of the previous layer, initially a zero vector ; =0, i.e., a zero vector of dimension ; denotes the uniform sequence length, represents the uniform embedding dimension; denotes the fusion encoder; is the joint modality representation vector of the layer fusion; is the index of the fusion layer; After layer iterations, the output of the layer fusion representation is aggregated into a uniform fusion semantic vector using Mean Pooling or Attention Pooling: ; where, is the fusion semantic vector; Pooling is the pooling operation; is the real space, and the vector set has a dimension of represents the uniform embedding dimension; is the fusion representation of the layer, is the index of the fusion layer; Mean Pooling is mean pooling, and Attention Pooling is attention pooling.

[0030] Preferably, the construction of the double-layer structure system of the rare disease knowledge graph including the ontology layer and the instance layer in step S5 specifically includes: The ontology layer is used to represent standard medical knowledge in the field of rare diseases, including disease concepts, clinical symptoms, pathogenic genes, and treatment plans, and the nodes come from authoritative medical ontology libraries, including Orphanet, UMLS, and HGNC. The instance layer is used to store specific case data and its corresponding multi-modal information, and the nodes include case instance nodes, text nodes, image nodes, and gene nodes, all of which have unique identifiers and modal attributes. The ontology layer and the instance layer are connected through semantic edges including instance_of, aligned_to, has_symptom, and has_gene, forming a hierarchical tree graph structure.

[0031] Among them, instance_of indicates that a specific entity is a specific instance of a certain general category, aligned_to indicates the correspondence of two entities in different systems or standards, has_symptom indicates the relationship between disease and symptoms, and has_gene indicates the relationship between disease, trait or organism and related genes.

[0032] Preferably, the edge relationships such as modal reference, semantic enhancement and traceability path established in the graph in step S5 specifically include: The modal fusion semantic vector generated in step S4 is used The instance nodes in the graph are embedded and represented, and the structural connection relationship between multiple modalities is constructed; the original modal data nodes are explicitly modeled, and a traceable path between the modal ontology and the instance is constructed; The structural connection relationship between multiple modalities is constructed by using the attention weight between modalities learned by the MultiAlign module to judge the semantic coupling strength between modalities; MultiAlign is an algorithm or software module for multiple sequence alignment (MSA) and is widely used in the fields of bioinformatics, genomics, protein structure prediction, etc. Its core goal is to align multiple biological sequences (DNA, RNA or protein), identify conserved regions, variant sites and evolutionary relationships.

[0033] If the weight exceeds the set threshold, a cross-modal semantic edge is automatically generated, specifically including: semantically_related_to: when the attention weight or fusion correlation between modal pairs exceeds the set threshold, a semantic association edge is automatically established between modal nodes; is_similar_to: when the cosine similarity of the fusion representation between two case nodes is lower than the pre-set distance threshold , it is determined that the semantic similar cases are generated is_similar_to edge; textual_equivalent_to: when the alignment confidence between the original text and the standardized UMLS term is higher than the set confidence threshold, the textual_equivalent_to edge is generated; the above edge relationship is automatically calculated and generated by the weight matrix and fusion vector space distance in the modal alignment module.

[0034] Preferably, the knowledge graph system supports multi-source heterogeneous query mode, respectively facing structured query and semantic vector retrieval, and the deployment structure includes: a graph database module for storing structured entity nodes, semantic edge relationships and modal reference paths, supporting graph query, path reasoning and subgraph extraction operations based on the Cypher query language, which is a declarative query language for Neo4j graph databases; a vector database module for storing the fusion semantic vectors of each instance node , supporting vector retrieval and case recall based on semantic similarity;

[0035] an object storage system module for storing raw modal data files, including structured clinical text, medical images and genomic data of rare disease patients, and providing path backtracking capability, supporting mapping and retrieval with graph node attributes.

[0036] The three modules cooperatively constitute the storage, indexing, retrieval and calling infrastructure of the multi-modal knowledge graph, supporting downstream applications for intelligent diagnosis, semantic question answering and clinical reasoning.

[0037] The knowledge graph system constructed by the present application integrates two types of heterogeneous query mechanisms, structured graph query and vectorized semantic retrieval, which are respectively deployed in the graph database and vector database modules, and have the following advantages: (1) Graph database (structured path query) - support semantic rule reasoning and relationship path search: Based on Neo4j + Cypher query language, multi-hop query, instance tracing and subgraph extraction are supported; complex rule queries can be implemented, such as: rare disease X → has_gene → GENE_Y → is_similar_to → patient P; case B → has_image → CT image → textual_equivalent_to → symptom S.

[0038] Corresponding to the habit of "rule-based diagnostic logic" of doctors, it is suitable for scenarios such as etiology tracking, gene candidate screening and disease clustering.

[0039] (2) Vector database (semantic similarity query) - support intelligent retrieval and recommendation in semantic vector space: Based on Milvus or Faiss, support vector Top-K retrieval and similar case recall; the query method can be: input the fusion vector of the case → retrieve similar disease; only input text / image / gene fragment → quickly complete similar graph node in modal collaboration.

[0040] The vector database improves the accessibility of the graph under unstructured input, and is suitable for intelligent question answering, auxiliary diagnosis and data-driven discovery tasks.

[0041] The above-mentioned scheme of the present application has the following beneficial effects: (1) The rare disease knowledge graph is constructed by selecting text, image and gene multi-modal medical information, the application range, coverage and accuracy of the knowledge graph are improved, the correspondence adaptation of rare disease clinical diagnosis and treatment can be solved, and the identification ability is strong; (2) When the complex multi-modal medical information is extracted and modality fusion is carried out, the modality injection means and cross-modal alignment module are selected, the representation alignment of different information is fully realized, and a unified fusion semantic vector is constructed through a multi-level semantic fusion step, so that the rare disease knowledge graph can be stably constructed, and the knowledge graph data obtained has strong discrimination ability; (3) The knowledge graph construction of the present application includes a double-layer structure system of ontology layer and instance layer, and establishes edge relationships such as modality reference, semantic enhancement and traceability path in the graph, which brings great convenience to graph query application, has high accessibility, strong practicability, and can be widely applied in clinical assistance of rare disease treatment; (4) The three-layer metadata structure system composed of disease type layer, case layer and modality layer is adopted, which effectively supports the full-link mapping and management from disease ontology, individual case to original modality data, enhances the traceability and quality control ability of modality data while realizing clear organization of graph structure, and has good expansibility and iterative training support; (5) The knowledge graph system constructed by the present application fuses the dual-channel query mechanism of graph database and vector database, supports both structured semantic path reasoning and similarity retrieval based on modality fusion vector, can meet the diversified needs of complex logical query and intelligent semantic recall, and shows high flexibility and practicability in application scenes such as clinical auxiliary diagnosis, semantic question answering and case recommendation.

[0042] Other beneficial effects of the present application will be described in detail in the subsequent specific embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 The overall flowchart of the modality injection and multi-modal fusion rare disease knowledge graph construction method of the present application is shown in the figure. Figure 2 The detailed flowchart of the multi-modal data acquisition and standardization processing in S1 step of the present application is shown in the figure. DETAILED DESCRIPTION

[0044] In order to make the technical problems, technical schemes and advantages of the present application clearer, specific embodiments will be described in detail below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor belong to the scope of protection of the present application.

[0045] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", "third" are only for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0046] In the description of the present application, it should be noted that unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connection" should be understood broadly, for example, it can be a locking connection, or a detachable connection, or an integral connection; it can be a mechanical connection, or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, or the internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0047] In addition, the technical features involved in the different embodiments of the application described below can be combined with each other as long as they do not conflict with each other.

[0048] The rare disease knowledge graph construction method based on modal injection and multi-modal fusion comprises the following steps: (1) Obtain multi-source heterogeneous medical data, collect structured clinical text of rare disease patients from hospital information system, medical image data of PACS system, and whole genome or targeted mutation data of genomic sequencing platform, and perform desensitization processing and modal format standardization; (2) Based on the three-layer metadata structure (disease layer Schema, case layer Instance, and modal layer Evidence) constructed, the above data is uniformly organized and semantically modeled to form a standardized multi-modal data system that can be parsed; (3) In view of the problem of missing modal data in the collected data in the three-layer metadata structure, an improved RareGAN-VAE generation model is used for conditional modal completion, the input is the embedding fragment feature of the existing modal in the three-layer metadata structure and the noise vector, and the output is the synthesized modal data, and a three-way verification mechanism is constructed to evaluate its credibility; The original feature representations of the three modalities, text, image, and gene, were extracted and mapped to a unified global semantic modality representation vector using Bio_ClinicalBERT, 3D ResNet + U-Net, and MLP models. A modality injection mechanism was introduced to enhance semantic interaction between modalities. The enhanced representation vector after modality injection was input into the MultiAlign alignment module to construct a bidirectional path-independent attention mapping mechanism. A path consistency loss function was introduced for alignment learning in a shared semantic space. (4) Use the multi-level residual fusion module FusionBLK to fuse the three types of aligned features, and construct the final unified fusion semantic vector through modality quality weighting, sequence length alignment, encoder fusion, and residual connection; (5) Construct a multimodal medical knowledge graph structure consisting of an ontology layer and an instance layer. Rely on authoritative knowledge bases such as Orphanet, UMLS, and HGNC to establish standard semantic nodes and their structural relationships, and use modality alignment output vectors to embed the instance layer; automatically generate semantically enhanced edges based on the attention weights and semantic similarity of modality fusion, such as semantically_related_to, textual_equivalent_to, is_similar_to, etc. edge relationships; deploy the graph on heterogeneous platforms of graph databases, vector databases, and object storage systems, and support dual-channel retrieval and reasoning services based on structural paths and semantic vectors.

[0049] Furthermore, step (1) is as follows: Figure 2 Detailed flowchart of multimodal data acquisition and standardization processing in step S1 of the present invention; Multimodal medical data originates from hospital clinical information systems (HIS), picture archiving and communication systems (PACS), and gene sequencing centers. Structured clinical text data includes basic patient information, diagnostic records, course of illness descriptions, diagnosis descriptions, treatment descriptions, treatment plans, and follow-up records. Medical imaging data includes typical examination results such as chest CT and brain MRI, stored in DICOM format. Genomic data includes information on variant sites in target regions or the entire genome, typically represented as VCF files. After collection, this multimodal data undergoes irrelevant information cleaning, named entity desensitization, and unified file format conversion. A physical-semantic mapping relationship is established through path indexing and semantic labeling, and stored in a multimodal metadata framework to facilitate cross-modal parsing and call-up in subsequent tasks.

[0050] All standardized data will be uniformly organized and stored in the following "three-layer metadata structure": the Schema layer defines disease ontology, standard disease coding, core symptom terminology, related pathogenic genes, treatment options, and other standard medical entities, achieving standardized expression and unified reference of upper-layer medical knowledge; the Instance layer carries specific individual cases, including patient basic information, diagnosis and treatment timeline, family medical history, etc., building a bridge from standard knowledge to actual cases; the Evidence layer records the multi-modal original data path and semantic label, feature vector, and quality score corresponding to each case, providing a basis for graph reasoning and traceability. This structure effectively avoids the problem of "flat structure and flat entity" in traditional graphs, enabling the graph to have three-layer semantic capabilities of ontology reusability, instance extensibility, and modality traceability.

[0051] The Evidence layer explicitly models each data path, supporting path mapping between original text, images, gene files, and semantic entities; combined with the modality alignment module to generate semantic enhanced edges (such as semantically_related_to, is_similar_to, textual_equivalent_to), enabling automatic generation of semantic links between heterogeneous data. Support doctors or systems to query the graph nodes to trace back the original data source and quality indicators.

[0052] Further, in step (3): The RareGAN-VAE model addresses the serious problem of modality missing in rare disease cases by using a conditional generation model to complete the missing modality. The input consists of a random noise vector and existing modality embedding fragments (such as gene sequences and text descriptions) from the three-layer metadata structure, and the output is synthetic data for the corresponding modality. For modality missing cases, the RareGAN-VAE conditional generation model is used to complete the missing modality data, and a three-level verification mechanism including automatic evaluation, semantic rule verification, and expert scoring is used to mark its credibility; The model takes the fragment features of existing modalities and noise vectors as input, and the synthetic data of the missing modality as output. Automatic evaluation uses Frechet Inception Distance to evaluate image similarity, and uses the ClinVar database to verify gene mutations and the matching of UMLS definitions between genes and disease phenotypes; expert scoring is performed by double-blind scoring by experts with clinical experience in rare diseases, with an average score not less than the pre-set threshold; To ensure the quality and clinical reliability of the generated data, a three-level verification mechanism is introduced: (1) For image modalities, the FID (Frechet Inception Distance) is used to calculate the similarity between the synthetic images and the distribution of real cases; (2) For synthetic mutation sites in the gene modality, match known pathogenic mutations in authoritative databases such as ClinVar to verify their clinical reasonableness; (3) For the semantic consistency of synthetic case texts and structural descriptions, three specialists with clinical experience in rare diseases conduct blind review and scoring, and only when the average score is greater than 4 points on a five-point scoring system is the data considered usable. All verification results and generation parameters are stored in the metadata to form a traceable synthetic data meta record.

[0053] The three-level verification mechanism is mainly based on the high complexity of the missing modalities of rare diseases and the key influence of the completion results on the accuracy and explainability of the knowledge graph construction. Compared with other diseases, rare disease case data is generally characterized by sample scarcity, incomplete modalities, and strong heterogeneity, and genomic or imaging information often plays a decisive role in diagnosis. Therefore, when using a generative model to complete modalities, more stringent quality assurance methods are needed to prevent low-quality synthetic data from negatively affecting the graph structure and diagnosis reasoning.

[0054] The three-level verification mechanism adopted by the present application starts from three dimensions of structural quality, semantic compliance, and clinical reasonableness, and constructs a verification system of "statistical evaluation + knowledge rules + manual judgment", effectively realizing the comprehensive screening and reliable control of the generated samples. The specific advantages are: first, to enhance the clinical safety of the generated data and avoid semantic bias; second, to improve the semantic consistency and structural matching degree of the completed samples; third, to realize the quantifiable and traceable management of synthetic data quality; fourth, to fully adapt to different modal characteristics and ensure the scientificity and effectiveness of the evaluation method.

[0055] Further, in step (3): To alleviate the problem of dimensional and semantic space deviation of multi-modal feature representation, first, the original representations of text, image, and gene modalities are mapped to a unified dimension, and a modality representation head (Modality Representation Head, MRH) is designed to extract a unified format of global semantic modality representation vector.

[0056] Text modality (Text-MRH): Bio_ClinicalBERT encoder is used to obtain the semantic representation of clinical text, and a 256-dimensional vector is output after aggregating the context vector; Img-MRH: Based on 3D ResNet to extract global anatomical features, combined with U-Net to segment the lesion area, and finally output a 256-dimensional vector through a lightweight convolutional layer for dimension reduction processing; Gene-MRH: The ClinVar annotated pathogenic mutation information is fused with the HGNC encoded gene site embedding vector, and is encoded through a multilayer perceptron (MLP) to output a 256-dimensional representation vector.

[0057] Respectively denoted as: Text modality representation vector : ; Image modality representation vector : ; Gene modality representation vector : ; Among them, , , respectively represent the text, image and gene three modalities.

[0058] In order to improve the perception ability of each modality to the semantics of other modalities, a modality injection mechanism is introduced. For any modality , its original representation vector , the enhanced representation vector is defined as: ; Among them, is the modality injection adjustment coefficient, which regulates the cross-modality semantic injection strength, represents the global vector of other modalities.

[0059] Although the text, image and gene three modalities are all mapped into 256-dimensional vectors, due to the significant differences in their original structures, information density and semantic expression methods, there are still many challenges in the joint training process. First, the semantic space distribution between the three types of modalities is offset, and direct fusion can easily cause semantic mismatch. Second, the model structures used by each type are quite different. Bio_ClinicalBERT has complex parameters and slow training, 3D ResNet has dense features and fast convergence, Gene-MLP model is simple and its gradient is easy to be suppressed, resulting in different training rhythms and unstable fusion effect. In addition, the gene modality itself has sparse semantics, which is easy to be "overwhelmed" by other modalities in the fusion, affecting the expression of its diagnostic value.

[0060] In order to solve the above problems, the present application makes targeted optimization on the training strategy. First, a normalization module (LayerNorm + ) to unify the feature scale; secondly, to set the learnable coefficient in the modal injection mechanism , dynamically adjust the semantic injection intensity; third, dynamic weighting of modal loss is used during joint training to keep the gradient of each modality balanced; fourth, a freeze-fine-tuning strategy is adopted for Bio_ClinicalBERT and 3D ResNet, and high-level parameters are updated only in the fusion stage to reduce the training-dominant bias.

[0061] Furthermore, in step (3): After completing the modality injection, each modality obtains the global semantic cues of other modalities; We further propose a cross-modal alignment module MultiAlign, which is used to map the enhanced representation vectors of three types of modalities: text, image, and gene, to a shared semantic space, to achieve feature alignment between modalities, and to realize semantic consistency and interoperability between modalities.

[0062] The inputs of the three types of modal features in the cross-modal alignment module are: Text Modality Enhanced Representation Vector : ; Image modality enhanced representation vector : ; Gene Modality Enhanced Representation Vector : ; in, , , Text Numbers and images Number, number of gene mutation sites, To unify the embedding dimension; For any mode pair , which is from the modal To Modal The attention weight map is defined as follows: ; ; ; in, is the learnable projection parameter, Indicates a The real matrix space of ; is modal The query matrix has the same dimension as consistent; yes The projection parameters of is modal The key matrix of consistent; yes The projection parameters of The T in represents the transpose of the matrix. In the attention mechanism, yes The transpose of and The dot product between represents the unified embedding dimension; and Represents the mode and modal The enhanced representation vector of Softmax represents the normalization function, is the attention weight matrix; Further calculate the cross-modal representation alignment under this mapping: ; ; in, is the learnable projection parameter; From the modal To Modal The weighted sum of the value vectors of ; is the attention weight matrix; is modal A vector of values ​​of ; Indicates modality The enhanced representation vector of is the projection matrix of the value vector; Indicates a The real matrix space of ; In order to improve the reversibility and stability of the aligned structure, a “path independence” loss term is introduced, requiring that the mode 𝑚 should be close to the original representation after being aligned with the mode 𝑛 and then back-projected. The loss is defined as follows: ; in, is modal No. indivual Mapping to modality the expression; After the corresponding reverse mapping, return to the modal The alignment result of Measuring modality and modal Alignment consistency loss, and The superscript represents the generalization variable of the modal, which is a generalization of the three modalities 、 、 , and are selected from 、 、 ; Finally, sum all the modal combinations to get the overall alignment consistency loss : ; Wherein, is the overall alignment consistency loss; is the alignment consistency loss of modal and modal ; is the alignment consistency loss of modal and modal ; is the alignment consistency loss of modal and modal .

[0063] Further, in step (4): The application adopts a multi-level fusion module FusionBLK to perform deep semantic fusion on the aligned modal features. The module is composed of fusion blocks, denotes the total number of fusion blocks or the number of fusion blocks, and each layer includes four steps of modal quality weighting, sequence alignment, encoder fusion and residual connection.

[0064] The aligned representation of each modal is dynamically adjusted according to its metadata quality weight: ; ; ; Wherein, is the aligned modal feature sequence, is the modal quality score (derived from metadata); is the aligned modal feature sequence, is the modal quality score (derived from metadata); is the aligned modal character sequence, is the modal quality score (derived from metadata); is the enhanced text modal representation vector; is the enhanced image modal representation vector; is the enhanced gene modal representation vector; Adjust the modal sequence of different lengths to a consistent size (such as by average pooling, upsampling, or zero padding), and unify it to a length : ; is the enhanced text modal representation vector; is the enhanced image modal representation vector; is the enhanced gene modal representation vector; represents a real matrix space; The meaning is to unify the sequence length, represents the unified embedding dimension; Concatenate the features of the three modalities and send them to the fusion encoder: ; where, " represents concatenation in dimension; is the layer fusion joint modal representation vector; is the index of the fusion layer; represents a real matrix space, represents the unified sequence length, represents the unified embedding dimension; is the enhanced text modal representation vector; is the enhanced image modal representation vector; is the enhanced gene modal representation vector; Send the concatenated multi-modal representation to a fusion encoder module (which can be a Transformer, Restormer, or Gated MLP): ; ; where, is the output of the previous layer, initially a zero vector ; =0, i.e., a zero vector of dimension ; denotes the uniform sequence length, represents the uniform embedding dimension; denotes the fusion encoder; is the joint modality representation vector of the th fusion layer; is the index of the fusion layer; After layer iterations, the output of the th fusion representation is aggregated into a uniform fusion semantic vector using Mean Pooling or Attention Pooling: ; where, is the fusion semantic vector; Pooling is the pooling operation; is the real space, and is a vector set of dimension represents the uniform embedding dimension; is the fusion representation of the th layer, is the index of the fusion layer; Mean Pooling is the mean pooling, and Attention Pooling is the attention pooling.

[0065] The multi-level fusion module (FusionBLK) proposed in the present application has obvious advantages compared with the existing conventional modal splicing or linear weighted fusion method. First, FusionBLK is based on multiple fusion units, and clearly divides the modal fusion process into four sub-steps: quality weighting, sequence alignment, splicing encoding, and uniform aggregation. This step-by-step strategy not only ensures the structural clarity of the fusion process, but also improves the perception and adaptation ability of each step to the differences between modalities.

[0066] Secondly, the modal quality weight , to realize adaptive adjustment of the contribution proportion of modal expression according to the quality of original data, and to avoid interference of low-quality modal on the overall expression. In the sequence alignment stage, the length of each modal sequence is unified through pooling or padding, so as to solve the semantic misplacement problem caused by inconsistent information length. In the encoding stage, the three-modal features are spliced and then sent to an optional deep fusion network (such as Transformer, Restormer or Gated MLP), and a residual connection structure is added to effectively alleviate the problems of gradient disappearance and information loss in deep fusion, and to enhance the expression stability and fusion depth.

[0067] Finally, the fusion output is aggregated into a unified vector through Mean Pooling or Attention Pooling, which not only improves the compactness of semantic expression, but also preserves the complementary semantic relationship between modalities.

[0068] Further, in step (5): A two-layer structure system of rare disease knowledge graph including an ontology layer and an instance layer is constructed. The ontology layer is used to represent standard medical knowledge in the rare disease field, including disease concepts, clinical symptoms, pathogenic genes and treatment plans. The nodes come from authoritative medical ontology libraries, including Orphanet, UMLS and HGNC. The instance layer is used to store specific case data and its corresponding multi-modal information. The nodes include case instance nodes, text nodes, image nodes and gene nodes, all of which have unique identifiers and modal attributes. The ontology layer and the instance layer are connected through semantic edges including instance_of, aligned_to, has_symptom and has_gene, forming a hierarchical tree graph structure. Among them, instance_of represents that a specific entity is a specific instance of a general category, aligned_to represents the corresponding relationship between two entities in different systems or standards, has_symptom represents the relationship between disease and symptom, and has_gene represents the relationship between disease, trait or organism and related genes.

[0069] The establishment of modal reference, semantic enhancement and traceability path edge relationships specifically includes: The knowledge graph supports the generation of semantic enhancement edge relationships based on modal fusion representation, specifically including: semantically_related_to: when the attention weight or fusion correlation between modal pairs exceeds a certain threshold, a semantic association edge is automatically established between modal nodes; is_similar_to: when the cosine similarity of the fusion representation between two case nodes is lower than the pre-set distance threshold , it is determined as a semantically similar case, and an is_similar_to edge is generated; textual equivalent to: when the alignment confidence between the original text and the standardized UMLS term is higher than the set confidence threshold, generate the textual equivalent to edge; the above edge relationship is automatically calculated and generated by the weight matrix and the fusion vector space distance in the modal alignment module.

[0070] The knowledge graph system supports a dual-channel query service oriented to structural relationships and semantic spaces, and is deployed on a heterogeneous system architecture. A graph database (Neo4j) is used to store ontology layer and instance layer nodes, entity attributes and semantic edge relationships, and supports graph query, path reasoning and subgraph extraction operations based on the Cypher query language. A vector database (Milvus) is used to store all fusion vectors , which realizes semantic similarity calculation and vector neighbor retrieval, and can be used for recommending similar cases or modal collaborative retrieval; An object storage system (Ceph) is used to save original modal data files, including medical images (such as DICOM), gene files (such as VCF), and text data (such as electronic medical record XML / JSON), which realizes fast positioning and data retrieval through path indexing in the graph node attribute. The three of them cooperatively constitute the storage, indexing, retrieval and calling infrastructure of the multi-modal knowledge graph, supporting downstream applications oriented to intelligent diagnosis, semantic question answering and clinical reasoning.

[0071] Embodiment 1 As Figure 1 shown, the embodiment of the present application provides a rare disease knowledge graph construction method based on modal injection and multi-modal fusion, comprising the following steps: S1 Multi-modal data acquisition and standardization processing For the construction of the rare disease multi-modal knowledge graph, the present embodiment first acquires and standardizes multi-source heterogeneous medical data, establishes a unified data expression system, and ensures the accuracy and consistency of subsequent graph construction.

[0072] S11 Multi-modal medical data acquisition In this step, by interfacing with various medical information systems, structured and unstructured data related to rare diseases are collected, mainly including the following three types of modal information: (1) Structured clinical text data Through the hospital information system (HIS), the core medical record information of the patient is obtained, including but not limited to: patient basic demographic information (gender, age, etc.); diagnosis record (initial diagnosis time, diagnosis conclusion, diagnosis code); treatment process record (treatment plan, drug use, follow-up summary, etc.).

[0073] (2) Medical image data Patient-related medical image data is collected through a picture archiving and communication system (PACS), mainly including: CT images; MRI images; all data is uniformly retained in DICOM format.

[0074] (3) Genomic sequencing data Genomic information of the patient is obtained from a gene sequencing center, and the data sources include: whole genome sequencing (WGS); targeted region sequencing (such as rare disease related gene panel); annotation content includes HGNC standard naming, ClinVar pathogenicity annotation, low frequency mutation site information, etc.

[0075] S12 Standardization processing of multi-modal data To ensure that data from different sources can be integrated and calculated, the multi-modal data described above is uniformly standardized in this embodiment, including the following aspects: (1) Text data processing: using regular expressions to clean up redundant content with disordered format and irrelevant semantics; desensitizing sensitive information such as patient name and ID number; retaining core clinical semantic content to ensure professionalism and data security.

[0076] (2) Image and gene data format standardization: medical image data is uniformly converted to a storage format that meets the DICOM protocol; gene data is converted to standard VCF / FASTA format, with HGNC naming and ClinVar annotation.

[0077] (3) Unified three-layer metadata structure storage: all standardized data will be uniformly organized and stored in the following "three-layer metadata structure": Schema Layer: used to define the standard medical knowledge system, covering disease ontology, clinical symptom terminology, pathogenic gene and treatment plan, etc. It includes standard disease name (ORPHA code, ICD-10 code), standardized symptom terminology (based on UMLS, HPO), pathogenic gene identification (HGNC standard symbol, ClinVar pathogenic annotation) and standard treatment plan (ATC classified drugs, operation code, etc.). This layer serves as the knowledge baseline for the case layer and the modal layer, ensuring data consistency and scalability across institutions and data sources.

[0078] Instance Layer: carries specific individual case instances, records the basic demographic information of the patient after desensitization (gender, age, family history), diagnosis and treatment timeline (initial diagnosis date, diagnosis date, follow-up record) and medical events (such as examination items, treatment behavior, drug use, etc.). The case layer serves as a bridge between standard knowledge and actual data, supporting dynamic expansion, asynchronous supplementation and time series modeling of cases, ensuring traceability of the association between instance data and standard knowledge.

[0079] Evidence Layer: Detailed records of each case corresponding to the multi-modal original data and its derived features, including structured text data (such as medical record description file path and entity labels), medical image data (DICOM image path, lesion mask file, image metadata such as resolution and device model), genomic data (VCF / FASTA format file, mutation annotation table), and feature vector representation extracted from each modality; The Evidence Layer synchronously maintains data quality indicators (such as text term coverage, image FID score, and genetic mutation pathogenicity score), and supports bidirectional mapping between original data path and graph nodes to meet the traceability and quality control requirements.

[0080] S2: Synthesis and verification mechanism for missing modal data. To address the common problem of missing multi-modal information in rare disease clinical data (such as missing image or genetic sequencing data), this embodiment proposes a multi-modal data synthesis method based on an improved RareGAN-VAE, and designs a multi-level verification mechanism to ensure the authenticity and clinical usability of the synthesized data.

[0081] S21 Conditional generation of multi-modal data. Due to the scarcity of case samples in actual rare disease data sets, there are often cases of missing partial modalities; To improve data integrity and model training effectiveness, this embodiment uses a conditional generation model RareGAN-VAE to complete the data based on a small number of labeled cases.

[0082] (1) Generator input design: The synthesis model takes random noise variable z and partial data of existing modalities as input, for example: Text description + partial gene sequence → synthesized image; Text description + image features → synthesized gene fragments.

[0083] (2) Cross-modal attention mechanism: The model introduces a cross-modal attention module to fully utilize the semantic information carried in the remaining modalities when generating the target modality. For example: When a case is missing a chest image, the model automatically synthesizes a CT image with characteristic rib cartilage lesions based on the genetic mutation information (such as UBA1 p.Met41Val) and clinical description (such as "recurrent chondritis") it carries.

[0084] S22 Three-level synthesis data verification mechanism: To ensure the clinical credibility of the generated data, this embodiment designs a three-level verification mechanism including automatic evaluation and manual review: (1) Automated similarity and consistency evaluation: Image quality evaluation: use Frechet Inception Distance (FID) index to evaluate the distribution difference of synthetic images and real images in feature space; Text semantic consistency detection: use medical pre-training language model (such as Med-BERT) to calculate the embedding similarity of original text and synthetic text, and eliminate samples with semantic deviation; Gene mutation rationality evaluation: analyze the generated gene mutation site, verify whether it exists in ClinVar or OMIM database, and ensure that the generated variation has known clinical relevance or pathogenic potential.

[0085] (2) Cross-modal semantic rule checking mechanism: Gene-disease association rule: the synthetic gene mutation should have explicit association (through mutation site or gene pathway) with the target disease in ClinVar / HGNC database; Image-pathological feature matching rule: if the disease clearly indicates a certain pathological sign (such as "costal cartilage thickening") in Radiopaedia or UMLS, the synthetic image should appear corresponding features in the target area (judged by deep model segmentation).

[0086] (3) Expert manual evaluation mechanism: randomly select 20% of synthetic samples; invite at least 3 rare disease doctors with clinical diagnosis experience to independently score the medical rationality and mutual coordination of text, image and gene data in synthetic samples; use 5-point scale score, get the average score, if the average is ≥ 4.0, mark as "clinically credible".

[0087] S23 Credibility labeling and iterative optimization: After the synthetic data is verified, it will be labeled and archived according to the quality grading mechanism, and will participate in the subsequent iterative optimization process.

[0088] Data labeling structure: Each synthetic data record is attached with metadata, including: synthetic_rank: quality level (A / B / C); fid_score: image generation quality index; expert_consensus: expert score average.

[0089] Model iteration mechanism: add the verified synthetic data to the training set; update the RareGAN-VAE model parameters; re-execute the synthesis and verification process after each iteration to gradually improve the generation quality.

[0090] Data archiving and storage: all synthetic data and verification metadata are stored in object storage system (such as Ceph); stored separately from real data, and through labeling to achieve traceability and controllability.

[0091] S31 Multi-modal data analysis and feature representation extraction; This step parses the three types of modal data, text, image, and gene, from the standardized three-layer metadata structure, extracts their semantic features, and represents them as fixed-length vectors.

[0092] (1) Text modality processing: Structured symptom descriptions and clinical text are extracted from the Schema layer. Context encoding is performed using the pre-trained medical language model Bio_ClinicalBERT. The UMLS clinical terminology database is combined to filter non-medical entities and retain standard terminology. To further enhance the medical expression of text semantics, a symptom keyword positioning mechanism and a synonym expansion strategy are introduced to improve term recognition rate and semantic clarity. The final text representation is a 256-dimensional semantic vector that combines contextual information and standardized medical term aggregation characteristics.

[0093] (2) Medical imaging modality processing: DICOM format medical imaging data is loaded according to the path recorded in the Evidence layer. Global anatomical features are extracted using 3D ResNet-50. Potential lesion regions are segmented based on the U-Net model, and a saliency mask is generated. This mask guides the network to focus on key lesion regions and generate local enhanced representations. The FID score is introduced as a quality scoring indicator to dynamically adjust feature credibility. After fusion, a 512-dimensional image vector is obtained.

[0094] (3) Gene Modality Processing: The HGNC codes and ClinVar annotations recorded in the Evidence layer are parsed, and the mutation sites are embedded using a multi-layer perceptron (MLP). The pathogenicity level weights and KEGG pathway context representations are integrated to enhance the biological semantic interpretation capability of the gene representation. To meet privacy protection requirements, a mild Gaussian perturbation is added to the output vector. The resulting gene representation is a 128-dimensional embedding vector.

[0095] (4) Modal representation unification and standardization: Modal representation heads (MRH) are introduced to perform unified conversion on modal vectors of different dimensions: Text modality → Text-MRH (BERT feature aggregation) Imaging modality → Img-MRH (ResNet+U-Net feature compression) Gene Modality → Gene-MRH (MLP embedding) After mapping, the three modalities are unified into 256-dimensional representation vectors as input for cross-modal alignment.

[0096] S32 cross-modal semantic alignment and fusion mechanism design; Modal representation unification and injection mechanism: To alleviate the dimension and spatial offset problems of each modality representation, first, the unified dimension mapping and modality semantic injection of different modalities are performed: (1) Modality representation unified design Text modality (Text-MRH): The Bio_ClinicalBERT encoder is used to obtain the semantic representation of the clinical text, and a 256-dimensional vector is output after aggregating the context vector; Image modality (Img-MRH): Based on 3D ResNet, the global anatomical features are extracted, combined with U-Net to segment the lesion area, and finally output a 256-dimensional vector through lightweight convolutional layer dimension reduction processing; Gene modality (Gene-MRH): The ClinVar annotated pathogenic mutation information is fused with the HGNC encoded gene site embedding vector, and the multi-layer perceptron (MLP) is used for encoding, and a 256-dimensional representation vector is output.

[0097] Denoted as: Text modality representation vector : ; Image modality representation vector : ; Gene modality representation vector : ; Wherein, , , represent the text, image and gene three modalities respectively.

[0098] (2) Modality injection mechanism To improve the perception ability of each modality to the semantics of other modalities, a modality injection mechanism is introduced. For the original representation vector of the modality , its enhanced representation vector is defined as follows: ; Where is the modality injection coefficient, which controls the cross-modality semantic injection strength, represents the global vector of other modalities.

[0099] Cross-modality alignment module (MultiAlign) After the modality injection mechanism is completed, each modality feature has obtained the global semantic prompt of other modalities, forming the enhanced representation vectors of text, image and gene three types of modality features , , ; Further proposed is a cross-modal alignment module MultiAlign, aiming to map the enhanced representations of the three modalities into a unified shared semantic space, so as to achieve semantic consistency and interoperability between modalities. The module adopts a bidirectional path-independent attention mechanism, and introduces an alignment consistency loss to ensure that the representations between different modalities tend to be symmetric in structure and semantically matched.

[0100] (1) The three modalities feature inputs are respectively: Text modality enhanced representation vector : ; Image modality enhanced representation vector : ; Gene modality enhanced representation vector : ; Wherein, , , are the number of text , the number of image , and the number of gene mutation sites, respectively, is the unified embedding dimension; For any pair of modalities , the attention weight mapping from modality to modality is denoted as , which is defined by the following formula: ; ; ; Wherein, are learnable projection parameters, represents a real matrix space; is the query matrix of modality , with the same dimension as ; is the projection parameter of ; is the key matrix of modality , with the same dimension as ; is the projection parameter of ; T in the above formula represents the transpose of the matrix, and in the attention mechanism, is the transpose of , used to calculate the dot product between and ; representing unified embedding dimensions; and represent enhanced representation vectors of modalities and respectively; softmax represents a normalization function, is an attention weight matrix; further calculate the cross-modality alignment representation under the mapping: ; ; wherein, is a learnable projection parameter; is a weighted sum of value vectors from modality to modality ; is an attention weight matrix; is a value vector of modality ; represent enhanced representation vectors of modalities ; is a projection matrix of value vectors; represent a real matrix space of ; (3) To enhance the stability and reversibility of the alignment process, this module introduces a "path independence" constraint, requiring consistent alignment between two modalities in both directions. The specific loss function is defined as follows: ; wherein, is the th mapping of modality to modality ; corresponds to the alignment result after the corresponding reverse mapping back to modality ; measures the consistency loss of modality and modality alignment, and the subscript represents the general variable of the modality, which is a general term for the three modalities , , , and are respectively selected from , , ; summing over all modality combinations, the overall alignment consistency loss is obtained: ; in, is the overall alignment consistency loss; is modal and modal Alignment consistency loss; is modal and modal Alignment consistency loss; is modal and modal Alignment consistency loss.

[0101] At the same time, to enhance the targeted modeling of each modality’s semantics in the alignment phase, the following strategies are introduced: Text modality: for medical terms with high matching Allocate higher attention to improve the accuracy of semantic alignment between symptoms and genes, and symptoms and images; Imaging modality: Use the saliency mask of the lesion area to generate a local guidance vector and participate in the construction of the attention matrix; Gene modality: The attention distribution is adjusted based on the pathogenicity level of the mutation site and the KEGG pathway participation, so that disease-related genes are preferentially matched with corresponding clinical manifestations and imaging signs.

[0102] Multi-level fusion module (FusionBLK) After completing the semantic alignment of multimodal features, the present invention further introduces a multi-level fusion module FusionBLK to integrate the alignment features from the three modalities of text, image and gene to generate a structured unified representation of the fusion semantic vector , to support subsequent entity embedding and reasoning tasks in the knowledge graph. This module adopts a multi-level residual aggregation structure and combines the modality quality score in the metadata to achieve weight control, ensuring that the fusion results have both structural integrity and semantic consistency.

[0103] The FusionBLK module consists of Fusion Block Indicates the total number of layers or fusion blocks of the fusion module. Layer by layer, modality aggregation and semantic enhancement are performed. Each layer includes the following steps: (1) Dynamically adjust the alignment representation of each modality according to its metadata quality weight: ; ; ; in, Is the aligned modal Feature sequence, is modal Quality score; (derived from metadata) Is the aligned modal Feature sequence, is modal Quality score; (derived from metadata) Is the aligned modal Feature sequence, is modal Quality score; (derived from metadata) is the enhanced text modality representation vector; is the enhanced image modality representation vector; is the enhanced gene modality representation vector; (2) Adjust the modal sequences of different lengths to a consistent size (e.g., by average pooling, upsampling, or zero padding) and unify them to a length of : ; is the enhanced text modality representation vector; is the enhanced image modality representation vector; is the enhanced gene modality representation vector; Indicates a The real matrix space of The meaning is to unify the sequence length, represents the unified embedding dimension; (3) Concatenate the features of the three modalities and send them to the fusion encoder: ; in," " means press Dimensional splicing; It is Joint modality representation vector of layer fusion; is the index of the fusion layer; Indicates a The real matrix space of represents the uniform sequence length, represents the unified embedding dimension; is the enhanced text modality representation vector; is the enhanced image modality representation vector; is the enhanced gene modality representation vector; (4) The spliced multi-modal representation is sent to a fusion encoder module (which can be a Transformer, Restormer, or Gated MLP): ; ; wherein, is the output of the previous layer, initially a zero vector ; = 0, i.e., a zero vector with a dimension of ; denotes a uniform sequence length, represents a uniform embedding dimension; denotes the fusion encoder; is the joint modality representation vector of the layer fusion; is the index of the fusion layer; After layer iterations, the fusion representation of the layer is output , which is aggregated into a unified fusion semantic vector using Mean Pooling or Attention Pooling: ; wherein, is the fusion semantic vector; Pooling is the pooling operation; is a real number space with a dimension of a vector set, represents a uniform embedding dimension; is the fusion representation of the layer, is the index of the fusion layer; Mean Pooling is mean pooling, and Attention Pooling is attention pooling.

[0104] At the same time, to perceive the individualized differences of the enhanced modality fusion expression, the following optimization mechanisms are designed: Text modality: structured term coverage and paragraph semantic confidence are introduced as attention factors during fusion; Image modality: dynamic weight correction is performed based on lesion area distribution density and FID quality score; Gene modality: information such as pathway involvement, disease severity, and site quantity density is fused to improve the semantic proportion of disease explanation feature vectors.

[0105] S4 Construction method of multi-modal medical knowledge graph, construct multi-modal medical knowledge graph for rare disease field, adopt double-layer hierarchical structure design, including ontology layer (Schema layer) and instance layer (Instance layer), realize structured semantic mapping from general medical knowledge to specific case data.

[0106] S41 Hierarchical structure design of ontology layer and instance layer (1) Ontology layer (Schema layer), the ontology layer is used for describing the general medical knowledge structure related to rare diseases, which has high stability and standardization characteristics, and the node types include: (2) Disease concept node: cited from Orphanet coding system, records standard disease name and ORPHA code; Clinical symptom node: cited from UMLS clinical terminology system, standardizes the representation of disease manifestations; Pathogenic gene node: cited from HGNC naming system, combined with ClinVar annotation to represent known pathogenic mutations; Therapeutic means node: such as drugs, surgical methods, reference WHO ATC classification; Relationship between nodes: including has_symptom (disease-symptom), related_gene (disease-gene), treated_by (disease-treatment), etc.

[0107] (2) Instance layer (Instance layer), the instance layer is used to express the specific case information actually collected or generated, which is independent of the ontology layer structure and dynamically expanded, each node has a unique identifier, and records key metadata such as source modality, data quality level, path information and whether it is a synthetic sample in the attribute. Node types include: Case instance node: represents a single patient sample, carries attributes such as gender, age, and visit time; Text modality node: structured medical record description, symptom text; Imaging modality node: index node of CT, MRI and other medical image files; Gene modality node: gene mutation record, such as VCF file or HGVS expression; Structural symptom node: parsed standardized UMLS term entries.

[0108] S42 Fusion vector embedding and instance mapping mechanism, use the modality fusion semantic vector generated in S3 step , embed the instance nodes in the graph, and construct the structural connection relationship between multi-modal: (1) Vector embedding identification: embed the fusion semantic vector Semantic representation as a case node; vectors are bound to vector databases (such as Milvus) through unique identification, supporting vectorized retrieval; when writing the graph, the vector ID is taken as an instance node attribute field.

[0109] (2) Semantic enhanced edge relationship automatic generation: use the inter-modal attention weight learned by the MultiAlign module to judge the semantic coupling strength between modes; if the weight exceeds the set threshold, automatically generate cross-modal semantic edges, such as: semantically_related_to: text symptom Gene mutation; is_similar_to: high semantic similarity relationship between case nodes.

[0110] S43 structure rule driven and semantic perception driven relationship construction, the strategy of the present application for constructing graph edge relationship is divided into two categories: (1) Structure rule driven path Based on the pre-defined entity relationship template of the medical knowledge base, the static semantic connection edge is constructed, such as: has_symptom: disease → symptom (according to the Orphanet symptom mapping table); has_gene: disease → gene (according to ClinVar mutation attribution); instance_of: case node → disease ontology node; All such relationships can be automatically generated by templates during graph writing stage, with fixed structure and stable logic.

[0111] (2) Semantic perception driven path Combine the semantic information in the modal fusion and alignment process to dynamically judge the potential semantic correlation and generate the following edge types: is_similar_to: fusion vector distance less than threshold Determine that it is a semantically similar case; aligns_with_image: the weight between gene features and image features is the largest; textual_equivalent_to: when the semantic alignment degree between unstructured text and standard UMLS terms is higher than the set confidence threshold, it is automatically generated.

[0112] S44 modal node and traceability path design, in order to enhance the explainability and data traceability of the graph, the present embodiment explicitly models the original modal data node and constructs the "traceable path" between modal entities and instances: (1) The modal node has the following characteristics: Has independent identifiers (such as text_id, image_id, gene_id); Recorded in node attributes: physical path, modal category, quality score, and whether it is synthetic; Establish modal edges with corresponding case nodes, such as has_image, has_gene, and has_symptom.

[0113] (3) Path organization logic: Modality trace path: such as case node → has_image → DICOM file node, which can be used for data retrieval; Symptom-gene-image path: through has_symptom → textual_equivalent_to → related_gene → aligns_with_image to realize modal collaborative semantic reasoning; Modality aggregation path: calculate the semantic aggregation degree of multiple modal nodes under the same instance, generate a fused_representation edge, and realize modal question answering and intelligent retrieval.

[0114] S45 Knowledge graph deployment and service interface, the multi-modal knowledge graph constructed in the embodiment supports two types of heterogeneous query methods, respectively facing the structure logic and the vector space, and the deployment structure is as follows: Graph database (Neo4j): store structured nodes and semantic relationships and modal reference paths; support Cypher query, path reasoning, subgraph extraction and other operations.

[0115] Vector database (Milvus): store fused semantic vectors Support vector-based similar case recall and semantic association reasoning.

[0116] Object storage system (Ceph): store original modal data files (CT images, gene sequences, structured text); provide the ability to quickly locate underlying data through graph paths.

[0117] S46 The query and auxiliary diagnosis method of the multi-modal rare disease knowledge graph constructed by the application is suitable for the case that the input forms of users are different in actual clinical scenarios, supports two paths of structured query and semantic vector retrieval, and has good adaptability and interpretability.

[0118] In practical applications, users may input standardized structured medical information, such as specific symptom terms, rare disease names, gene sites, etc., at which time the system enters a structured query mode. This mode is based on a Neo4j graph database for graph access. The system first performs entity recognition and standardization mapping on user input, mapping keywords to entity node types in the graph, such as Disease, Gene, or Symptom. The system then automatically generates a Cypher query statement based on the graph relationship structure and calls the graph database interface to perform retrieval operations. For example, if the user input keyword is “IKBKG”, the system will construct the query instruction: “MATCH (d:Disease)-[:ASSOCIATED_WITH]->(g:Gene{name:’IKBKG’}) RETURN d.name, g.name”, and return all disease information associated with the gene in the database. The query result includes the basic attributes of the hit node, the relationship type between nodes and its semantic weight, and the front-end interface supports graph structure visualization, path highlighting, node detail expansion, edge attribute viewing, and other interactive functions, making it easy for professional users to trace and verify and provide clinical knowledge support.

[0119] When the user input is unstructured text in natural language description, the system automatically switches to a semantic vector retrieval mode. The system first performs semantic analysis and embedding coding on the input text, using Bio_ClinicalBERT and other medical field pre-training models to map the text into a fixed-length semantic vector representation. During the graph construction stage, the system has pre-generated embedding vector representations for all key nodes such as diseases, symptoms, and genes and stored them. The query vector will be compared with the embedding vectors of all entity nodes in the graph for similarity, usually using cosine similarity as the similarity measure function, and returning the Top-K similar nodes as the matching results. Then, the system will perform path backtracking based on the hit nodes, extract their upstream and downstream knowledge structures, construct multi-hop semantic chains containing symptoms-diseases-genes, and provide visual display and explanation of the path graph.

[0120] In the query result output stage, the system outputs a structured result set regardless of the path used, including: matching entity name, entity type, semantic matching confidence, path structure relationship, node attribute description, data source reference (such as ClinVar, Orphanet ID), etc. The system supports path export, result download, and original data traceable access.

[0121] To evaluate the adaptability and effectiveness of the method in this embodiment under actual clinical conditions, a test set consisting of 50 real or simulated rare disease cases was constructed, covering a variety of input formats, including text input, gene input, and natural language questions. The system achieved 86% accuracy for top-1 queries and 94% accuracy for top-3 queries. Physicians with experience in rare disease diagnosis also conducted a blind review, rating the output results on their rationality, semantic interpretation, and clinical reference value. The average score was above 4.6 / 5, validating the system's high adaptability and promising application in rare disease assisted search tasks.

[0122] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for constructing a rare disease knowledge graph based on modality injection and multimodal fusion, characterized in that: The following steps are involved: S1. Collect multimodal medical information including text, images and genes; S2, standardize multimodal medical information and build a three-layer metadata structure; S3. Complete missing modal data by performing feature extraction and unified dimensionality conversion on the modal data to obtain a global semantic modal representation vector in a unified format. Use the modal injection mechanism to process the modal representation vector to enhance cross-modal semantic perception and obtain an enhanced representation vector. Cross-modal representation alignment in the shared semantic space is achieved through the cross-modal alignment module. S4, performing multi-level semantic fusion on the modal features after cross-modal representation alignment to obtain a unified fused semantic vector; S5. Based on the unified fusion semantic vector, a rare disease knowledge graph with a two-layer structure system including an ontology layer and an instance layer is constructed, and side relationships such as modal reference, semantic enhancement and traceability path are established in the graph.

2. The method according to claim 1, wherein The texts, images and genes include structured clinical texts, medical images and genomic data from patients with rare diseases, and the genomic data include low-frequency mutation sites and their pathogenicity annotation information.

3. The method according to claim 1, wherein The three-layer metadata structure described in S2 includes a disease layer, a case layer, and a modality layer. The disease layer is used to record standard disease codes, core symptom terms, and related pathogenic gene identifiers; the case layer is used to record basic patient information, diagnosis and treatment timeline, and family medical history; and the modality layer is used to record modality data paths, semantic labels, feature vectors, and quality scores.

4. The method according to claim 3, wherein The missing modality data completion described in S3 specifically includes: For missing modalities, the missing modality data is completed based on the RareGAN-VAE conditional generative model, and its credibility is marked using a three-level verification mechanism including automatic evaluation, semantic rule verification, and expert scoring. RareGAN-VAE is a hybrid generative model that combines variational autoencoders and generative adversarial networks. The model takes segment features and noise vectors of existing modalities as input and synthesized data of missing modalities as output. The automatic evaluation uses Frechet Inception Distance to assess image similarity. Frechet Inception Distance is an indicator used to evaluate the quality of images generated by the generative model. The ClinVar database is used to verify the matching of gene mutations and UMLS definitions of genes and disease phenotypes. UMLS is the Unified Medical Language System. The semantic rule verification includes gene-disease association rules and image-pathology feature matching rules. The gene-disease association rule stipulates that the synthesized gene mutation should be explicitly associated with the target disease in the ClinVar / HGNC database. The image-pathology feature matching rule stipulates that if a disease clearly indicates a certain pathological sign in Radiopaedia or UMLS, the synthesized image should show the corresponding feature in the target area. Radiopaedia is an encyclopedia of medical imaging and a teaching platform. The expert scoring is performed double-blind by experts with clinical experience in rare diseases, and the average score is no less than a preset threshold. The feature extraction and unified dimension conversion of the modal data specifically include: Feature extraction is performed on text, image, and genetic data respectively. The original representations of the three modalities are mapped to a unified dimension. A modality representation head is designed to extract a global semantic modality representation vector in a unified format. Text modality: The Bio_ClinicalBERT encoder is used to obtain the semantic representation of clinical text, aggregating the context vector and outputting a 256-dimensional vector. Bio_ClinicalBERT is a BERT model pre-trained specifically for biomedical and clinical text. Imaging modality: 3D ResNet is used to extract global anatomical features, combined with U-Net to segment lesion areas, and lightweight convolutional layers are used for dimensionality reduction, ultimately outputting a 256-dimensional vector. 3D ResNet is a three-dimensional extension of the traditional residual network, while U-Net is a classic convolutional neural network used for medical image segmentation. Genetic modality: The pathogenic mutation information annotated by ClinVar is fused with the HGNC coding gene locus embedding vector, encoded through a multi-layer perceptron, and outputs a 256-dimensional representation vector; Denoted as: Text modality representation vector : ; Image modality representation vector : ; Genetic modality representation vector : ; in, 、 、 They represent the three modalities of text, image and gene respectively.

5. The method according to claim 1, wherein The modal injection mechanism described in S3 specifically includes: The global semantic representation of other modalities is introduced into each modality representation vector to enhance cross-modal semantic perception capabilities; For modal The original representation vector , whose enhanced representation vector It is defined by the following formula: ; in, is the modal injection adjustment coefficient, Represents the global vector of other modes.

6. The method according to claim 5, wherein The cross-modal alignment module in S3 specifically includes: A representation mapping method based on a bidirectional attention mechanism is used to map the enhanced representation vectors of three modalities: text, image, and gene, into a shared semantic space, achieving feature alignment between the modalities. The inputs of the three types of modal features in the cross-modal alignment module are: Text Modality Enhanced Representation Vector : ; Image modality enhanced representation vector : ; Gene Modality Enhanced Representation Vector : ; in, , , Text Numbers and images Number, number of gene mutation sites, To unify the embedding dimension; For any mode pair , which is modal To Modal The attention weight map is defined as follows: ; ; ; in, is the learnable projection parameter, Indicates a The real matrix space of ; is modal The query matrix has the same dimension as consistent; yes The projection parameters of is modal The key matrix of consistent; yes The projection parameters of The T in represents the transpose of the matrix. In the attention mechanism, yes The transpose of and The dot product between represents the unified embedding dimension; and Represents the mode and modal The enhanced representation vector of Softmax represents the normalization function, is the attention weight matrix; Calculate the cross-modal alignment representation under this mapping: ; ; in, is the learnable projection parameter; From the modal To Modal The weighted sum of the value vectors of ; is the attention weight matrix; is modal A vector of values ​​of ; Indicates modality The enhanced representation vector of is the projection matrix of the value vector; Indicates a The real matrix space of ; To enhance the consistency and reversibility of alignment, path-independent consistency loss is introduced : ; in, is modal No. indivual Mapping to modality the expression; After the corresponding reverse mapping, return to the modal The alignment result of Measuring modality and modal Alignment consistency loss, and The subscript represents the generalized variable of the mode, which is the generalization variable of the three modes. 、 、 A general term for and Selected from 、 、 One of the following; Finally, all modal combinations are summed to obtain the overall alignment consistency loss : ; in, is the overall alignment consistency loss; is modal and modal Alignment consistency loss; is modal and modal Alignment consistency loss; is modal and modal Alignment consistency loss.

7. The method according to claim 1, wherein The multi-level semantic fusion in step S4 specifically includes: The multi-level fusion module FusionBLK is used to perform deep semantic fusion on the aligned modal features; this module consists of The fusion blocks are composed of Indicates the total number of layers or fusion blocks of the fusion module. Each layer includes four steps: modality quality weighting, sequence alignment, encoder fusion, and residual connection: Dynamically adjust the alignment representation of each modality according to its metadata quality weight: ; ; ; in, Is the aligned modal Feature sequence, is modal Quality rating; Is the aligned modal Feature sequence, is modal Quality rating; Is the aligned modal Feature sequence, is modal Quality rating; is the enhanced text modality representation vector; is the enhanced image modality representation vector; is the enhanced gene modality representation vector; Adjust modal sequences of different lengths to a consistent size and unify them to the same length : ; is the enhanced text modality representation vector; is the enhanced image modality representation vector; is the enhanced gene modality representation vector; Indicates a The real matrix space of The meaning is to unify the sequence length, represents the unified embedding dimension; The features of the three modalities are concatenated and fed into the fusion encoder: ; in," " means press Dimensional splicing; It is Joint modality representation vector of layer fusion; is the index of the fusion layer; Indicates a The real matrix space of represents the uniform sequence length, represents the unified embedding dimension; is the enhanced text modality representation vector; is the enhanced image modality representation vector; is the enhanced gene modality representation vector; The concatenated multimodal representation is fed into a fusion encoder module: ; ; in, Output of the previous layer, initially a zero vector ; =0, that is, an all-zero vector with dimension ; represents the uniform sequence length, represents the unified embedding dimension; represents the fusion encoder; It is Joint modality representation vector of layer fusion; is the index of the fusion layer; go through After layer iteration, the output Layer fusion representation , use Mean Pooling or AttentionPooling to converge into a unified fusion semantic vector: ; in, It is the fusion semantic vector; Pooling is the pooling operation; is a real number space with dimension Vector collection of represents the unified embedding dimension; It is The fusion representation of the layers, is the index of the fusion layer; Mean Pooling is mean pooling, and Attention Pooling is attention pooling.

8. The method according to claim 1, wherein The construction of a rare disease knowledge graph with a two-layer structure including an ontology layer and an instance layer in step S5 specifically includes: The ontology layer represents standard medical knowledge in the field of rare diseases, including disease concepts, clinical symptoms, pathogenic genes, and treatment options. Nodes are sourced from authoritative medical ontology libraries, including Orphanet, UMLS, and HGNC. The instance layer stores specific case data and its corresponding multimodal information. Nodes include case instance nodes, text nodes, image nodes, and gene nodes, all with unique identifiers and modality attributes. The ontology layer and instance layer are connected by semantic edges, including instance_of, aligned_to, has_symptom, and has_gene, forming a hierarchical fir tree graph structure. Among them, instance_of indicates that a specific entity is a specific instance of a general category, aligned_to indicates the corresponding relationship between two entities in different systems or standards, has_symptom indicates the relationship between disease and symptom, and has_gene indicates the relationship between disease, trait or organism and related genes.

9. The method according to claim 1, wherein The establishment of side relationships such as modal reference, semantic enhancement, and traceability path in the graph in step S5 specifically includes: Use the modality fusion semantic vector generated in step S4 , embed the instance nodes in the graph and build the structural connection relationship between multiple modalities; explicitly model the original modal data nodes and build a traceable path between the modal ontology and the instance; The construction of the structural connection relationship between multiple modalities uses the inter-modal attention weights learned by the MultiAlign module to determine the semantic coupling strength between modalities. MultiAlign is a multiple sequence alignment algorithm or software module. If the weight exceeds a set threshold, cross-modal semantic edges are automatically generated, specifically including: semantically_related_to: Automatically establish semantically related edges between modality nodes when the attention weight or fusion correlation between the modality pairs exceeds the set threshold; is_similar_to: When the cosine similarity of the fusion representation between two case nodes is lower than the preset distance threshold , it is determined to be a semantically similar case and an is_similar_to edge is generated; textual_equivalent_to: When the alignment confidence between the original text and the standardized UMLS term is higher than the set confidence threshold, a textual_equivalent_to edge is generated; the above edge relationship is automatically calculated by the weight matrix and fused vector space distance in the modality alignment module.

10. The method according to claim 1, wherein The knowledge graph supports multi-source heterogeneous query methods, targeting structured queries and semantic vector retrieval respectively. The deployment structure includes: The graph database module is used to store structured entity nodes, semantic edge relationships, and modal reference paths. It supports graph query, path reasoning, and subgraph extraction operations based on the Cypher query language; Cypher is the declarative query language of the Neo4j graph database; Vector database module, used to store the fused semantic vector of each instance node , supports vector retrieval and case recall based on semantic similarity; The object storage system module is used to store raw modality data files, including structured clinical texts, medical images, and genomic data of patients with rare diseases. It also provides path tracing capabilities and supports mapping and retrieval of graph node attributes.

Citation Information

Patent Citations

  • Inner ear disease diagnosis model construction method and system, medium, product and terminal

    CN118571495A

  • Autism spectrum disorder knowledge graph construction method and system

    CN120163222A

  • Drug-disease relation prediction method and system based on dual-channel fusion knowledge graph

    CN120372560A

  • Intelligent urinary surgery diagnosis and treatment data processing system based on artificial intelligence

    CN120452830A

  • System and method for medical disease diagnosis by enabling artificial intelligence

    US20250069744A1

Cited By

  • Multi-modal health medical data management method and system based on cloud platform

    CN121075530A

  • Heart and cerebral vessel scientific research multi-modal data semantic alignment method and medium

    CN121191795A

  • Medical data processing system and processing method

    CN121393703A

  • A medical data processing system and processing method

    CN121393703B

  • Alzheimer's disease data completion method based on multi-modal generation and fusion

    CN121439269A