Mapping knowledge domain-based oral cavity multi-modal data fusion method

By constructing a unified semantic framework based on knowledge graphs, deep semantic integration of structured medical records, unstructured text, and 3D image data was achieved, solving the problems of semantic fragmentation and rigid integration in existing oral AI systems, and improving the reliability and security of intelligent assisted diagnosis and treatment.

CN122050673APending Publication Date: 2026-05-15CHINA UNIV OF MINING & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIV OF MINING & TECH
Filing Date
2026-04-17
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing oral AI systems suffer from semantic fragmentation, rigid integration, and evolutionary deficiencies in high-risk scenarios such as orthodontics, implantology, and maxillofacial surgery, making it difficult to achieve reliable, traceable, and evolvable intelligent assisted diagnosis and treatment based on multimodal data.

Method used

A unified semantic framework based on knowledge graphs is constructed, and dynamic comprehensive data views are generated through cross-modal alignment, confidence-weighted fusion, and conflict resolution, including deep semantic integration of structured medical records, unstructured text, and 3D image data.

Benefits of technology

It significantly improves the semantic consistency, traceability, and clinical credibility of multi-source information fusion, ensuring the stability and security of output results and supporting precise and safe diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122050673A_ABST
    Figure CN122050673A_ABST
Patent Text Reader

Abstract

The invention discloses an oral cavity multi-modal data fusion method based on a knowledge graph, and belongs to the technical field of oral medicine artificial intelligence and medical big data fusion. The method comprises the following steps: firstly, constructing a stomatology knowledge graph with a'top-layer multiplexing-specialized extension 'double-layer architecture, and forming a standardized semantic framework; then, multi-modal data such as structured medical records, unstructured texts and CBCT images of the patient are collected and subjected to quality control; then, analyzing and mapping each modal data into a standardized semantic representation associated with the knowledge graph by applying a natural language processing and computer vision technology; and finally, under a unified semantic framework, deep integration of multi-source information is realized through cross-modal semantic alignment, confidence weighted fusion and conflict resolution, and a patient dynamic comprehensive data view containing a three-dimensional panorama, a time axis and a semantic network is generated. According to the method, deep semantic understanding and collaborative utilization of oral multi-source heterogeneous data are realized, and a comprehensive and credible data basis is provided for precise diagnosis and treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and medical information processing technology, specifically involving a knowledge graph-based method for fusion of oral multimodal data. It is particularly suitable for orthodontics, implantology, and maxillofacial surgery scenarios, and is an intelligent auxiliary diagnosis and treatment system that performs semantic alignment, confidence-weighted fusion, and conflict resolution on structured medical records, unstructured text, and three-dimensional image data to generate a dynamic comprehensive data view. Background Technology

[0002] With the development of artificial intelligence and medical information processing technologies, knowledge graphs are being widely explored for high-risk oral clinical decision-making scenarios such as orthodontics, implantology, and maxillofacial surgery. Leveraging their structured semantic expression and logical reasoning capabilities, knowledge graphs can integrate diverse data such as medical records, images, and guidelines to support precision diagnosis and treatment. However, oral clinical decision-making involves multimodal heterogeneous data, interdisciplinary knowledge, and dynamic temporal sequences. Traditional methods often suffer from semantic fragmentation, quality control issues, and rigid evolution, making it difficult to build reliable, traceable, and evolvable intelligent assistance systems.

[0003] Existing dental AI methods mostly rely on single-modality or task models, such as rule-based engine mapping and encoding, CNN image segmentation, and keyword-based text retrieval. While effective for local tasks, they generally suffer from three major drawbacks when facing complex decisions: First, semantic anchors are lacking, and text entities cannot be aligned with image regions, leading to "different names for the same disease"; second, the fusion mechanism is crude, simply weighting and merging results without considering modality confidence differences and clinical logic constraints; and third, conflict resolution is lacking, and when text descriptions contradict image features, there is a lack of intelligent judgment mechanisms based on knowledge graph context and clinical rules, raising questions about the reliability of the output.

[0004] In recent years, technologies such as multimodal alignment, contrastive learning, and graph neural networks have been introduced into the medical field, making progress in improving representation consistency. However, they still have significant limitations in the oral cavity setting: they lack a unified semantic framework covering "patient-examination-entity-relationship-feature," achieving only local modality pairing; confidence weights are mostly fixed values, without dynamic adjustment based on data quality scores and entity link confidence; more importantly, a continuous optimization loop from fusion results → clinical validation → atlas update → model retraining has not yet been established, resulting in stagnant system capabilities that cannot evolve with real-world data.

[0005] Therefore, there is an urgent need for an intelligent assistance method that can achieve multimodal semantic unification, adaptive confidence fusion, intelligent conflict resolution, and continuous knowledge evolution. This method should use a knowledge graph as the core semantic hub, breaking down data barriers between structured medical records, free text, and 3D images. Through cross-modal alignment, weighted fusion, and rule verification, it should generate a dynamic comprehensive view that includes spatiotemporal dimensions and semantic relationships, ultimately meeting the core requirements of modern precision medicine for data credibility, reasoning interpretability, system security, and personalized services. Summary of the Invention

[0006] This invention provides a knowledge graph-based method for fusion of multimodal oral data. The method constructs a cross-modal alignment mechanism within a unified semantic framework, implements a confidence-driven weighted fusion strategy, and establishes a conflict resolution and dynamic verification closed loop. This supports deep semantic integration of structured medical records, unstructured text, and 3D image data in high-risk oral care scenarios such as orthodontics, implantology, and maxillofacial surgery, generating a traceable, interpretable, and evolvable comprehensive patient data view.

[0007] To achieve the above objectives, this invention provides a knowledge graph-based method for fusing oral multimodal data, comprising the following steps:

[0008] S1. Construct a knowledge graph in the field of oral medicine. The knowledge graph is based on a standardized medical terminology system and adopts a two-layer architecture of "top-level reuse - specialty expansion". It includes entities such as teeth, periodontal tissues, diseases, treatments, and imaging examinations, as well as semantic relationships describing the clinical logical relationships between entities.

[0009] S2. Collect multimodal oral medical data of the target patient and perform quality control. The data includes at least structured electronic medical records, unstructured clinical text, and medical imaging data including CBCT images.

[0010] S3. The multimodal oral medical data is parsed, and the information in the structured electronic medical records and clinical texts, as well as the visual features and structural information extracted from the medical image data, are mapped into standardized semantic representations associated with entities and relationships in the knowledge graph.

[0011] S4. Based on the patient's unique identifier, examination events, and the unified semantic framework provided by the knowledge graph, the standardized semantic representations from different modalities are associated and fused through cross-modal semantic alignment, confidence-weighted fusion, and conflict resolution to generate a comprehensive patient data view containing cross-modal semantic associations.

[0012] Furthermore, the steps in S1 for constructing a knowledge graph in the field of oral medicine specifically include:

[0013] S1.1 Construct a knowledge graph with a two-layer architecture of "top-level reuse - specialty extension". The top-level reuse layer reuses the international standardized medical terminology system as the basic framework through semantic mapping and inheritance mechanisms to ensure semantic interoperability between the graph and the general medical field. Under the constraints of the top-level framework, the specialty extension layer transforms implicit clinical experience into computable rules and relationships by structurally integrating authoritative specialty knowledge sources and collaborating deeply with clinical experts, thereby realizing the dynamic expansion and updating of in-depth specialty knowledge in oral medicine.

[0014] S1.2 Under the aforementioned two-layer architecture, define and instantiate the core entities constituting the knowledge graph; the entities include at least: anatomical structure entities, disease and pathology entities, treatment and operation entities, examination and imaging entities, and patient and time entities; each entity is identified by a globally unique identifier and assigned a set of attributes describing its characteristics; among them, the attributes of key anatomical structure entities include three-dimensional spatial coordinates and morphological quantitative parameters, the attributes of disease entities include standard codes, typical clinical manifestations, and imaging features, and the attributes of treatment entities include operation codes, indications, contraindications, and standard procedure sequences; the assignment of entity attributes must follow predefined clinical constraint rules, among which, for tooth mobility grades... The assessment is based on clinically measured buccal-lingual movement. and vertical motion Through a piecewise function The mapping is performed as follows:

[0015] ;

[0016] S1.3 Define and construct a set of semantic relations connecting the entities to form a knowledge network that conforms to clinical logic. The relations are directional semantic links and include at least: suffersFrom representing the disease condition, occurrenceIn representing the disease location, targets representing the treatment purpose, employees representing the examination method, depicts representing the image content, isPartOf representing the composition structure, and mayCause representing the causal risk.

[0017] Relationships are constructed through multi-modal collaboration, directly mapping from structured data. For unstructured text, an automated extraction model based on a pre-trained language model, using a joint entity and relation extraction model, is employed. The training objective of this model is a total loss. It minimizes the entity span recognition loss. Relationship classification loss Weighted sum:

[0018] ;

[0019] in, It is a process of balancing hyperparameters. All automatically extracted relationships must be reviewed by clinical experts and logical consistency is verified using descriptive logic axioms to ensure the correctness of the knowledge network.

[0020] S1.4. A dual strategy is adopted to represent and store the entities and relationships. On the one hand, entities are stored as nodes and relationships as edges, along with all their attributes, in a high-performance graph database to support complex topology queries and path analysis. On the other hand, knowledge graph embedding technology is used to map discrete entities and relationships to a continuous low-dimensional vector space to support semantic similarity calculation and link prediction. A rotation model is used for knowledge graph embedding, and its core relationship modeling formula is:

[0021] ;

[0022] in, Let be the vector representation of the head entity in the embedding space. For relation vectors, For the tail entity vector, Represents the Hadamard product between complex vectors;

[0023] The training process employs a self-adversarial negative sampling strategy to optimize the embedding model, and its loss function is defined as follows:

[0024] ;

[0025] in, It is a head entity Through relationships After rotation and tail entity The Euclidean distance between them It is the sigmoid activation function. It is the interval parameter. These are negative sample triples, obtained by randomly replacing the original triples. Generate the head or tail entity in the middle. This represents the number of negative samples in each round of sampling.

[0026] Furthermore, the step of collecting multimodal oral medical data in S2 specifically includes:

[0027] S2.1 Planning and quality control for multi-source data acquisition; First, based on the oral clinical diagnosis and treatment process and AI research and development needs, a unified data acquisition standard is formulated. This standard clarifies the acquisition standards for each modality of data: For CBCT images, the spatial resolution is specified to be no less than 0.2mm×0.2mm×0.2mm, and the radiation dose must comply with the ALARA principle; For intraoral scan data, the accuracy error is required to be less than 20 micrometers; For text medical records, the use of structured templates for key fields is mandatory.

[0028] Secondly, establish a quantitative evaluation mechanism for data quality, and calculate a quality score for each piece of collected data. For CBCT images, their quality score Contrast noise ratio Signal-to-noise ratio and motion artifact level The following was decided jointly:

[0029] ;

[0030] in, and For reference standard values, For artifact scoring calculated based on image gradient distribution, The weighting coefficients are satisfied. Only when At that time, the image data is incorporated into the subsequent processing flow;

[0031] For intraoral scan data, its quality score From surface integrity Noise density and geometric distortion rate The following was decided jointly:

[0032] ;

[0033] in, This is a standard completeness reference value. This is a reference value for the maximum allowable noise density. The geometric distortion rate is calculated by detecting the deviation of a standard sphere or a known distance. The weighting coefficients are satisfied. ;

[0034] For text data, its quality score Field fill rate Terminology standardization Number of logical contradictions The following was decided jointly:

[0035] ;

[0036] in, The key field fill rate is calculated by dividing the number of filled key fields by the total number of key fields. To determine the terminology standardization, the semantic similarity between the text and a standard medical dictionary is calculated using an NLP algorithm. The number of logical contradictions is the number of logical errors detected by the rule engine and measured after normalization. The weighting coefficients are satisfied. ;

[0037] S2.2. Data acquisition and enhancement of structured electronic medical records: Specifically, through the standardized interface of the hospital information system, batch acquisition of patients' demographic information, medical history, allergy history, diagnostic codes, surgical procedure codes, medication records, and laboratory test results; to solve the problem of coding differences between different hospital systems, a knowledge graph-based coding mapping and normalization algorithm is adopted. This algorithm utilizes the synonym expansion and hierarchical relationships of entities in the knowledge graph to construct a mapping matrix from local codes to standard terms. For a local encoding record Its corresponding standard encoding vector Calculated using the following formula:

[0038] ;

[0039] Among them, matrix Optimization is achieved by minimizing the cosine distance loss of semantic vectors before and after mapping to ensure semantic consistency;

[0040] S2.3. Data collection and privacy anonymization of unstructured clinical text: Specifically, free text fields such as chief complaint, present medical history, specialist examinations, and treatment plans from outpatient medical records are collected. Secondly, to protect patient privacy, a two-stage anonymization pipeline is designed. The first stage uses a rule-based and dictionary-based method to quickly identify and replace explicit identifiers. The second stage uses a fine-tuned biomedical BERT model for context-sensitive entity recognition, identifying and generalizing implicit identifiers and sensitive clinical events. The anonymized text... Privacy risk Defined in a given set of background knowledge In this case, the attacker uses a combination of quasi-identifiers to uniquely re-identify a specific patient, the probability of which... The k-anonymity of the dataset is estimated using the following formula:

[0041] ;

[0042] in, This represents the minimum number of records with the same combination of quasi-identifiers in the anonymized dataset.

[0043] S2.4. Acquisition and multimodal registration of medical image data: First, image data from different devices are acquired in parallel, including 3D volume data of the jawbone provided by CBCT, 3D mesh data of the dentition and soft tissue surfaces provided by an intraoral scanner, and 3D facial appearance models provided by a facial scanner. Second, a coarse-to-fine hybrid registration algorithm is used to achieve spatial alignment of multimodal images. This algorithm consists of two stages: The first stage is global rigid alignment, which, for hard tissues such as the jawbone, automatically detects stable anatomical landmarks through a pre-trained deep learning model and calculates the rigid transformation matrix. The second stage is local non-rigid fine-tuning, which, for the gingiva and facial soft tissues, minimizes the surface distance between the source and target images based on rigid alignment, and introduces bending energy constraints to solve the non-rigid deformation field to compensate for the geometric deformation of soft tissues under different imaging states. The overall optimization objective function is defined as follows:

[0044]

[0045] );

[0046] in, Represents a spatial transformation function. Let N be the set of anatomical landmarks detected in the source image. For the target image and The corresponding set of N paired marker points, It is a random sample from the surface of the source image. A set of points; Indicates the point The point closest to the Euclidean distance projected onto the surface of the target image. To balance the weighting coefficients, The distance is represented by Euclidean distance, and the optimization process uses the L-BFGS algorithm for iterative solution. This represents the deformation regularization term, used to constrain the smoothness of non-rigid transformations. These are the regularization weight coefficients;

[0047] S2.5 Collect time-series information and device metadata associated with the core data mentioned above. This information is associated with the core data through a globally unique three-level identifier of "patient-examination-sequence" to form a multimodal data package containing timestamps and collection context, providing support for subsequent dynamic analysis and tracing.

[0048] Furthermore, the step of parsing and acquiring multimodal oral medical data in S3 specifically includes:

[0049] S3.1. For the collected structured electronic medical record data, perform standardized mapping based on the entities and relationships defined in the knowledge graph: For standardized coding fields such as ICD / CPT, directly map them to the corresponding disease or treatment entities by querying the "coding-entity" correspondence table of the knowledge graph; For numerical and categorical fields, design a rule-based parser to convert them into attribute assertions defined in the knowledge graph, and finally output a structured set of attribute assertions as part of the standardized semantic representation.

[0050] S3.2 Deep semantic parsing and entity linking of unstructured clinical text: First, a pre-trained oral medicine language model is used to identify named entity references in the text. Subsequently, entity mentions are calculated based on contextual semantic similarity, clinical prior probability, and graph topological association strength. With candidate entities in the knowledge graph Link score, the formula for calculating link score is:

[0051]

[0052] ;

[0053] in, Mentioning entities in the text that are to be linked. Candidate entities in the knowledge graph. For entity mention The context vector of the sentence. For the semantic embedding vector of the candidate entity, The cosine similarity function is used. Candidate entities are obtained based on statistics from a large-scale medical corpus. The prior probability of occurrence, The set of disambiguated neighboring entities within the current text context window. These are the weighting coefficients;

[0054] Ultimately, if max( () greater than the preset confidence threshold If the candidate entity with the highest score is selected as the link result, then the mention is marked as an unregistered entity and not linked. Simultaneously, the pre-trained language model is used to encode the entire clinical text, and the hidden layer states corresponding to the [CLS] tags are extracted as the global semantic feature vector of the text. This is used for subsequent alignment with image features in the shared semantic space;

[0055] S3.3, Visual feature and structural information analysis of medical image data: First, an anatomical structure segmentation network based on 3D U-Net is constructed to extract voxel-level masks of the target anatomical structure. The training uses a joint loss function of Dice and cross-entropy, and the training loss formula is as follows:

[0056] ;

[0057] in, The Dice loss measures the spatial overlap between the predicted mask and the true mask. To optimize pixel-level classification accuracy using cross-entropy loss, For balancing weighting coefficients;

[0058] Secondly, a 3D Faster R-CNN network is used to detect pathological features within the region of interest defined by the mask, outputting the three-dimensional bounding box coordinates of the pathological lesions and their category probability distribution. Simultaneously, the detected pathological lesion regions are recorded as image entity regions. The ROI Align layer extracts feature map patches corresponding to each lesion region from the feature map, and then maps them into local visual feature vectors through a fully connected layer. It is used to characterize the fine-grained visual properties of lesion areas;

[0059] Subsequently, the segmented anatomical structure voxel data are input into a 3D residual network 3D-ResNet image encoder to extract global semantic feature vectors. Through a cross-modal contrastive learning mechanism, minimize The corresponding clinical text feature vector generated in step S3.2 The cosine distance between them, while maximizing the distance to negative samples within the batch, is the contrastive loss function. Defined as:

[0060] ;

[0061] in, This is the global semantic feature vector of the current sample image. To and Paired clinical text feature vectors, The total number of samples in a training batch. This is the text feature vector of the j-th sample in the batch. When j corresponds to the current sample, it is a positive sample; otherwise... All samples are considered negative samples. The cosine similarity function is used. This is a temperature coefficient used to adjust the smoothness of the similarity distribution, making the model pay more attention to the hard-to-bear samples;

[0062] Finally, the geometric parameters of the segmentation mask, the category labels of the pathological tests, and the aligned global feature vectors are concatenated to form a standardized multimodal semantic representation associated with the entities in the knowledge graph.

[0063] S3.4 Unify the encoding of the parsing results of structured medical records, clinical texts and medical images into standardized semantic representations: For each examination event, construct a structured data framework that includes a set of entity nodes, a set of attribute assertions and a set of relation triples;

[0064] Specifically, the text global feature vector generated in S3.2 Image global semantic feature vector extracted from S3.3 Modal alignment and fusion are performed to form a multimodal feature vector set, which is then attached as an attribute to the corresponding knowledge graph entity node. Finally, the data is encapsulated to form a structured patient-examination level multimodal data object, which serves as the basic input for cross-modal semantic reasoning and fusion in the subsequent step S4.

[0065] Furthermore, the step in S4 of generating a comprehensive patient data view that includes cross-modal semantic associations specifically includes:

[0066] S4.1. Based on the knowledge graph embedding space, achieve cross-modal semantic alignment and association: Project the entity feature vectors obtained from text and image parsing onto the unified vector space trained by the knowledge graph, calculate their weighted similarity with candidate entity vectors, and perform joint linking decisions; for entity mentions that have both text and image features, calculate their cross-modal coreference confidence. When the threshold is exceeded, they are determined to be the same real object, and their multimodal attributes are merged and integrated in the comprehensive view. Specifically, the cross-modal co-reference confidence is defined as:

[0067]

[0068] ;

[0069] in, For entity references in the text defined in S3.2, The image pathological lesion area detected in S3.3, for The context embedding vector, for The corresponding local visual feature vector, This represents the vector representation of candidate entities in the knowledge graph. The cosine similarity function is used. For the weighting coefficients, satisfying ;

[0070] like If the confidence level is greater than the preset confidence threshold, it is determined that the two refer to the same entity, and their multimodal attributes are merged to generate a unified representation.

[0071] S4.2 Confidence-weighted fusion and conflict resolution of cross-modal parsing results: For attribute assertions extracted from the same entity in different modalities, weighted aggregation is performed based on the confidence weight of their source modality; when attribute values ​​in different modalities conflict, knowledge graph constraints and clinical rules are introduced to resolve the conflict, ultimately generating a comprehensive attribute representation with enhanced consistency;

[0072] Specifically, for entities Attributes If the attribute is numeric, its merged value Defined as:

[0073] ;

[0074] If the attribute is categorical, a weighted voting mechanism is used to select the category with the largest sum of weights as the fusion result;

[0075] in, The attribute value provided for the k-th modality The confidence weights for this mode satisfy the following conditions: Weight The link confidence score is determined by both the modal quality score and the link confidence score calculated for the entity in S4.1. To provide the number of valid modalities for this attribute;

[0076] If a conflict exists, the conflict resolution module is triggered: first, it queries the context relationship of the entity in the knowledge graph, then it judges the priority in combination with the predefined clinical rule base, and finally outputs attribute values ​​that are consistent and in line with medical logic.

[0077] S4.3 Construct a comprehensive patient data view: Using the patient's unique identifier and examination event as keys, encapsulate the fused entity instances, attribute assertions, relation assertions, and multimodal feature vectors into a structured data packet header and inject it into the knowledge graph; at the same time, based on the relational paths defined in the knowledge graph, automatically deduce implicit semantic associations and generate a visual comprehensive view containing cross-modal semantic links;

[0078] Specifically, the integrated data packet header is defined as a five-tuple structure:

[0079] ;

[0080] in, As a globally unique identifier for the patient, This serves as the unique identifier for this inspection event. This is the set of attribute assertions after fusion by S4.2. For the set of assertions of relationships between entities, A multimodal feature vector set for key entities, including text context vectors. With image local feature vector This data package serves as the basic unit for subsequent clinical decision support and dynamic tracking, and supports interactive visualization based on time axis or spatial dimension.

[0081] S4.4 Output and validate the comprehensive patient data view: Write the encapsulated structured data package into the clinical decision support system and trigger the consistency verification module; this module performs logical consistency verification on the fusion result according to the constraint rules defined in the knowledge graph; if an anomaly is found, it is marked for manual review to ensure that the final output comprehensive view is reliable and usable in both semantics and clinical logic;

[0082] Specifically, the packet consistency check score is defined as:

[0083] ;

[0084] in, For the i-th knowledge graph constraint rule, This is an indicator function that indicates when a data packet meets the rules. The value is 1 if the condition is met, and 0 otherwise. For the importance weight of the rule, satisfying , The total number of rules participating in the verification;

[0085] like If the score is less than the preset verification score, the data packet is determined to have semantic conflicts or clinical logic errors, triggering a manual review process to ensure that the final output comprehensive view is reliable and usable in both semantics and clinical logic.

[0086] Beneficial Effects: This invention constructs a unified semantic framework centered on a knowledge graph, mapping structured medical records, unstructured text, and 3D image data into standardized semantic representations. It enforces a three-stage processing flow of cross-modal entity alignment, confidence-weighted fusion, and conflict resolution, significantly improving the semantic consistency, traceability, and clinical credibility of multi-source information fusion. Through a dynamic weighting mechanism, based on the data quality scores of each modality, entity link confidence, and clinical rule priorities, it adaptively adjusts the contribution ratio of different modalities in attribute aggregation, suppressing noise interference while retaining high-confidence information, effectively ensuring the stability and security of output results in complex cases. Furthermore, it establishes a closed-loop optimization mechanism from fusion results → clinical validation → knowledge graph update → model retraining, transforming physician correction behavior, image annotation feedback, and patient treatment outcomes into multi-dimensional verification signals, driving continuous system evolution. This ensures that while absorbing new guidelines and cases, it avoids forgetting historical knowledge and guarantees that semantic boundaries do not degrade. In summary, this invention effectively addresses the three major bottlenecks faced by existing oral AI systems in high-risk scenarios such as orthodontics, implantology, and maxillofacial surgery: semantic fragmentation, rigid fusion, and lack of evolution. It significantly improves the interpretability, security, and robustness of intelligent decision-making systems, especially under challenging conditions such as contradictions between text and images, conflicting multimodal evidence, or insufficient experience at the primary care level. It can still output a consistent, compliant, and traceable comprehensive view, supporting accurate and safe diagnosis and treatment. This method can be widely applied to intelligent diagnosis, surgical planning, multidisciplinary consultation, and primary care empowerment scenarios, demonstrating outstanding clinical practical value and promising prospects for industrial promotion. Attached Figure Description

[0087] Figure 1 This is a schematic diagram of the overall process of the present invention;

[0088] Figure 2 It is a flowchart for constructing a knowledge graph in oral medicine;

[0089] Figure 3 This is a flowchart of multimodal data acquisition and quality control;

[0090] Figure 4 This is a flowchart of the multimodal data parsing and standardized semantic representation generation process;

[0091] Figure 5 This is a flowchart of cross-modal semantic alignment, fusion, and the generation of a comprehensive patient data view; Detailed Implementation

[0092] The invention will now be further described with reference to the accompanying drawings.

[0093] Example

[0094] Furthermore, such as Figure 1 As shown, a knowledge graph-based method for fusing oral multimodal data includes the following steps:

[0095] S1. Construct a knowledge graph in the field of oral medicine. The knowledge graph is based on a standardized medical terminology system and adopts a two-layer architecture of "top-level reuse - specialty expansion". It includes entities such as teeth, periodontal tissues, diseases, treatments, and imaging examinations, as well as semantic relationships describing the clinical logical relationships between entities.

[0096] S2. Collect multimodal oral medical data of the target patient and perform quality control. The data includes at least structured electronic medical records, unstructured clinical text, and medical imaging data including CBCT images.

[0097] S3. The multimodal oral medical data is parsed, and the information in the structured electronic medical records and clinical texts, as well as the visual features and structural information extracted from the medical image data, are mapped into standardized semantic representations associated with entities and relationships in the knowledge graph.

[0098] S4. Based on the patient's unique identifier, examination events, and the unified semantic framework provided by the knowledge graph, the standardized semantic representations from different modalities are associated and fused through cross-modal semantic alignment, confidence-weighted fusion, and conflict resolution to generate a comprehensive patient data view containing cross-modal semantic associations.

[0099] Furthermore, such as Figure 2 As shown, the steps in S1 for constructing a knowledge graph in the field of oral medicine specifically include:

[0100] S1.1 Construct a knowledge graph with a two-layer architecture of "top-level reuse - specialty extension". The top-level reuse layer builds a basic framework based on SNOMED CT through semantic mapping and inheritance mechanisms to ensure semantic interoperability between the graph and the general medical field. Under the constraints of the top-level framework, the specialty extension layer integrates authoritative specialty knowledge sources in a structured manner and collaborates deeply with clinical experts to transform implicit clinical experience into computable rules and relationships, thereby realizing the dynamic expansion and updating of in-depth specialty knowledge in oral medicine.

[0101] S1.2 Under the aforementioned two-layer architecture, define and instantiate the core entities constituting the knowledge graph; these entities include: anatomical structure entities, disease and pathology entities, treatment and operation entities, examination and imaging entities, and patient and time entities; each entity is identified by a globally unique identifier and assigned a set of attributes describing its characteristics; among them, the attributes of key anatomical structure entities include three-dimensional spatial coordinates and morphological quantitative parameters, the attributes of disease entities include standard codes, typical clinical manifestations, and imaging features, and the attributes of treatment entities include operation codes, indications, contraindications, and standard procedure sequences; the assignment of entity attributes must follow predefined clinical constraint rules, including those for tooth mobility grades. The assessment is based on clinically measured buccal-lingual movement. and vertical motion Through a piecewise function The mapping is performed as follows:

[0102]

[0103] in, This indicates that the tooth mobility is normal; This indicates a degree of looseness (Level I). This indicates a degree of loosening (Level II). This indicates a degree of looseness (Grade III).

[0104] S1.3 Define and construct a set of semantic relations connecting the entities to form a knowledge network that conforms to clinical logic. The relations are directional semantic links and include at least: suffersFrom representing the disease condition, occurrenceIn representing the disease location, targets representing the treatment purpose, employees representing the examination method, depicts representing the image content, isPartOf representing the composition structure, and mayCause representing the causal risk.

[0105] Relationships are constructed through multi-modal collaboration, directly mapping from structured data. For unstructured text, a joint entity and relation extraction model is built using a pre-trained language model based on the Transformer architecture for automated extraction. This model extracts the contextual semantic features of the text through a shared encoding layer. The training objective of this model is a total loss. Loss identification for entity span Relationship classification loss Weighted sum:

[0106] ;

[0107] in, It is a balancing hyperparameter greater than 0, used to balance the learning weights of the two tasks and prevent the loss value of one task from being too large and dominating the training process. It is the entity span recognition loss, which uses the cross-entropy loss function to measure the entity boundary predicted by the model; It is a relation classification loss, which uses the binary cross-entropy loss function to measure the difference between the relation type predicted by the model and the actual relation type;

[0108] S1.4. A dual strategy is adopted to represent and store the entities and relationships. On the one hand, entities are stored as nodes and relationships as edges, along with all their attributes, in a high-performance graph database to support complex topology queries and path analysis. On the other hand, knowledge graph embedding technology is used to map discrete entities and relationships to a continuous low-dimensional vector space to support semantic similarity calculation and link prediction. A rotation model is used for knowledge graph embedding, and its core relationship modeling formula is:

[0109] ;

[0110] in, Let be the vector representation of the head entity in the embedding space. For relation vectors, For the tail entity vector, Represents the Hadamard product between complex vectors;

[0111] The training process employs a self-adversarial negative sampling strategy to optimize the embedding model, and its loss function is defined as follows:

[0112] ;

[0113] in, It is a head entity Through relationships After rotation and tail entity The Euclidean distance between them It is the sigmoid activation function. It is the interval parameter. These are negative sample triples, obtained by randomly replacing the original triples. Generate the head or tail entity in the middle. This represents the number of negative samples in each round of sampling.

[0114] Furthermore, such as Figure 3 As shown, the step of collecting multimodal oral medical data in S2 specifically includes:

[0115] S2.1 Planning and quality control for multi-source data acquisition; First, based on the oral clinical diagnosis and treatment process and AI research and development needs, a unified data acquisition standard is formulated. This standard clarifies the acquisition standards for each modality of data: For CBCT images, the spatial resolution is specified to be no less than 0.2mm×0.2mm×0.2mm, and the radiation dose must comply with the ALARA principle; For intraoral scan data, the accuracy error is required to be less than 20 micrometers; For text medical records, the use of structured templates for key fields is mandatory.

[0116] Secondly, establish a quantitative evaluation mechanism for data quality, and calculate a quality score for each piece of collected data. For CBCT images, their quality score Contrast noise ratio Signal-to-noise ratio and motion artifact level The following was decided jointly:

[0117] ;

[0118] in, and For reference standard values, For artifact scoring calculated based on image gradient distribution, The weighting coefficients are satisfied. Only when At that time, the image data is incorporated into the subsequent processing flow;

[0119] For intraoral scan data, its quality score From surface integrity Noise density and geometric distortion rate The following was decided jointly:

[0120] ;

[0121] in, This is a standard completeness reference value. This is a reference value for the maximum allowable noise density. The geometric distortion rate is calculated by detecting the deviation of a standard sphere or a known distance. The weighting coefficients are satisfied. ;

[0122] For text data, its quality score Field fill rate Terminology standardization Number of logical contradictions The following was decided jointly:

[0123] ;

[0124] in, The key field fill rate is calculated by dividing the number of filled key fields by the total number of key fields. To determine the terminology standardization, the semantic similarity between the text and a standard medical dictionary is calculated using an NLP algorithm. The number of logical contradictions is the number of logical errors detected by the rule engine and measured after normalization. The weighting coefficients are satisfied. ;

[0125] S2.2. Data acquisition and enhancement of structured electronic medical records: Specifically, through the standardized interface of the hospital information system, batch acquisition of patients' demographic information, medical history, allergy history, diagnostic codes, surgical procedure codes, medication records, and laboratory test results; to solve the problem of coding differences between different hospital systems, a knowledge graph-based coding mapping and normalization algorithm is adopted. This algorithm utilizes the synonym expansion and hierarchical relationships of entities in the knowledge graph to construct a mapping matrix from local codes to standard terms. For a local encoding record Its corresponding standard encoding vector Calculated using the following formula:

[0126] ;

[0127] Among them, matrix Optimization is achieved by minimizing the cosine distance loss of semantic vectors before and after mapping to ensure semantic consistency;

[0128] S2.3. Data collection and privacy desensitization of unstructured clinical text are performed, specifically: collecting free text fields such as chief complaint, present medical history, specialist examinations, and treatment plans from outpatient medical records; secondly, to protect patient privacy, a two-stage desensitization pipeline is designed. The first stage uses a rule-based and dictionary-based method to quickly identify and replace explicit identifiers, which at least include: patient name, ID number, phone number, and detailed address; the second stage uses a fine-tuned biomedical BERT model for context-sensitive entity recognition, identifying and generalizing implicit identifiers, which include: descriptions of rare diseases, specific surgery dates, and highly characteristic clinical manifestations; the desensitized text... Privacy risk Defined in a given set of background knowledge In this case, the attacker uses a combination of quasi-identifiers to uniquely re-identify a specific patient, the probability of which... The k-anonymity of the dataset is estimated using the following formula:

[0129] ;

[0130] in, This represents the minimum number of records with the same quasi-identifier combination in the anonymized dataset. The security threshold is set as follows: ;

[0131] S2.4. Acquisition and multimodal registration of medical image data: First, image data from different devices are acquired in parallel, including 3D volume data of the jawbone provided by CBCT, 3D mesh data of the dentition and soft tissue surfaces provided by an intraoral scanner, and 3D facial appearance models provided by a facial scanner. Second, a coarse-to-fine hybrid registration algorithm is used to achieve spatial alignment of multimodal images. This algorithm consists of two stages: The first stage is global rigid alignment, which, for hard tissues such as the jawbone, automatically detects stable anatomical landmarks through a pre-trained deep learning model and calculates the rigid transformation matrix. The second stage is local non-rigid fine-tuning, which, for the gingiva and facial soft tissues, minimizes the surface distance between the source and target images based on rigid alignment, and introduces bending energy constraints to solve the non-rigid deformation field to compensate for the geometric deformation of soft tissues under different imaging states. The overall optimization objective function is defined as follows:

[0132]

[0133] );

[0134] in, Represents a spatial transformation function. Let N be the set of anatomical landmarks detected in the source image. For the target image and The corresponding set of N paired marker points, where N must be greater than or equal to 5. It is a random sample from the surface of the source image. A set of points; Indicates the point The point closest to the Euclidean distance projected onto the surface of the target image. The distance is represented by Euclidean distance. The optimization process uses the L-BFGS algorithm for iterative solution, with a convergence threshold set to 1e-6. This represents the deformation regularization term, used to constrain the smoothness of non-rigid transformations. To balance the weighting coefficients, These are the regularization weight coefficients. and Pre-defined through experiments based on the validation set;

[0135] S2.5 Collect time-series information and device metadata associated with the core data mentioned above. This information is associated with the core data through a globally unique three-level identifier of "patient-examination-sequence" to form a multimodal data package containing timestamps and collection context, providing support for subsequent dynamic analysis and tracing.

[0136] Furthermore, such as Figure 4 As shown, the steps in S3 for parsing and acquiring multimodal oral medical data specifically include:

[0137] S3.1. For the collected structured electronic medical record data, perform standardized mapping based on the entities and relationships defined in the knowledge graph: For standardized coding fields such as ICD / CPT, directly map them to the corresponding disease or treatment entities by querying the "coding-entity" correspondence table of the knowledge graph; For numerical and categorical fields, design a rule-based parser to convert them into attribute assertions defined in the knowledge graph, and finally output a structured set of attribute assertions as part of the standardized semantic representation.

[0138] S3.2 Deep semantic parsing and entity linking of unstructured clinical text: First, a pre-trained oral medicine language model is used to identify named entity references in the text. Subsequently, entity mentions are calculated based on contextual semantic similarity, clinical prior probability, and graph topological association strength. With candidate entities in the knowledge graph Link score, the formula for calculating link score is:

[0139]

[0140] ;

[0141] in, Mentioning entities in the text that are to be linked. Candidate entities in the knowledge graph. For entity mention The context vector of the sentence. For the semantic embedding vector of the candidate entity, The cosine similarity function is used. Candidate entities are obtained based on statistics from a large-scale medical corpus. The prior probability of occurrence, The set of disambiguated neighboring entities within the current text context window. These are weighting coefficients, with values ​​of 0.5, 0.2, and 0.3 respectively.

[0142] Ultimately, if max( () greater than the preset confidence threshold If the candidate entity with the highest score is selected as the link result, then the mention is marked as an unregistered entity and not linked. Simultaneously, the pre-trained language model is used to encode the entire clinical text, and the hidden layer states corresponding to the [CLS] tags are extracted as the global semantic feature vector of the text. This is used for subsequent alignment with image features in the shared semantic space;

[0143] S3.3, Visual feature and structural information analysis of medical image data: First, an anatomical structure segmentation network based on 3D U-Net is constructed to extract voxel-level masks of the target anatomical structure. The training uses a joint loss function of Dice and cross-entropy, and the training loss formula is as follows:

[0144] ;

[0145] in, The Dice loss measures the spatial overlap between the predicted mask and the true mask. To optimize pixel-level classification accuracy using cross-entropy loss, For balancing weighting coefficients;

[0146] Secondly, a 3D Faster R-CNN network is used to detect pathological features within the region of interest defined by the mask, outputting the three-dimensional bounding box coordinates of the pathological lesions and their category probability distribution. Simultaneously, the detected pathological lesion regions are recorded as image entity regions. The ROI Align layer extracts feature map patches corresponding to each lesion region from the feature map, and then maps them into local visual feature vectors through a fully connected layer. It is used to characterize the fine-grained visual properties of lesion areas;

[0147] Subsequently, the segmented anatomical structure voxel data are input into a 3D residual network 3D-ResNet image encoder to extract global semantic feature vectors. Through a cross-modal contrastive learning mechanism, minimize The corresponding clinical text feature vector generated in step S3.2 The cosine distance between them, while maximizing the distance to negative samples within the batch, is the contrastive loss function. Defined as:

[0148] ;

[0149] in, This is the global semantic feature vector of the current sample image. To and Paired clinical text feature vectors, The total number of samples in a training batch. This is the text feature vector of the j-th sample in the batch. When j corresponds to the current sample, it is a positive sample; otherwise... All samples are considered negative samples. The cosine similarity function is used. This is a temperature coefficient used to adjust the smoothness of the similarity distribution, making the model pay more attention to the hard-to-bear samples;

[0150] Finally, the geometric parameters of the segmentation mask, the category labels of the pathological tests, and the aligned global feature vectors are concatenated to form a standardized multimodal semantic representation associated with the entities in the knowledge graph.

[0151] S3.4 Unify the encoding of the parsing results of structured medical records, clinical texts and medical images into standardized semantic representations: For each examination event, construct a structured data framework that includes a set of entity nodes, a set of attribute assertions and a set of relation triples;

[0152] Specifically, the text global feature vector generated in S3.2 Image global semantic feature vector extracted from S3.3 Modal alignment and fusion are performed to form a multimodal feature vector set, which is then attached as an attribute to the corresponding knowledge graph entity node. Finally, the data is encapsulated to form a structured patient-examination level multimodal data object, which serves as the basic input for cross-modal semantic reasoning and fusion in the subsequent step S4.

[0153] Furthermore, such as Figure 5 As shown, the step in S4 that generates a comprehensive patient data view containing cross-modal semantic associations specifically includes:

[0154] S4.1. Based on the knowledge graph embedding space, achieve cross-modal semantic alignment and association: Project the entity feature vectors obtained from text and image parsing onto the unified vector space trained by the knowledge graph, calculate their weighted similarity with candidate entity vectors, and perform joint linking decisions; for entity mentions that have both text and image features, calculate their cross-modal coreference confidence. When the threshold is exceeded, they are determined to be the same real object, and their multimodal attributes are merged and integrated in the comprehensive view. Specifically, the cross-modal co-reference confidence is defined as:

[0155]

[0156] ;

[0157] in, For entity references in the text defined in S3.2, The image pathological lesion area detected in S3.3, for The context embedding vector, for The corresponding local visual feature vector, This represents the vector representation of candidate entities in the knowledge graph. The cosine similarity function is used. For the weighting coefficients, satisfying ;

[0158] like If the two refer to the same entity, their multimodal attributes are merged to generate a unified representation.

[0159] S4.2 Confidence-weighted fusion and conflict resolution of cross-modal parsing results: For attribute assertions extracted from the same entity in different modalities, weighted aggregation is performed based on the confidence weight of their source modality; when attribute values ​​in different modalities conflict, knowledge graph constraints and clinical rules are introduced to resolve the conflict, ultimately generating a comprehensive attribute representation with enhanced consistency;

[0160] Specifically, for entities Attributes If the attribute is numeric, its merged value Defined as:

[0161] ;

[0162] If the attribute is categorical, a weighted voting mechanism is used to select the category with the largest sum of weights as the fusion result;

[0163] in, To provide the effective number of modalities for this attribute, The attribute value provided for the k-th modality The confidence weights for this mode satisfy the following conditions: Weight The modal quality score obtained in S2.1 and the link confidence score calculated for the entity in S4.1 are jointly determined, as follows:

[0164] ;

[0165] in, To score the data quality of the k-th modality, The confidence level of entity links in this modality. This is a balancing factor used to adjust the relative importance of quality scores and link confidence.

[0166] If a conflict exists, the conflict resolution module is triggered: First, it queries the context relationship of the entity in the knowledge graph, then combines the predefined priority rule library, logical constraint rule library and time decay rule library to determine the priority, and finally outputs attribute values ​​that are consistent and conform to medical logic.

[0167] S4.3 Construct a comprehensive patient data view: Using the patient's unique identifier and examination event as keys, encapsulate the fused entity instances, attribute assertions, relation assertions, and multimodal feature vectors into a structured data packet header and inject it into the knowledge graph; at the same time, based on the relational paths defined in the knowledge graph, automatically deduce implicit semantic associations and generate a visual comprehensive view containing cross-modal semantic links;

[0168] Specifically, the integrated data packet header is defined as a five-tuple structure:

[0169] ;

[0170] in, As a globally unique identifier for the patient, This serves as the unique identifier for this inspection event. This is the set of attribute assertions after fusion by S4.2. For the set of assertions of relationships between entities, A multimodal feature vector set for key entities, including text context vectors. With image local feature vector This data package serves as the basic unit for subsequent clinical decision support and dynamic tracking, and supports interactive visualization based on time axis or spatial dimension.

[0171] S4.4 Output and validate the comprehensive patient data view: Write the encapsulated structured data package into the clinical decision support system and trigger the consistency verification module; this module performs logical consistency verification on the fusion result according to the constraint rules defined in the knowledge graph; if an anomaly is found, it is marked for manual review to ensure that the final output comprehensive view is reliable and usable in both semantics and clinical logic;

[0172] Specifically, the packet consistency check score is defined as:

[0173] ;

[0174] in, For the i-th knowledge graph constraint rule, This is an indicator function that indicates when a data packet meets the rules. The value is 1 if the condition is met, and 0 otherwise. For the importance weight of the rule, satisfying , The total number of rules participating in the verification;

[0175] like If the data packet score is less than the preset verification score of 0.85, it is determined that the data packet has a semantic conflict or clinical logic error, triggering a manual review process. First, the system highlights the specific attribute fields that caused the deduction and the rules for the conflict, suspends the data group and pushes it to the manual review queue. Second, doctors or data annotators correct or confirm the conflicting fields based on their clinical expertise, generating corrected standard data. Finally, only data that has passed verification or been manually corrected and confirmed will be written into the comprehensive view to ensure that it is reliable and usable in both semantics and clinical logic.

[0176] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. The scope of protection of the present invention should be determined by the scope of protection of the appended claims.

Claims

1. A knowledge graph-based method for fusing oral multimodal data, characterized in that, Includes the following steps: S1. Construct a knowledge graph in the field of oral medicine. The knowledge graph is based on a standardized medical terminology system and adopts a two-layer architecture of "top-level reuse - specialty expansion". It includes entities such as teeth, periodontal tissues, diseases, treatments, and imaging examinations, as well as semantic relationships describing the clinical logical relationships between entities. S2. Collect multimodal oral medical data of the target patient and perform quality control. The data includes at least structured electronic medical records, unstructured clinical text, and medical imaging data including CBCT images. S3. The multimodal oral medical data is parsed, and the information in the structured electronic medical records and clinical texts, as well as the visual features and structural information extracted from the medical image data, are mapped into standardized semantic representations associated with entities and relationships in the knowledge graph. S4. Based on the patient's unique identifier, examination events, and the unified semantic framework provided by the knowledge graph, the standardized semantic representations from different modalities are associated and fused through cross-modal semantic alignment, confidence-weighted fusion, and conflict resolution to generate a comprehensive patient data view containing cross-modal semantic associations.

2. The oral multimodal data fusion method based on knowledge graph according to claim 1, characterized in that, The specific steps in S1 for constructing a knowledge graph in the field of oral medicine include: S1.1 Construct a knowledge graph with a two-layer architecture of "top-level reuse - specialty extension". The top-level reuse layer reuses the international standardized medical terminology system as the basic framework through semantic mapping and inheritance mechanisms to ensure semantic interoperability between the graph and the general medical field. Under the constraints of the top-level framework, the specialty extension layer integrates authoritative specialty knowledge sources in a structured manner and collaborates deeply with clinical experts to transform implicit clinical experience into computable rules and relationships, thereby realizing the dynamic expansion and updating of in-depth specialty knowledge in oral medicine. S1.2 Under the aforementioned two-layer architecture, define and instantiate the core entities constituting the knowledge graph; the entities include at least: anatomical structure entities, disease and pathology entities, treatment and operation entities, examination and imaging entities, and patient and time entities; each entity is identified by a globally unique identifier and assigned a set of attributes describing its characteristics; among them, the attributes of key anatomical structure entities include three-dimensional spatial coordinates and morphological quantitative parameters, the attributes of disease entities include standard codes, typical clinical manifestations, and imaging features, and the attributes of treatment entities include operation codes, indications, contraindications, and standard procedure sequences; the assignment of entity attributes must follow predefined clinical constraint rules, among which, for tooth mobility grades... The assessment is based on clinically measured buccal-lingual movement. and vertical motion Through a piecewise function The mapping is performed as follows: ; S1.3 Define and construct a set of semantic relations connecting the entities to form a knowledge network that conforms to clinical logic. The relations are directional semantic links and include at least: suffersFrom representing the disease condition, occurrenceIn representing the disease location, targets representing the treatment purpose, employees representing the examination method, depicts representing the image content, isPartOf representing the composition structure, and mayCause representing the causal risk. Relationships are constructed through multi-modal collaboration, directly mapping from structured data. For unstructured text, an automated extraction model based on a pre-trained language model, using a joint entity and relation extraction model, is employed. The training objective of this model is a total loss. It minimizes the entity span recognition loss. Relationship classification loss Weighted sum: ; in, It is a process of balancing hyperparameters. All automatically extracted relationships must be reviewed by clinical experts and logical consistency is verified using descriptive logic axioms to ensure the correctness of the knowledge network. S1.

4. A dual strategy is adopted to represent and store the entities and relationships. On the one hand, entities are stored as nodes and relationships as edges, along with all their attributes, in a high-performance graph database to support complex topology queries and path analysis. On the other hand, knowledge graph embedding technology is used to map discrete entities and relationships to a continuous low-dimensional vector space to support semantic similarity calculation and link prediction. A rotation model is used for knowledge graph embedding, and its core relationship modeling formula is: ; in, Let be the vector representation of the head entity in the embedding space. For relation vectors, For the tail entity vector, Represents the Hadamard product between complex vectors; The training process employs a self-adversarial negative sampling strategy to optimize the embedding model, and its loss function is defined as follows: ; in, It is a head entity Through relationships After rotation and tail entity The Euclidean distance between them It is the sigmoid activation function. It is the interval parameter. These are negative sample triples, obtained by randomly replacing the original triples. Generate the head or tail entity in the middle. This represents the number of negative samples in each round of sampling.

3. The oral multimodal data fusion method based on knowledge graph according to claim 1, characterized in that, The steps for collecting multimodal oral medical data in S2 specifically include: S2.1 Planning and quality control for multi-source data acquisition; First, based on the oral clinical diagnosis and treatment process and AI research and development needs, a unified data acquisition standard is formulated. This standard clarifies the acquisition standards for each modality of data: For CBCT images, the spatial resolution is specified to be no less than 0.2mm×0.2mm×0.2mm, and the radiation dose must comply with the ALARA principle; For intraoral scan data, the accuracy error is required to be less than 20 micrometers; For text medical records, the use of structured templates for key fields is mandatory. Secondly, establish a quantitative evaluation mechanism for data quality, and calculate a quality score for each piece of collected data. For CBCT images, their quality score Contrast noise ratio Signal-to-noise ratio and motion artifact level The following was decided jointly: ; in, and For reference standard values, For artifact scoring calculated based on image gradient distribution, The weighting coefficients are satisfied. Only when At that time, the image data is incorporated into the subsequent processing flow; For intraoral scan data, its quality score From surface integrity Noise density and geometric distortion rate The following was decided jointly: ; in, This is a standard completeness reference value. This is a reference value for the maximum allowable noise density. The geometric distortion rate is calculated by detecting the deviation of a standard sphere or a known distance. The weighting coefficients are satisfied. ; For text data, its quality score Field fill rate Terminology standardization Number of logical contradictions The following was decided jointly: ; in, The key field fill rate is calculated by dividing the number of filled key fields by the total number of key fields. To determine the terminology standardization, the semantic similarity between the text and a standard medical dictionary is calculated using an NLP algorithm. The number of logical contradictions is the number of logical errors detected by the rule engine and measured after normalization. The weighting coefficients are satisfied. ; S2.

2. Data acquisition and enhancement of structured electronic medical records: Specifically, through the standardized interface of the hospital information system, batch acquisition of patients' demographic information, medical history, allergy history, diagnostic codes, surgical procedure codes, medication records, and laboratory test results; to solve the problem of coding differences between different hospital systems, a knowledge graph-based coding mapping and normalization algorithm is adopted. This algorithm utilizes the synonym expansion and hierarchical relationships of entities in the knowledge graph to construct a mapping matrix from local codes to standard terms. For a local encoding record Its corresponding standard encoding vector Calculated using the following formula: ; Among them, matrix Optimization is achieved by minimizing the cosine distance loss of semantic vectors before and after mapping to ensure semantic consistency; S2.

3. Data collection and privacy anonymization of unstructured clinical text: Specifically, free text fields such as chief complaint, present medical history, specialist examinations, and treatment plans from outpatient medical records are collected. Secondly, to protect patient privacy, a two-stage anonymization pipeline is designed. The first stage uses a rule-based and dictionary-based method to quickly identify and replace explicit identifiers. The second stage uses a fine-tuned biomedical BERT model for context-sensitive entity recognition, identifying and generalizing implicit identifiers and sensitive clinical events. The anonymized text... Privacy risk Defined in a given set of background knowledge In this case, the attacker uses a combination of quasi-identifiers to uniquely re-identify a specific patient, the probability of which... The k-anonymity of the dataset is estimated using the following formula: ; in, This represents the minimum number of records with the same combination of quasi-identifiers in the anonymized dataset. S2.

4. Acquisition and multimodal registration of medical image data: First, image data from different devices are acquired in parallel, including 3D volume data of the jawbone provided by CBCT, 3D mesh data of the dentition and soft tissue surfaces provided by an intraoral scanner, and 3D facial appearance models provided by a facial scanner. Second, a coarse-to-fine hybrid registration algorithm is used to achieve spatial alignment of multimodal images. This algorithm consists of two stages: The first stage is global rigid alignment, which, for hard tissues such as the jawbone, automatically detects stable anatomical landmarks through a pre-trained deep learning model and calculates the rigid transformation matrix. The second stage is local non-rigid fine-tuning, which, for the gingiva and facial soft tissues, minimizes the surface distance between the source and target images based on rigid alignment, and introduces bending energy constraints to solve the non-rigid deformation field to compensate for the geometric deformation of soft tissues under different imaging states. The overall optimization objective function is defined as follows: ); in, Represents a spatial transformation function. Let N be the set of anatomical landmarks detected in the source image. For the target image and The corresponding set of N paired marker points, It is a random sample from the surface of the source image. A set of points; Indicates the point The point closest to the Euclidean distance projected onto the surface of the target image. To balance the weighting coefficients, The distance is represented by Euclidean distance, and the optimization process uses the L-BFGS algorithm for iterative solution. This represents the deformation regularization term, used to constrain the smoothness of non-rigid transformations. These are the regularization weight coefficients; S2.5 Collect time-series information and device metadata associated with the core data mentioned above. This information is associated with the core data through a globally unique "patient-examination-sequence" three-level identifier, and together they form a multimodal data package containing timestamps and collection context, providing support for subsequent dynamic analysis and tracing.

4. The oral multimodal data fusion method based on knowledge graph according to claim 1, characterized in that, The steps for parsing and acquiring multimodal oral medical data in S3 specifically include: S3.

1. For the collected structured electronic medical record data, perform standardized mapping based on the entities and relationships defined in the knowledge graph: For standardized coding fields such as ICD / CPT, directly map them to the corresponding disease or treatment entities by querying the "coding-entity" correspondence table of the knowledge graph; For numerical and categorical fields, design a rule-based parser to convert them into attribute assertions defined in the knowledge graph, and finally output a structured set of attribute assertions as part of the standardized semantic representation. S3.2 Deep semantic parsing and entity linking of unstructured clinical text: First, a pre-trained oral medicine language model is used to identify named entity references in the text. Subsequently, entity mentions are calculated based on contextual semantic similarity, clinical prior probability, and graph topological association strength. With candidate entities in the knowledge graph Link score, the formula for calculating link score is: ; in, Mentioning entities in the text that are to be linked. Candidate entities in the knowledge graph. For entity mention The context vector of the sentence. For the semantic embedding vector of the candidate entity, The cosine similarity function is used. Candidate entities are obtained based on statistics from a large-scale medical corpus. The prior probability of occurrence, The set of disambiguated neighboring entities within the current text context window. These are the weighting coefficients; Ultimately, if max( () greater than the preset confidence threshold If the candidate entity with the highest score is selected as the link result, then the mention is marked as an unregistered entity and not linked. Simultaneously, the pre-trained language model is used to encode the entire clinical text, and the hidden layer states corresponding to the [CLS] tags are extracted as the global semantic feature vector of the text. This is used for subsequent alignment with image features in the shared semantic space; S3.3, Visual feature and structural information analysis of medical image data: First, an anatomical structure segmentation network based on 3D U-Net is constructed to extract voxel-level masks of the target anatomical structure. The training uses a joint loss function of Dice and cross-entropy, and the training loss formula is as follows: ; in, The Dice loss measures the spatial overlap between the predicted mask and the true mask. To optimize pixel-level classification accuracy using cross-entropy loss, For balancing weighting coefficients; Secondly, a 3D Faster R-CNN network is used to detect pathological features within the region of interest defined by the mask, outputting the three-dimensional bounding box coordinates of the pathological lesions and their category probability distribution. Simultaneously, the detected pathological lesion regions are recorded as image entity regions. The ROI Align layer extracts feature map patches corresponding to each lesion region from the feature map, and then maps them into local visual feature vectors through a fully connected layer. It is used to characterize the fine-grained visual properties of lesion areas; Subsequently, the segmented anatomical structure voxel data are input into a 3D residual network 3D-ResNet image encoder to extract global semantic feature vectors. Through a cross-modal contrastive learning mechanism, minimize The corresponding clinical text feature vector generated in step S3.2 The cosine distance between them, while maximizing the distance to negative samples within the batch, is the contrastive loss function. Defined as: ; in, This is the global semantic feature vector of the current sample image. To and Paired clinical text feature vectors, The total number of samples in a training batch. This is the text feature vector of the j-th sample in the batch. When j corresponds to the current sample, it is a positive sample; otherwise... All samples are considered negative samples. The cosine similarity function is used. This is a temperature coefficient used to adjust the smoothness of the similarity distribution, making the model pay more attention to the hard-to-bear samples; Finally, the geometric parameters of the segmentation mask, the category labels of the pathological tests, and the aligned global feature vectors are concatenated to form a standardized multimodal semantic representation associated with the entities in the knowledge graph. S3.4 Unify the encoding of the parsing results of structured medical records, clinical texts and medical images into standardized semantic representations: For each examination event, construct a structured data framework that includes a set of entity nodes, a set of attribute assertions and a set of relation triples; Specifically, the text global feature vector generated in S3.2 Image global semantic feature vector extracted from S3.3 Modal alignment and fusion are performed to form a multimodal feature vector set, which is then attached as an attribute to the corresponding knowledge graph entity node. Finally, the data is encapsulated to form a structured patient-examination level multimodal data object, which serves as the basic input for cross-modal semantic reasoning and fusion in the subsequent step S4.

5. The oral multimodal data fusion method based on knowledge graph according to claim 1, characterized in that, The step in S4 that generates a comprehensive patient data view containing cross-modal semantic associations specifically includes: S4.

1. Based on the knowledge graph embedding space, achieve cross-modal semantic alignment and association: Project the entity feature vectors obtained from text and image parsing onto the unified vector space trained by the knowledge graph, calculate their weighted similarity with candidate entity vectors, and perform joint linking decisions; for entity mentions that have both text and image features, calculate their cross-modal coreference confidence. When the threshold is exceeded, they are determined to be the same real object, and their multimodal attributes are merged and integrated in the comprehensive view. Specifically, the cross-modal co-reference confidence is defined as: ; in, For entity references in the text defined in S3.2, The image pathological lesion area detected in S3.3, for The context embedding vector, for The corresponding local visual feature vector, This represents the vector representation of candidate entities in the knowledge graph. The cosine similarity function is used. For the weighting coefficients, satisfying ; like If the confidence level is greater than the preset confidence threshold, it is determined that the two refer to the same entity, and their multimodal attributes are merged to generate a unified representation. S4.2 Confidence-weighted fusion and conflict resolution of cross-modal parsing results: For attribute assertions extracted from the same entity in different modalities, weighted aggregation is performed based on the confidence weight of their source modality; when attribute values ​​in different modalities conflict, knowledge graph constraints and clinical rules are introduced to resolve the conflict, ultimately generating a comprehensive attribute representation with enhanced consistency; Specifically, for entities Attributes If the attribute is numeric, its merged value Defined as: ; If the attribute is categorical, a weighted voting mechanism is used to select the category with the largest sum of weights as the fusion result; in, The attribute value provided for the k-th modality The confidence weights for this mode satisfy the following conditions: Weight The link confidence score is determined by both the modal quality score and the link confidence score calculated for the entity in S4.

1. To provide the number of valid modalities for this attribute; If a conflict exists, the conflict resolution module is triggered: first, it queries the context relationship of the entity in the knowledge graph, then it judges the priority in combination with the predefined clinical rule base, and finally outputs attribute values ​​that are consistent and in line with medical logic. S4.3 Construct a comprehensive patient data view: Using the patient's unique identifier and examination event as keys, encapsulate the fused entity instances, attribute assertions, relation assertions, and multimodal feature vectors into a structured data packet header and inject it into the knowledge graph; at the same time, based on the relational paths defined in the knowledge graph, automatically deduce implicit semantic associations and generate a visual comprehensive view containing cross-modal semantic links; Specifically, the integrated data packet header is defined as a five-tuple structure: ; in, As a globally unique identifier for the patient, This serves as the unique identifier for this inspection event. This is the set of attribute assertions after fusion by S4.

2. For the set of assertions of relationships between entities, A multimodal feature vector set for key entities, including text context vectors. With image local feature vector This data package serves as the basic unit for subsequent clinical decision support and dynamic tracking, and supports interactive visualization based on time axis or spatial dimension. S4.4 Output and validate the comprehensive patient data view: Write the encapsulated structured data package into the clinical decision support system and trigger the consistency verification module; this module performs logical consistency verification on the fusion result according to the constraint rules defined in the knowledge graph; if an anomaly is found, it is marked for manual review to ensure that the final output comprehensive view is reliable and usable in both semantics and clinical logic; Specifically, the packet consistency check score is defined as: ; in, For the i-th knowledge graph constraint rule, This is an indicator function that indicates when a data packet meets the rules. The value is 1 if the condition is met, and 0 otherwise. For the importance weight of the rule, satisfying , The total number of rules participating in the verification; like If the score is less than the preset verification score, the data packet is determined to have semantic conflicts or clinical logic errors, triggering a manual review process to ensure that the final output comprehensive view is reliable and usable in both semantics and clinical logic.