Medical multi-modal data fusion analysis method and system based on cross-modal alignment

CN122025111BActive Publication Date: 2026-08-21HANGZHOU QUADRANT DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610485915.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-14
Publication Date
2026-08-21
Estimated Expiration
2046-04-14

AI Technical Summary

Technical Problem

[0003]本申请提供基于跨模态对齐的医疗多模态数据融合分析方法及系统,解决了现有技术数据质量不均、语义对齐步骤以及融合权重分配不灵活的技术问题

Benefits of technology

[0070] This application improves the accuracy and reliability of medical data fusion through a systematic multimodal data quality grading assessment and differentiated processing strategy. Specifically, the quality scoring system based on multi-dimensional indicators can automatically identify the quality differences of data in each modality and implement differentiated processing for different levels, effectively avoiding the interference of low-quality data on the fusion results, while optimizing the allocation of computing resources. This grading processing mechanism not only improves the intelligence level of data preprocessing, but also ensures the standardization and consistency of input data in subsequent analysis stages, laying a solid foundation for the overall fusion effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122025111B_ABST
    Figure CN122025111B_ABST
Patent Text Reader

Abstract

The application provides a medical multi-modal data fusion analysis method and system based on cross-modal alignment, relates to the field of artificial intelligence, and solves the technical problems of uneven data quality, inflexible semantic alignment steps and fusion weight distribution in the prior art. The method comprises the following steps: acquiring multi-modal medical data and performing quality grading evaluation to obtain a modal quality score of each modality; using a two-stage alignment algorithm of intra-modal feature enhancement and cross-modal semantic calibration, performing semantic alignment on the multi-modal data after quality grading evaluation to obtain a semantic alignment result; based on a knowledge graph correlation degree, a clinical guideline priority and the modal quality score, using a dynamic weight model to perform feature fusion on the semantic alignment result; generating a structured diagnosis report through an artificial intelligence generated content (AIGC) model on the fusion result, and performing clinical compliance verification and sensitivity analysis, and feeding back the verification and analysis results to the feature fusion step for optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence, specifically a method and system for medical multimodal data fusion and analysis based on cross-modal alignment. Background Technology

[0002] With the rapid development of medical informatization, medical institutions generate massive amounts of multimodal medical data during daily diagnosis and treatment, which is of great value in clinical diagnosis, treatment planning, and disease prediction. However, current methods for fusing multimodal medical data have significant shortcomings. First, due to the diverse sources and formats of medical data, traditional methods lack unified quality assessment standards, making it difficult to effectively identify and distinguish data of different quality levels. This results in high-quality and low-quality data being treated equally during the fusion process, seriously affecting the reliability of the final analysis results. Second, existing technologies often employ simple feature splicing or shallow fusion strategies when processing different modalities, making it difficult to establish deep semantic relationships. This leads to significant semantic gaps between different modalities, affecting the accuracy of clinical diagnosis. Furthermore, existing multimodal fusion methods mostly use static weight allocation strategies, which cannot be dynamically adjusted according to specific disease types, changes in data quality, and clinical guideline requirements, resulting in a lack of clinical adaptability in the fusion results. Summary of the Invention

[0003] This application provides a method and system for medical multimodal data fusion analysis based on cross-modal alignment, which solves the technical problems of uneven data quality, semantic alignment steps, and inflexible fusion weight allocation in existing technologies.

[0004] To achieve the above objectives, this application adopts the following technical solution:

[0005] Firstly, it provides a method for medical multimodal data fusion and analysis based on cross-modal alignment, including:

[0006] Acquire multimodal medical data and perform quality grading assessment to obtain modality quality scores for each modality; wherein, the multimodal medical data includes at least medical imaging data, electronic medical record text data, and physiological signal time series data;

[0007] A two-stage alignment algorithm, consisting of intra-modal feature enhancement and cross-modal semantic calibration, is used to perform semantic alignment on multimodal data after quality grading assessment, yielding semantic alignment results.

[0008] Based on the knowledge graph relevance, clinical guideline priority, and modality quality score, a dynamic weighting model is used to perform feature fusion on the semantic alignment results.

[0009] The AIGC (Artificial Intelligence Generates Content) model generates a structured diagnostic report from the fusion results, performs clinical compliance verification and sensitivity analysis, and feeds the verification and analysis results back to the feature fusion step for optimization.

[0010] Based on the above technical solutions, the medical multimodal data fusion and analysis method based on cross-modal alignment provided in this application automatically identifies and quantifies the quality differences of each modality through quality grading assessment of multimodal medical data, ensuring the reliability and consistency of input data and laying a solid foundation for subsequent processing. Secondly, a two-stage alignment algorithm combining intra-modal feature enhancement and cross-modal semantic calibration effectively addresses the semantic gap between multi-source heterogeneous data, enhancing the correlation between different modalities and improving the accuracy of semantic alignment. Then, a dynamic weight model based on knowledge graph correlation, clinical guideline priority, and modality quality score achieves intelligent fusion through multi-factor collaboration, ensuring that the fusion process conforms to the medical knowledge system and adapts to specific clinical scenarios, improving the scientific nature of decision-making. Finally, a structured diagnostic report is generated through an AIGC model, and combined with clinical compliance verification and sensitivity analysis, the compliance and robustness of the output results are ensured. Simultaneously, a feedback optimization mechanism forms a closed-loop system, continuously improving overall performance. This method comprehensively optimizes the entire process of multimodal data from acquisition to output, providing an efficient and reliable solution for smart healthcare applications.

[0011] Furthermore, the acquisition of multimodal medical data and the performance of quality grading assessment include:

[0012] For medical image data, spatial coverage is calculated based on the ratio of the number of pixels actually covered by the lesion area to the total number of pixels in the entire medical image. The first information density score is calculated based on the information entropy algorithm. The spatial coverage and the first information density score are weighted and summed to obtain the comprehensive score of the image data.

[0013] For electronic medical record text data, a first text index is calculated based on the ratio of the number of core fields filled in to the total number of core fields. A second text index is calculated based on the ratio of the number of text terms in the medical record that match preset medical standard terms to the total number of all relevant terms in the medical record. A third text index is calculated based on the ratio of the number of fields in the electronic medical record text format that conform to the HL7 medical data exchange standard to the total number of fields in the medical record. A second information density score is calculated based on the information entropy algorithm. The first, second, and third text indices and the second information density score are then weighted and summed to obtain a comprehensive score for the text data. The preset medical standard terms include the medical system nomenclature—the clinical terminology SNOMED CT standard terminology.

[0014] For physiological signal time series data, the time series continuity score is calculated based on the ratio of continuous signal duration to total signal acquisition duration within the physiological signal acquisition period. The third information density score is calculated based on the information entropy algorithm. The time series continuity score and the third information density score are then weighted and summed to obtain the comprehensive score of the time series data.

[0015] Based on the comprehensive score of each modality, the data of each modality are classified into different quality levels according to the level threshold, and a differentiated post-processing strategy is set for each level.

[0016] Furthermore, the provision of differentiated post-processing strategies for each level includes:

[0017] For the first-level high-quality modes, directly input them into the subsequent two-stage alignment algorithm;

[0018] For qualified modalities at the second level, an enhancement strategy based on AIGC (Artificial Intelligence Generated Content) is adopted to optimize the data of each modality through generative models.

[0019] For the third level of low-quality modalities, an AIGC-based alternative strategy is adopted, which uses a generative model to generate the missing content.

[0020] For the extremely low modes of the fourth level, they are directly discarded.

[0021] Furthermore, the two-stage alignment algorithm includes an intra-modal feature enhancement stage and a cross-modal semantic calibration stage; wherein,

[0022] The intramodal feature enhancement stage is used to receive medical image data, electronic medical record text data, and physiological signal time series data after quality grading assessment, and to enhance the features of each modality through a pre-trained neural network model to obtain enhanced feature blocks for each modality.

[0023] The cross-modal semantic calibration stage is used to perform clinical semantic association mapping and consistency verification on the enhanced feature blocks combined with medical knowledge graph embedding, thereby enhancing the semantic association between different modalities and obtaining a unified semantic vector that conforms to clinical logic.

[0024] Furthermore, the internal workflow of the intra-modal feature enhancement stage is as follows:

[0025] Input multimodal data after quality grading; among them, medical imaging data is in standardized DICOM 3.0 format, electronic medical record text data is in structured XML data, and physiological signal time series data is in standardized JSON format;

[0026] For medical image data, a pre-trained MedViT model is used, and the following steps are performed in sequence: extracting lesion features through multi-layer convolutional blocks, enhancing the size / density / boundary features of lesions through an attention module, encoding global features through multi-layer Transformers, and outputting multi-dimensional image feature blocks.

[0027] For electronic medical record text data, a pre-trained MedicalBERT model is used, which sequentially performs word segmentation and mapping with SNOMEDCT standard terms, medical entity recognition, and model semantic encoding to output multi-dimensional text feature blocks;

[0028] For physiological signal time series data, a pre-trained multilayer perceptron is used, and the following steps are performed in sequence: the original time series values ​​are encoded by the multilayer perceptron, the physiological index features are enhanced by the attention module, and multidimensional numerical feature blocks are output.

[0029] The effectiveness score of each modality feature block is calculated by weighted summation of the proportion of key feature dimensions and the feature signal-to-noise ratio of each modality. When the effectiveness score is greater than the preset judgment threshold, the enhanced feature block of each modality is output. The proportion of key feature dimensions is the proportion of feature dimensions representing clinical core indicators in each modality to the total feature dimensions, and the feature signal-to-noise ratio is the ratio of effective feature variance to noise feature variance.

[0030] Furthermore, in the process of enhancing feature extraction for each modality, the feature dimension of the clinical core indicator refers to the feature dimension corresponding to the indicator directly related to the diagnosis, treatment and prognosis of the target disease, which is obtained through clinical guideline screening and expert consensus confirmation; the effective feature is obtained through a joint algorithm of mutual information and variance analysis; and the noise feature is obtained through threshold screening and dynamic anomaly detection.

[0031] Furthermore, the cross-modal semantic calibration stage includes a medical memory generation layer, a cross-modal attention layer, and a semantic consistency verification layer, wherein,

[0032] The medical memory generation layer receives medical knowledge graphs and medical pairing data, learns clinical association patterns through a generator of a generative adversarial network (GAN), and obtains a three-dimensional medical memory matrix containing "modality-terminology-disease". The matrix element values ​​of the three-dimensional medical memory matrix are used to characterize the association strength between modality features and medical entities, providing medical constraints for semantic mapping. The medical pairing data consists of clinical pairing triplets of image data, electronic medical record text data, and physiological signal time series data, with each triplet containing a diagnostic association label confirmed by a parent clinical expert.

[0033] The cross-modal attention layer receives enhanced feature blocks from each modality and the three-dimensional medical memory matrix. It constructs key and value vectors for the attention mechanism using the enhanced feature blocks, embeds target disease candidate entities from the medical knowledge graph to construct query vectors, and uses the three-dimensional medical memory matrix as the basis for attention weight constraints to achieve cross-modal semantic mapping and obtain a preliminary unified semantic vector. The target disease candidate entities are determined through semantic matching between clinical information in electronic medical record text data and disease entities in the medical knowledge graph.

[0034] The semantic consistency verification layer is used to receive the preliminary unified semantic vector and the ideal alignment vector of each modality, calculate the semantic consistency score and dynamically calibrate to obtain the final unified semantic vector, i.e. the semantic alignment result.

[0035] Furthermore, the construction process of the medical knowledge graph includes:

[0036] Medical data sources were collected, including the National Comprehensive Cancer Network (NCCN) clinical guidelines, Chinese clinical practice guidelines, medical journal articles, clinical practice data, and expert consensus.

[0037] Multiple entities and their clinical relationships are extracted from the medical data source; wherein, the entities include, but are not limited to, diseases, symptoms, test indicators, imaging features, drugs, and physiological signals.

[0038] Entities were uniformly labeled using SNOMED CT standard terminology, ICD-10 disease codes, and LOINC test index codes, and clinical associations were assigned levels of clinical evidence; the clinical evidence refers to the amount of clinical research data and the degree of expert consensus that support the association.

[0039] A medical knowledge graph is obtained by storing entities and clinical relationships using a graph database.

[0040] Furthermore, the process of determining the target disease candidate set entities includes:

[0041] Based on clinical information in electronic medical record text data, semantic similarity is calculated with the "disease" entity in the knowledge graph. Disease entities with similarity greater than a preset threshold are selected to form a target disease candidate set. The clinical information includes, but is not limited to, chief complaint and present medical history. The semantic similarity is calculated using the cosine similarity formula.

[0042] Furthermore, the ideal alignment vectors for each modality are obtained based on entities and clinical relationships in the medical knowledge graph, including:

[0043] Extract entities related to the target disease from the medical knowledge graph to form a core set of related entities;

[0044] The core associated entity set is bound to the corresponding modality to form a modality-entity mapping table;

[0045] The enhanced features of each modality obtained from historical multimodal medical data are labeled as the basic feature vectors of each modality;

[0046] The medical knowledge graph is pre-trained using the TransE algorithm to obtain the knowledge graph embedding vector of each entity in the core associated entity set;

[0047] Extract the association levels between the core associated entity set and the target disease, and convert them into training weights;

[0048] Constructing the loss function ;in, The modal base feature vector bound to the i-th core entity. Let be the knowledge graph embedding vector of the i-th core entity, and n be the number of entities in the core associated entity set. Let be the training weights for the i-th core entity;

[0049] The modal basic feature vector, knowledge graph embedding vector, and training weights are input to a modal feature alignment training network built on a fully connected layer. The loss function is minimized by gradient descent to perform alignment training on the modal basic feature vectors, thereby obtaining the modal alignment vectors corresponding to each core entity.

[0050] According to the formula Calculate the weight percentage of each core entity. The ideal alignment vector for modality i is obtained by weighted fusion of the alignment vectors of all core entities in the same modality. .

[0051] Furthermore, the calculation process of the cross-modal attention layer is as follows:

[0052] According to the formula Calculate the basic attention weights ;in, Embed query vectors into the target disease candidate set entities. The key vector is the concatenated enhanced feature blocks from each modality. For feature dimensions;

[0053] Based on the aforementioned basic attention weights and the three-dimensional medical memory matrix The weight adjustment is performed using the following formula: ;in, This is the attention weight matrix after medical constraints.

[0054] A preliminary unified semantic vector is calculated based on the attention weight matrix after the aforementioned medical constraints. The calculation formula is: ;in, This is the value vector concatenated from the enhanced feature blocks of each modality.

[0055] Furthermore, the semantic consistency score The calculation formula is: ;in, Let be the ideal alignment vector for the i-th mode. For vector dot product operation, Let be the vector magnitude.

[0056] Furthermore, the step of using a dynamic weighting model to perform feature fusion on the semantic alignment result includes:

[0057] Obtain the unified semantic vector corresponding to the semantic alignment result. Simultaneously, extract the medical knowledge graph correlation vector R and the clinical guideline priority vector. And the modality quality score vector Q; where R is the clinical association strength between each modality and the target disease, extracted from the medical knowledge graph. Prioritize guideline recommendations for each modality, based on clinical practice guidelines;

[0058] Set the initial weight vector Based on dynamic weight formula The final mode weights are calculated. ;

[0059] Feature fusion is performed using an attention network with a fusion gating mechanism, through a formula. Calculate the gating coefficient G, and then use the formula Output the fused feature vector; where, , representing the knowledge graph embedding vector of each core entity based on the clinical association strength between the target disease and the core associated entities in the medical knowledge graph. The knowledge graph reasoning vector obtained by weighted aggregation The association level between the i-th core entity and the target disease; This indicates an attention calculation operation.

[0060] Furthermore, the process of generating a structured diagnostic report from the fusion results using the AIGC model, and performing clinical compliance verification and sensitivity analysis, includes:

[0061] The fused feature vector is input into a medical-specific AIGC model, and clinical semantic mapping, feature extraction, and structured report generation are performed on the fused features. The report includes a multimodal data summary, diagnostic results, feature contribution, and clinical recommendations.

[0062] The report content is corrected by traversing the clinical guideline rule base of the target disease using the Drools-Med rule engine; and the semantic consistency of the report is verified by reasoning through a medical knowledge graph.

[0063] Monte Carlo simulation is used to calculate the fluctuation range of the diagnostic results under data disturbance. When the fluctuation value is less than the preset threshold, it is determined to be stable and the sensitivity analysis is completed; otherwise, the feature fusion step is returned to be executed again.

[0064] Secondly, this application provides a medical multimodal data fusion and analysis system based on cross-modal alignment, including: a quality grading module, a semantic alignment module, a feature fusion module, and a report generation module; wherein,

[0065] The quality grading module is used to acquire multimodal medical data and perform quality grading assessment to obtain a modality quality score for each modality; wherein, the multimodal medical data includes at least medical imaging data, electronic medical record text data, and physiological signal time series data;

[0066] The semantic alignment module is used to perform semantic alignment on multimodal data after quality grading assessment using a two-stage alignment algorithm of intramodal feature enhancement and cross-modal semantic calibration, and obtain semantic alignment results.

[0067] The feature fusion module is used to perform feature fusion on the semantic alignment results based on knowledge graph relevance, clinical guideline priority, and modality quality score, using a dynamic weight model.

[0068] The report generation module is used to generate a structured diagnostic report from the fusion results through an AIGC (Artificial Intelligence Generated Content) model, and to perform clinical compliance verification and sensitivity analysis. The verification and analysis results are then fed back to the feature fusion step for optimization.

[0069] Compared with the prior art, the beneficial effects of this application are:

[0070] This application improves the accuracy and reliability of medical data fusion through a systematic multimodal data quality grading assessment and differentiated processing strategy. Specifically, the quality scoring system based on multi-dimensional indicators can automatically identify the quality differences of data in each modality and implement differentiated processing for different levels, effectively avoiding the interference of low-quality data on the fusion results, while optimizing the allocation of computing resources. This grading processing mechanism not only improves the intelligence level of data preprocessing, but also ensures the standardization and consistency of input data in subsequent analysis stages, laying a solid foundation for the overall fusion effect.

[0071] In the feature processing stage, this application achieves deep semantic integration of multi-source heterogeneous data through a two-stage alignment algorithm consisting of intra-modal feature enhancement and cross-modal semantic calibration. The intra-modal feature enhancement stage utilizes a pre-trained model to strengthen the core clinical features of each modality and dynamically selects high-value feature blocks based on effectiveness scores. The cross-modal semantic calibration stage combines the three-dimensional memory matrix and attention mechanism of a medical knowledge graph to map multi-modal features to a unified semantic space, and then dynamically optimizes the alignment results through consistency checks. This method effectively solves the semantic gap problem in traditional multimodal fusion, improving the clinical relevance of feature representation and alignment accuracy. Furthermore, the ideal alignment vector generation mechanism based on the knowledge graph provides a reference standard for semantic calibration, further ensuring the scientific rigor and interpretability of the alignment process.

[0072] The dynamic weighting model and closed-loop optimization mechanism in this application further enhance the system's practicality and adaptability. By integrating multiple dimensions such as knowledge graph relevance, clinical guideline priority, and modality quality scores, the dynamic weighting model can adaptively adjust the contribution of each modality in the fusion, making the results more consistent with clinical logic. Furthermore, the structured diagnostic reports generated by AIGC, combined with compliance verification and sensitivity analysis, not only improve the reliability of the output results but also form a continuously improving closed-loop system through feedback optimization. The overall solution improves diagnostic accuracy while strengthening clinical compliance and robustness, providing an efficient and secure decision support tool for smart healthcare applications. Attached Figure Description

[0073] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0074] Figure 1 A system architecture diagram of a medical multimodal data fusion and analysis system based on cross-modal alignment provided in an embodiment of this application;

[0075] Figure 2 A flowchart illustrating the medical multimodal data fusion and analysis method based on cross-modal alignment provided in this application embodiment;

[0076] Figure 3 A flowchart illustrating another medical multimodal data fusion and analysis method based on cross-modal alignment provided in this application embodiment;

[0077] Figure 4 This is a flowchart illustrating another medical multimodal data fusion and analysis method based on cross-modal alignment, provided as an embodiment of this application. Detailed Implementation

[0078] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.

[0079] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0080] The medical multimodal data fusion and analysis method based on cross-modal alignment provided in this application can be applied to, for example... Figure 1 In the medical multimodal data fusion and analysis system shown, based on cross-modal alignment, such as... Figure 1 As shown, the communication system includes: a quality grading module, a semantic alignment module, a feature fusion module, and a report generation module; among which,

[0081] The quality grading module is used to acquire multimodal medical data and perform quality grading assessment to obtain the modality quality score for each modality; among which, multimodal medical data includes at least medical imaging data, electronic medical record text data, and physiological signal time series data;

[0082] The semantic alignment module is used to perform semantic alignment on multimodal data after quality grading assessment using a two-stage alignment algorithm that combines intramodal feature enhancement and cross-modal semantic calibration, and obtain semantic alignment results.

[0083] The feature fusion module is used to perform feature fusion on semantic alignment results based on knowledge graph relevance, clinical guideline priority, and modality quality score, using a dynamic weight model.

[0084] The report generation module is used to generate structured diagnostic reports from the fusion results through an AIGC (Artificial Intelligence Generate Content) model, and to perform clinical compliance verification and sensitivity analysis. The verification and analysis results are then fed back to the feature fusion step for optimization.

[0085] To address the technical problems in existing medical multimodal data, such as uneven quality, semantic gaps in heterogeneous data, static and rigid fusion weight allocation, and lack of standardization and closed-loop optimization mechanisms in diagnostic report generation, this application provides a method and system for medical multimodal data fusion analysis based on cross-modal alignment. The method includes:

[0086] Acquire multimodal medical data and conduct quality grading assessment to obtain modality quality scores for each modality; among which, multimodal medical data includes at least medical imaging data, electronic medical record text data, and physiological signal time series data;

[0087] A two-stage alignment algorithm, consisting of intra-modal feature enhancement and cross-modal semantic calibration, is used to perform semantic alignment on multimodal data after quality grading assessment, yielding semantic alignment results.

[0088] Based on knowledge graph relevance, clinical guideline priority, and modality quality score, a dynamic weighting model is used to perform feature fusion on semantic alignment results;

[0089] The AIGC (Artificial Intelligence Generates Content) model generates a structured diagnostic report from the fusion results, performs clinical compliance verification and sensitivity analysis, and feeds the verification and analysis results back to the feature fusion step for optimization.

[0090] Based on this, the method in this application achieves efficient integration and reliable application of multimodal medical data through standardized processing and intelligent optimization mechanisms throughout the entire process. It not only adapts to the differentiated needs of different clinical scenarios, but also ensures the scientific and compliant nature of diagnostic results, thereby improving the accuracy and robustness of smart healthcare decision-making.

[0091] like Figure 2 As shown in the embodiments of this application, the medical multimodal data fusion and analysis method based on cross-modal alignment provided includes:

[0092] S1. Acquire multimodal medical data and conduct quality grading assessment to obtain modal quality scores for each modality.

[0093] Multimodal medical data includes at least medical imaging data, electronic medical record text data, and physiological signal time series data.

[0094] In some implementation methods, quality grading assessment methods include the proportion of missing values ​​in statistical data, verification of data format standardization, analysis of data signal integrity, sampling verification of data accuracy, and assessment of the relevance of data to clinical scenarios. The role of quality grading assessment is to screen out high-quality and highly available data, eliminate or mark low-quality and redundant data, avoid low-quality data from interfering with subsequent fusion analysis, and at the same time match appropriate processing strategies for data of different quality levels to optimize the efficiency of computing resource allocation.

[0095] It should be noted that multimodal medical data comes from a wide range of sources, including imaging equipment, electronic medical record systems, and physiological monitoring instruments in medical institutions. Data from different sources vary significantly in terms of collection standards, storage formats, and completeness. Quality grading assessment is a prerequisite for ensuring the reliability of subsequent analysis results.

[0096] For example, for medical imaging data, quality can be graded by judging the proportion of images without artifacts and the clarity of lesion areas; for electronic medical record text data, quality can be graded by statistically analyzing the completeness of core diagnostic and treatment fields and the accuracy of terminology; for physiological signal time series data, quality can be graded by analyzing the duration of uninterrupted signal and the rationality of numerical fluctuations.

[0097] S2. A two-stage alignment algorithm combining intra-modal feature enhancement and cross-modal semantic calibration is used to perform semantic alignment on the multimodal data after quality grading assessment, and the semantic alignment results are obtained.

[0098] Among them, intramodal feature enhancement refers to the process of strengthening key features related to clinical diagnosis and treatment and suppressing redundant noise information for the characteristics of a single modality of data; cross-modal semantic calibration refers to the process of establishing semantic associations between different modalities of data and eliminating differences in the expression of heterogeneous data; semantic alignment refers to mapping data from different modalities to a unified semantic space, so that the data from each modality are comparable in the same logical dimension.

[0099] In some implementations, intramodal feature enhancement methods include local feature extraction based on convolutional neural networks (CNNs), temporal feature enhancement based on recurrent neural networks (RNNs), key information focusing based on attention mechanisms, and data standardization based on feature normalization; cross-modal semantic calibration methods include modal association modeling based on canonical correlation analysis (CCA), semantic mapping based on contrastive learning, and cross-modal feature transformation based on dictionary learning; semantic alignment methods include vector space mapping, semantic similarity matching, entity association alignment, and feature dimension unification.

[0100] It should be noted that the value of the two-stage alignment algorithm lies in first improving the feature quality of single-modal data, and then establishing deep semantic associations between modalities, avoiding the loss of key information or meaningless associations caused by direct fusion, which can effectively solve the semantic gap problem of multi-source heterogeneous medical data.

[0101] For example, for medical image data, lesion edge features can be extracted using CNN, and then key information such as lesion size and density can be enhanced through attention mechanism to complete intramodal enhancement; for electronic medical record text data, text can be converted into vector features through word embedding technology, and then intramodal enhancement can be completed by combining it with a medical terminology dictionary; subsequently, contrastive learning is used to align the image feature vector and the text feature vector in a unified space to achieve cross-modal semantic calibration.

[0102] S3. Based on knowledge graph relevance, clinical guideline priority, and modality quality score, a dynamic weight model is used to perform feature fusion on semantic alignment results.

[0103] Among them, knowledge graph correlation refers to the degree of correlation between the features corresponding to each modality of data and entities such as target diseases and clinical indicators, which is quantified by the correlation between entities in the medical knowledge graph; clinical guideline priority refers to the definition of the importance of different modalities of data in disease diagnosis and treatment assessment in clinical practice guidelines; dynamic weight model refers to a model that can adaptively adjust the contribution ratio of each modality feature in the fusion process based on multiple dimensions such as data quality, clinical scenario, and knowledge correlation, which is different from the fixed weight allocation method.

[0104] In some implementations, common forms of dynamic weight models include weight calculation models based on statistical rules, weight prediction models based on machine learning, weight optimization models based on reinforcement learning, and comprehensive decision-making models based on multi-factor weighting; feature fusion methods include weighted summation fusion, attention mechanism fusion, gating fusion, feature concatenation fusion, and probabilistic model fusion.

[0105] It should be noted that the dynamic weight model, which combines knowledge graph relevance, clinical guideline priority, and modality quality score, can make the fusion process conform to the medical knowledge system, adapt to differences in data quality and clinical scenario needs, and make the fusion result more clinically applicable.

[0106] For example, in the context of lung cancer diagnosis, if pathological data has a higher priority than imaging data in clinical guidelines, the dynamic weight model will increase the fusion weight of pathological modalities; if a patient's imaging data quality score is significantly higher than that of the test data, the model will increase the contribution ratio of imaging features accordingly; at the same time, combined with the high correlation between "ground-glass nodules" and "lung cancer" in the knowledge graph, the weights of the corresponding features will be further adjusted.

[0107] S4. The AIGC model generates a structured diagnostic report from the fusion results, performs clinical compliance verification and sensitivity analysis, and feeds the verification and analysis results back to the feature fusion step for optimization.

[0108] AIGC stands for Artificial Intelligence Generated Content. The AIGC model learns the structured format of clinical reports, medical terminology, and diagnostic logic, and transforms the fused high-dimensional feature vectors into structured diagnostic reports that conform to clinical standards. The reports typically include modules such as basic patient information, multimodal data summaries, diagnostic conclusions, core evidence, and clinical recommendations.

[0109] In some implementations, optimization methods include adjusting the parameters of the dynamic weight model based on clinical compliance verification results, optimizing the threshold setting for semantic alignment based on sensitivity analysis results, revising the report generation rules of the AIGC model in conjunction with expert feedback, and improving the matching degree between feature fusion and report generation through iterative training.

[0110] It should be noted that clinical compliance verification is used to ensure that the report complies with medical industry standards and clinical practice guidelines, avoiding irregular statements or logical contradictions; sensitivity analysis is used to verify the stability of diagnostic results under small data perturbations, ensuring the robustness of the method; and the feedback optimization mechanism forms a closed loop of "fusion-reporting-verification-optimization", continuously improving the performance of the overall solution.

[0111] For example, the AIGC model can transform the fused feature vector into a structured statement based on a pre-trained medical language model: "The patient has ground-glass nodules in the lungs. Combined with elevated CEA levels and a history of smoking, the lung cancer risk assessment is intermediate to high risk, and further pathological biopsy is recommended." Clinical compliance verification is performed by comparing with NCCN guidelines and correcting statements in the report that do not meet the risk grading standards. Sensitivity analysis is performed by slightly adjusting the imaging feature parameters and observing the fluctuations in the diagnostic risk assessment results. If the fluctuations exceed the threshold, the process returns to the fusion step to adjust the weight allocation.

[0112] Based on the above technical solutions, the medical multimodal data fusion and analysis method based on cross-modal alignment provided in this application employs a quality grading assessment stage that adapts to the quality characteristics of different modalities with targeted assessment logic, ensuring the reliability and standardization of input data from the source. The two-stage alignment algorithm, through the collaborative design of intra-modal enhancement and cross-modal calibration, addresses the semantic gap in heterogeneous data, strengthening the clinical core relevance of the alignment results. The dynamic weight model and feature fusion method deeply integrate the association patterns of medical knowledge graphs with the core requirements of clinical guidelines, improving the clinical adaptability and scientific rigor of the fusion results. The AIGC report generation and closed-loop optimization mechanism comprehensively ensure the clinical compliance, diagnostic robustness, and continuous iteration capability of the output results. The overall method constructs a full-process intelligent optimization system covering data preprocessing, semantic alignment, feature fusion, and report generation, flexibly adapting to the diagnostic and treatment scenarios of different medical institutions and the multimodal data fusion needs of various diseases, providing an efficient, universal, and reliable decision support solution for smart healthcare.

[0113] In one possible implementation of this application embodiment, the above-mentioned S1 can be specifically implemented by the following S101, S102 and S103, which are described in detail below:

[0114] S101. Calculate the comprehensive score according to the modality.

[0115] The comprehensive score is a quantitative score obtained by weighted summation based on the core quality indicators of each modality. It is used to objectively reflect the quality level of single-modality data. The selection of core indicators needs to be in line with the essential characteristics of each modality data.

[0116] In some implementations, the comprehensive score calculation for different modalities adopts a targeted index system, as follows:

[0117] 1. Comprehensive scoring of medical imaging data The calculation formula is: ;in,

[0118] Spatial coverage, with a value range of 0-10, is calculated as "the number of pixels actually covered by the lesion area ÷ the total number of pixels in the entire medical image × 10", and is used to characterize the completeness of the image's coverage of the lesion.

[0119] The first information density score, ranging from 1 to 10, is calculated using the information entropy algorithm, and the formula is as follows: denoted as , where is the probability of occurrence of the i-th clinically valid feature in the image (such as lesion size, density, boundary, calcification, etc.), and k is the total number of clinically valid feature categories in the image; information density score. , It represents the maximum standard information entropy of similar images, used to quantify the richness of effective clinical information in images.

[0120] 2. Comprehensive scoring of electronic medical record text data The calculation formula is: ;in,

[0121] The first text indicator has a value range of 0-10 points and is calculated as "the number of core fields that are filled in completely and effectively ÷ the total number of core fields × 10". Core fields include essential fields for clinical diagnosis and treatment such as chief complaint, present illness, past medical history, diagnosis results, medication history, and examination indicators.

[0122] The second text indicator has a value range of 0-10 points and is calculated as "the number of matches between the text terms in the medical record and the standard terms of Systematized Nomenclature of Medicine - Clinical Terms (SNOMED CT) ÷ the total number of all relevant terms in the medical record × 10". The matching process is automatically compared through a term mapping dictionary to ensure the standardization of terms such as disease, symptoms, and procedures.

[0123] The third text indicator has a value range of 0-10 points and is calculated as "the number of fields in the electronic medical record text format that conform to the Health Level 7 (HL7) medical data exchange standard ÷ the total number of fields in the medical record × 10". The HL7 standard ensures data interoperability between different systems.

[0124] The second information density score, ranging from 1 to 10, is calculated using the information entropy algorithm, and the formula is as follows: ,in represents the probability of occurrence of the i-th type of valid medical information in the medical record (such as symptom description details, duration of medical history, degree of abnormality in examination results, etc.), where m is the total number of categories of valid medical information; information density score. , This represents the maximum standard information entropy for medical records of the same type.

[0125] 3. Comprehensive scoring of physiological signal time series data The calculation formula is:

[0126] ;in,

[0127] The score represents the continuity of the time sequence, ranging from 0 to 10. It is calculated as "the length of the time period without missing or abnormal interruption within the physiological signal acquisition period ÷ the total signal acquisition time × 10". No abnormal interruption means that the signal value fluctuation is within the clinically reasonable range. A single missing time exceeding 10 seconds is considered an interruption. Among them, a signal without missing or abnormal interruption is a continuous signal.

[0128] The third information density score, ranging from 1 to 10, is calculated using the information entropy algorithm. The formula is as follows: ,in Let be the probability of occurrence of the i-th type of effective physiological indicator (such as heart rate, blood oxygen saturation, tumor marker concentration, blood pressure, etc.) in the physiological signal, and p be the total number of categories of effective physiological indicators; information density score. , It represents the maximum standard information entropy of physiological signals of the same type.

[0129] It should be noted that all three types of modal data must meet the standardized format requirements: medical imaging data should be in DICOM 3.0 format, electronic medical record text data should be in structured XML format, and physiological signal time series data should be in standardized JSON format. If the format does not meet the requirements, it must be converted first; otherwise, it will be directly judged as an extremely low-quality modality. All indicator scores are mapped to the 0-10 score range through linear scaling to ensure the comparability of the comprehensive scores.

[0130] S102. Quality grading based on comprehensive scores.

[0131] Quality grading is a process of dividing data into different quality levels based on the comprehensive score of each modality and a preset level threshold, with the aim of providing a basis for subsequent differentiated processing.

[0132] In some implementations, a four-level grading standard can be adopted. The preset grade thresholds are determined based on statistical analysis of multi-center clinical data and expert consensus among multiple senior physicians. The specific grading rules are as follows:

[0133] High-quality modality: A comprehensive score of ≥8 points indicates extremely high data quality with no obvious defects, and can be directly used for subsequent analysis;

[0134] Qualified mode: 6 ≤ comprehensive score < 8 points, which means that the data quality is good, but there are minor defects, such as a few missing non-core fields or short-term signal fluctuations. After simple optimization, it can meet the analysis requirements.

[0135] Low-quality modality: 3 ≤ overall score < 6 points, indicating poor data quality and obvious defects, such as missing core fields, incomplete coverage of lesions in the image, and multiple signal interruptions, requiring in-depth processing or supplementation of key information;

[0136] Extremely low quality modality: The overall score is <3 points, which means that the data quality is extremely poor, the core information is seriously missing or distorted, and there is no value in processing it.

[0137] S103. Set differentiated post-processing strategies for each quality level.

[0138] Among them, the differentiated post-processing strategy is to design appropriate optimization, supplementation or discard schemes for modal data of different quality levels. The core goal is to maximize the use of high-quality data, repair medium and low-quality data, and remove invalid data to ensure the reliability of subsequent analysis.

[0139] In some implementations, the specific post-processing strategies for each level are as follows:

[0140] High-quality modalities: No additional processing is required; they can be directly input into the subsequent two-stage alignment algorithm to retain the core clinical information of the original data. At the same time, data quality labels are recorded for use in assigning modal quality score vectors in the subsequent dynamic weight model.

[0141] Qualified modality: Employs a lightweight enhancement strategy based on Artificial Intelligence Generated Content (AIGC).

[0142] Medical imaging data: Enhanced super-resolution generative adversarial network (ESRGAN) optimizes image clarity, removes minor artifacts, and maintains the authenticity of lesion features;

[0143] Electronic medical record text data: The medical AIGC model is used to correct ambiguities in terminology and supplement missing non-core fields without changing the core diagnosis and treatment information;

[0144] Physiological signal time series data: A small amount of interpolated data is generated by generative adversarial network (GAN) to fill in short time gaps and ensure the continuity of time series.

[0145] Low-quality modalities: Employing a "data completion + deep augmentation" strategy:

[0146] Medical imaging data: The missing areas of the images are filled by the mean lesion features of patients with the same disease and the same stage, and the edge features of the lesions are enhanced by the U-Net denoising network to improve the lesion recognition.

[0147] Electronic medical record text data: Based on medical knowledge graphs and similar medical records, core diagnosis and treatment information is reconstructed using the MedicalBERT+AIGC model. The reconstructed content needs to be verified through terminology standardization.

[0148] Physiological signal time series data: Long-term missing segments are filled in using a Long Short-Term Memory (LSTM) network, and abnormal fluctuation values ​​are corrected using a signal smoothing algorithm to ensure that the data conforms to clinical physiological laws.

[0149] Extremely low quality modes: discard them directly, record the reason for discarding, and feed it back to the data acquisition end.

[0150] It should be noted that AIGC-driven enhancement or replacement processes must comply with clinical compliance requirements, the generated data must meet medical data standards, and the generated diagnostic and treatment information must be within the clinically reasonable range. The quality of low-quality modal processing must be reassessed, and only those with a comprehensive score of ≥6 after reassessment can proceed to the next step; otherwise, they will still be processed as low-quality modal or discarded.

[0151] Based on the above technical solutions, S1 achieves a systematic assessment and optimization of the quality of multimodal medical data. This process takes into account the differences in characteristics of different modalities of data, ensuring the objectivity of the scoring through targeted indicators and formulas, and achieving precise control of data quality through grading and post-processing strategies. This effectively avoids the interference of low-quality data on subsequent semantic alignment and feature fusion, laying a reliable data foundation for the entire multimodal data fusion and analysis process.

[0152] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 3 As shown, the above S2 can be implemented through the following S201, S202 and S203, which are explained in detail below:

[0153] S201. Perform intra-modal feature enhancement to generate enhanced feature blocks for each modality.

[0154] Intramodal feature enhancement is a process that strengthens the core clinical features and suppresses redundant noise in a single modality by using a pre-trained neural network model for data after quality grading. The generated enhanced feature blocks (i.e., tokens) are the basic inputs for subsequent cross-modal semantic calibration.

[0155] In some implementations, each modality employs a dedicated pre-trained model and a standardized processing flow, as detailed below:

[0156] First, the input multimodal data must meet the preset format requirements. For example, medical imaging data is in DICOM 3.0 format, electronic medical record text data is in structured XML format, and physiological signal time series data is in standardized JSON format. If the formats do not match, they need to be unified using medical data conversion tools.

[0157] Next, modal feature enhancement is performed:

[0158] For medical image data, a pre-trained MedViT (Medical Vision Transformer) model was used for intramodal feature enhancement. The specific process is as follows: First, basic features were extracted through a 4-layer multi-layer convolutional block. The kernel size of this convolutional block was set to 3×3, stride 1, and padding 1, which can accurately capture basic visual features such as local texture and grayscale distribution of lesions, laying the foundation for subsequent feature enhancement. Next, a lesion attention mechanism was constructed through CBAM (Convolutional Block Attention Module). This mechanism automatically identifies and highlights areas with high diagnostic value in the image through the synergistic effect of channel attention and spatial attention, focusing on enhancing key features closely related to clinical diagnosis, such as lesion size, density, boundary clarity, and calcification, while suppressing interference from irrelevant background information. Finally, global feature encoding was performed through a 6-layer Transformer encoder, effectively integrating the spatial correlation information between the lesion and surrounding tissues to form a high-dimensional feature representation that combines local details and global context, ultimately outputting a 256-dimensional image feature block.

[0159] For electronic medical record text data, feature enhancement is performed using a pre-trained MedicalBERT (Medical Bidirectional Encoder Representations from Transformers) model. The specific process is as follows: First, word segmentation and standardized terminology mapping are performed. The medical record text is first segmented, and then the segmented natural language terms are accurately matched and mapped with the standard terms of the Medical System Nomenclature - Clinical Terminology (SNOMED CT). For example, "lung cancer" is standardized as "lung cancer" to ensure the standardization and consistency of text terminology. Subsequently, BiLSTM-CRF (Bidirectional Long Short-Term Memory-Conditional Random Access Memory) is used. The Field (Bidirectional Long Short-Term Memory-Conditional Random Field) sub-model is used for medical entity recognition. This sub-model can accurately extract core medical entities from medical records, including key diagnostic and treatment information such as disease name, symptoms, past medical history, and medication history, providing core elements for semantic encoding. Finally, the model performs semantic encoding through a 12-layer Transformer structure, deeply mining the semantic associations and logical relationships between entities, transforming text information into structured high-dimensional semantic features, and finally outputting 128-dimensional text feature blocks.

[0160] For time-series physiological signal data, a pre-trained multilayer perceptron (MLP) combined with a clinical indicator attention mechanism is used for feature enhancement. The specific process is as follows: First, the original time-series values ​​are encoded and transformed through a 3-layer MLP. The hidden layer dimensions of the MLP are set to 256, 128, and 64 respectively, and the activation function is ReLU. This transforms the discrete original physiological signal values ​​into a high-dimensional feature representation with a hierarchical structure. Next, a clinical indicator attention mechanism based on the weights initialized according to the NCCN clinical guidelines is introduced. This mechanism automatically enhances the feature contribution of core physiological indicators such as heart rate, blood oxygen saturation, and tumor marker concentration according to the clinical diagnosis and treatment priority, while effectively suppressing the interference of non-critical indicators, ensuring that the feature expression conforms to clinical needs. Finally, through feature dimensionality reduction, the feature dimensions are simplified while retaining the core effective information, and a 32-dimensional numerical feature block is finally output.

[0161] It should be noted that each modality pre-trained model is fine-tuned and optimized based on a multi-center medical dataset to ensure that the model is adapted to the distribution of clinical data. Furthermore, during the data collection process, strict adherence to relevant laws and regulations such as the *Personal Information Protection Law of the People's Republic of China*, the *Guidelines for Medical Data Security*, and the *Measures for the Management of Cybersecurity in Medical and Health Institutions*, as well as international privacy protection standards (such as HIPAA), is maintained, implementing full-process privacy protection measures. Specifically, sensitive patient information (including name, ID number, contact information, home address, etc.) is thoroughly de-identified, removing all fields that can directly or indirectly link to personal identity; a professional anonymization algorithm is used to separate the unique mapping between data and patients, retaining only feature information related to diagnosis and treatment; end-to-end encryption technology is used in the data transmission stage, and a hierarchical authorization access and encrypted storage mechanism is implemented in the storage stage to strictly prevent data leakage, tampering, or misuse; all data collection work is reviewed and approved in advance by the ethics committee of the multi-center medical institution, and written informed consent is obtained from patients in accordance with the law, clearly informing them of the scope and purpose of data use, ensuring the legality and compliance of data collection and use, and ensuring that no patient privacy information is collected, retained, or disclosed throughout the process.

[0162] S202. Verify the effectiveness of each modal enhancement feature and select qualified enhancement feature blocks.

[0163] Among them, feature validity verification is to evaluate the clinical value of enhanced feature blocks through quantitative indicators, and to remove low-quality feature blocks to avoid interfering with subsequent semantic calibration. The core indicators are the proportion of key feature dimensions and feature signal-to-noise ratio. The verification results directly determine whether a feature block enters the next stage.

[0164] In some implementations, the formula for calculating feature validity verification is as follows: ;in,

[0165] The feature effectiveness score is used to quantify the clinical value and quality level of the enhanced feature block. The value ranges from 0 to 1. The higher the score, the more the feature block meets the requirements of subsequent semantic calibration.

[0166] w1 is the weighting coefficient of the key feature dimension, with a value range of 0.6-0.8 (default is 0.7). It is determined by the consensus of multi-center clinical experts in combination with different disease diagnosis and treatment scenarios, and is used to highlight the dominant role of the feature dimension corresponding to the core clinical indicators in the effectiveness evaluation.

[0167] The key feature dimension proportion refers to the proportion of feature dimensions representing core clinical indicators in each modality feature block to the total feature dimensions, with a value ranging from 0 to 1; among which, The number of feature dimensions corresponding to core clinical indicators is determined by screening relevant clinical guidelines for the target disease and confirming them in conjunction with expert consensus, focusing on the feature dimensions corresponding to indicators directly related to disease diagnosis, treatment and prognosis. The total dimension of the feature block is, for example, 256 dimensions for medical images, 128 dimensions for electronic medical record text, and 32 dimensions for physiological signal time series data, which is determined by the model output specifications in the intramodal feature enhancement stage.

[0168] The feature signal-to-noise ratio refers to the effective feature variance. With noise characteristic variance The ratio of ; where is the effective feature variance, obtained through a combined algorithm of mutual information (MI) and analysis of variance (ANOVA)—first, calculate the mutual information value between each feature and the core clinical indicators, screen candidate features with MI≥0.6, then perform ANOVA test on the candidate features, remove features with no statistical difference (P>0.05), and calculate the variance of the remaining effective features; The feature data is obtained by combining threshold screening and dynamic anomaly detection. First, a feature threshold is set based on a clinically reasonable range to remove abnormal features that exceed the range. Then, the isolated forest algorithm is used to model the feature data, and features with anomaly scores ≥0.8 are selected as noise features. Finally, the variance is calculated.

[0169] lg is a logarithmic function with base 10, used to compress the range of feature signal-to-noise ratio values ​​to ensure the balance of evaluation results.

[0170] It should be noted that the core clinical indicators need to be dynamically adjusted according to the target disease. For example, tumor diagnosis focuses on lesion characteristics, and cardiovascular disease focuses on cardiac function indicators, to ensure that the selection of key feature dimensions is appropriate for the specific diagnosis and treatment scenario.

[0171] S203. Perform cross-modal semantic calibration to generate a unified semantic vector.

[0172] Among them, cross-modal semantic calibration eliminates the semantic gap between different modalities through medical knowledge graph constraints and multi-module collaborative processing, maps qualified enhanced feature blocks to a unified semantic space, and finally generates a unified semantic vector that conforms to clinical logic, i.e., semantic alignment result.

[0173] In some implementations, cross-modal semantic calibration is performed in a three-tiered chain: a medical memory generation layer, a cross-modal attention layer, and a semantic consistency verification layer. The specific operations are as follows:

[0174] The input to the medical memory generation layer includes a medical knowledge graph and medical pairing data. Through a generative adversarial network (GAN) generator driven by artificial intelligence-generated content (AIGC), it learns the clinical association patterns between various modal features and disease entities in the medical pairing data, and outputs a three-dimensional medical memory matrix of "modality-terminology-disease". (Dimension 512×512), the element values ​​of this matrix It represents the association strength between the modal features of type a and the medical entities of type b. The closer the value is to 1, the stronger the association, providing a solid medical constraint basis for subsequent semantic mapping.

[0175] The medical pairing data represents a clinical pairing of image-text-physiological signal triplets, with each triplet accompanied by a clinically confirmed diagnostic association label. For example, the image data is “chest CT image (8mm ground-glass nodule in the upper lobe of the right lung, with blurred borders)”, the text data is “electronic medical record (cough with hemoptysis for 1 month, smoking history for 20 years)”, and the physiological signal data is “time-series detection data (serum CEA 6.8 ng / mL)”. The corresponding clinically confirmed diagnostic association label is “lung adenocarcinoma (IA1 stage): ground-glass nodule + elevated CEA + smoking history”.

[0176] After constructing the clinical semantic constraint matrix, the cross-modal attention layer is used to achieve cross-modal semantic mapping. Its inputs are qualified enhanced feature blocks for each modality and a three-dimensional medical memory matrix. And the embedding of target disease candidate set entities in medical knowledge graphs;

[0177] First, the core vector of the attention mechanism is constructed by concatenating the enhanced feature blocks of images (256-dimensional), text (128-dimensional), and physiological signals (32-dimensional) to form a key vector. (Dimensions 316×512) and value vector (Dimensions 316×512), query vector (Dimension 512×512) is composed of the embedding of the target disease candidate set entities, and the target disease candidate set is determined by cosine similarity matching (similarity > 0.75) between the chief complaint, present medical history and other information in the electronic medical record text and the "disease" entity in the knowledge graph.

[0178] Next, the basic attention weights are calculated using the following formula: ,in The basic attention weight matrix (dimension 512×316). =512 is the feature dimension, used to alleviate the curse of dimensionality. For matrix multiplication, This is the transpose of the key vector V_k; subsequently, medical constraint weights are adjusted using the formula... Attention weight matrix after obtaining medical constraints (Dimensions 512×316), where For element-wise matrix multiplication, using Filter out cross-modal associations that have no clinical significance;

[0179] Finally, through the formula Generate preliminary unified semantic vectors (Dimensions 512×1).

[0180] After obtaining the initial unified semantic vector, dynamic calibration is performed through a semantic consistency verification layer, the input of which includes the initial unified semantic vector. and ideal alignment vectors for each mode (i=1 corresponds to image, i=2 corresponds to text, i=3 corresponds to physiological signal);

[0181] First, calculate the semantic consistency score. The formula is ,in, For vector dot product operation, It is the L2 norm (magnitude) of the vector. The ideal alignment vector generated based on entities and relationships in the medical knowledge graph serves as a standard reference for semantic calibration.

[0182] Then the dynamic calibration logic is executed: when When the value is ≥0.85, the semantic alignment is considered acceptable, and the result is output directly. As the final unified semantic vector; when 0.7 ≤ When <0.85, adjust Elements with a correlation strength ≤ 0.3 are set to 0, and the cross-modal attention layer calculation is re-executed; when If the value is less than 0.7, return to S201 to re-execute intramodal feature enhancement, or delete the low-quality modal data from the data quality grading assessment stage that cannot be optimized by secondary enhancement, and then perform intramodal feature enhancement and subsequent processes.

[0183] It should be noted that the process of constructing a medical knowledge graph includes:

[0184] First, medical data sources are collected, covering a wide range of authoritative medical resources, including core industry guidelines such as the National Comprehensive Cancer Network (NCCN) Clinical Guidelines and Chinese Clinical Practice Guidelines. In addition, medical journal articles, real clinical practice data accumulated by medical institutions, and expert consensus in the clinical field are incorporated to fill the gaps in empirical medical knowledge not covered in guidelines and literature, forming a multi-dimensional and comprehensive data source system.

[0185] Secondly, based on technologies such as natural language processing and machine learning, multiple core medical entities are systematically extracted from the aforementioned multi-source data, covering entity types directly related to clinical diagnosis and treatment, such as diseases, symptoms, laboratory indicators, imaging features, drugs, and physiological signals. Then, the clinical relationships between entities are explored, including but not limited to the correspondence between "disease-symptom", the treatment relationship between "disease-drug", the diagnostic relationship between "disease-laboratory indicator", and the relationship between "imaging feature-disease". All relationships are based on clinical diagnosis and treatment logic to ensure that they conform to medical common sense and actual application scenarios.

[0186] Next, the extracted entities are uniformly labeled using industry-standardized coding and terminology systems. For example, the medical system nomenclature-clinical terminology (SNOMED CT) is used to standardize and unify entity terminology; ICD-10 (International Classification of Diseases, 10th Revision) disease codes are used to standardize the identification of disease-related entities; and LOINC (Logical Observation Identifiers Names and Codes) test indicator codes are used to uniformly name test-related entities. At the same time, the extracted clinical associations are assigned a level of clinical evidence, which is specifically reflected in the amount of clinical research data supporting the association (a higher level is assigned if supported by a large-sample, multi-center study) and the degree of expert consensus (a higher level is assigned if validated by an authoritative expert team). The credibility of the association is clarified by quantifying the level.

[0187] Finally, a graph database is used as the storage medium, standardized and labeled entities are used as graph nodes, and clinical relationships with hierarchical assignments are used as edges between nodes to construct a structured medical knowledge graph.

[0188] In some implementations, the ideal alignment vector for each modality is obtained through knowledge graph entity association constraints and modality feature alignment training processes. The core is to deeply bind modality features with the clinical logic of the medical knowledge graph to form a standardized semantic alignment reference. The specific steps are as follows:

[0189] Step 1: Extract the core associated entity set of the target disease and bind it to the modality.

[0190] First, entities directly related to the target disease are screened from the constructed medical knowledge graph to form a core set of associated entities. These entities include, but are not limited to, typical symptoms, key test indicators, characteristic imaging manifestations, and core physiological signals corresponding to the target disease. The target disease refers to the specific disease to be diagnosed, assessed, or predicted in the multimodal data fusion analysis process of this application. Its determination depends on the core clinical information in the input multimodal medical data—specifically, it is the disease to be analyzed obtained by extracting core clues from key diagnostic and treatment information such as the chief complaint and present medical history from electronic medical record texts, and then performing semantic similarity matching (e.g., cosine similarity > 0.75) with the "disease" entity in the medical knowledge graph. The target disease is the core guideline for the entire multimodal data fusion analysis; subsequent core associated entity screening, ideal alignment vector construction, and cross-modal semantic calibration all revolve around it, ensuring that all processing procedures align with specific diagnostic and treatment needs.

[0191] Subsequently, the core associated entity set is precisely bound to the corresponding modality to form a "modality-entity" mapping table (such as "image modality-ground-glass nodules", "text modality-coughing and hemoptysis", "physiological signal modality-elevated CEA"), ensuring that each core entity can correspond to a specific modality, laying the foundation for the subsequent alignment of modality features with knowledge graph entities.

[0192] Step 2: Obtain the basic feature vectors for each modality.

[0193] Historical multimodal medical data is collected, and intramodal feature enhancement is performed to obtain enhanced features for each modality. These enhanced features are then labeled as the basic feature vectors for each modality. , where the subscript i corresponds to the i-th entity in the core associated entity set;

[0194] Step 3: Obtain the knowledge graph embedding vector of the core entity.

[0195] The TransE algorithm was used to pre-train the medical knowledge graph. This algorithm learns the semantic association patterns between entities by mapping entities and relations to a low-dimensional vector space. Through pre-training, the knowledge graph embedding vector of each entity in the core associated entity set was extracted from the knowledge graph. This vector is a quantitative representation of entity semantics, which can reflect the position and relationship of the entity in the medical knowledge system.

[0196] Step 4: Determine the training weights of the core entities.

[0197] Extract the association level between each entity in the core associated entity set and the target disease. This level has been assigned a value based on clinical evidence during the medical knowledge graph construction stage, with a value range of 1-5 points. The closer the association and the more sufficient the clinical evidence, the higher the level. For example, the association level of "ground-glass nodules-lung cancer" is 5 points, and the association level of "mild cough-lung cancer" is 2 points.

[0198] The association level is directly converted into training weights. That is, the training weights of the i-th core entity. This is equivalent to the level of its association with the target disease;

[0199] Step 5: Construct the modality feature alignment training network and loss function.

[0200] A modal feature alignment training network is constructed based on fully connected layers. The input of this network is the basic feature vector of each core entity. Knowledge graph embedding vectors and training weights The goal is to align modal feature vectors with knowledge graph embedding vectors through training, so that the modal features contain standardized medical semantics.

[0201] Simultaneously, a loss function is constructed, with the following formula:

[0202] Where L represents the loss function value, used to quantify the difference between the modality base feature vector and the knowledge graph embedding vector; the smaller the value, the higher the degree of alignment between the two; n is the total number of entities in the core associated entity set; This represents the square of the L2 norm, used to calculate the squared Euclidean distance between two vectors, measuring the degree of difference between the vectors.

[0203] It should be noted that this loss function minimizes the difference between the modal feature vector corresponding to each core entity and the knowledge graph embedding vector through a weighted summation, and the training weights... It will amplify the alignment priority of entities with high correlation levels, ensuring that the alignment results conform to clinical logic.

[0204] Step 6: Perform modal feature alignment training to obtain the core entity alignment vector.

[0205] Will , as well as The input modality feature alignment training network employs gradient descent to iteratively optimize the training process, continuously minimizing the loss function L. Training stops when the loss function converges to a preset threshold. At this point, the network output vector is the modality alignment vector corresponding to the i-th core entity. This vector retains the original information of the modality features while incorporating the standardized semantic constraints of the knowledge graph.

[0206] Step 7: Weighted fusion generates ideal alignment vectors for each modality.

[0207] First, calculate the weight percentage of each core entity according to the formula. : ;in, This represents the weight percentage of the i-th core entity;

[0208] Subsequently, the alignment vectors of all core entities under the same modality are weighted and fused, and the weighting coefficients are the weight proportions of each core entity. Finally, the ideal alignment vector for this mode is obtained. This vector is a comprehensive representation of the alignment vectors of all core entities under the same modality. It integrates the clinical association logic of the knowledge graph with the essential attributes of modal features, and provides a standardized semantic reference benchmark for cross-modal semantic calibration.

[0209] Based on the above technical solution, S2 achieves deep semantic alignment of multimodal data through intramodal feature enhancement, validity verification, and cross-modal semantic calibration. Intramodal feature enhancement provides a high-quality feature foundation for semantic calibration, validity verification ensures the clinical value of input features, and cross-modal semantic calibration effectively eliminates the semantic gap through medical knowledge graph constraints and multi-module collaboration.

[0210] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 4 As shown, the above S3 can be implemented through the following S301, S302 and S303, which are explained in detail below:

[0211] S301. Obtain the multi-dimensional input vector required for feature fusion.

[0212] In some implementations, the definitions, acquisition methods, and quantization standards of each input vector are as follows:

[0213] 1. Unified Semantic Vector As the core output of semantic alignment, it is a multi-dimensional vector obtained after cross-modal semantic calibration, which integrates the deep clinical semantic features of three modalities: medical images, electronic medical record text, and physiological signals, and is extracted from the semantic alignment results of S2.

[0214] 2. Medical Knowledge Graph Relevance Vector R: Represents the degree of clinical association between each modality and the target disease. The vector dimension is 3 (corresponding to the three modalities), and the value range is 1-5 (5 points for strong association, 1 point for weak association). It is obtained by extracting the association level between the entities corresponding to the core features of each modality and the entities of the target disease from the medical knowledge graph, and taking the average of the association levels of all core entities under that modality as the vector element value. For example, the core entity "ground-glass nodule" in the medical imaging modality has an association level of 5 with "lung adenocarcinoma," and an association level of 4 with "pleural traction sign." Therefore, the relevance vector element of the imaging modality is (5+4)÷2=4.5.

[0215] 3. Clinical Guideline Priority Vector This reflects the definition of the importance of each modality of data in disease diagnosis by clinical practice guidelines. The vector dimension is 3, and the values ​​are standardized priority coefficients based on NCCN clinical guidelines and Chinese clinical practice guidelines. Specific values ​​are based on the following criteria: medical imaging modality, which directly presents lesion morphology, has a priority coefficient of 1.8; electronic medical record text modality, which carries core medical history and symptom information, has a priority coefficient of 1.5; and physiological signal temporal sequence modality, which provides auxiliary diagnostic indicators, has a priority coefficient of 1.2. These values ​​have been confirmed by consensus among multiple clinical experts.

[0216] 4. Modal Quality Score Vector Q: This vector reflects the quality level of each modality's data. It has a dimension of 3 and a value range of 1-10. It is directly extracted from the quality grading assessment results of S1, i.e., the comprehensive score of each modality. , , These are respectively used as the three elements of the vector.

[0217] It should be noted that all input vectors must ensure dimensional consistency and value standardization: the unified semantic vector must be standardized to the range of [-1,1], and the relevance vector and quality score vector must be linearly mapped to the range of [1,10] to avoid deviations in weight calculation due to differences in numerical ranges.

[0218] S302. Calculate the final modal weights based on the dynamic weight model.

[0219] Among them, the dynamic weight model is a core module that adaptively adjusts the contribution ratio of each modality in the fusion through multi-factor collaborative operation. Its value lies in breaking through the rigidity of static weights and making the weight allocation adapt to the dynamic changes of disease type, data quality and clinical standards.

[0220] In some implementations, the calculation process, formulas, and constraint rules for dynamic weights are as follows:

[0221] First, set the initial weight vector. This vector sets an initial proportion based on the clinical importance of each modality to ensure that the fusion starting point conforms to general diagnostic and treatment logic. Specifically, the values ​​are 0.35 for medical imaging modality, 0.4 for electronic medical record text modality, and 0.25 for physiological signal modality. =[0.35,0.4,0.25];

[0222] Next, preliminary weights are obtained through multi-factor element-level multiplication and normalization operations, as shown in the formula: , ,in, Provides a basic percentage, and R incorporates knowledge association constraints. This approach reflects guideline priorities, uses Q to reflect data quality, and employs a four-factor synergistic adjustment process. Normalization ensures the total weight of each factor is equal to 1, achieving a three-dimensional fit between data characteristics, knowledge constraints, and clinical standards. This is the initial weight vector. The kth element of the initial weight vector (k=1 corresponds to image, k=2 corresponds to text, k=3 corresponds to physiological signal). This is the final modal weight vector;

[0223] Finally, a weight constraint verification is performed. To avoid fusion imbalance caused by excessively high weights for a single modality or excessively low weights for core modalities, clear constraint rules are set: the weight of a single modality ≤ 0.5 to prevent a single modality from dominating the fusion result; medical images and electronic medical record texts are considered core modalities, with a weight ≥ 0.15; the weight of physiological signal modalities ≥ 0.1 to avoid excessive weakening of auxiliary indicators. If the weight fails the constraint verification, the weights exceeding the threshold are first truncated (e.g., if a modality's weight is calculated to be 0.55, it is truncated to 0.5), and then all elements are renormalized to a sum of 1 to ensure that the final weight distribution is balanced and conforms to clinical logic.

[0224] It should be noted that the calculation of dynamic weights responds in real time to changes in the input vector: when data quality fluctuates, disease types change, or clinical guidelines are updated, the weights are automatically recalculated without manual intervention.

[0225] S303, Execution feature fusion based on gating attention mechanism.

[0226] Feature fusion is a key step in integrating unified semantic vectors, dynamic weights, and medical knowledge constraints. It balances data-driven features and knowledge-driven constraints through gating mechanisms and highlights the core contributions of high-weight modalities through attention operations, ultimately generating a fused feature vector that combines data characteristics with clinical reliability.

[0227] In some implementations, the fusion process is as follows:

[0228] First, calculate the knowledge graph reasoning vector. This vector, serving as a quantifiable carrier of medical knowledge, provides rigid knowledge constraints for fusion. Its calculation is based on the embedding vectors of core entities in the medical knowledge graph, and the formula is as follows: This formula transforms the association strength between the target disease and the core entity into weights, and performs weighted aggregation on the knowledge graph embedding vectors of the core entities, so that... It contains the clinical correlation logic of "disease-entity"; among which, This is a 512-dimensional knowledge graph reasoning vector, where m is the total number of core entities associated with the target disease. The association level between the i-th core entity and the target disease is 1-5. The knowledge graph embedding vector for the i-th core entity;

[0229] Next, the gating coefficient G is calculated, which is used to balance the contribution ratio of the unified semantic vector (data-driven) and the knowledge graph reasoning vector (knowledge-driven). The formula is as follows: By using the Sigmoid activation function, the fused features are mapped to the [0,1] interval. The closer G is to 1, the greater the contribution of the unified semantic vector; the closer it is to 0, the stronger the constraint on the knowledge graph inference vector, thus achieving a dynamic balance between the two. Regarding the parameters... It is the Sigmoid activation function. For the gated weight matrix, This represents a vector concatenation operation;

[0230] Then, attention-weighted aggregation is performed, based on the final modality weights. The feature components corresponding to each modality in the unified semantic vector are weighted and enhanced, i.e., attention operations are performed. ;

[0231] Finally, a final feature fusion is performed, integrating the gating coefficient, attention-weighted features, and knowledge graph inference vectors, using the following formula: By using linear fusion to achieve deep integration of data features and knowledge constraints, a 512-dimensional fused feature vector is ultimately generated. This approach preserves the personalized characteristics of multimodal data while ensuring that the results conform to the medical knowledge system.

[0232] Based on the above technical solutions, S3 achieves deep integration of multimodal features through dynamic weight adaptive calculation and gated attention fusion. Specifically, the dynamic weight model balances data quality, medical knowledge, and clinical standards, enabling scenario-adaptive weight allocation; the gated attention mechanism balances data-driven and knowledge-driven approaches, preserving personalized data characteristics while avoiding clinically meaningless fusion results; the final generated fusion feature vector possesses accuracy, clinical relevance, and reliability, laying a solid foundation for subsequent structured diagnostic report generation.

[0233] In one possible implementation of this application embodiment, the above-mentioned S4 specifically includes the following S401 to S403:

[0234] S401. Generate structured diagnostic reports based on the AIGC model.

[0235] The AIGC model used in this step is an existing large-scale AIGC model specifically for medical use. Its input is a multimodal fusion multidimensional vector obtained after S1 multimodal medical data quality grading assessment, S2 two-stage alignment of intra-modal feature enhancement and cross-modal semantic calibration, and S3 dynamic weight feature fusion. That is, the entire process of S1-S3 completes the deep fusion processing of heterogeneous medical multimodal raw data such as medical images, electronic medical record text, and physiological signals. It transforms the raw multimodal data with different formats, dimensions, and semantic dispersion into a standardized fusion multidimensional vector that integrates medical knowledge graph association constraints, clinical guideline priorities, and modal quality features. The large-scale AIGC model specifically for medical use can complete a series of operations based on the fusion multidimensional vector, including accurate clinical semantic mapping, core diagnostic feature extraction, multimodal feature contribution quantification, and diagnosis and treatment logic sorting. Based on its own medical professional training foundation, it transforms the quantitative feature information contained in the fusion vector into professional natural language expressions that conform to clinical diagnosis and treatment standards, and then automatically generates a standardized structured diagnostic report containing core modules such as multimodal data summary, diagnostic results, feature contribution, and clinical suggestions.

[0236] It should be noted that compared to directly inputting unprocessed multimodal raw data into a large medical AIGC model, inputting it into the model after S1-S3 fusion processing has many significant advantages: First, at the data processing level, the fused standardized multidimensional vectors achieve semantic unification of heterogeneous multimodal data, fundamentally eliminating the problems of inconsistent raw data formats and semantic gaps between modalities. This significantly reduces the model's cost of parsing and processing multi-source heterogeneous data, improving the efficiency and accuracy of structured diagnostic report generation. Simultaneously, the fusion process extracts core information and removes invalid features, avoiding the waste of computational resources caused by the model processing large amounts of redundant raw data. At the data storage and transmission level… The fusion of multidimensional vectors results in a significantly lower dimensionality than the original multimodal data, greatly saving on storage space costs and cross-process transmission bandwidth costs. Furthermore, the standardized vector format offers enhanced retrieval and reusability, facilitating subsequent clinical data analysis and model iteration optimization. At the diagnostic report generation level, the fusion process incorporates medical knowledge graph correlation and clinical guideline priorities, ensuring the fusion vectors input to the model have a clear clinical logical orientation. This makes the generated diagnostic reports more aligned with actual clinical needs, reducing issues such as information fragmentation, scattered diagnostic evidence, and missing feature associations caused by individual input of raw data, thus enhancing the clinical reference value of the report content.

[0237] In some implementations, the specific operation process and key details of S401 are as follows:

[0238] First, the input data is preprocessed, and the 512-dimensional fused feature vector V generated by S3 is... fusion Standardization is performed to map it to the [-1,1] interval, and key information such as the modal quality score vector Q of S1 and the core feature associated entities in the semantic alignment result of S2 is associated with it.

[0239] Then, a medical-specific AIGC model is used. This model is built upon MedCLIP (Medical Contrastive Language-Image Pre-training) and a medically-tuned GPT (Generative Pre-trained Transformer), and consists of a three-level processing flow:

[0240] The first level is clinical semantic mapping, which uses MedCLIP to map V... fusion The high-dimensional feature elements in the graph are mapped one by one to a clinically understandable concrete expression. For example, "feature dimension 128 corresponds to a value of 0.85" is mapped to "the diameter of the ground-glass nodule in the lung is about 8mm and the boundary is blurred". This mapping rule is based on the "feature-term" association relationship in the medical knowledge graph that is homologous with the ideal alignment vector of S2.

[0241] The second level involves extracting core evidence. By fine-tuning the entity association module of GPT, it deeply mines the inherent logical relationship between various modal features and diagnostic results. For example, it achieves the clinical logical deduction of "imaging features (ground-glass nodules) + physiological signal features (elevated CEA) + textual features (20-year smoking history) → intermediate to high risk of lung cancer". At the same time, it directly reuses the final modal weight W of S3. final The contribution ratio of each modal feature is accurately quantified;

[0242] The third level is structured report generation. The model automatically organizes the various types of clinical information processed above according to clinical documentation standards. The generated diagnostic report may include: basic patient information (basic data such as patient ID, age, and gender from the associated medical system), multimodal data summary (extracting core abnormal information from images, text, and physiological signals), diagnostic results (including disease name, probability of benign or malignant / disease classification, and diagnostic confidence level), feature contribution (with contribution percentage of each modality and a list of key features), clinical recommendations (targeted recommendations for examination, treatment, or follow-up based on the diagnostic results), and compliance statement (marking the model version used to generate the report and a compliance statement for the data source). The basic patient information involved in the report generated here is all basic data that has been compliantly linked from the medical system during actual application. During the model training process, no patient privacy information or personally identifiable data is involved, and the entire process complies with relevant laws and standards for medical data privacy protection.

[0243] It should be noted that the medical-specific AIGC model is a generative artificial intelligence model specifically designed for medical scenarios, optimized and trained with medical data, and compliant with medical compliance standards. It can understand medical professional data such as medical images, electronic medical records, and physiological signals, generating content that aligns with clinical logic and is based on evidence. It also strictly adheres to privacy compliance requirements such as HIPAA and China's Personal Information Protection Law. Currently, publicly available versions are available, such as the commercially available and HIPAA-compliant OpenAI ChatGPT Healthcare Edition, and the open-source Google ChatGPT that supports multimodal processing and local deployment. The MedGemma series and the open-source model Baichuan-M1-14B, natively optimized for medical scenarios, are examples of existing large models. However, when these existing large models are applied to this application, the input is not the original medical data, but a fused feature vector obtained after quality grading assessment, two-stage semantic alignment, and dynamic weight feature fusion. Therefore, targeted improvements are needed: First, optimize the input adaptation module to accurately parse the multimodal clinical semantic relationships contained in the fused feature vector, rather than adapting to the original data format. Second, strengthen the clinical semantic mapping capability by accurately binding the high-dimensional elements of the fused feature vector with standardized medical terminology and diagnostic logic based on the medical knowledge graph constructed in this application, ensuring that the report fully restores the core multimodal information and relationships. Third, adapt to closed-loop optimization requirements by converting the results of clinical compliance verification and sensitivity analysis into feedback signals to continuously optimize the mapping accuracy from feature vector to clinical text. This ensures that the generated structured diagnostic report not only meets compliance requirements but also maintains consistency with the aforementioned fusion logic and knowledge graph constraints, efficiently outputting a professional report containing multimodal data summaries, diagnostic results, feature contribution, and clinical recommendations.

[0244] S402. Perform clinical compliance verification and sensitivity analysis.

[0245] Clinical compliance verification is used to ensure that reports comply with clinical guidelines and industry standards, while sensitivity analysis is used to verify the robustness of diagnostic results. Together, they constitute a dual quality control for report output, preventing irregular or unstable results from entering clinical practice.

[0246] First, guideline rules are validated. A clinical guideline rule base for the target disease is built based on Drools-Med (a medical-specific rule engine). The data sources for the rule base include NCCN clinical guidelines, Chinese clinical practice guidelines, and expert consensus. Each rule has a "trigger condition-correction logic," such as "If the diameter of the ground-glass nodule is <5mm and CEA ≤5ng / mL, then the probability of malignancy is ≤8%" or "If the pathology result is benign, then the diagnosis result must exclude malignancy." During validation, the rule engine traverses the report content, and after triggering the matching rule, it automatically corrects the non-compliant statements in the report. The correction magnitude is expressed by the formula P. corrected =P original ×α+P standard Calculate using ×(1-α), where P corrected P is the corrected probability value. original The original probability value P generated for AIGC standard The standard probability value is defined in the guidelines, and α is a correction coefficient with a range of 0.3-0.5, determined by clinical expert consensus.

[0247] Then, semantic verification of the knowledge graph is performed. The medical knowledge graph constructed in stage S2 is invoked, and the semantic consistency of "feature-diagnosis-recommendation" in the report is inferred through graph neural network (GNN). For example, if the report diagnoses "lung adenocarcinoma" but does not mention related imaging lesion features or tumor marker indicators, it is determined that the diagnostic evidence is insufficient, and the content of "suggesting supplementary detailed description of lesions and related test indicators" will be automatically added. If the report gives a clinical recommendation of "surgical treatment", and the diagnosis result is early lung cancer with no contraindications to surgery, then the three are determined to be semantically consistent.

[0248] After completing the clinical compliance verification, a sensitivity analysis was conducted using Monte Carlo simulation to verify the stability of the diagnostic results when there are minor perturbations in the fused feature vector, thus preventing the reversal of diagnostic conclusions due to minor data fluctuations. The specific simulation process is as follows:

[0249] First, the fused feature vector V... fusion Each dimension is randomly perturbed with a perturbation amplitude of ±5%, determined based on the normal fluctuation range of clinical data. This perturbation operation is then repeated 1000 times, generating 1000 perturbed fusion feature vectors V. fusion,k (k=1,2,...,1000);

[0250] Finally, each perturbed feature vector is input into the AIGC model to obtain 1000 corresponding diagnostic result probabilities P. k .

[0251] Then through the formula The relative error of the diagnostic results is calculated to quantify the range of fluctuation in the diagnostic results, where ΔP represents the relative range of fluctuation in the diagnostic results. This represents the maximum probability value across 1000 simulations. This represents the minimum probability value across 1000 simulations. This is the average probability of 1000 simulations.

[0252] At the same time, a clear stability judgment rule is set: the preset threshold is 10%. If the calculated ΔP ≤ 10%, the diagnosis result is determined to be stable and the sensitivity analysis is successfully passed; if it is greater than the threshold, the diagnosis result is determined to be unstable and the feature fusion operation needs to be re-executed in the S3 stage.

[0253] S403. Implement a feedback optimization mechanism to iteratively improve feature fusion performance.

[0254] Among them, feedback optimization transforms the clinical compliance verification results, sensitivity analysis results, and clinical expert feedback into optimization instructions, and adjusts the feature fusion parameters of S3 in reverse to continuously improve the clinical adaptability of the overall method.

[0255] In some implementation methods, the specific operation process is as follows:

[0256] First, collect quality control feedback and extract the fusion feature vector V corresponding to the "rule-triggered correction" records in compliance verification and the "fluctuation exceeding threshold" cases in sensitivity analysis. fusion Feature dimensions and dynamic weight parameters for the S3 stage; in addition, clinical expert feedback is collected, and the evaluation data of clinicians on the report is obtained through the structured feedback interface of the doctor's workstation, including "result compliance" on a 1-5 scale (1 point is completely inconsistent, 5 points is completely consistent), "omitted feature annotation" and "correction suggestions", and valid feedback with "result compliance ≤ 3 points + clear correction suggestions" or "compliance 4 points but with key omissions" is selected, and the sample size of valid feedback in a single batch must be ≥ 10 cases.

[0257] Next, corresponding optimization strategies are adopted for feedback from different sources: When optimizing based on compliance verification feedback, if a report shows a violation due to "unreasonable allocation of feature contribution," the dynamic weight constraint rules of S3 are adjusted; if semantic inconsistency is caused by "insufficient knowledge constraints," the knowledge graph inference vector V in S3 is adjusted. kgThe aggregation weights are increased to increase the proportion of entities with high correlation strength. When optimizing based on sensitivity analysis feedback, if the diagnostic results fluctuate beyond a threshold, it is determined that the fused features are too sensitive to a certain type of modal feature. Therefore, the gating coefficient calculation parameters of S3 are adjusted, such as increasing the gating weight matrix W. g The weighting of inference vectors in the knowledge graph is adjusted to reduce the impact of data disturbances on diagnostic results; when optimizing based on clinical expert feedback, the filtered effective feedback is transformed into "original V". fusion - Corrected V fusion The training samples of "-feedback labels" are used to fine-tune the initial weights W0 of the S3 dynamic weight model and the parameters of the gated attention network through incremental learning.

[0258] Finally, the fine-tuned S3 stage parameters are applied to the new multimodal data fusion process, repeating the entire process from S3 feature fusion to S4 report generation and verification. The three indicators of the diagnostic report are verified: compliance pass rate, sensitivity analysis stability rate, and clinical conformity. If the indicators do not meet the preset targets, the above feedback data collection and parameter optimization process is repeated until all indicators meet the clinical application requirements, forming a closed-loop system of continuous optimization.

[0259] Based on the above technical solution, S4S serves as the terminal output and closed-loop optimization core of the multimodal data fusion analysis in this application. Through four main steps—customized construction of a medical-specific AIGC model, dual-dimensional clinical compliance verification, Monte Carlo simulation sensitivity analysis, and multi-source feedback closed-loop iteration—it creates a diagnostic report generation system that combines clinical professionalism, content compliance, result stability, and scenario adaptability. S4 employs a dedicated AIGC model based on MedCLIP and a medical-domain fine-tuned GPT. It inputs fused standardized feature vectors to efficiently transform high-dimensional features into clinical semantics and standardized structured reports, improving generation efficiency and accuracy. The reports are clearly based and modularly standardized. Dual-dimensional compliance verification is performed using the Drools-Med engine and knowledge graph GNN inference, achieving dual quality control of guideline compliance and semantic logic, enhancing the professional rigor of the reports. Sensitivity analysis is conducted through Monte Carlo simulation to quantitatively assess result stability, preventing diagnostic reversals caused by minor data fluctuations and ensuring result reliability. A multi-source feedback system is constructed for parameter optimization and incremental learning, forming a closed-loop iteration that drives overall process optimization, improving overall clinical adaptability while also considering technological application and data security. Ultimately, it provides clinicians with an efficient, reliable, and compliant diagnostic support tool.

[0260] How this application works:

[0261] This application focuses on cross-modal alignment to achieve end-to-end fusion analysis of multimodal medical data. First, it acquires multimodal medical data including medical images, electronic medical record texts, and physiological signals. Quality grading assessments are conducted for each modality from multiple dimensions, and differentiated post-processing strategies are implemented based on the assessment levels to select high-quality, compliant data sources, laying a reliable foundation for subsequent processing. Then, a two-stage alignment algorithm combining intra-modal feature enhancement and cross-modal semantic calibration is employed. First, a pre-trained neural network model is used to strengthen the core features of each modality and select effective feature blocks. Then, combined with a constructed medical knowledge graph, cross-modal semantic mapping is completed through the generation of a medical memory matrix and an attention mechanism. Simultaneously, semantic consistency verification is performed based on the ideal alignment vector to obtain a unified semantic vector that conforms to clinical logic. Next, a dynamic weight model is constructed based on the relevance of the medical knowledge graph, clinical guideline priority, and modality quality score to intelligently fuse the semantic alignment results. This fusion process considers both the medical knowledge system and the actual clinical scenario, generating a fused feature vector that integrates the core information of multiple modalities. The fused feature vector is then input into a medically customized and optimized AIGC model to generate a standardized, structured diagnostic report. Simultaneously, a two-dimensional clinical compliance quality control process is implemented through guideline rule verification and knowledge graph semantic verification. Monte Carlo simulation is used for sensitivity analysis to ensure the compliance of the report content and the stability of the diagnostic results. Finally, the quality control results and multi-source effective feedback from clinical experts are collected to specifically optimize the relevant parameters of the initial feature fusion and the model parameters, forming an end-to-end closed-loop iterative optimization system. This achieves end-to-end optimization of medical multimodal data from acquisition and processing to report output, continuously improving the clinical adaptability, accuracy, and reliability of the overall analysis results.

[0262] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0263] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely illustrative descriptions of the application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and variations.

Claims

1. A medical multimodal data fusion and analysis method based on cross-modal alignment, characterized in that, include: Acquire multimodal medical data and perform quality grading assessment to obtain modality quality scores for each modality; wherein, the multimodal medical data includes at least medical imaging data, electronic medical record text data, and physiological signal time series data; A two-stage alignment algorithm, consisting of intra-modal feature enhancement and cross-modal semantic calibration, is used to perform semantic alignment on multimodal data after quality grading assessment, yielding semantic alignment results. Based on the knowledge graph relevance, clinical guideline priority, and modality quality score, a dynamic weighting model is used to perform feature fusion on the semantic alignment results. An AIGC (Artificial Intelligence Generated Content) model is used to generate a structured diagnostic report from the fusion results, and clinical compliance verification and sensitivity analysis are performed. The verification and analysis results are then fed back to the feature fusion step for optimization. The two-stage alignment algorithm includes an intra-modal feature enhancement stage and a cross-modal semantic calibration stage; the intra-modal feature enhancement stage is used to generate enhanced feature blocks for each modality; the cross-modal semantic calibration stage includes a medical memory generation layer, a cross-modal attention layer, and a semantic consistency verification layer, wherein, The medical memory generation layer is used to generate a three-dimensional medical memory matrix; the matrix elements of the three-dimensional medical memory matrix include modality type, standard terminology and disease name, which are used to characterize the association strength between modality features and medical entities, and provide medical constraints for semantic mapping. The cross-modal attention layer is used to receive the enhanced feature blocks of each modality and the three-dimensional medical memory matrix, and combine them with the medical knowledge graph to generate a preliminary unified semantic vector based on the attention mechanism. The semantic consistency verification layer is used to receive the preliminary unified semantic vector and the ideal alignment vector of each modality, calculate the semantic consistency score and dynamically calibrate to obtain the semantic alignment result. The ideal alignment vectors for each modality are obtained based on entities and clinical relationships in the medical knowledge graph, including: Extract entities related to the target disease from the medical knowledge graph to form a core set of related entities; The core associated entity set is bound to the corresponding modality to form a modality-entity mapping table; The enhanced features of each modality obtained from historical multimodal medical data are labeled as the basic feature vectors of each modality; The medical knowledge graph is pre-trained using the TransE algorithm to obtain the knowledge graph embedding vector of each entity in the core associated entity set; Extract the association levels between the core associated entity set and the target disease, and convert them into training weights; Constructing the loss function ;in, The modal base feature vector bound to the i-th core entity. Let be the knowledge graph embedding vector of the i-th core entity, and n be the number of entities in the core associated entity set. Let be the training weights for the i-th core entity; The modal basic feature vector, knowledge graph embedding vector, and training weights are input to a modal feature alignment training network built on a fully connected layer. The loss function is minimized by gradient descent to perform alignment training on the modal basic feature vectors, thereby obtaining the modal alignment vectors corresponding to each core entity. According to the formula Calculate the weight percentage of each core entity. The ideal alignment vector for modality i is obtained by weighted fusion of the alignment vectors of all core entities in the same modality. .

2. The medical multimodal data fusion and analysis method based on cross-modal alignment according to claim 1, characterized in that, The acquisition of multimodal medical data and the assessment of its quality grading include: For medical image data, spatial coverage is calculated based on the ratio of the number of pixels actually covered by the lesion area to the total number of pixels in the entire medical image. The first information density score is calculated based on the information entropy algorithm. The spatial coverage and the first information density score are weighted and summed to obtain the comprehensive score of the image data. For electronic medical record text data, the first text index is calculated based on the proportion of the number of core fields filled in to the total number of core fields. The second text index is calculated based on the proportion of the number of text terms in the medical record that match the preset medical standard terms to the total number of all relevant terms in the medical record. The third text index is calculated based on the proportion of the number of fields in the electronic medical record text format that conform to the HL7 medical data exchange standard to the total number of fields in the medical record. The second information density score is calculated based on the information entropy algorithm. The first text index, the second text index, the third text index, and the second information density score are weighted and summed to obtain the comprehensive score of the text data. For physiological signal time series data, the time series continuity score is calculated based on the ratio of continuous signal duration to total signal acquisition duration within the physiological signal acquisition period. The third information density score is calculated based on the information entropy algorithm. The time series continuity score and the third information density score are then weighted and summed to obtain the comprehensive score of the time series data. Based on the comprehensive score of each modality, the data of each modality are classified into different quality levels according to the level threshold, and a differentiated post-processing strategy is set for each level.

3. The medical multimodal data fusion and analysis method based on cross-modal alignment according to claim 1, characterized in that, The two-stage alignment algorithm includes an intra-modal feature enhancement stage and a cross-modal semantic calibration stage; wherein... The intramodal feature enhancement stage is used to receive medical image data, electronic medical record text data, and physiological signal time series data after quality grading assessment, and to enhance the features of each modality through a pre-trained neural network model to obtain enhanced feature blocks for each modality. The cross-modal semantic calibration stage is used to perform clinical semantic association mapping and consistency verification on the enhanced feature blocks combined with medical knowledge graph embedding, thereby enhancing the semantic association between different modalities and obtaining a unified semantic vector that conforms to clinical logic.

4. The medical multimodal data fusion and analysis method based on cross-modal alignment according to claim 3, characterized in that, The cross-modal semantic calibration phase includes a medical memory generation layer, a cross-modal attention layer, and a semantic consistency verification layer, wherein... The medical memory generation layer receives medical knowledge graphs and medical pairing data, learns clinical association patterns through a generator of a generative adversarial network (GAN), and obtains a three-dimensional medical memory matrix. The matrix elements of the three-dimensional medical memory matrix include modality types, standard terms, and disease names, which are used to characterize the association strength between modality features and medical entities, providing medical constraints for semantic mapping. The medical pairing data consists of clinical pairing triplets of image data, electronic medical record text data, and physiological signal time series data, with each triplet containing a diagnostic association label confirmed by a parent clinical expert. The cross-modal attention layer receives enhanced feature blocks from each modality and the three-dimensional medical memory matrix. It constructs key and value vectors for the attention mechanism using the enhanced feature blocks, embeds target disease candidate entities from the medical knowledge graph to construct query vectors, and uses the three-dimensional medical memory matrix as the basis for attention weight constraints to achieve cross-modal semantic mapping and obtain a preliminary unified semantic vector. The target disease candidate entities are determined through semantic matching between clinical information in electronic medical record text data and disease entities in the medical knowledge graph. The semantic consistency verification layer is used to receive the preliminary unified semantic vector and the ideal alignment vector of each modality, calculate the semantic consistency score and dynamically calibrate to obtain the final unified semantic vector, i.e. the semantic alignment result.

5. The medical multimodal data fusion and analysis method based on cross-modal alignment according to claim 4, characterized in that, The construction process of the medical knowledge graph includes: Medical data sources were collected, including the National Comprehensive Cancer Network (NCCN) clinical guidelines, Chinese clinical practice guidelines, medical journal articles, clinical practice data, and expert consensus. Multiple entities and their clinical relationships are extracted from the medical data source; the entities include diseases, symptoms, laboratory indicators, imaging features, drugs, and physiological signals. Entities are uniformly labeled using pre-defined medical standard terminology, ICD-10 disease codes, and LOINC test index codes, and clinical associations are assigned a level of clinical evidence; the clinical evidence refers to the amount of clinical research data and the degree of expert consensus that support the association. A medical knowledge graph is obtained by storing entities and clinical relationships using a graph database.

6. The medical multimodal data fusion and analysis method based on cross-modal alignment according to claim 4, characterized in that, The calculation process for the cross-modal attention layer is as follows: According to the formula Calculate the basic attention weights ;in, Embed query vectors into the target disease candidate set entities. The key vector is the concatenated enhanced feature blocks from each modality. For feature dimensions; Based on the aforementioned basic attention weights and the three-dimensional medical memory matrix The weight adjustment is performed using the following formula: ;in, This is the attention weight matrix after medical constraints. A preliminary unified semantic vector is calculated based on the attention weight matrix after the aforementioned medical constraints. The calculation formula is: ;in, This is the value vector concatenated from the enhanced feature blocks of each modality.

7. The medical multimodal data fusion and analysis method based on cross-modal alignment according to claim 4, characterized in that, The semantic consistency score The calculation formula is: ;in, Let be the ideal alignment vector for the i-th mode. For vector dot product operation, Let be the vector magnitude.

8. The medical multimodal data fusion and analysis method based on cross-modal alignment according to claim 1, characterized in that, The step of using a dynamic weighting model to perform feature fusion on the semantic alignment results includes: Obtain the unified semantic vector corresponding to the semantic alignment result. Simultaneously, extract the medical knowledge graph correlation vector R and the clinical guideline priority vector. And the modality quality score vector Q; where R is the clinical association strength between each modality and the target disease, extracted from the medical knowledge graph. Prioritize guideline recommendations for each modality, based on clinical practice guidelines; Set the initial weight vector Based on dynamic weight formula The final mode weights are calculated. ; Feature fusion is performed using an attention network with a fusion gating mechanism, through a formula. Calculate the gating coefficient G, and then use the formula Output the fused feature vector; where, , representing the knowledge graph embedding vector of each core entity based on the clinical association strength between the target disease and the core associated entities in the medical knowledge graph. The knowledge graph reasoning vector obtained by weighted aggregation The association level between the i-th core entity and the target disease; This indicates an attention calculation operation.

9. A medical multimodal data fusion and analysis system based on cross-modal alignment, characterized in that, include: The module includes a quality grading module, a semantic alignment module, a feature fusion module, and a report generation module; among them, The quality grading module is used to acquire multimodal medical data and perform quality grading assessment to obtain a modality quality score for each modality; wherein, the multimodal medical data includes at least medical imaging data, electronic medical record text data, and physiological signal time series data; The semantic alignment module is used to perform semantic alignment on multimodal data after quality grading assessment using a two-stage alignment algorithm of intramodal feature enhancement and cross-modal semantic calibration, and obtain semantic alignment results. The feature fusion module is used to perform feature fusion on the semantic alignment results based on knowledge graph relevance, clinical guideline priority, and modality quality score, using a dynamic weight model. The report generation module is used to generate a structured diagnostic report from the fusion results using an AIGC (Artificial Intelligence Generated Content) model, and to perform clinical compliance verification and sensitivity analysis. The verification and analysis results are then fed back to the feature fusion step for optimization. The two-stage alignment algorithm includes an intra-modal feature enhancement stage and a cross-modal semantic calibration stage; the intra-modal feature enhancement stage is used to generate enhanced feature blocks for each modality; the cross-modal semantic calibration stage includes a medical memory generation layer, a cross-modal attention layer, and a semantic consistency verification layer, wherein, The medical memory generation layer is used to generate a three-dimensional medical memory matrix; the matrix elements of the three-dimensional medical memory matrix include modality type, standard terminology and disease name, which are used to characterize the association strength between modality features and medical entities, and provide medical constraints for semantic mapping. The cross-modal attention layer is used to receive the enhanced feature blocks of each modality and the three-dimensional medical memory matrix, and combine them with the medical knowledge graph to generate a preliminary unified semantic vector based on the attention mechanism. The semantic consistency verification layer is used to receive the preliminary unified semantic vector and the ideal alignment vector of each modality, calculate the semantic consistency score and dynamically calibrate to obtain the semantic alignment result. The ideal alignment vectors for each modality are obtained based on entities and clinical relationships in the medical knowledge graph, including: Extract entities related to the target disease from the medical knowledge graph to form a core set of related entities; The core associated entity set is bound to the corresponding modality to form a modality-entity mapping table; The enhanced features of each modality obtained from historical multimodal medical data are labeled as the basic feature vectors of each modality; The medical knowledge graph is pre-trained using the TransE algorithm to obtain the knowledge graph embedding vector of each entity in the core associated entity set; Extract the association levels between the core associated entity set and the target disease, and convert them into training weights; Constructing the loss function ;in, The modal base feature vector bound to the i-th core entity. Let be the knowledge graph embedding vector of the i-th core entity, and n be the number of entities in the core associated entity set. Let be the training weights for the i-th core entity; The modal basic feature vector, knowledge graph embedding vector, and training weights are input to a modal feature alignment training network built on a fully connected layer. The loss function is minimized by gradient descent to perform alignment training on the modal basic feature vectors, thereby obtaining the modal alignment vectors corresponding to each core entity. According to the formula Calculate the weight percentage of each core entity. The ideal alignment vector for modality i is obtained by weighted fusion of the alignment vectors of all core entities in the same modality. .

Citation Information

Patent Citations

  • Multi-source heterogeneous medical data fusion and intelligent diagnosis method

    CN121545724A

  • Multi-modal heterogeneous medical data dynamic weighting intelligent disease analysis system

    CN121682497A