A chronic disease classification and prediction method and system based on knowledge and data fusion driving

CN122474237BActive Publication Date: 2026-08-28YANTAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610956784.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-08-28
Estimated Expiration
2046-06-30

AI Technical Summary

Technical Problem

然而,上述研究主要依赖通用医学知识,对患者个性化知识与多模态医疗数据融合的关注极少

Benefits of technology

[0030] (1) In this invention, artificial intelligence and deep learning algorithms are introduced, medical multimodal data and prior knowledge of patient health profiles are integrated, and a series of technologies such as Transformer, BERT, residual attention network, multi-head attention mechanism and multimodal fusion are used to achieve accurate, efficient identification and dynamic monitoring of abnormal health characteristics of patients with chronic diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122474237B_ABST
    Figure CN122474237B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of chronic disease classification prediction, in particular to a chronic disease classification prediction method and system based on knowledge and data fusion driving. The method comprises constructing a patient health portrait knowledge based on the obtained clinical chronic multi-source data; performing high-density semantic coding on the patient health portrait knowledge based on an improved BERT encoder to obtain a portrait knowledge representation matrix; performing directional extraction based on the portrait knowledge representation matrix to obtain a portrait knowledge embedding vector; extracting multi-modal data features from the clinical chronic multi-source data based on an improved Transform encoder; and performing data feature fusion based on a CMD multi-modal alignment mechanism; and finally generating an enhanced representation with knowledge guiding capability through the deep interaction and fusion of the fused features, physiological experimental features and portrait knowledge representation by the multi-layer CA mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chronic disease classification and prediction technology, and in particular to a chronic disease classification and prediction method and system driven by knowledge and data fusion. Background Technology

[0002] With the rapid development of smart healthcare, wearable devices, and multi-omics technologies, clinical data has expanded from single laboratory indicators to multimodal and heterogeneous formats such as electronic medical record texts, medical images, laboratory tests, physiological signals, and follow-up records, providing a data foundation for intelligent diagnosis of chronic diseases. Compared to single-modal data, multimodal data can provide complementary and comprehensive patient medical information, effectively mitigating misdiagnosis and missed diagnosis caused by single-modal data, thereby significantly improving diagnostic robustness and accuracy. Therefore, multimodal intelligent chronic disease diagnosis is gradually becoming an important direction in the field of medical artificial intelligence.

[0003] In recent years, multimodal AI has been widely applied in the diagnosis of chronic diseases. For example, some studies have proposed DeepDR-LLM, which integrates fundus images and large language models to achieve auxiliary diagnosis and personalized management of diabetic retinopathy; other studies have proposed a multimodal deep learning urodynamic diagnostic model that integrates image and text data, achieving a diagnostic accuracy of over 90% for lower urinary tract dysfunction; and some studies have built AI decision-making systems based on multidimensional data, reducing the incidence of postoperative complications in colorectal cancer by 32%–36%. However, most methods rely solely on data-driven approaches, seriously neglecting the integration of relevant prior medical knowledge, resulting in insufficient accuracy and weak interpretability in scenarios such as similar symptom differentiation, comorbidity assessment, diagnosis of low-resource populations, and risk stratification.

[0004] Current research has attempted to incorporate medical knowledge into chronic disease diagnosis. For example, studies have proposed the KG4Diagnosis framework, integrating knowledge graphs and large language models to diagnose 362 common diseases; the REKG-MDP model, combining representation learning and knowledge graph reasoning, is used for diabetes and complication prediction; and even the MedGraphNet network model, integrating electronic medical record text knowledge, accurately infers the associations between diseases, genes, and drugs, maintaining excellent performance even in isolated nodes and sparse data scenarios. However, these studies primarily rely on general medical knowledge, with minimal attention paid to the integration of personalized patient knowledge and multimodal medical data. Therefore, effectively integrating personalized patient knowledge (personal information, medical records, and genetic history, etc.) with multimodal medical data to achieve more accurate intelligent classification and prediction of chronic diseases remains a pressing challenge. Summary of the Invention

[0005] To address the aforementioned problems, this invention provides a chronic disease classification and prediction method and system driven by knowledge and data fusion.

[0006] Firstly, the present invention provides a chronic disease classification and prediction method based on knowledge and data fusion, which adopts the following technical solution:

[0007] A chronic disease classification and prediction method driven by knowledge and data fusion includes:

[0008] Acquire multi-source clinical chronic disease data;

[0009] Based on the acquired clinical chronic multi-source data, construct patient health profile knowledge;

[0010] Based on the improved BERT encoder, high-density semantic encoding is performed on patient health profile knowledge to obtain a profile knowledge representation matrix;

[0011] The portrait knowledge embedding vector is obtained by targeted extraction based on the portrait knowledge representation matrix;

[0012] Multimodal data features are extracted from multi-source clinical chronic disease data based on an improved Transformer encoder;

[0013] Data feature fusion based on CMD multimodal alignment mechanism;

[0014] A knowledge-aware interactive fusion network is used to perform deep interactive fusion of the fused data features and the profile knowledge embedding vectors.

[0015] A knowledge-guided augmented representation method is used to classify and predict the features of the fused data.

[0016] Output the prediction results.

[0017] Secondly, a chronic disease classification and prediction system driven by knowledge and data fusion includes:

[0018] The data acquisition module is configured to acquire multi-source clinical chronic disease data.

[0019] The profiling module is configured to construct patient health profile knowledge based on acquired clinical chronic multi-source data;

[0020] The matrix module is configured to perform high-density semantic encoding on patient health profile knowledge based on an improved BERT encoder to obtain a profile knowledge representation matrix.

[0021] The embedding module is configured to perform targeted extraction based on the profile knowledge representation matrix to obtain the profile knowledge embedding vector;

[0022] The feature module is configured to extract multimodal data features from multi-source clinical chronic disease data based on an improved Transformer encoder;

[0023] The fusion module is configured to perform data feature fusion based on the CMD multimodal alignment mechanism;

[0024] The interaction module is configured to use a knowledge-aware interaction fusion network to perform deep interactive fusion of the fused data features and profile knowledge embedding vectors.

[0025] The prediction module is configured to perform classification prediction on the fused data features using a knowledge-guided augmented representation method.

[0026] The output module is configured to output the prediction results.

[0027] Thirdly, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the aforementioned chronic disease classification and prediction method based on knowledge and data fusion.

[0028] Fourthly, the present invention provides a terminal device, including a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a knowledge and data fusion-driven chronic disease classification and prediction method.

[0029] In summary, the present invention has the following beneficial technical effects:

[0030] (1) In this invention, artificial intelligence and deep learning algorithms are introduced, medical multimodal data and prior knowledge of patient health profiles are integrated, and a series of technologies such as Transformer, BERT, residual attention network, multi-head attention mechanism and multimodal fusion are used to achieve accurate, efficient identification and dynamic monitoring of abnormal health characteristics of patients with chronic diseases.

[0031] (2) In order to effectively learn patients' personal knowledge, this invention proposes a patient health profile learning method. The method first systematically constructs a structured health profile of the patient, and then learns the representation of the patient's health profile through an improved BERT encoder, thereby obtaining an effective user profile embedding representation, providing reliable prior guidance for subsequent chronic disease diagnosis.

[0032] (3) To effectively integrate user knowledge and multimodal data, this invention proposes a knowledge-aware interactive fusion network. This network deeply interacts and integrates fusion features, physiological experimental features and profile knowledge representation through a multi-layer Cross-Attention mechanism, ultimately generating an enhanced representation with knowledge guidance capabilities. Attached Figure Description

[0033] Figure 1This is a schematic diagram of a chronic disease classification and prediction method based on knowledge and data fusion driven by Embodiment 1 of the present invention;

[0034] Figure 2 This is a schematic diagram of an improved Transformer encoder module according to Embodiment 1 of the present invention;

[0035] Figure 3 This is a schematic diagram of the knowledge perception and interaction fusion network of Embodiment 1 of the present invention;

[0036] Figure 4 This is a schematic diagram illustrating the performance analysis of the training dataset with different proportions in Embodiment 1 of the present invention. Detailed Implementation

[0037] The present invention will be further described in detail below with reference to the accompanying drawings.

[0038] Example 1

[0039] Reference Figure 1 This embodiment of a chronic disease classification and prediction method based on knowledge and data fusion includes:

[0040] S1: The health profile learning module systematically learns and structures the patient's overall health profile knowledge. First, it comprehensively collects multi-source data, including compliant basic static information, long-term follow-up dynamic health records, family history of genetic diseases, and archived health records from all previous medical visits. This data is standardized and organized according to the specific guidelines for chronic disease diagnosis and treatment, constructing standardized patient-specific health profile knowledge. Then, using an improved BERT encoder adapted to the semantic scenarios of medical text, the constructed full-data patient health profile is semantically encoded dimension by dimension with high density, generating a unique profile knowledge representation matrix that is dimensionally unified, semantically aligned, and pathologically correlated. Finally, based on the patient's unique target ID as the index key, a one-to-one correspondence index table between IDs and profile knowledge vectors is established to accurately retrieve and extract the unique personal health profile knowledge representation corresponding to the current target patient for diagnosis.

[0041] Step S1 specifically includes:

[0042] S1.1: Pre-define all target patients to be assessed (total sample set), and standardize the formalized form as shown in formula (1):

[0043] (1),

[0044] in, This represents the total number of chronic disease patients diagnosed in the batch. Under the premise of strictly adhering to medical data compliance, cross-platform multi-source heterogeneous original health data sources corresponding to each patient are collected, including: patient basic data, archived data of physical examination indicators over the years, past outpatient and inpatient electronic health record data, family genetic background tracing ledger, daily behavior and chronic disease medication adherence follow-up registration data, and other types of data sources. Standardized and organized data are obtained by summarizing and organizing, and a unique health profile knowledge is independently constructed for each patient to form a global patient structured health profile set, as shown in formula (2).

[0045] (2),

[0046] in, Indicates the patient Personalized health profile This represents the semantic alignment and normalization processing function. Structured data representing patients, This represents the patient's unstructured data. Each personalized profile comprehensively includes the patient's relevant health background information, ensuring that the profile as a whole aligns with the targeted needs of chronic disease diagnosis.

[0047] To improve the semantic consistency and information integrity of the profile, this invention strictly follows the original key-value pair construction approach, simultaneously strengthening internal semantic alignment constraints, and uniformly decomposing each health attribute into standardized attribute labels and corresponding semantic description values. Therefore, each patient's profile consists of attribute labels and corresponding semantic values, formally represented as shown in formula (3):

[0048] (3),

[0049] in, This refers to profile attribute tags (such as age, past medical records, family history of genetic diseases, and long-term lifestyle risk records). The standardized natural language semantic description content generated after each attribute tag undergoes a unified medical semantic transformation.

[0050] S1.2: Based on the completion of the patient health profile construction, in order to avoid the loss of chronic disease correlation between attributes due to isolated encoding of a single profile, this invention conducts global unified semantic space modeling on the batch patient profiles of the entire domain, transforming the discrete and independent single profile text into a structured topological representation set with internal implicit correlation constraints, ensuring that subsequent encoding can simultaneously capture the explicit information of attributes and the implicit pathological dependencies. The overall modeling architecture completely follows the original core definition, and a global patient profile structured representation space is constructed. The global unified sequence of patient profiles is shown in (4):

[0051] (4),

[0052] in, A collection of text sequences representing all patient profiles. This represents the implicit semantic association mapping between various attributes within a portrait, used to characterize the dependency relationships and semantic coupling between attribute text. Through the above construction method, It can systematically characterize patients' information in terms of individual attributes, background disease records, and family genetic history, providing a unified and semantic input foundation for subsequent patient health profile representation learning. Meanwhile, As a carrier reflecting the patient's external knowledge, it helps the model to introduce personalized prior information in the diagnosis of chronic diseases, thereby enhancing the accuracy and adaptability of the diagnosis.

[0053] S1.3: Obtaining a unified set of images through modeling Subsequently, this invention revolves around the collection of images. Pre-processing work is carried out.

[0054] Step S1.3 specifically includes:

[0055] S1.3.1: Collection of calls to this invention Complete set of internal text sequences Simultaneously, retrieve the global implicit semantic association mapping. As a structural prior constraint, each patient is individually extracted and partitioned according to the one-to-one correspondence rule of patient sample number. The corresponding original portrait text sequence Each sequence fully reproduces the semantic content of the multi-source heterogeneous health profile key-value pairs constructed earlier, resulting in a global health profile sequence set for patients. ,in This indicates the patient's ID number.

[0056] S1.3.2: For each image text sequence Based on the unified semantic standard for chronic disease diagnosis and treatment, character compliance verification, invalid noise attribute masking, partial semantic completion of time-series health information, and alignment of attribute labels with semantic descriptions are completed. The length distribution of the input sequence is standardized and fixed, forming a standardized medical text tensor sequence group that can be encoded in parallel batches. This avoids encoding gradient distortion caused by heterogeneous image length variations. The general preprocessing formula is:

[0057] (5),

[0058] in, It provides a unified semantic standard constraint set for the diagnosis and treatment of chronic diseases, including four types of prior medical rules: medical stop word list, compliant character set, chronic disease-specific semantic knowledge base, and time sequence information alignment rules, providing professional semantic judgment standards for preprocessing.

[0059] S1.4: Obtaining the preprocessed image set Subsequently, this invention completes the learning of deep semantic representations specific to health profiles through an improved BERT encoder, outputting a profile knowledge representation with unified dimensions, semantic alignment, and direct compatibility with downstream fusion tasks. In this step, to adapt the basic BERT encoder to the heterogeneous characteristics of chronic disease profile text, a lightweight improvement to the attention layer is achieved by relying on global association priors. By linking the encoding with the patient ID number, the semantic relationship of isolated profiles is not fully established, and the computational knowledge transformation of the entire patient profile is efficiently completed.

[0060] Step S1.4 specifically includes:

[0061] S1.4.1: This step will normalize the batch portrait text sequence. It is fed into the BERT encoding backbone for standardized embedding encoding, as shown in formula (6):

[0062] (6),

[0063] in, BERT has built-in embedding encoding functions; For the first The basic embedding feature vectors corresponding to each patient profile.

[0064] At the same time, the globally implicit semantic association mapping that has been solidified within the set is directly retrieved. With the patient's ID number ,Will and Preprocessed into a pathological association bias matrix that perfectly matches the attention calculation dimension. The specific process is shown in the announcement (7);

[0065] (7),

[0066] in, Transformation function for dimension adaptation; This is a bias matrix built offline to carry the pathological correlation relationships of the whole-domain image.

[0067] Within each layer of bidirectional self-attention, the Query, Key, and Value base matrices are generated normally based on the input profile text sequence. The original attention similarity scores of the native context are calculated and normalized to obtain the score matrix. The specific process is shown in formulas (8) and (9):

[0068] (8),

[0069] (9),

[0070] in, , , These are the self-attention projection weights; The original attention score matrix without incorporating medical priors. For hierarchical adaptive Softmax normalized mapping function, This is a mask matrix for masking invalid semantics in medical contexts.

[0071] S1.4.2: After multi-layer improved self-attention iterative computation and output of multi-scale fused hidden layer semantic features, the semantic features are first subjected to global pooling and... After standardization, a regularization term for the weight of chronic disease clinical attributes is introduced to perform global spatial alignment processing on the multi-scale hidden layer features, and finally a patient-specific profile knowledge embedding vector with strong pathological correlation is generated, representing the learning process as shown in formula (10):

[0072] (10),

[0073] in, This is a pooling operation for the top-level standard global temporal features of BERT. This is a standard global dimension normalization correction function used to unify the vector space distribution of all patients; Regularization terms for clinical attributes of chronic diseases; To standardize the hidden feature dimension globally, the following representation learning process is performed, resulting in a unified and semantically consistent global portrait knowledge embedding representation as shown in Equation (11):

[0074] (11),

[0075] S1.5: After completing the semantic encoding of the global patient health profile through an improved BERT encoder, a global profile knowledge embedding with unified dimensions and semantic alignment is obtained. Subsequently, this invention further relies on the patient's unique target ID to achieve precise positioning and targeted extraction of target patient profile knowledge.

[0076] Step S1.5 specifically includes:

[0077] S1.5.1: This invention relies on global semantic association mapping. Simultaneously, a structured relational index table is established with the patient's unique target ID as the primary key. This index table uses semantic relational mapping... Each patient's unique identifier ID and the embedded vector of the personalized profile knowledge output by the improved BERT are bound one by one to form a stable, ordered, and fast-addressable bidirectional association mapping relationship, as shown in formula (12):

[0078] (12),

[0079] in, Embed a bidirectional association index table for patient ID-profile; Functions for building structured indexes; This is a global implicit semantic association mapping.

[0080] S1.5.2: For the target patient to be diagnosed that needs to be assessed, extract their corresponding unique target ID as the search keyword, and rely on global semantic association mapping. A pre-established one-to-one correspondence index table is used to perform precise key-value matching and addressing, accurately retrieving the patient's unique personalized profile knowledge embedding representation. The structured index retrieval process is shown in formula (13):

[0081] (13),

[0082] in, This indicates a search operation based on patient ID, using an index table. Return to target patient Image knowledge embedding This embedding vector can then be fused with the multimodal fusion features of the target sample, providing the model with personalized prior knowledge about the patient.

[0083] S2: For three types of heterogeneous medical multimodal data—chief complaint texts, physiological experimental data, and clinical notes—this paper abandons the traditional, simplistic feature extraction method using a single native Transformer multi-head attention approach and designs a modality-adaptive hierarchical refined feature extraction module. Combining the semantic sparsity and professional specificity of medical data, modality-differentiated embedding correction and hierarchical multi-head attention weight calibration are introduced to achieve deep, fine-grained, and differentiated feature mining of multimodal raw data. This addresses the problems of poor generalization and lack of modality specificity in native Transformer feature extraction. Figure 2 As shown.

[0084] S2.1: This invention focuses on multimodal data preprocessing and adaptive embedding optimization. It designs a modality-specific adaptive embedding method to perform differentiated purification and dimensionality mapping on different types of raw data, filtering out invalid noise and retaining modality-specific effective features. This provides high-quality input for subsequent attention feature extraction and slightly optimizes the original simplified feature extraction logic, balancing model rationality and appropriate innovation. The differentiated adaptive embedding formula is as follows:

[0085] =LN ,

[0086] =LN (14)

[0087] =LN ,

[0088] in, , , Learnable embedding mapping layers are designed for three modalities, respectively adapting to feature mappings for text semantics, physiological signals, and professional medical notes; LN This is a layer normalization operation used to stabilize feature distribution; , , The modality-adaptive noise suppression coefficient is dynamically updated through model backpropagation, balancing the preservation of original information with the effect of feature purification. , , This refers to the refined modal input features after preprocessing.

[0089] S2.2: To avoid the problems of indiscriminate feature modeling and redundant feature interference in traditional multi-head attention modeling, this invention introduces a hierarchical gated multi-head attention feature extraction mechanism based on adaptive embedding preprocessing. This mechanism breaks down feature extraction into modal feature projection, gated attention weight calculation, multi-head feature splicing and fusion, and residual feature calibration, progressively completing the extraction of deep features for each modality.

[0090] Step S2.2 specifically includes:

[0091] S2.2.1: To adapt to the dimensional characteristics of different modal features, a modality-specific projection matrix is ​​first constructed. Linear transformations of the query, key, and value are then performed on the input features to complete feature dimensionality adaptation and preliminary feature reconstruction. The formula is as follows:

[0092] ,

[0093] (15),

[0094] ,

[0095] in, These correspond to three modalities: chief complaint, physiological experiment, and clinical notes. , , The projective weight matrix is ​​an independent and learnable projection weight matrix for each modality; , , These are the query matrix, key matrix, and value matrix corresponding to each modality.

[0096] S2.2.2: Traditional attention mechanisms distribute weights equally across all features, failing to filter key features from medical data. This invention introduces a modality-adaptive gating factor to filter and correct the original attention weights, suppressing redundant feature weights and strengthening core feature weights. The calculation formula is as follows:

[0097] (16),

[0098] in, This is a scaling factor to prevent the gradient from saturating due to excessively large vector dot products. For Hadamard product operations; It is a modal adaptive gating matrix, which is adaptively generated from the global statistical features of modal features and is used to dynamically select effective features.

[0099] S2.2.3: After obtaining the corrected attention weights, the corrected attention weights are weighted and fused with the value matrix to obtain single-head attention features. Then, multiple sets of single-head features are concatenated and fused to achieve complementary multi-dimensional feature information. The formula is as follows:

[0100] (17),

[0101] (18),

[0102] in, For the number of attention heads; For the first Single-granularity modal features extracted by the attention head; This is a multi-head feature stitching operation to achieve multi-scale feature fusion.

[0103] S2.2.4: After completing the multi-head feature splicing and fusion, the final feature output of each modality is completed through linear mapping, and the extracted modal features are obtained, as shown in formula (19):

[0104] (19),

[0105] in, This is used to balance the information ratio between deep attention features and original preprocessed features, thereby improving the stability of feature representation.

[0106] S3: Multimodal feature fusion is a crucial process for achieving multimodal disease diagnosis. This paper utilizes a Transformer-based multimodal fusion module to perform multimodal fusion on subject description features, physiological experiment features, and clinical note features, providing more information for model decision-making.

[0107] Step S3 specifically includes:

[0108] S3.1: The self-attention mechanism can consider the global information of the input sequence, not just local regions. Using a self-attention module separately for each modality helps avoid information confusion. Therefore, after each modality feature enters the multimodal fusion module, three self-attention mechanisms are first used to capture and learn the global information of each modality, as shown in formula (20):

[0109] ,

[0110] in, Representing each mode, , , , These represent the query matrix, key matrix, and value matrix, respectively. This represents the parameters learned during model training. This represents the dimension of the key matrix.

[0111] S3.2: Considering that textual modalities are crucial for multimodal intent recognition, this paper conducts similarity learning on text-audio and text-image feature pairs. This paper uses the central moment difference (CMD) as the similarity loss function, and then calculates the similarity between text-audio and text-image modalities. By minimizing CMD, image and audio modalities are guided to better align with text modalities at the distribution level, thereby accelerating more coordinated fusion of multimodalities. The calculation process follows equation (21):

[0112] ,

[0113] in, It is the empirical expectation vector of sample x, and It is a vector of the central moments of all k-order samples at the x-coordinate. For example, this paper computes the central moments between text-audio and text-image modalities. As shown in equation (22):

[0114] (twenty two),

[0115] in, It is a feature vector extracted from the self-attention mechanism. .

[0116] S3.3: Secondly, physiological test indicators are the core diagnostic criteria that are objectively quantified, possessing stronger disease specificity and clinical diagnostic value, and are key support for chronic disease classification and abnormality identification; while patient reports and clinical notes are unstructured text information, highly subjective, and discrete in expression, serving only as supplementary information. Based on this, we did not calculate cross-modal features from patient report to clinical notes and from clinical notes to patient report. Therefore, cross-modal attention was used to learn the relevant features from physiological tests to clinical notes. Clinical notes on physiological experimental features Physiological experiments to the main description of relevant characteristics and the main description of physiological experiment-related characteristics A total of four cross-modal transformers are needed to obtain four feature vectors. Each cross-modal attention mechanism consists of n layers of cross-modal attention modules. Information is obtained from modal... Transition to mode For example, the cross-modal attention modules for i=1, 2, ..., n are shown in equations (23), (24), and (25):

[0117] (twenty three),

[0118] (twenty four),

[0119] (25),

[0120] in, It is by The parameterized position-feedforward sublayer, CT is the multi-head cross-modal attention module, and LN is layer normalization.

[0121] S3.4: Then, we combine physiological experiments with clinical notes to identify relevant features. Clinical notes on physiological experiment-related features This data is then input into the clinical gating system to obtain the final clinical note features. This combines physiological experiments with subject-specific features. and the main description of physiological experiment-related characteristics And input them into the subject gating to obtain the final subject features, as shown in formula (26) and formula (27):

[0122] (26),

[0123] (27),

[0124] in, This indicates a splicing operation. Indicates a convolutional layer. The Sigmoid function represents a tensor.

[0125] S3.5: Finally, the feature vectors of all modalities are combined to obtain the final fused feature representation, as shown in Equation (28):

[0126] (28),

[0127] in, This indicates a splicing operation. These are physiological experimental features extracted through the self-attention mechanism. Clinical note features. This is the result of clinical gating fusion processing. (Main characteristic) This is the result of the main gating fusion processing.

[0128] S4: The heterogeneity of knowledge and data often leads to distribution gaps and information redundancy, resulting in task-independent or semantically ambiguous joint representations. Furthermore, the question of "who takes the lead" during the fusion process presents a significant challenge. To address these issues, this invention designs a knowledge-aware interactive fusion network to achieve deep interactive fusion of fusion features, physiological experimental features, and profile knowledge representations.

[0129] Step S4 specifically includes:

[0130] S4.1: Feature-driven interaction: In feature-driven interaction, the fusion features are first... Image knowledge representation splicing to obtain spliced ​​representation Then fuse the features Physiological experimental characteristics splicing to obtain spliced ​​representation To enhance the stability of feature interactions, residual connections and layer normalization (Add & Norm) are introduced in the Cross-Attention module. Then, the features are fused. As a query As the key and value, they are calculated through Cross-Attention (CA) and then obtained through residual connections and layer normalization (Add & Norm). Similarly, the fusion features As a query. As the key and value, they are obtained through Cross-Attention (CA), residual connections, and layer normalization (Add & Norm). Secondly, and Summing the elements yields the representation. The specific process is shown in formulas (29), (30) and (31).

[0131] (29)

[0132] (30)

[0133] (31),

[0134] in, Indicates the dominant mode, and Indicates the fused mode. , , These are mapping matrices for queries, keys, and values, respectively. Let be the dimension of the key vector. Then, to obtain a more stable and uniformly distributed feature representation, this paper will... The input is fed into the feedforward network FFN and Add&Norm layers, thereby generating an enhanced representation with certain structural robustness and good semantic expressiveness. As shown in formula (32).

[0135] (32),

[0136] in, To integrate features An enhanced representation of the final output of the dominant branch.

[0137] S4.2: Interaction Based on Physiological Experiment Features: Considering that the subject-matter modality data and clinical note modality data are not as important as the physiological experiment modality data, this paper will also use physiological experiment features as the primary focus for interaction. In the dominant interaction, the physiological experimental characteristics are first presented. With fusion features splicing to obtain spliced ​​representation Then, physiological experimental characteristics Image knowledge representation splicing to obtain spliced ​​representation Next As a query As the key and value, they are calculated through Cross-Attention (CA) and then obtained through residual connections and layer normalization (Add & Norm). Similarly, As a query As the key and value, they are obtained through Cross-Attention (CA), residual connections, and layer normalization (Add & Norm). Secondly, and Summing the elements yields the representation. The specific process is shown in formulas (33), (34) and (35).

[0138] (33),

[0139] (34)

[0140] (35),

[0141] in, Indicates the dominant mode, and Indicates the fused mode. , , These are mapping matrices for queries, keys, and values, respectively. Let be the dimension of the key vector. Then, to obtain a more stable and uniformly distributed feature representation, this paper will... The input is fed into the feedforward network FFN and Add&Norm layers, thereby generating an enhanced representation with certain structural robustness and good semantic expressiveness. As shown in formula (36).

[0142] (36),

[0143] in, To use physiological experimental characteristics An enhanced representation of the final output of the dominant branch.

[0144] S4.3: Interaction led by portrait knowledge representation: First, the portrait knowledge is represented... With fusion features splicing to obtain spliced ​​representation Then represent the portrait knowledge Physiological experimental characteristics splicing to obtain spliced ​​representation Next, the knowledge of the portrait will be represented. As a query, As the key and value, they are calculated through Cross-Attention (CA) and then obtained through residual connections and layer normalization (Add & Norm). Similarly, As a query As the key and value, they are obtained through Cross-Attention (CA), residual connections, and layer normalization (Add & Norm). Secondly, and Summing the elements yields the representation. The specific process is shown in formulas (37), (38) and (39).

[0145] (37),

[0146] (38),

[0147] (39),

[0148] in, Indicates the dominant mode, and Indicates the fused mode. , , These are mapping matrices for queries, keys, and values, respectively. Let be the dimension of the key vector. Then, to obtain a more stable and uniformly distributed feature representation, this paper will... The input is fed into the feedforward network FFN and Add&Norm layers, thereby generating an enhanced representation with certain structural robustness and good semantic expressiveness. The specific process is shown in formula (40).

[0149] (40),

[0150] in, To represent with portrait knowledge An enhanced representation of the final output of the dominant branch.

[0151] S4.4: Finally, this paper will present the enhanced representation of the three-stage branches. , and The images are then spliced ​​together to obtain the final enhanced representation. The specific engineering process is shown in formula (41).

[0152] (41),

[0153] in, This indicates a splicing operation. It is a branch-enhanced representation dominated by fusion features. It is a branched enhancement representation dominated by physiological experimental characteristics. It is a branch of augmented representation that is dominated by image knowledge representation.

[0154] S5: To achieve accurate identification and intelligent assisted diagnostic classification output for chronic diseases, this invention will ultimately enhance the representation with knowledge-guided capabilities. The data is then fed into the hierarchical projection classification module to complete high-order feature purification and nonlinear discriminant mapping.

[0155] Step S5 specifically includes:

[0156] S5.1: First, the final enhanced representation of the output. The first-level hidden layer projection fully connected unit is fully input, and cross-dimensional feature ordered compression mapping is completed based on the learnable weight matrix. The adaptive bias compensation vector is superimposed to correct the baseline deviation of the representation. Then, the ReLU nonlinear activation function is embedded to complete the global feature sparsity denoising process, effectively eliminating the invalid redundant noise components in the fused representation and highlighting the core discriminative pathological features of chronic diseases. Then, the second-level output mapping fully connected layer is connected to complete the feature dimension adaptation and secondary accurate projection calibration, and finally mapped to the chronic disease diagnosis category probability output space to generate the model's refined classification prediction logic score vector. The whole hierarchical progressive nonlinear inference operation process is strictly as shown in formula (42):

[0157] (42),

[0158] in, It is the first-level hidden layer high-dimensional feature compression learnable weight projection matrix, which is responsible for completing the cross-channel feature weight redistribution; This is the adaptive bias correction vector for the hidden layer nonlinear mapping; For piecewise nonlinear activation constraint functions; To accurately adapt the weight mapping matrix to the second-level output layer and adapt it to the spatial dimension of chronic disease diagnosis categories; This is the global equilibrium bias compensation vector for the output layer. This is the confidence score vector for the global prediction of chronic disease diagnosis categories in the final output of the model.

[0159] S5.2: Based on this, combined with real chronic disease labels, a refined cross-entropy classification loss function with smooth regularization constraint is constructed. The difference between the predicted confidence distribution and the real label distribution is compared sample by sample to quantify the classification error in a targeted manner. Simultaneously, the training bias problem of small sample chronic disease categories is suppressed to enhance the model's global diagnostic adaptability. The calculation logic of refined classification supervision loss is shown in formula (43):

[0160] (43),

[0161] in, This represents the main loss term in the batch-normalized chronic disease diagnosis classification across the entire domain; The total number of patient samples for single-batch synchronous inference training; Preset the total number of categories for chronic disease subtype diagnosis; For the first patient, the corresponding number is... Distribution of unique heat markers in the true clinical gold standard for chronic diseases; The corresponding category prediction probability distribution of the model's hierarchical inference output is quantified by the global log-likelihood bias to precisely constrain the gradient update direction of the model's classification. Finally, the overall training loss of the model is defined as shown in formula (44):

[0162] (44),

[0163] Experimental verification

[0164] To verify the performance of the method proposed in this application (denoted as PHKF-Net), this paper conducts extensive experimental validation on a private dataset. This section first introduces the dataset, experimental environment, and parameter settings used in the experiments, then describes the nine baseline models used for comparison experiments, and finally introduces the specific details of the comparison and ablation experiments, and provides an in-depth analysis of the experimental results.

[0165] 1. Dataset

[0166] This paper uses a private dataset—PHMD—to verify the effectiveness of the proposed model, PHKF-Net. The following sections will provide a detailed description of this dataset.

[0167] PHMD: The PHMD dataset contains 5560 real patient samples, focusing on the construction of disease diagnosis scenarios. Each sample is a combination of multimodal heterogeneous medical data, comprehensively covering multi-dimensional clinical information of patients, mainly including four core data dimensions: patient complaint text information, physiological test experimental index data, outpatient clinical notes text records, and accurately labeled chronic disease classification tags. PHMD includes 4 disease label classifications, and is divided into training set, validation set, and test set in a 3:1:1 ratio.

[0168] 2. Parameter settings

[0169] In this paper, the experiments were conducted on a personal computer equipped with the following hardware: Windows 10 operating system, Intel® Core™ i9-10900K processor, NVIDIA 3090 graphics card, and 96GB of RAM. The model was implemented using PyTorch 1.4.0, with Python 3.6 as the programming language. This paper uses the Adam optimizer to minimize the total loss function; the parameters are summarized in Table 1.

[0170] Table 1 Experimental parameter settings

[0171]

[0172] 3. Baseline Model

[0173] To verify the performance of the proposed model PHKF-Net, experiments were conducted to test it against six state-of-the-art multimodal learning models and three unimodal learning models. A detailed description of the nine benchmark models is as follows:

[0174] BERT[#]: BERT is a pre-trained language model built on the Transformer encoder, designed to model text using bidirectional contextual information to more accurately understand semantic content.

[0175] Wav2vec[#]: Wav2Vec is an end-to-end speech representation learning model based on a deep learning framework, which can extract high-level speech feature representations from raw audio signals.

[0176] R-CNN[#]: This is an intent prediction framework that combines TalkNet with Faster R-CNN. The method first uses the TalkNet module to capture the user in each frame of the image, then uses Faster R-CNN to extract user features, and finally imports the extracted feature data into a classifier to achieve intent recognition.

[0177] MulT[#]: MulT integrates a bidirectional cross-modal attention mechanism into an end-to-end framework to capture the interaction relationships between unaligned multimodal data.

[0178] MAG-B[#]: MAG-B introduces a multimodal adaptive gating mechanism (MAG) on the basis of BERT, enabling BERT to process multimodal data other than language during fine-tuning.

[0179] IF-MMIN[#]: IF-MMIN integrates invariant feature learning to address intermodal differences and predict missing data, maintaining robust performance even when modalities are incomplete or uncertain.

[0180] MISA[#]: MISA employs a multi-task learning framework that projects modalities to invariant and specific subspaces to generate representations that combine shared characteristics with modality specificity.

[0181] EMRFM[#]: EMRFM constructs modality sharing and modality-specific encoders to learn shared and unique features, and employs an adaptive fusion mechanism to reduce noise in multimodal intent recognition in real-world scenarios.

[0182] SDIF-DA[#]: SDIF-DA employs a data-enhanced shallow-to-deep interaction framework to progressively align and fuse multiple modal features.

[0183] 4. Performance Verification

[0184] To verify the performance of the proposed model PHKF-Net, experiments were conducted on the PHMD dataset to compare PHKF-Net with nine contrast models, as shown in Table 2. The performance metrics evaluated in this paper mainly include F1 score, accuracy, precision, and recall. As shown in Table 2, compared to models based on unimodal data, the multimodal data-based model significantly outperforms the unimodal model on the PHMD dataset. Furthermore, PHKF-Net surpasses the contrast models in all metrics on the PHMD dataset. PHKF-Net's F1 score improved by 0.65% to 3.91%, accuracy by 1.04% to 2.76%, precision by 0.85% to 3.8%, and recall by 1.01% to 2.98%.

[0185] Table 2 shows the comparative experimental results in the PHMD dataset.

[0186]

[0187] To compare the performance of PHKF-Net with four benchmark models—SDIF-DA, EMRFM, MISA, and MAG_B—on training sets of varying sizes, this paper uses a proportional selection method for the training set to test the performance metrics of the five models. For example, the training set is partitioned into stepwise subsets ranging from 10% to 100%, and the F1 scores of the five models are tested. Figure 3 The test results show that increasing the size of the training samples directly improves the model's performance. However, sometimes, as the proportion of the training dataset increases, the measured F1 score tends to decrease. This is because the selected training set is random, resulting in fewer samples for some classes, leading to data imbalance and affecting the model's recognition performance. Furthermore, compared with other models, the performance improvement of PHKF-Net is relatively stable, and in most cases, it outperforms other models. Therefore, the proposed PHKF-Net can effectively alleviate the sample imbalance problem and has good stability.

[0188] 5. Ablation test

[0189] To systematically evaluate the impact of each component of the PHKF-Net model on the overall recognition performance, this study conducted ablation experiments on the modules, and the specific data are listed in Table 3.

[0190] Table 3. Results of PHKF-Net ablation experiments

[0191]

[0192] As shown in Table 3, scheme (1) is the complete PHKF-Net model; schemes (2) and (3) are variants generated after removing the corresponding innovation modules; the experimental results show that, compared with the complete model, the absence of any innovation module leads to a significant decrease in the overall performance of the model, which verifies the rationality of the PHKF-Net architecture design. The following detailed analysis is performed: After removing the UF module, the F1 score, accuracy, precision, and recall of the model decreased by 0.52%, 1.07%, 0.71%, and 1.00%, respectively. After introducing the patient health profile, this module can provide key patient background prior knowledge for chronic disease diagnosis. The removal of the KDF module also has the most significant impact on the model performance, with the accuracy decreasing by 0.98% and the F1 score decreasing by 1.13%. As the main component for knowledge and multimodal data fusion, KDF achieves semantic alignment and interaction between knowledge and multimodal features through a multi-layer cross-modal attention mechanism, a layer normalization (LN) layer, and a feedforward network (FFN).

[0193] Example 2

[0194] This embodiment provides a chronic disease classification and prediction system driven by knowledge and data fusion.

[0195] A computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the aforementioned knowledge and data fusion-driven chronic disease classification and prediction method.

[0196] A terminal device includes a processor and a computer-readable storage medium, the processor being configured to implement various instructions; the computer-readable storage medium being configured to store multiple instructions adapted for loading and execution by the processor of the aforementioned knowledge and data fusion-driven chronic disease classification and prediction method.

[0197] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A chronic disease classification and prediction method based on knowledge and data fusion, characterized in that, include: Acquire multi-source clinical chronic disease data; Based on the acquired clinical chronic multi-source data, construct patient health profile knowledge; Based on the improved BERT encoder, high-density semantic encoding is performed on patient health profile knowledge to obtain a profile knowledge representation matrix; The portrait knowledge embedding vector is obtained by targeted extraction based on the portrait knowledge representation matrix; Multimodal data features are extracted from multi-source clinical chronic disease data based on an improved Transformer encoder; Data feature fusion based on CMD multimodal alignment mechanism; A knowledge-aware interactive fusion network is used to perform deep interactive fusion of the fused data features and the profile knowledge embedding vectors. A knowledge-guided augmented representation method is used to classify and predict the features of the fused data. Output the prediction results; The high-density semantic encoding of patient health profile knowledge based on the improved BERT encoder includes first encoding a batch of profile text sequences. The data is fed into the BERT encoding backbone for standardized embedding encoding, while simultaneously retrieving the globally implicit semantic association mappings that have been solidified within the set. With the patient's ID number ,Will and Preprocessed into a pathological association bias matrix that perfectly matches the attention calculation dimension. Within each layer of bidirectional self-attention, a Query, Key, and Value matrix is ​​generated normally based on the input profile text sequence, completing the calculation of the original attention similarity score of the native context. Subsequently, before the Softmax normalization operation node, the preprocessed pathological bias matrix is ​​element-wise weighted and superimposed to compensate the original attention score matrix. External medical prior constraints are used to correct literal semantic bias, forcibly increasing the correlation response weights between heterogeneous health attributes related to chronic diseases, and strengthening the static baseline features. Then, after multi-layer improved self-attention iterative operation and output of multi-scale fused hidden layer semantic features, the semantic features are globally pooled and... Standardized operations are then implemented, followed by the introduction of weighted regularization terms for chronic disease clinical attributes. Layered fusion, bias correction, and global spatial alignment are then performed on multi-scale hidden layer features to unify and standardize the vector distribution range of all patient profile representations, ultimately generating patient-specific profile knowledge embedding vectors with strong pathological correlation. The method for extracting multimodal data features from multi-source clinical chronic disease data based on the improved Transformer encoder includes performing differentiated purification and dimension mapping on different types of original data based on modality-specific adaptive embedding methods, filtering invalid noise and retaining modality-specific effective features; Then, based on the adaptive embedding preprocessing, a hierarchical gated multi-head attention feature extraction mechanism is introduced, which breaks down the feature extraction into modal feature projection, gated attention weight calculation, multi-head feature splicing and fusion, and residual feature calibration, and completes multimodal deep feature extraction layer by layer. To adapt to the dimensional characteristics of different modal features, a modality-specific projection matrix is ​​constructed to perform linear transformations on the input features (query, key, and value) to complete feature dimensional adaptation and preliminary feature reconstruction. Finally, a modality-adaptive gating factor is introduced to filter and correct the original attention weights, suppressing redundant feature weights and strengthening core feature weights. After obtaining the corrected attention weights, the corrected attention weights are weighted and fused with the value matrix to obtain single-head attention features. Then, multiple sets of single-head features are concatenated and fused to achieve multi-dimensional feature information complementarity. After completing the concatenation and fusion of multi-head features, the final multimodal deep features are obtained through linear mapping.

2. The chronic disease classification and prediction method based on knowledge and data fusion as described in claim 1, characterized in that, The knowledge of constructing patient health profiles based on acquired clinical chronic multi-source data includes pre-identifying all target patients to be assessed and standardizing the formalized format: ,in, This represents the total number of chronic disease patients diagnosed in a batch. Standardized and organized data is then aggregated, and a unique health profile is independently constructed for each patient, forming a global set of structured health profiles for all patients. ,in, Indicates the patient Personalized health profile This represents the semantic alignment and normalization processing function. Structured data representing patients, This represents unstructured patient data, where each patient's profile consists of attribute labels and corresponding semantic values, formally represented as follows: ,in, Indicates the image attribute tags. Standardized content for each attribute tag after semantic transformation; to avoid the loss of chronic disease correlation between attributes due to isolated encoding of a single profile, a globally unified semantic space model is carried out for batch patient profiles across the entire domain. Discrete and independent single profile texts are transformed into a set of structured topological representations with internal implicit correlation constraints, thus constructing a global patient profile structured representation space, and a unified sequence of global patient profiles: ,in, A collection of text sequences representing all patient profiles. This represents the implicit semantic association mapping between various attributes within a portrait; finally, it represents the portrait set. Perform preprocessing, call the collection Complete set of internal text sequences Simultaneously, retrieve the global implicit semantic association mapping. As a structural prior constraint, each patient is extracted one by one according to the patient sample number correspondence rule. The corresponding original portrait text sequence Each sequence fully reproduces the semantic content of the multi-source heterogeneous health profile key-value pairs constructed earlier, thus obtaining a global health profile sequence set for patients. ,in Indicate the patient's ID number, and then for each entry Based on the unified semantic standard for chronic disease diagnosis and treatment, character compliance verification, invalid noise attribute masking, partial semantic completion of time-series health information, and alignment of attribute labels with semantic descriptions are completed. The length distribution of the input sequence is unified and fixed, forming a standardized medical text tensor sequence that can be encoded in parallel batches. This avoids distortion of the encoding gradient caused by heterogeneous images of varying lengths.

3. The chronic disease classification and prediction method based on knowledge and data fusion as described in claim 2, characterized in that, The high-density semantic encoding of patient health profile knowledge based on the improved BERT encoder includes first encoding a batch of profile text sequences. The data is fed into the BERT encoding backbone for standardized embedding encoding, while simultaneously retrieving the globally implicit semantic association mappings that have been solidified within the set. With the patient's ID number ,Will and Preprocessed into a pathological association bias matrix that perfectly matches the attention calculation dimension. Within each layer of bidirectional self-attention, the Query, Key, and Value base matrices are generated normally based on the input profile text sequence, completing the calculation of the original attention similarity score of the native context. Subsequently, before the Softmax normalization operation node, the preprocessed pathological bias matrix is ​​element-wise weighted and superimposed to compensate the original attention score matrix. External medical prior constraints are used to correct literal semantic bias, forcibly increasing the correlation response weights between heterogeneous health attributes related to chronic diseases, and strengthening static baseline features. This layer-by-layer improvement operation process is represented as follows: ,in, These are the query matrix, key matrix, and value matrix generated in real time within the encoding layer for the i-th image text sequence; This is the inherent top-level feature dimension of the encoder model; An adaptive attention weight allocation matrix for clinical attributes of chronic diseases; For global prior balanced hyperparameters; To construct and reuse the global profiling pathology association bias matrix for front-end offline construction; A mask matrix for masking invalid medical semantics; It is a hierarchical adaptive Softmax normalization mapping function; The semantic features are then processed through multiple layers of improved self-attention iterative computation to output multi-scale fused hidden semantic features. First, global pooling is performed on these semantic features. Standardized operations are then implemented, followed by the introduction of regularization terms based on chronic disease clinical attributes. This involves hierarchical fusion, bias correction, and global spatial alignment of multi-scale hidden features to unify and standardize the vector distribution range of all patient profile representations. Ultimately, this generates patient-specific profile knowledge embedding vectors with strong pathological correlations, representing the learning process. ,in, This is a pooling operation for the top-level standard global temporal features of BERT. This is the standard global dimension normalization correction function; To standardize the hidden feature dimensions globally, a representation learning process is conducted to ultimately obtain a global profile knowledge embedding representation with unified dimensions and semantic consistency. ,in, An embedded representation of a single patient's profile. A global image knowledge embedding representation.

4. The chronic disease classification and prediction method based on knowledge and data fusion as described in claim 3, characterized in that, The process of extracting portrait knowledge embedding vectors based on the portrait knowledge representation matrix includes, firstly, relying on global semantic association mapping... Simultaneously, a structured association index table is established with the patient's unique target ID as the primary key, and semantic association mapping is used to achieve this. Each patient's unique identifier ID is bound to a personalized profile knowledge embedding vector output by an improved BERT algorithm, forming a fast-addressable bidirectional association mapping relationship. This achieves positional synchronization between the ID and the profile knowledge vector. Then, for the target patient to be diagnosed, the corresponding unique target ID is extracted as the search keyword, relying on global semantic association mapping. Pre-established corresponding associated index tables in the middle are used to carry out precise key-value matching and addressing; By using ID-based forward fast matching of index ledgers, global profile knowledge is embedded. The system directly locates the vector storage location corresponding to the target patient, accurately retrieves the personalized profile knowledge embedding representation unique to that patient, without needing to traverse and calculate all global profile vectors, thus avoiding cross-patient profile feature confusion and interference. The structured index retrieval process is represented as follows: ,in, This indicates a patient ID-based retrieval operation, and is performed through a global semantic association mapping. Return to target patient Image knowledge embedding .

5. The chronic disease classification and prediction method based on knowledge and data fusion as described in claim 4, characterized in that, The method for extracting multimodal data features from multi-source clinical chronic disease data based on an improved Transformer encoder includes differential purification and dimensionality mapping of different types of raw data using a modality-specific adaptive embedding method, filtering out invalid noise, and retaining modality-specific effective features. The differential adaptive embedding formula is as follows: =LN , =LN , =LN , in, , , Learnable embedding mapping layers are designed for three modalities, adapting feature mappings for text semantics, physiological signals, and professional medical notes, respectively; LN For layer normalization operation; , , The modality-adaptive noise suppression coefficient is dynamically updated through model backpropagation, balancing the preservation of original information with the effect of feature purification. , , The input features are refined after preprocessing. Then, based on adaptive embedding preprocessing, a hierarchical gated multi-head attention feature extraction mechanism is introduced, breaking down feature extraction into modal feature projection, gated attention weight calculation, multi-head feature concatenation and fusion, and residual feature calibration, progressively completing multimodal deep feature extraction. To adapt to the dimensional characteristics of different modal features, a modality-specific projection matrix is ​​constructed, performing linear transformations on the input features (query, key, and value) to complete feature dimensionality adaptation and preliminary feature reconstruction. The formula is: , , ,in, These correspond to three modalities: chief complaint, physiological experiment, and clinical notes. , , The projective weight matrix is ​​an independent and learnable projection weight matrix for each modality; , , These are the query matrix, key matrix, and value matrix corresponding to each modality. Finally, a modality-adaptive gating factor is introduced to filter and correct the original attention weights, suppressing redundant feature weights and strengthening core feature weights. The calculation formula is as follows: ,in, This is the scaling factor; For Hadamard product operations; As the modality adaptive gating matrix, after obtaining the corrected attention weights, the corrected attention weights are weighted and fused with the value matrix to obtain single-head attention features. Then, multiple sets of single-head features are concatenated and fused to achieve multi-dimensional feature information complementarity. The formula is as follows: , ,in, For the number of attention heads; For the first Single-granularity modal features extracted by the attention head; This involves a multi-head feature concatenation operation to achieve multi-scale feature fusion. After the multi-head feature concatenation and fusion are completed, the final multimodal deep features are obtained through linear mapping, represented as follows: ,in, This is used to balance the information ratio between deep attention features and original preprocessed features.

6. The chronic disease classification and prediction method based on knowledge and data fusion as described in claim 5, characterized in that, The data feature fusion based on the CMD multimodal alignment mechanism includes performing multimodal fusion on the subject description features, physiological experiment features, and clinical note features using a multimodal fusion module based on an improved Transformer. After each modal feature enters the multimodal fusion module, three self-attention mechanisms are first used to capture and learn the global information of each modality, as shown below: , in, Representing each mode, , , , These represent the query matrix, key matrix, and value matrix, respectively. This represents the parameters learned during model training. The dimension of the key matrix is ​​represented; then, similarity learning is performed on the feature pairs of text-audio and text-image, using the central moment difference (CMD) as the similarity loss function to calculate the similarity between the text-audio and text-image modalities. By minimizing CMD, the image and audio modalities are guided to align with the text modal at the distribution level, as shown below: , in, It is the empirical expectation vector of sample x, and It is a vector of the central moments of all k-order samples at the x-coordinate, used to calculate the central moments between text-audio and text-image modalities. The final representation is: ,in, It is a feature vector extracted from the self-attention mechanism. Then, cross-modal attention was used to learn features related to physiological experiments and clinical notes. Clinical notes on physiological experimental characteristics Physiological experiments to the main description of relevant characteristics and the main description of physiological experiment-related characteristics A total of four cross-modal transformers are needed to obtain four feature vectors. Each cross-modal attention mechanism consists of n layers of cross-modal attention modules, with information transferred from different modalities. Transition to mode The formula for cross-modal attention is: , , , in, It is by The parameterized location-feedforward sublayer, CT is a multi-head cross-modal attention module, and LN is layer normalization; then, features related to physiological experiments and clinical notes are combined. Clinical notes on physiological experiment-related features The data is then input into the clinical gating system to obtain the final clinical note features, which are then combined with physiological experiments to obtain subject-related features. and the main description of physiological experiment-related characteristics And input it into the subject gating to obtain the final subject features, represented as: , , in, This indicates a splicing operation. Indicates a convolutional layer. The sigmoid function represents the tensor; finally, the feature vectors of all modalities are combined to obtain the final fused feature representation. ,in, This indicates a splicing operation. These are physiological experimental features and clinical note features extracted through self-attention mechanisms. This is the result of clinical gating fusion processing, with the main characteristic described. This is the result of the main gating fusion processing.

7. The chronic disease classification and prediction method based on knowledge and data fusion as described in claim 6, characterized in that, The method utilizes a knowledge-aware interactive fusion network to perform deep interactive fusion of the fused data features and profile knowledge embedding vectors. This includes, in an interaction dominated by fused features, firstly, fused features... Image knowledge representation splicing to obtain spliced ​​representation Then fuse the features Physiological experimental characteristics splicing to obtain spliced ​​representation To enhance the stability of feature interactions, residual connections and layer normalization (Add&Norm) are introduced into the cross-modal attention (CA) module. Then, the fused features are... As a query As the key and value, the result is obtained through CA calculation, residual join, and layer normalization (Add&Norm). Similarly, the fusion features As a query As the key and value, they are obtained through the CA and Add&Norm modules. Secondly, and Summing the elements yields the representation. , represented as: , , ,in, Indicates the dominant mode, and This represents the splicing representation that has been merged. , , These are mapping matrices for queries, keys, and values, respectively. Let be the dimension of the key vector, and then, to obtain a more stable and uniformly distributed feature representation, The input is fed into the feedforward network FFN and Add&Norm layers to produce an enhanced representation that is structurally robust and has good semantic expressiveness. , represented as: ,in To integrate features An enhanced representation of the final output of the dominant branch.

8. The chronic disease classification and prediction method based on knowledge and data fusion as described in claim 7, characterized in that, The method of utilizing a knowledge-aware interactive fusion network to perform deep interactive fusion of the fused data features and profile knowledge embedding vectors also includes prioritizing physiological experiment features in the interaction process, given that the importance of the subject-narrative modality data and clinical note modality data is weaker than that of the physiological experiment modality data. In the dominant interaction, the physiological experimental characteristics are first presented. With fusion features splicing to obtain spliced ​​representation Then, physiological experimental characteristics Image knowledge representation splicing to obtain spliced ​​representation Next As a query As the key and value, they are obtained through cross-modal attention (CA) calculation followed by residual connection and layer normalization (Add&Norm). Similarly, As a query As keys and values, they are obtained through CA, residual join, and layer normalization (Add&Norm). Secondly, and Summing the elements yields the representation. The process is represented as follows: , , in, Indicates the dominant mode, and This represents the splicing representation that has been merged. , , These are mapping matrices for queries, keys, and values, respectively. Let be the dimension of the key vector; then, to obtain a more stable and uniformly distributed feature representation, The input is fed into the feedforward network FFN and Add&Norm layers, thereby generating an enhanced representation with structural robustness and good semantic expressiveness. : ,in To use physiological experimental characteristics An enhanced representation of the final output of the dominant branch.

9. A chronic disease classification and prediction method based on knowledge and data fusion as described in claim 8, characterized in that, The method of using knowledge-guided augmented representation to classify and predict the features of fused data includes augmenting the representations of three stages. , and The images are then spliced ​​together to obtain the final enhanced representation. : ,in, This indicates a splicing operation. It is a branch-enhanced representation dominated by fusion features. It is a branched enhancement representation dominated by physiological experimental characteristics. This is a branch of augmented representation primarily based on image knowledge representation; to achieve accurate identification and intelligent assisted diagnostic classification output for chronic diseases, the final augmented representation with knowledge guidance capabilities will be used. The data is fed into a hierarchical projection classification module to complete high-order feature extraction and nonlinear discriminative mapping. The final enhanced representation of the output is then processed. The first-level hidden layer projection fully connected unit is fully input, and cross-dimensional feature ordered compression mapping is completed based on the learnable weight matrix. An adaptive bias compensation vector is superimposed to correct the baseline deviation of the representation. Then, the ReLU nonlinear activation function is embedded to complete the global feature sparsity denoising process, effectively eliminating invalid and redundant noise components in the fused representation and highlighting the core discriminative pathological features of chronic diseases. Then, it is connected to the second-level output mapping fully connected layer to complete the second accurate projection calibration of feature dimension adaptation. Finally, it is mapped to the chronic disease diagnosis category probability output space to generate the model's refined classification prediction logic score vector. ,in, The first-level hidden layer high-dimensional feature compression learnable weight projection matrix; This is the adaptive bias correction vector for the hidden layer nonlinear mapping; For piecewise nonlinear activation constraint functions; To accurately adapt the weight mapping matrix to the second-level output layer; This is the global equilibrium bias compensation vector for the output layer. The final output of the model is the global prediction confidence score vector for chronic disease diagnosis categories. Finally, combining the real chronic disease labels, a refined cross-entropy classification loss function with smoothing regularization constraints is constructed. This function compares the difference between the predicted confidence distribution and the real label distribution sample by sample, quantifies the classification error, and simultaneously suppresses the training bias problem of small-sample chronic disease categories, thus enhancing the model's global diagnostic adaptability. The refined classification supervision loss is expressed as follows: ,in, This represents the main loss term in the batch-normalized chronic disease diagnosis classification across the entire domain; The total number of patient samples for single-batch synchronous inference training; Preset the total number of categories for chronic disease subtype diagnosis; For the first patient, the corresponding number is... Distribution of unique heat markers in the true clinical gold standard for chronic diseases; This is the corresponding category prediction probability distribution output by the hierarchical inference model.

10. A chronic disease classification and prediction system based on knowledge and data fusion, executing the chronic disease classification and prediction method based on knowledge and data fusion as described in claim 1, characterized in that, include: The data acquisition module is configured to acquire multi-source clinical chronic disease data. The profiling module is configured to construct patient health profile knowledge based on acquired clinical chronic multi-source data; The matrix module is configured to perform high-density semantic encoding on patient health profile knowledge based on an improved BERT encoder to obtain a profile knowledge representation matrix. The embedding module is configured to perform targeted extraction based on the profile knowledge representation matrix to obtain the profile knowledge embedding vector; The feature module is configured to extract multimodal data features from multi-source clinical chronic disease data based on an improved Transformer encoder; The fusion module is configured to perform data feature fusion based on the CMD multimodal alignment mechanism; The interaction module is configured to use a knowledge-aware interaction fusion network to perform deep interactive fusion of the fused data features and profile knowledge embedding vectors. The prediction module is configured to perform classification prediction on the fused data features using a knowledge-guided augmented representation method. The output module is configured to output the prediction results.

Citation Information

Patent Citations

  • Text enhanced multi-mode spectrum occupancy prediction system and method

    CN121486834A

  • Mapping knowledge domain-based oral cavity multi-modal data fusion method

    CN122050673A