Multi-modal deep learning fused early-stage intelligent screening system for rheumatic diseases
By integrating clinical phenotype, laboratory test and imaging information through a multimodal deep learning fusion system, and utilizing topological sensing graph convolutional networks and continuous homology feature extraction technology, the problem of insufficient early symptom identification of rheumatic diseases is solved, and early accurate screening and diagnosis are achieved.
Patent Information
- Application Number
- CN202511796150.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-03-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies struggle to effectively integrate multimodal medical data, resulting in insufficient ability to identify early symptoms of rheumatic diseases, often leading to misdiagnosis and delayed treatment.
A multimodal deep learning fusion system is adopted, which integrates clinical phenotype, laboratory test and imaging information by using topology-aware graph convolutional network and continuous homology feature extraction technology. The information is then represented in a structured manner through knowledge graph module, and multi-scale feature extraction and cross-modal feature fusion are performed by spectral topology fusion network to generate rheumatic disease risk prediction results.
It achieves precision and comprehensiveness in early screening of rheumatic diseases, identifying early signs up to 6 months in advance, with a diagnostic accuracy of 85%, specificity of 78%, and sensitivity of 88%, providing reliable support for early intervention.
Smart Images

Figure CN121641397A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical artificial intelligence, and in particular to an early intelligent screening system for rheumatic diseases based on multimodal deep learning fusion, which is used to achieve early intelligent screening of rheumatic diseases by integrating clinical phenotypic information, laboratory test information and imaging information. Background Technology
[0002] Rheumatic diseases are a group of autoimmune diseases characterized primarily by inflammation of the joints and connective tissues, including rheumatoid arthritis, systemic lupus erythematosus, and ankylosing spondylitis. Early symptoms of these diseases are often atypical and are frequently overlooked or misdiagnosed, leading to delayed treatment and irreversible joint damage and organ dysfunction.
[0003] Currently, the diagnosis of rheumatic diseases mainly relies on the experience and judgment of specialists, combined with laboratory tests and imaging examinations. However, the following problems exist: First, there is a lack of specialists in primary healthcare institutions, resulting in limited diagnostic capabilities; second, it is difficult to integrate multi-source heterogeneous medical data, making it difficult to comprehensively assess the patient's condition; and third, traditional diagnostic methods are insufficient in recognizing early, subtle symptoms, often requiring waiting for more obvious symptoms before a diagnosis can be made.
[0004] In recent years, artificial intelligence technology has made significant progress in the application of medical technology. Existing technologies mainly focus on single-modal data analysis, such as analyzing only imaging data or only laboratory test data, lacking deep fusion of multimodal data. Furthermore, existing methods often employ standard deep learning models, neglecting the structured representation of medical knowledge and failing to capture complex disease correlation patterns, especially weak signals in the early stages.
[0005] Therefore, there is an urgent need for an intelligent system that can effectively integrate multimodal medical data and has early screening capabilities to provide decision support for the early intervention of rheumatic diseases. Summary of the Invention
[0006] The purpose of this invention is to provide an early intelligent screening system for rheumatic diseases based on multimodal deep learning fusion. By introducing topologically perceptual graph convolutional networks and continuous homology feature extraction technology, it effectively integrates clinical phenotype, laboratory test and imaging information to achieve early and accurate screening of rheumatic diseases.
[0007] This invention proposes a multimodal deep learning fusion-based intelligent early screening system for rheumatic diseases, comprising:
[0008] The intelligent question-and-answer module is used to collect patients' clinical phenotype information, laboratory test information, and imaging information;
[0009] The knowledge graph module is communicatively connected to the intelligent question-answering module. It is used to receive information collected by the intelligent question-answering module, convert the clinical phenotype information, laboratory test information and imaging information into a simple complex knowledge representation structure, and establish a medical knowledge topology space.
[0010] The auxiliary diagnosis module is communicatively connected to the knowledge graph module. Based on the topology-aware graph convolutional network and self-attention mechanism, it is used to extract multi-scale continuous homology features from the simple complex knowledge representation structure. Based on the spectral topology fusion network, it realizes multi-modal feature fusion and generates rheumatic disease risk prediction results.
[0011] The stratified diagnostic module, which is communicatively connected to the auxiliary diagnostic module, is used to receive the risk prediction results of the rheumatic disease, classify the patient's risk, and generate diagnostic suggestions.
[0012] Preferably, the intelligent question-answering module inputs the clinical phenotype information, laboratory test information, and imaging information in a structured manner using unified annotation rules, and uses BERT and Word2vec to encode the input information into a vector representation, which is then transmitted to the knowledge graph module.
[0013] Preferably, the process by which the knowledge graph module converts the clinical phenotype information, laboratory test information, and imaging information into a simple complex knowledge representation structure includes:
[0014] Medical entities are represented as 0-simplexes, binary relations between entities are represented as 1-simplexes, and complex relations between multiple entities are represented as higher-order simplexes.
[0015] Construct boundary mappings to describe the topological relationships between simplexes;
[0016] Establish nested complex structures to capture complex medical association patterns in the diagnosis of rheumatic diseases.
[0017] Preferably, the process of multi-scale continuous homology feature extraction in the auxiliary diagnostic module includes:
[0018] Construct multi-scale filtered complex sequences and control the filtering of relation strength by setting multi-level thresholds;
[0019] Calculate topological features at different scales to generate a persistent graph;
[0020] Extract the feature vectors of the persistent graph and quantify the importance of the features using persistent entropy;
[0021] Enhancement processing is performed on features with persistent high signal but weak signal to achieve early detection of subtle symptoms.
[0022] Preferably, the spectral topology fusion network in the auxiliary diagnostic module includes:
[0023] The spectral domain transformation unit is used to construct the simple complex Laplacian operator to map medical features to the spectral domain;
[0024] The multi-spectral filtering unit is used to extract low-frequency, mid-frequency, and high-frequency features, which correspond to the global disease pattern, main manifestations, and individual differences, respectively.
[0025] Topology-aware attention units are used to assign attention weights based on the importance of topological locations;
[0026] The cross-modal fusion unit is used to integrate features from different modalities, resolve data conflicts, and generate a comprehensive feature vector.
[0027] Preferably, the hierarchical diagnostic module includes:
[0028] A risk assessment unit is used to classify patients into high-risk and low-risk groups;
[0029] Disease typing unit, used to identify specific rheumatic disease types in high-risk patients;
[0030] The diagnostic suggestion generation unit is used to generate clinical diagnostic suggestions and interventions based on risk assessment results and disease classification results.
[0031] Preferably, the multi-level threshold settings in the multi-scale continuous homology feature extraction process include:
[0032] Set 10 gradient thresholds, from high correlation strength (0.9) to low correlation strength (0.1).
[0033] A long window (6 months) was used to assess chronic progression characteristics.
[0034] A short time window (7 days) is used for acute features;
[0035] Threshold sensitivity is dynamically adjusted based on patient baseline data.
[0036] Preferably, the topology-aware attention unit employs a multi-head attention mechanism, including:
[0037] Eight independent attention heads, each focusing on a different type of medical association;
[0038] Calculate the topological position score based on node centrality and the importance of continuous homology;
[0039] The weights of each head are dynamically balanced according to the distribution of training data.
[0040] Generate an attention weight matrix to guide the multimodal feature fusion process.
[0041] Preferably, the cross-modal fusion unit includes:
[0042] Modality-specific encoders are designed with specialized feature extractors for clinical phenotypes, laboratory tests, and imaging information;
[0043] A shared representation mapper maps features from different modalities to a unified topological space;
[0044] The conflict resolver handles data inconsistencies based on a confidence-weighted average.
[0045] Information completer, used to handle feature inference in the case of missing modalities.
[0046] Preferably, the system connects to the hospital information system, laboratory information system, and medical imaging system to achieve automatic data acquisition; the system can identify early signs of rheumatic diseases up to 6 months in advance, with a diagnostic accuracy rate of 85%, providing decision support for clinical intervention.
[0047] The present invention has the following beneficial effects:
[0048] 1. Early diagnosis time: This system can identify early signs of rheumatic diseases up to 6 months in advance, which is significantly earlier than traditional diagnostic methods, creating conditions for early intervention and preventing disease progression and irreversible joint damage.
[0049] 2. Multi-source data integration: The system innovatively integrates clinical phenotypes, laboratory tests, and imaging features. Through simple complex knowledge representation and spectral topology fusion network, it preserves the intrinsic correlation of multi-source data, thereby improving the comprehensiveness and accuracy of diagnosis.
[0050] 3. Enhancement of weak signals: By extracting the topological stability features of the disease through the theory of continuous cohomology, the system can amplify the weak but persistent early manifestations that are easily overlooked by traditional methods, providing a new perspective for early diagnosis.
[0051] 4. High diagnostic accuracy: The system achieves an 85% detection rate in the early diagnosis of rheumatic diseases, with a specificity of 78% and a sensitivity of 88%, which is superior to existing screening methods and provides reliable support for clinical decision-making.
[0052] 5. Flexible clinical application: The system supports various application scenarios such as primary healthcare screening, specialist auxiliary diagnosis, and multi-center clinical research, and has good promotional value. Attached Figure Description
[0053] Figure 1 This is a diagram of the overall system architecture of the present invention;
[0054] Figure 2 This is a structural diagram of the intelligent question-answering module of the present invention;
[0055] Figure 3 This is a structural diagram of the knowledge graph module of the present invention;
[0056] Figure 4 This is a structural diagram of the auxiliary diagnostic module of the present invention;
[0057] Figure 5 This is a structural diagram of the hierarchical diagnostic module of the present invention;
[0058] Figure 6 Flowchart for multi-scale continuous homology feature extraction;
[0059] Figure 7 This is a diagram of the spectrum topology fusion network architecture. Detailed Implementation
[0060] Please refer to Figures 1-7 The present invention will now be described in further detail with reference to the accompanying drawings.
[0061] Reference Figure 1 The multimodal deep learning fusion-based early intelligent screening system for rheumatic diseases of the present invention includes an intelligent question answering module 1, a knowledge graph module 2, an auxiliary diagnosis module 3, and a hierarchical diagnosis module 4.
[0062] The intelligent question-answering module 1 collects patients' clinical phenotype, laboratory test, and imaging information. The knowledge graph module 2 communicates with the intelligent question-answering module 1, receives the collected information, and converts the clinical phenotype, laboratory test, and imaging information into a simple complex knowledge representation structure, establishing a medical knowledge topology space. The auxiliary diagnosis module 3 communicates with the knowledge graph module 2, and based on a topology-aware graph convolutional network and self-attention mechanism, extracts multi-scale continuous homology features from the simple complex knowledge representation structure. It then uses a spectral topology fusion network to achieve multi-modal feature fusion, generating rheumatic disease risk prediction results. The hierarchical diagnosis module 4 communicates with the auxiliary diagnosis module 3, receives the rheumatic disease risk prediction results, classifies patients by risk, and generates diagnostic suggestions.
[0063] Reference Figure 2 In a preferred embodiment of the present invention, the intelligent question answering module 1 inputs clinical phenotype information, laboratory test information and imaging information in a structured manner using unified annotation rules, and uses BERT and Word2vec to encode the input information into a vector representation, which is then transmitted to the knowledge graph module 2.
[0064] Specifically, the intelligent question-answering module 1 first collects patient information through standardized electronic questionnaires or natural language interactive interfaces. For clinical phenotype information, such as symptoms like joint pain, morning stiffness, and rash, standardized symptom naming conventions and severity grading (0-4 levels) are used for recording. For laboratory test information, such as RF, ESR, and CRP, specific test values and their reference ranges are recorded. For imaging information, such as X-rays and MRIs, key lesion areas and characteristic descriptions are marked.
[0065] Then, the intelligent question answering module 1 uses the BERT model to perform contextual semantic encoding on the text information. BERT is a pre-trained bidirectional Transformer encoder that can capture deep semantic relationships in text. In this system, a BERT model specifically fine-tuned for medical text (such as BioBERT or ClinicalBERT) is used, with standardized medical text descriptions as input and a 768-dimensional feature vector as output.
[0066] Simultaneously, the system uses the Word2vec model to represent medical terms as word vectors. Word2vec employs a Skip-gram model with a word vector dimension of 300, a context window size of 5, and a negative sampling number of 10, obtained through pre-training on a large-scale medical corpus. This method can map semantically similar medical terms to neighboring positions in the vector space, effectively capturing the semantic relationships between terms.
[0067] Finally, the intelligent question answering module 1 concatenates the context semantic vector generated by BERT with the word vector generated by Word2vec, fuses them through a fully connected layer to generate a unified feature vector representation, and transmits it to the knowledge graph module 2.
[0068] Reference Figure 3 In another preferred embodiment of the present invention, the process by which the knowledge graph module 2 converts clinical phenotype information, laboratory test information and imaging information into a simple complex knowledge representation structure includes: representing medical entities as 0-simplexes, binary relationships between entities as 1-simplexes, and complex relationships between multiple entities as higher-order simplexes; constructing boundary mappings to describe the topological relationships between simplexes; and establishing nested complex structures to capture complex medical association patterns in the diagnosis of rheumatic diseases.
[0069] Specifically, Knowledge Graph Module 2 first maps various types of medical information into simplex structures. In algebraic topology, a simplex is a mathematical structure capable of expressing higher-order relations. A 0-simplex corresponds to a point, a 1-simplex to a line segment, a 2-simplex to a triangle, and so on. In this system, medical entities (such as symptoms, detection indicators, and imaging features) are represented as 0-simplexes, relations between two entities are represented as 1-simplexes, and combinations of three or more entities are represented as higher-order simplexes.
[0070] For example, joint pain is a 0-simplex, elevated erythrocyte sedimentation rate (ESR) is another 0-simplex, and the correlation between the two is represented as a 1-simplex. The combined feature of symmetrical joint pain, morning stiffness >1 hour, and elevated ESR, which all point to rheumatoid arthritis, can be represented as a 2-simplex.
[0071] Then, the system constructs a boundary map to describe the topological relationships between simplexes. The boundary map is defined as follows:
[0072] ,
[0073] in: Let be a k-dimensional boundary operator, representing the mapping from a k-dimensional simplex to a k-1-dimensional simplex; Let be the free Abelian group of k-dimensional simplexes, that is, the linear combination of all k-dimensional simplexes; It is a simple complex form, representing the entire knowledge structure; This indicates a mapping relationship. Boundary mapping describes how a higher-dimensional simplex is composed of lower-dimensional simplexes. For example, the boundary of a 2-simplex (triangle) is composed of three 1-simplexes (edges).
[0074] Finally, the system establishes a nested complex structure. By setting different association strength thresholds, it constructs multi-level simple complexes to capture association patterns of varying intensities in rheumatic diseases. The association strength is determined based on clinical data statistics and expert knowledge, with thresholds typically set between 0.1 and 0.9 and a step size of 0.1, forming a 10-level nested structure.
[0075] Reference Figure 4 In another preferred embodiment of the present invention, the process of multi-scale continuous homology feature extraction by the auxiliary diagnostic module 3 includes: constructing a multi-scale filtered complex sequence and controlling the screening of relationship strength by setting multi-level thresholds; calculating topological features at different scales to generate a continuous graph; extracting feature vectors of the continuous graph and quantifying feature importance by continuous entropy; and enhancing features with high persistence but weak signal to achieve early detection of weak symptoms.
[0076] Specifically, the auxiliary diagnostic module 3 first constructs a multi-scale filtered complex sequence. The system sets a series of threshold parameters. (Typical settings) , , The strength of relations in simplex complexes is filtered to generate a series of nested complex sequences. ,in Inclusion strength The simplex.
[0077] in: The i-th threshold parameter represents the screening criterion for relation strength, with a value range of [0,1]. The total number of threshold parameters is set to 10 in this embodiment; For the i-th filtering complex, the inclusion relation strength is not less than All simplexes; Representing a subset relation, i.e. yes Subsets of , and so on, form a nested structure.
[0078] Then, the system calculates topological features at different scales. Continuous cohomology theory focuses on the "lifecycle" of topological features (such as connected components, cycles, and holes) at different scales. The system calculates each Homogroups And the sequence of Betty numbers, where Betty numbers This indicates the number of k-dimensional holes.
[0079] in: Let * represent the homology group of the i-th filtered complex, where * represents the dimension (e.g., ...). , , wait); Let be the k-th Betty number, representing the number of k-dimensional topological features (holes): Indicates the number of connected components. Indicates the number of 1-dimensional rings. This indicates the number of 2D voids, and so on.
[0080] Next, the system tracks the "birth" and "death" of topological features to construct a persistence graph. A persistence graph is a set of points on a two-dimensional plane, where each point... This indicates that a topological feature appears at parameter value b and disappears at parameter value d. Durability is defined as follows: , indicating the stability of the feature.
[0081] Where: b is the birth parameter value of the topological feature, representing the threshold for the first appearance of the feature; d is the "death" parameter value of the topological feature, representing the threshold for the disappearance of the feature; p is the persistence of the feature, calculated as... A larger value indicates a more stable feature.
[0082] The system quantifies feature importance using persistence entropy:
[0083] ,
[0084] in: Persistent entropy measures the distribution of the importance of topological features; To determine the persistence of the i-th topological feature, calculate the difference between its birth parameter and death parameter; For total persistence, it is calculated as That is, the sum of the persistence of all characteristics; This represents summing over all topological features; It is the natural logarithm; Let be the normalized persistence of the i-th feature, representing its proportion in the population. A higher persistence entropy indicates a more uniform distribution of features; a lower value indicates that some features are particularly significant.
[0085] Finally, the system enhances features with high persistence but weak signals. For features with long duration (exceeding 50% of the total scale range) but low intensity, a nonlinear enhancement function is applied:
[0086] ,
[0087] in The enhanced feature strength; The original feature strength represents the original saliency of the feature in the data; The persistence of the feature is calculated as the difference between birth parameters and death parameters; It represents the maximum persistence among all features; This is the enhancement factor, typically 3-5, used to control the degree of enhancement; This is a normalized persistence ratio. This treatment can amplify early, subtle but persistent symptom characteristics, improving early diagnostic capabilities.
[0088] Reference Figure 5 In another preferred embodiment of the present invention, the spectral topology fusion network in the auxiliary diagnosis module 3 includes: a spectral domain transformation unit for constructing a simple complex Laplacian operator to map medical features to the spectral domain; a multi-spectral filtering unit for extracting low-frequency, mid-frequency, and high-frequency features, corresponding to the global disease pattern, main manifestations, and individual differences, respectively; a topology-aware attention unit for allocating attention weights based on the importance of topological position; and a cross-modal fusion unit for integrating features from different modalities, resolving data conflicts, and generating a comprehensive feature vector.
[0089] Specifically, the spectral domain transformation unit first constructs a simplicial complex Laplacian operator. Unlike the traditional graph Laplacian, which only considers nodes and edges, the simplicial complex Laplacian considers higher-order structures. For For a 3D simple complex, the Laplace operator is defined as:
[0090] ,
[0091] in: The Laplace operator for a k-dimensional simplex complex is a square matrix whose size depends on the number of k-dimensional simplex complexes. The matrix representation of the boundary operator describes the relationship between the k-dimensional simplex and the (k-1)-dimensional simplex. for The transpose of the matrix; For the matrix representation of the k+1 dimensional boundary operator; for The transpose of the matrix; "+" denotes matrix addition. The eigendecomposition of the Laplace operator is:
[0092] ,
[0093] in: It is an eigenvalue diagonal matrix, where each diagonal element corresponds to an eigenvalue; This is an eigenvector matrix, with each column corresponding to an eigenvector; for The transpose of .
[0094] The multi-spectral filtering unit extracts features from different frequency bands by designing multiple sets of frequency response functions.
[0095] ,
[0096] ,
[0097] ,
[0098] in: These are low-frequency features, corresponding to the global disease pattern; The mid-frequency characteristic corresponds to the main clinical manifestations; These are high-frequency features, corresponding to individual differences; The input feature vector; , , These are low-pass, band-pass, and high-pass filter functions, respectively, acting on the Laplace eigenvalues; and These are the eigenvector matrix and its transpose, used to achieve transformation between the spectral and spatial domains. The low-frequency channel (0%–20% frequency band) captures the global disease pattern, the mid-frequency channel (20%–70% frequency band) captures typical clinical manifestations, and the high-frequency channel (70%–100% frequency band) captures individual differences.
[0099] The topology-aware attention unit uses a multi-head attention mechanism, and the calculation formula is as follows:
[0100] ,
[0101] in: The query matrix has dimensions (n, ...). ), where n is the number of features; K is the key matrix with dimensions (n, ..., n). V is a value matrix with dimensions (n, ...). ); The dimension of the key; Dimensions of the value; This represents matrix multiplication of the query matrix and the transpose of the key matrix, calculating the correlation between features; is a scaling factor to prevent the dot product value from becoming too large; softmax is the softmax function that converts the dot product value into attention weights; This represents matrix multiplication, applying attention weights to the value matrix. In this system, the attention weights are further adjusted based on the importance of topological positions.
[0102] ,
[0103] in: This is a topology-aware attention matrix; This is the output of the standard attention mechanism; This represents the Hadamard product (element-wise product). This is a topology weight matrix calculated based on node centrality and continuous homology importance, where each element represents the importance of the corresponding feature in the topology.
[0104] The cross-modal fusion unit integrates features from different modalities and resolves data conflicts. The formula is as follows:
[0105] ,
[0106] Where Z is the fused feature vector, and its dimension is consistent with that of each modality feature; Let i be the feature vector of the i-th mode; These are the topological attention weights corresponding to the i-th modality; This represents a weighted summation over all modes; This represents element-wise multiplication, implementing attention-based weighting.
[0107] Reference Figure 6 In another preferred embodiment of the present invention, the stratified diagnosis module 4 includes: a risk assessment unit for classifying patients into high-risk groups and low-risk groups; a disease classification unit for identifying specific rheumatic disease types in high-risk patients; and a diagnosis suggestion generation unit for generating clinical diagnosis suggestions and intervention measures based on the risk assessment results and disease classification results.
[0108] Specifically, the risk assessment unit uses a binary classification model to divide patients into a high-risk group and a low-risk group for rheumatic diseases. The model input is the comprehensive feature vector generated by the auxiliary diagnostic module 3, and the output is a risk score (a probability value between 0 and 1). The risk threshold is set at 0.5. Patients with a value higher than this are classified as high-risk and require further subtyping; patients with a value lower than this are classified as low-risk, and the system recommends regular follow-up.
[0109] The disease classification unit only performs further analysis on high-risk patients, using a multi-classification model to identify specific rheumatic disease types, including common rheumatic diseases such as rheumatoid arthritis, systemic lupus erythematosus, and ankylosing spondylitis. The model outputs the probability distribution for each disease type, using the disease type with the highest probability as the prediction result.
[0110] The diagnostic recommendation generation unit generates personalized diagnostic recommendations and interventions based on risk assessment and disease classification results, combined with clinical guidelines (such as the ACR / EULAR criteria). For high-risk patients, recommendations include further examination, specialist consultation, and treatment plans; for low-risk patients, recommendations include regular follow-up and lifestyle modifications.
[0111] Reference Figure 6 In another preferred embodiment of the present invention, the multi-level threshold setting in the multi-scale continuous coherence feature extraction process includes: setting 10 gradient thresholds, from high correlation strength (0.9) to low correlation strength (0.1); using a long time window (6 months) for chronic progression features; using a short time window (7 days) for acute features; and dynamically adjusting the threshold sensitivity based on patient baseline data.
[0112] Specifically, the system sets 10 gradient thresholds. = {0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1, 0.05}, decreasing from high to low correlation strength. High thresholds (e.g., 0.9) retain only strong correlations for identifying salient features; low thresholds (e.g., 0.1) allow weak correlations for capturing early, weak signals.
[0113] in: This is a set of threshold parameters containing 10 elements, each representing a screening criterion; the value range is [0,1], 0.9 represents the highest association strength, retaining only extremely strong correlations; 0.05 represents the lowest association strength, allowing weaker associations to be included in the analysis.
[0114] For chronic progression features (such as joint erosion and loss of function), a long time window (6 months) was used for analysis to capture slow-changing trends; for acute features (such as joint swelling and fever), a short time window (7 days) was used for analysis to capture rapidly changing features.
[0115] The system also dynamically adjusts the threshold sensitivity based on patient baseline data. For high-risk individuals (such as those with a family history or HLA-B27 positivity), the threshold sensitivity is reduced (e.g., from 0.5 to 0.4) to improve early detection capabilities; for low-risk individuals, the standard threshold setting is maintained to reduce the false positive rate.
[0116] Reference Figure 7 In another preferred embodiment of the present invention, the topology-aware attention unit adopts a multi-head attention mechanism, including: 8 independent attention heads, each attention head focusing on different types of medical associations; calculating topological position scores based on node centrality and continuous cohomology importance; dynamically balancing the weights of each head according to the distribution of training data; and generating an attention weight matrix to guide the multimodal feature fusion process.
[0117] Specifically, the system is designed with eight independent attention heads, each focusing on a different type of medical association. For example, head 1 focuses on symptom-symptom associations, head 2 focuses on symptom-test associations, head 3 focuses on test-test associations, head 4 focuses on symptom-image associations, head 5 focuses on test-image associations, head 6 focuses on image-image associations, head 7 focuses on temporal associations, and head 8 focuses on individualized features.
[0118] The system calculates a topological position score based on node centrality and persistent homohomology importance. Node centrality measures the importance of a node in the topology, and the calculation formula is as follows:
[0119] ,
[0120] in: The centrality score for node v; The node to be evaluated; Let be the set of all nodes in the graph; Let be any node in the graph; For the node To the node The number of shortest paths; For the node The total number of shortest paths from the starting point; This indicates that for all nodes in the graph Summation is performed. The importance of persistence cohomology is calculated based on the persistence graph, and the importance score is:
[0121] ,
[0122] in: Let be the importance score of the continuous cohomology of node v; It is a continuous graph, containing multiple points representing topological features; For a point in a continuous graph, it represents a topological feature derived from parameter values. Appearing in parameter value disappear; The contribution weight of node v to this topological feature has a value range of [0,1]. For the persistence of this feature; This indicates summing over all points in the continuous graph.
[0123] The system dynamically balances the weights of each head based on the distribution of training data. The importance of different heads may vary depending on the disease type. For example, for rheumatoid arthritis, symptom-test association (head 2) may be more important; for ankylosing spondylitis, image-image association (head 6) may be more important. The system learns the weight coefficients of each head through a backpropagation algorithm. :
[0124] ,
[0125] in: An attention matrix for multi-head fusion; Let be the attention matrix for the i-th head; Let be the weight coefficient of the i-th head, initially set to 1 / 8, and dynamically adjusted during training to satisfy . ; This indicates a weighted summation of 8 attention heads; "." indicates a scalar multiplication by a matrix.
[0126] Finally, the system generates an attention weight matrix to guide the multimodal feature fusion process. Each element of the attention matrix... Representation of features Features The level of attention is used for weighted fusion of multimodal features.
[0127] In another preferred embodiment of the present invention, the cross-modal fusion unit includes: a modality-specific encoder, which designs a special feature extractor for clinical phenotypes, laboratory tests and imaging information; a shared representation mapper, which maps features of different modalities to a unified topological space; a conflict resolver, which handles data inconsistencies based on confidence-weighted averaging; and an information completer, which handles feature inference in the case of modality missing.
[0128] Specifically, modality-specific encoders are designed with specialized feature extractors for different types of medical data. For clinical phenotypic information, BERT and Word2vec are used for text encoding; for laboratory test information, multilayer perceptrons are used for feature extraction; and for imaging information, CNNs and ResNets are used to extract image features. The output dimension of each modality encoder is uniformly 256 to facilitate subsequent fusion.
[0129] The shared representation mapper maps features from different modalities to a unified topological space. The mapping function is defined as:
[0130] ,
[0131] in: Let be the original feature vector of the i-th modality, with a dimension of 256; This is the weight matrix for the corresponding modality, with dimensions of 256×256; This is the bias vector, with a dimension of 256; The mapped feature vectors have a dimension of 256; "." represents matrix multiplication; "+" represents vector addition. The goal of the mapping is to minimize the representation differences between different modalities while preserving modality-specific information.
[0132] The conflict resolver handles inconsistencies between data from different modalities. For example, a patient may present with joint pain based on clinical symptoms, but imaging studies may reveal no obvious abnormalities. The system resolves conflicts based on a confidence-weighted average.
[0133] ,
[0134] in: The feature vector after conflict resolution has a dimension of 256; Let be the mapping feature vector of the i-th mode; Let be the confidence coefficient of the i-th mode, with a value range of [0,1], and satisfying . ; The expression "+" indicates a weighted summation of all modalities; "." indicates a scalar multiplication by a vector. Confidence coefficients are based on historical statistics and expert knowledge presets, and can also be learned through model training. Typically, laboratory test information has a higher confidence level (0.4-0.5), followed by imaging information (0.3-0.4), and clinical phenotype information has a lower confidence level (0.2-0.3), but this will be dynamically adjusted according to the specific disease type.
[0135] The information completion tool handles modality missing cases. In clinical practice, patients may not complete all examinations, resulting in missing data for certain modalities. The system employs a feature inference method based on graph neural networks, utilizing existing modalities and knowledge graphs to complete the missing information:
[0136] ,
[0137] in: The inferred feature vector for the missing modality; The set of feature vectors for available modes; It is a knowledge graph that contains relationships between medical entities; This is a graph neural network model used to infer missing information based on known information and knowledge structures. This method can reduce the impact of missing data while ensuring diagnostic accuracy.
[0138] In another preferred embodiment of the present invention, the system connects to the hospital information system, the laboratory information system and the medical imaging system to achieve automatic data acquisition; the system can identify early signs of rheumatic diseases 6 months in advance with a diagnostic accuracy of 85%, providing decision support for clinical intervention.
[0139] Specifically, this system connects to the Hospital Information System (HIS), Laboratory Information System (LIS), and Picture Archiving and Communication System (PACS) via standard interfaces to achieve automatic acquisition of patient data. The system supports medical data standards such as HL7 and DICOM, ensuring compatibility with existing medical information systems.
[0140] In practical clinical applications, the system can identify early signs of rheumatic diseases an average of 6 months in advance. In a prospective study involving 500 high-risk individuals, the system achieved an accuracy of 85%, a specificity of 78%, and a sensitivity of 88% in the early diagnosis of rheumatic diseases, significantly outperforming traditional diagnostic methods. Early diagnosis allows patients to receive intervention and treatment as early as possible, reducing the risk of disease progression, minimizing irreversible joint damage, and improving quality of life.
[0141] The system provides decision support for clinical interventions, including risk warnings, disease classification, and treatment recommendations. For high-risk patients, the system recommends personalized plans for further examination and specialist referrals; for diagnosed patients, the system provides treatment recommendations and prognostic assessments based on disease subtype and severity, and with reference to the latest clinical guidelines.
[0142] The system is applicable to various clinical scenarios, including initial screening for rheumatic diseases in primary healthcare institutions, assisted diagnosis by rheumatologists, data analysis in multi-center clinical studies, and long-term health management for high-risk populations. By providing standardized and objective diagnostic support, the system helps reduce diagnostic bias and improve diagnostic consistency, especially in areas with limited medical resources.
[0143] In summary, this invention innovatively introduces algebraic topology theory to construct a complete multimodal deep learning-integrated intelligent early screening system for rheumatic diseases. The system integrates three core technologies: simple complex knowledge representation, continuous homology feature extraction, and spectral topology fusion network. This enables the effective capture and accurate identification of weak early signals of rheumatic diseases, providing an innovative solution for early screening and precise diagnosis of these diseases.
[0144] The embodiments described above are merely illustrative of specific implementations of the present invention, and while the descriptions are detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A multi-modal deep learning fusion based early intelligent screening system for rheumatic diseases, characterized in that, The application relates to a medical diagnosis system based on a simplicial complex knowledge representation structure, which comprises the following modules: An intelligent question-answering module for collecting clinical phenotype information, laboratory examination information and imaging information of a patient; A knowledge graph module in communication connection with the intelligent question-answering module, for receiving the information collected by the intelligent question-answering module, converting the clinical phenotype information, laboratory examination information and imaging information into a simplicial complex knowledge representation structure, and establishing a medical knowledge topological space; An auxiliary diagnosis module in communication connection with the knowledge graph module, based on a topological perception graph convolution network and a self-attention mechanism, for performing multi-scale persistent homology feature extraction on the simplicial complex knowledge representation structure, realizing multi-modal feature fusion based on a spectral topological fusion network, and generating a rheumatic disease risk prediction result; A hierarchical diagnosis module in communication connection with the auxiliary diagnosis module, for receiving the rheumatic disease risk prediction result, classifying the patient according to the risk, and generating a diagnosis suggestion.
2. The multi-modal deep learning fusion based early intelligent screening system for rheumatic diseases as claimed in claim 1, wherein, The intelligent question-answering module structures the clinical phenotype information, laboratory examination information and imaging information by using a unified labeling rule, encodes the input information by using BERT and Word2vec, generates a vector representation, and transmits the vector representation to the knowledge graph module.
3. The multi-modal deep learning fusion based rheumatic disease early intelligent screening system according to claim 1, wherein, The process of converting the clinical phenotype information, laboratory examination information and imaging information into a simplicial complex knowledge representation structure by the knowledge graph module comprises the following steps: Medical entities are represented as 0-simplices, binary relationships between entities are represented as 1-simplices, and complex relationships among multiple entities are represented as high-order simplices; A boundary mapping is constructed to describe the topological relationship between simplices; A nested complex structure is established to capture complex medical association patterns in rheumatic disease diagnosis.
4. The multi-modal deep learning fusion based early intelligent rheumatic disease screening system as claimed in claim 1, wherein, The process of multi-scale persistent homology feature extraction by the auxiliary diagnosis module comprises the following steps: A multi-scale filtering complex sequence is constructed, and the relationship strength is controlled by setting a multi-level threshold; Topological features under different scales are calculated to generate a persistence graph; A persistence graph feature vector is extracted, and the feature importance is quantified by using a persistence entropy; Features with high persistence but weak signals are enhanced to realize the detection of early weak symptoms.
5. The multi-modal deep learning fusion based early intelligent rheumatic disease screening system as claimed in claim 1, wherein, The spectral topological fusion network in the auxiliary diagnosis module comprises the following units: A spectral domain transformation unit for constructing a simplicial complex Laplacian to map medical features to a spectral domain; A multi-spectrum filtering unit for extracting low-frequency, medium-frequency and high-frequency features, which correspond to disease global patterns, main manifestations and individual differences, respectively; A topological perception attention unit for assigning attention weights based on the importance of topological positions; A cross-modal fusion unit for integrating different modal features, solving data conflicts, and generating a comprehensive feature vector.
6. The multi-modal deep learning fusion based early intelligent rheumatic disease screening system as claimed in claim 1, wherein, The hierarchical diagnosis module comprises the following units: A risk assessment unit for dividing patients into a high-risk group and a low-risk group; A disease typing unit for identifying specific rheumatic disease types for high-risk patients; A diagnosis suggestion generation unit for generating clinical diagnosis suggestions and intervention measures based on the risk assessment result and the disease typing result.
7. The multi-modal deep learning fusion based rheumatic disease early intelligent screening system according to claim 4, wherein, The multi-level threshold setting in the multi-scale persistent homology feature extraction process comprises the following steps: Ten gradient threshold values are set from high correlation strength to low correlation strength; A long time window is used for chronic progression features. Short time window is adopted for acute features; Threshold sensitivity is dynamically adjusted based on patient baseline data.
8. The multi-modal deep learning fusion based rheumatic disease early intelligent screening system according to claim 5, wherein, The topology-aware attention unit adopts a multi-head attention mechanism, including: 8 independent attention heads, each focusing on different types of medical associations; Topological position scores are calculated based on node centrality and persistent homology importance; Head weights are dynamically balanced according to training data distribution; An attention weight matrix is generated to guide the multi-modal feature fusion process.
9. The multi-modal deep learning fusion based rheumatic disease early intelligent screening system according to claim 5, wherein, The cross-modal fusion unit includes: Modality-specific encoders, designed with specialized feature extractors for clinical phenotypes, laboratory tests, and imaging information; Shared representation mapper, mapping different modalities to a unified topological space; Conflict resolver, handling data inconsistencies based on confidence-weighted averaging; Information completer, used for feature inference in the case of missing modalities.
10. The multi-modal deep learning fusion based early intelligent screening system for rheumatic diseases of claim 1, wherein, The system connects with hospital information systems, laboratory information systems, and medical imaging systems to achieve automatic data acquisition; The system identifies early signs of rheumatic diseases 6 months in advance, with a diagnosis accuracy rate of 85%, providing decision support for clinical intervention.