Multimodal computing-based early intelligent graded screening system for brain disease

Through the multimodal calculation of early intelligent hierarchical screening system, multimodal data collection, feature generation completion and fusion, and knowledge-based intelligent screening, the problem of inaccurate early screening of neurodegenerative diseases in the existing technology is solved, and high-precision and low-cost personalized screening is achieved, reducing misdiagnosis and misdiagnosis.

WO2025175424A1PCT designated stage Publication Date: 2025-08-28SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Patent Information

Application Number
PCT/CN2024/077582
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-19
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

The existing neurodegenerative disease screening system is difficult to conduct precise early screening, causing patients to miss the best treatment opportunity.

Method used

The early intelligent hierarchical screening system for brain diseases based on multimodal computing is adopted. Through multimodal data acquisition, feature generation completion and fusion, and knowledge-based intelligent screening, the neural network model trained by expert experience and knowledge is used to calculate to obtain accurate screening results.

Benefits of technology

It improves the accuracy and accuracy of early screening of neurodegenerative diseases, reduces screening costs, enhances personalization and adaptability, reduces the risks of misdiagnosis and missed diagnosis, and provides a more comprehensive assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024077582_28082025_PF_FP_ABST
    Figure CN2024077582_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a multimodal computing-based early intelligent graded screening system for a brain disease. The system comprises: a multimodal data acquisition unit, configured to acquire multimodal data of a target patient under a specified screening grade to form a screening data set; a multimodal feature generation, completion, fusion and calculation unit, configured to generate and complete feature data of modal features in the screening data set to obtain complete modal features, and extract pathological features for calculation to obtain a first screening result; a knowledge-based intelligent screening unit, configured to encode the feature data of the modal features in the screening data set into corresponding graph structure data features, use a multimodal association graph optimized by expert knowledge to match the graph structure data features to obtain knowledge-based association features, and obtain a second screening result on the basis of the knowledge-based association features; and an intelligent graded screening unit, configured to carry out weighted calculation on the first screening result and the second screening result to obtain a final screening result. The present disclosure achieves accurate early screening of brain diseases of patients.
Need to check novelty before this filing date? Find Prior Art

Description

Early intelligent grading screening system for brain diseases based on multimodal computing Technical Field

[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and in particular to an early intelligent grading screening system for brain diseases based on multimodal computing. Background Art

[0002] The pathogenesis of neurodegenerative diseases is highly complex and the causes are diverse. The existing neurodegenerative disease screening systems all analyze a type of patient data. This screening system is difficult to accurately screen patients for neurodegenerative diseases early based on a single type of data, causing patients to miss the best treatment opportunities.

[0003] How to accurately screen patients for neurodegenerative diseases at an early stage is an urgent problem that needs to be solved.

[0004] Summary of the Invention

[0005] To address the problem of being unable to accurately perform early screening for neurodegenerative diseases in patients, the embodiments of this specification provide an early intelligent graded screening system for brain diseases based on multimodal computing. The early intelligent graded screening system for brain diseases based on multimodal computing can extract features from the target patient's multimodal data, generate and complete missing modal features, and then fuse these modal features to obtain fused features. Finally, the fused features are input into a neural network model trained based on expert experience knowledge for calculation to obtain the patient's neurodegenerative disease screening results.

[0006] The specific technical solutions of the embodiments of this specification are as follows:

[0007] The embodiment of this specification provides an early intelligent grading screening system for brain diseases based on multimodal computing, including:

[0008] a multimodal data acquisition unit, configured to acquire multimodal data of a target patient at a specified screening level to form a screening data set; the screening level corresponds to the number of modal features in the multimodal data, and the multimodal data includes data of at least one modal feature from a predetermined plurality of modal features;

[0009] a multimodal feature generation, completion, and fusion calculation unit, configured to generate and complete feature data of each modality feature in the screening data set to obtain a complete modality feature, extract pathological features from the complete modality feature for calculation, and obtain a first screening result;

[0010] a knowledge-based intelligent screening unit, configured to encode the feature data of each modal feature in the screening data set into corresponding graph structure data features, match the graph structure data features using a multimodal association graph optimized by expert knowledge to obtain knowledge-based association features, and obtain a second screening result based on the knowledge-based association features;

[0011] The intelligent hierarchical screening unit is used to perform weighted calculation on the first screening result and the second screening result to obtain the final screening result of the target patient.

[0012] Utilizing the intelligent early-stage brain disease graded screening system based on multimodal computing of the embodiment of this specification, the multimodal data acquisition unit first obtains the multimodal data of the target patient at the specified screening level. Compared with the existing technology, the embodiment of this specification does not need to obtain data under all modalities of the target patient for screening, thereby effectively reducing the cost of large-scale screening of neurodegenerative diseases, providing a more personalized screening process, adapting to user needs and various input modes, improving user experience, and improving the accuracy of early screening of brain diseases.

[0013] Then, the multimodal feature generation, completion, and fusion calculation module of the embodiment of this specification generates and completes the feature data of each modal feature in the screening data set, which solves the limitation of the existing technology that accurate screening cannot be achieved when the multimodal data is incomplete, and solves the problem that the modality of the multimodal data of early brain diseases is missing and cannot be accurately screened. At the same time, the multimodal feature generation, completion, and fusion calculation module extracts the specific pathological features of each modality through different early brain disease pathological feature encoders, while ensuring the consistency between different modalities, making full use of the information of the early multimodal data of brain diseases, and using the training strategy of the pathological feature center moment difference distance constraint to reduce the difference between modalities, improve the efficiency of the fusion between neurodegenerative disease modalities, and obtain the first screening result corresponding to the multimodal data.

[0014] In addition, the knowledge-based intelligent screening unit of the embodiment of this specification uses the neural network of the early pathological graph of brain diseases to deeply mine the high-order topological features of multimodal data, effectively distinguishes and expresses the relevant information of the early pathological characteristics of brain diseases, effectively improves the accuracy of early screening of brain diseases, promotes the exploration of the early pathogenesis of brain diseases and the identification of biomarkers, and combines the specific graph matching of early pathological characteristics of brain diseases and the pathological subgraph mining algorithm to accurately extract pathological knowledge rules, thereby improving the pathological knowledge embedding capability of the multimodal auxiliary diagnosis framework for early brain diseases.

[0015] Finally, the intelligent hierarchical screening unit of the embodiment of this specification performs a weighted calculation on the first screening result and the second screening result to obtain the final screening result, so that the final screening result is consistent. The system can evaluate the target patient from different angles, different feature extraction methods, and other aspects. This complementarity can increase the system's ability to detect problems. Using a multimodal feature generation, completion, and fusion calculation unit and a knowledge-based intelligent screening unit for screening can increase the reliability and accuracy of the results. If both methods give similar results, the patient's condition can be determined more confidently. In addition, if there are differences between the screening results of the two units, the system can analyze these differences more deeply and provide a more comprehensive assessment. The two units have different sensitivities and specificities for different types of diseases or abnormalities. By using these two units for screening at the same time, the risk of misdiagnosis and missed diagnosis can be reduced. If the two screening results are inconsistent, the system can remind the doctor to conduct further examinations or assessments to reduce potential errors and risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for describing the embodiments or the prior art.

[0017] FIG1 is a schematic diagram showing the structure of an early intelligent grading screening system for brain diseases based on multimodal computing in an embodiment of this specification;

[0018] FIG2 is a detailed structural diagram of an early intelligent grading screening system for brain diseases based on multimodal computing in an embodiment of this specification;

[0019] FIG3 is a schematic diagram showing the process of modal feature generation, completion and binding alignment in an embodiment of this specification;

[0020] FIG4 is a schematic diagram showing a process of feature-level fusion of multimodal data in an embodiment of this specification;

[0021] FIG5 is a schematic diagram showing the process of constructing a multimodal association graph and regularizing knowledge in an embodiment of this specification;

[0022] FIG6 is a schematic diagram showing the process of constructing a knowledge-based intelligent module and early diagnosis of neurodegenerative diseases in an embodiment of this specification;

[0023] FIG7 is a schematic diagram showing a process for obtaining high-quality and highly available multimodal data according to an embodiment of this specification;

[0024] FIG8 is a schematic diagram showing a process of generating pseudo labels for multimodal data in an embodiment of this specification;

[0025] FIG9 is a schematic diagram showing the flow of the intelligent hierarchical screening system according to an embodiment of this specification;

[0026] FIG10 is a schematic diagram showing the structure of a computer device according to an embodiment of the present specification.

[0027] [Explanation of accompanying symbols]: 1. Multimodal data acquisition unit; 11. Multimodal data preprocessing module; 111. Questionnaire feature data preprocessing submodule; 112. Behavioral feature data preprocessing submodule; 113. EEG and imaging feature data preprocessing submodule; 12. Pseudo-label generation module; 2. Multimodal feature generation, completion and fusion calculation unit; 21. Multimodal feature generation, completion module; 22. Multimodal data feature-level fusion module; 23. Multimodal feature generation, completion and fusion calculation training module; 3. Knowledge-based intelligent screening unit; 31. Multimodal association graph training module; 311. Multimodal association graph construction submodule; 312. Expert knowledge rule extraction submodule; 313. Graph neural network model training submodule; 4. Intelligent hierarchical screening unit; 5. Dynamic incremental learning unit; 1002. Computer equipment; 1004. Processing equipment; 1006. Storage resources; 1008. Driving mechanism; 1010. Input / output module; 1012. Input device; 1014. Output device; 1016. Presentation device; 1018. Graphical user interface; 1020. Network interface; 1022. Communication link; 1024. Communication bus. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.

[0029] FIG1 is a schematic diagram of the results of an intelligent graded screening system for early brain diseases based on multimodal computing according to an embodiment of the present specification, which may include: a multimodal data acquisition unit 1, configured to acquire multimodal data of a target patient at a specified screening level to form a screening data set; the screening level corresponds to the number of modal features in the multimodal data, and the multimodal data includes data of at least one modal feature among a plurality of predetermined modal features;

[0030] A multimodal feature generation, completion and fusion calculation unit 2 is used to generate and complete the feature data of each modality feature in the screening data set to obtain a complete modality feature, extract the pathological feature from the complete modality feature for calculation, and obtain a first screening result;

[0031] The knowledge-based intelligent screening unit 3 is used to encode the feature data of each modal feature in the screening data set into corresponding graph structure data features, match the graph structure data features using the multimodal association graph optimized by expert knowledge to obtain knowledge-based association features, and obtain a second screening result based on the knowledge-based association features;

[0032] The intelligent hierarchical screening unit 4 is configured to perform weighted calculation on the first screening result and the second screening result to obtain a final screening result of the target patient.

[0033] In an embodiment of the present specification, an early intelligent graded screening system for brain diseases based on multimodal computing can be deployed on a processor, and the multimodal data acquisition unit 1, the multimodal feature generation, completion and fusion calculation unit 2, the knowledge-based intelligent screening unit 3 and the intelligent graded screening unit 4 can be functional modules implemented by computer programs. The processor obtains the multimodal data of the target patient, and then obtains the final screening result of the target patient through calculation by the multimodal data acquisition unit 1, the multimodal feature generation, completion and fusion calculation unit 2, the knowledge-based intelligent screening unit 3 and the intelligent graded screening unit 4.

[0034] All steps of the embodiments described in this specification can be implemented by computers and other devices.

[0035] In the embodiments of this specification, the types of brain diseases may include neurodegenerative diseases, and the predetermined multiple modal features may include questionnaire features, behavioral features, EEG features, and imaging features. Questionnaire feature data may be obtained by sending a questionnaire to the target patient, behavioral feature data may include the behavior of the target patient, and the behavioral feature data of the target patient may be obtained by video recording, EEG feature data may be resting-state EEG data, and imaging feature data may be MRI / PET imaging data of the target patient. Positron emission tomography (PET) technology may obtain information about the activity and metabolism of biomolecules, and the combination of brain MRI / PET images may provide a more comprehensive and multi-dimensional disease diagnosis.

[0036] The main contents of the multimodal data acquisition unit 1 include three parts: (1) multimodal data acquisition stage, (2) data preprocessing stage, and (3) data enhancement stage.

[0037] First, (1) multimodal data acquisition stage: the most advanced Peterson, MoCA, and MMSE scales in the mild cognitive impairment screening method were used for evaluation, and behavioral gait videos of the normal group and the patient group were recorded; secondly, the multi-channel acquisition device Neuroelectrics was used to collect EEG data; then, for some target patients, the MR-PET integrated scanning device was used to collect imaging data.

[0038] Secondly, (2) the data preprocessing stage is used to preprocess the collected multimodal data. As shown in FIG2 , the multimodal data acquisition unit 1 further includes: a multimodal data preprocessing module 11, which is used to preprocess the multimodal data and form the screening data set based on the preprocessed multimodal data. Specifically, the multimodal data preprocessing module 11 further includes:

[0039] The questionnaire feature data preprocessing submodule 111 is used to perform data cleaning and data missing processing on the questionnaire feature data, and vectorize the questionnaire feature data after data cleaning and data missing processing to obtain word vectors;

[0040] The behavior feature data preprocessing submodule 112 is used to extract action features from the behavior feature data;

[0041] The EEG and image feature data preprocessing submodule 113 is used to denoise, eliminate artifacts, and standardize the EEG feature data and the image feature data.

[0042] Finally, (3) the data enhancement stage is mainly used in the process of training the multimodal feature generation, completion and fusion calculation unit 2 and the knowledge-based intelligent screening unit. The correlation between the multimodal data of the same patient is used to generate pseudo labels based on the self-supervised learning algorithm. The training steps are explained in the subsequent content of the embodiment of this specification.

[0043] Through the three stages (1) to (3) above, complete multimodal data characterizing the early stages of neurodegenerative diseases are effectively obtained, thereby learning multimodal effective features globally and improving the accuracy of early diagnosis: the early development mechanism of neurodegenerative diseases is complex, the inducing factors are diverse, and the data is prone to missing and inaccurate labels, which makes it difficult for existing methods to learn multimodal effective features related to early neurodegenerative diseases from a global perspective. The present invention is based on multidisciplinary cross-integration and collaboration, fully utilizing the complementary characteristics of features between different modal data, and using multimodal features from multiple perspectives to accurately depict the early pathological mechanism of Alzheimer's disease, thereby facilitating the analysis of the pathogenesis of neurodegenerative diseases and accurate early detection. The present invention combines the correlation between the multimodal data of the subjects and proposes a self-supervised generative learning module to overcome the label missing problem in large-scale community screening, thereby improving the quality and availability of data. By proposing a multidisciplinary cross-integration and multimodal information complementary strategy, the present invention can learn multimodal effective features related to early neurodegenerative diseases from a global perspective, effectively improving the accuracy of early diagnosis of neurodegenerative diseases. Learn more feature data.

[0044] The multimodal feature generation, completion, and fusion calculation unit 2 and the knowledge-based intelligent screening unit 3 are the core algorithm modules for early screening of neurodegenerative diseases. To address the problem of missing modalities in multimodal data, the multimodal feature generation, completion, and fusion calculation unit 2 designs a modal feature generation and completion algorithm based on a collaborative diffusion model to achieve the generation, completion, and alignment of missing modal features, thereby achieving feature-level fusion of multimodal data. Driven by a "data + knowledge" fusion approach, the knowledge-based intelligent screening unit 3 addresses the individual differences in the early manifestations of neurodegenerative diseases and the difficulty in effectively embedding existing expert knowledge into data-driven diagnostic models. It designs an effective multimodal knowledge embedding algorithm to improve the model's universality and robustness.

[0045] Specifically, as shown in Figure 2, to address the problem of missing modalities in multimodal data, a modal feature generation and completion algorithm is designed based on the diffusion model to design a multimodal feature generation and completion module 21 to achieve the generation, completion and alignment of missing modal features; to address the challenge of difficult fusion of highly heterogeneous and heterogeneous multimodal data, a multimodal data feature-level fusion module 22 is designed based on the Transformer model to achieve multimodal feature-level fusion.

[0046] Specifically, the multimodal feature generation, completion and fusion calculation unit 2 further includes:

[0047] a multimodal feature generation and completion module 21 for extracting modal features of the multimodal data in the screening dataset using a variational autoencoder, completing the missing modal features in the screening dataset using a generative artificial intelligence algorithm, and aligning the completed modal features with multiple modal features in the screening dataset using a binding model to obtain complete modal features;

[0048] The multimodal data feature-level fusion module 22 is used to extract specific pathological features based on the complete modal features, where the specific pathological features are valid pathological features in the complete modal features, and to splice the specific pathological features to obtain invariant pathological features. The specific pathological features and the invariant pathological features are input into the classifier through a modal fusion encoder with a cross-modal multi-head attention mechanism for calculation to obtain the first screening result.

[0049] As shown in FIG3 , the main contents of the multimodal feature generation and completion module 21 include: (1) pathological feature extraction, (2) pathological feature generation and completion, and (3) pathological feature binding and alignment.

[0050] (1) Pathological feature extraction: In order to uniformly vectorize the pathological features of neurodegenerative disease data in order to address the challenge of high heterogeneity between different modalities, it is necessary to construct corresponding pathological feature encoders based on the characteristics of different neurodegenerative disease data modalities. First, the pathological feature encoder is used to extract multimodal data features and obtain a vector representation in the embedding space. Secondly, in order to ensure the accurate extraction of effective pathological related features and remove redundant information, it is necessary to jointly train the pathological feature encoders of each modality to achieve a reasonable expression of the relevant pathological information between different modalities of neurodegenerative diseases.

[0051] (2) Pathological feature generation and completion: To address the challenge of modality missing in multimodal data of neurodegenerative diseases, a collaborative diffusion module based on generative artificial intelligence is proposed to complete the missing modal features.

[0052] The first step is to use a pre-trained pathological feature conditional diffusion model based on data from different neurodegenerative disease modalities. This pre-trained pathological feature diffusion model is used to reduce model convergence time, improve training efficiency, and reduce unnecessary computing resource consumption.

[0053] In the second step, starting from Gaussian noise, denoising is performed through the pathological feature diffusion module. At each step of the denoising process, the collaborative diffusion model is used to dynamically coordinate the cooperation between the diffusion modules of different neurodegenerative disease modalities, thereby generating intermediate neurodegenerative disease pathological features. At the same time, during the coordination process of the pathological feature collaborative diffusion model, the dynamic pathological feature diffuser is used to predict the degree of influence of the diffusion module of each single neurodegenerative disease modality on the overall pathological feature generation result in the embedding space, thereby measuring the contribution of each pathological feature diffusion module.

[0054] In the third step, based on the contribution of different pathological feature diffusion modules, the predicted impact degree is used to perform pathological convolution with the neurodegenerative disease pathological feature generation results of each diffusion module and calculate the expectation to obtain the next pathological feature generation result.

[0055] Fourth, in the early stages of the pathological feature denoising process, the module will focus more on constructing coarse-grained information of the pathological features of neurodegenerative diseases, while in the later stages, the module will focus more on constructing fine-grained information of the pathological features of neurodegenerative diseases. This approach can enhance the accuracy of generating pathological features.

[0056] In the fifth step, the collaborative diffusion model coordinates the diffusion modules of different modalities in the early stages of neurodegenerative diseases to obtain the final fully modal pathological signature. Furthermore, the generated pathological signature is constrained by the difference ratio between neurodegenerative disease modalities, improving the reliability of the generated results and reducing the modality gap between different neurodegenerative disease modal signatures.

[0057] The loss function guiding the collaborative diffusion model is as follows:

[0058] Among them, L is the loss function value of the collaborative diffusion model, L rec is the reconstruction loss function, which is L1 loss here; L vgg is the perceptual loss, L dl is the collaborative diffusion loss.

[0059] x and y represent samples of multimodal data, λ l is the weight set, Φ k represents the kth convolutional layer, D represents the sample batch, Represents the original modal features.

[0060] (3) Pathological feature binding alignment: To address the bias caused by label imbalance and modality imbalance in early multimodal data of neurodegenerative diseases on the generation and completion of pathological features, a pathological feature multimodal binding model is used to align the multimodal joint representations of early neurodegenerative diseases.

[0061] The first step is to design different encoders based on the pathological characteristics of different modalities in the early stages of neurodegenerative diseases. Since most of the multimodal data of neurodegenerative diseases are questionnaires, scales, and video gait behavior data, pre-trained large-scale language models for neurodegenerative diseases and large-scale visual models for neurodegenerative diseases are used as encoders for the pathological characteristics of early text data and behavioral data of neurodegenerative diseases.

[0062] In the second step, linear projection is performed on the encoded embedded pathological features to unify the dimensions of the embedded pathological features generated by different pathological feature encoders. At the same time, the embedded pathological features of different modalities of neurodegenerative diseases are projected into a unified pathological feature space to achieve alignment between the pathological features of different modal data in the early stages of neurodegenerative diseases.

[0063] The third step is to align the distribution of neurodegenerative disease pathological features to minimize the differences in pathological features across different modalities. Therefore, based on the principle of maximizing the entropy of neurodegenerative disease pathological features, we maximize the entropy of the joint distribution of neurodegenerative disease pathological features while simultaneously ensuring that the pathological features of different modalities are as close as possible.

[0064] Fourth, integrating the above points, as the amount of collected neurodegenerative disease-related data reaches a certain level, a large-scale neurodegenerative disease pathology model will undergo an emergent phenomenon. This allows the model to more comprehensively explain the relevant content of neurodegenerative disease pathology, allowing different neurodegenerative disease modalities to communicate with each other, identifying connections between neurodegenerative disease pathological features, and performing emergent alignment of neurodegenerative disease data. This strengthens the connection between different neurodegenerative disease modalities and finds deeper intrinsic connections between pathological features. This model can understand the pathological features of different modalities in the early stages of neurodegenerative diseases without requiring a large amount of fully modal early neurodegenerative disease data, and has a high degree of robustness to multimodal joint representation.

[0065] To summarize, the loss function guiding the multimodal binding model is as follows:

[0066] Let’s use image binding as an example to explain: we use a modality pair (I,M), where I represents an image and M is another modality, to learn a single joint embedding, L I,M Denotes the loss function value between the image and the modality in the modality pair. Consider a modality pair (I, M) with aligned observations. Given an image I i and its other mode M i corresponding observations in , encoding them into normalized embeddings: q i =f(I i ) and k i =g(M i ), where f and g are deep networks. τ is a scalar that controls the smoothness of the softmax distribution, and j represents irrelevant observations, also known as "negative examples." Each sample j≠i in the mini-batch is considered a negative example. This loss makes the embedding vector g i and k i They are closer in the joint embedding space, thus aligning I and M.

[0067] As shown in FIG4 , the main contents of the multimodal data feature-level fusion module 22 include:

[0068] First, for the complete modal features representing the early stages of neurodegenerative diseases processed by the multimodal feature generation and completion module 21, unified encoding of the early pathological features of neurodegenerative diseases is required before fusion, so as to further reduce the modal gap between the multimodal highly heterogeneous and heterogeneous data of early neurodegenerative diseases.

[0069] In the first step, to learn the invariant pathological features across modalities of early neurodegenerative diseases, a specificity encoder is used to extract high-level neurodegenerative disease pathological features from the raw pathological features of each modality to represent the specific pathological features of early neurodegenerative diseases, preserving the effective pathological information in the pathological features of each modality. The specificity encoder for each modality is primarily composed of a neurodegenerative disease early pathology self-attention module, a pathology feedforward neural network, and a series of pathology feature normalization operations.

[0070] In the second step, the invariance encoder uses the specific pathological features of early neurodegenerative disease modalities as input to extract invariant pathological features of early neurodegenerative disease modalities. The invariant pathological features are composed of the concatenation of high-level pathological features from all early neurodegenerative disease modalities. The invariance encoder consists of a fully connected layer with pathological features and a dropout layer with pathological features.

[0071] The third step is to use a constrained training strategy based on the neuropathological center moment difference distance to learn modal invariance between various modal pathological features of early neurodegenerative diseases. This strategy fully explores the available multimodal pathological features of early neurodegenerative diseases, alleviates the modality gap problem in cross-modal early neurodegenerative disease, and thus improves the robustness of the multimodal joint representation of neurodegenerative diseases. The loss function based on the neuropathological center moment difference distance is as follows:

[0072] The invariant encoder consists of a fully connected layer, an activation function, and a dropout layer. Its purpose is to incorporate modality-specific features (h a ,h v ,h t ) is mapped into the shared subspace to obtain high-level features (H a ,H v ,H t Then, these three high-level features are concatenated into the modal invariant feature H. The distance constraint based on CMD aims to reduce the high-level features of the modality (H a ,H v ,H t ). Note that CMD is a state-of-the-art distance metric that measures the difference between two feature distributions by matching the order moment difference of the features. By minimizing the loss value L cmd To learn modality invariant features. Where E(H) is the empirical expectation vector of the input sample, C k (H)=E((HE(H))k) is the vector of central moments of all k-order samples of H coordinates, where K represents the order.

[0073] In the fourth step, the early fully modal pathological features of neurodegenerative diseases are uniformly encoded, and the modality-specific pathological features and modality-invariant pathological features are processed by the early multimodal fusion pathological feature encoder of neurodegenerative diseases with a cross-modal multi-head attention mechanism, and finally the corresponding first screening results are output.

[0074] In the fifth step, since early multimodal data on neurodegenerative diseases primarily consists of questionnaires, questionnaires, and behavioral data, a staged model training strategy for early neurodegenerative diseases is proposed. The pathology feature encoders described above are trained using large-scale questionnaires and video gait behavior data. Because the pathology features initially generated from the missing modality for neurodegenerative diseases are difficult to align with those from other modalities, the specificity pathology feature encoder is pre-trained to further align the pathology features of the multimodal data for early neurodegenerative diseases. Pre-training textual pathology features on questionnaires and questionnaires for early neurodegenerative diseases and visual pathology features on video gait behavior data simultaneously learns more pathology feature representations relevant to early neurodegenerative diseases. For visual pathology feature pre-training, the pathology attention module in the visual specificity pathology feature encoder is primarily considered. The pathology attention module of the visual specificity pathology feature encoder is initialized using pre-trained parameters from the BEIT for early neurodegenerative diseases. During textual pathology feature pre-training, the parameters of the visual specificity pathology feature encoder are frozen, and masked language modeling with pathology features is used to optimize the textual specificity pathology feature encoder on textual data from early neurodegenerative diseases.

[0075] In the embodiment of this specification, the multimodal feature generation and completion module 21 and the multimodal data feature-level fusion module 22, as well as the multimodal association graph optimized by expert knowledge in the knowledge-based intelligent screening unit 3, are all pre-trained. Specifically, as shown in FIG2 , the multimodal feature generation and completion and fusion calculation unit further includes:

[0076] The multimodal feature generation, completion and fusion calculation training module 23 is used to train the multimodal feature generation and completion module and the multimodal data feature-level fusion module using the historical multimodal data of multiple patients collected by the multimodal data acquisition unit and the clinical diagnosis results of brain diseases corresponding to the historical multimodal data.

[0077] During the training phase, historical multimodal data of multiple patients and the clinical diagnosis results corresponding to the historical multimodal data are input. Then, the multimodal feature generation and completion module 21 to be trained generates, completes and aligns the multimodal features in the input historical multimodal data according to the above steps to obtain complete modal features. Then, the multimodal data feature-level fusion module 22 to be trained is used to calculate the complete modal features according to the above steps to obtain the prediction results. Then, the loss value is calculated based on the prediction results and the clinical diagnosis results for iterative training.

[0078] Furthermore, as shown in FIG2 , the knowledge-based intelligent screening unit 3 includes:

[0079] The multimodal association graph training module 31 is used to train a graph neural network model using the historical multimodal data and the clinical diagnosis results corresponding to the historical multimodal data, and to constrain the training process of the graph neural network model based on expert knowledge rules to obtain the multimodal association graph after expert knowledge optimization.

[0080] The embodiments of this specification address the challenge of purely data-driven medical artificial intelligence methods being unable to universally learn multimodal representations and provide biological explanations. By constructing a "scale-behavior-EEG-brain imaging" association map, this approach provides effective inference rules and knowledge embedding for a multimodal assisted diagnosis framework. By integrating inference rules and expert knowledge to design an interpretable early diagnosis model for neurodegenerative diseases, this paper constructs an expert system-assisted diagnosis model driven by the "data + knowledge" fusion approach, improving the universality and interpretability of intelligent early diagnosis of neurodegenerative diseases.

[0081] Specifically, as shown in FIG2 , the multimodal association graph training module 31 further includes:

[0082] A multimodal association map construction submodule 311 is used to quantify the correlation between the modalities in the historical multimodal data of multiple patients, mine the common features between the features of the modalities based on the correlation, and construct the association map based on the common features;

[0083] The expert knowledge rule extraction submodule 312 is used to regularize the expert knowledge of diagnosing brain diseases to obtain a knowledge rule base;

[0084] The graph neural network model training submodule 313 is used to perform graph structure learning on different modal features in the association graph, aggregate the information within each modality through multiple graph representation learning and graph convolution layer channels, and then use the knowledge rule base to interact and aggregate the information between different modalities, extract valid graph features, and send them to the classification model to obtain prediction results. The loss value is calculated based on the prediction results and the corresponding clinical diagnosis results, and iterative training is performed based on the loss value to obtain the multimodal association graph after the expert knowledge is optimized.

[0085] Specifically, the main contents of the multimodal association map training module 31 include: (1) construction of the "scale-behavior-EEG-brain image" association map and knowledge rule extraction, and (2) multimodal association map integrating reasoning rules and expert knowledge.

[0086] (1) Construction of “scale-behavior-EEG-brain imaging” association map and knowledge rule extraction: As shown in Figure 5, in order to address the problem of individual differences in the early manifestations of neurodegenerative diseases and the limitation that existing expert knowledge is difficult to effectively embed and integrate into the multimodal auxiliary diagnosis framework, a “scale-behavior-EEG-brain imaging” association map is jointly constructed to obtain effective knowledge reasoning rules for the early manifestations of neurodegenerative diseases, provide effective knowledge embedding for the multimodal auxiliary diagnosis framework, improve the universality and robustness of the model, provide auxiliary support for the graded screening results, and improve the understanding of the early pathogenesis of neurodegenerative diseases and the efficiency of early diagnosis.

[0087] The multi-modal features are used from multiple perspectives to comprehensively characterize the early stages of neurodegenerative diseases, and the pathological mechanisms of early neurodegenerative diseases are accurately portrayed in the form of association maps. The knowledge reasoning rules of the early manifestations of neurodegenerative diseases are effectively obtained, which helps to analyze the pathogenesis of neurodegenerative diseases and accurately detect them in the early stages.

[0088] ① Graph-based multi-scale homogeneous-heterogeneous decoupling of inter-modal complementary features. This approach decomposes the high-order topological features of early-stage multimodal data on neurodegenerative diseases into a peculiar way, enabling efficient separation of complementary information related to various early-stage neurodegenerative diseases provided by multimodal data features. This approach comprehensively mines comprehensive disease representations, including pathological phenotypes, physiological signals, and behavioral characteristics, revealing the inherent correlations among multimodal heterogeneous information. This approach provides effective support for the construction of a correlation graph for early detection of neurodegenerative diseases based on multimodal features.

[0089] ② Construction of a multimodal association map for the early detection of neurodegenerative diseases. Quantify the correlations between modalities, explore the shared features between modal features, and characterize the relationships between the corresponding modal features to construct a multimodal feature association map. Targeting the multimodal features of early neurodegenerative diseases, a multi-level heterogeneous feature fusion module is used to fuse features from different modal data to establish a "scale-behavior-EEG-brain imaging" association map. This will help explore new mechanisms of early neurodegenerative disease and identify biomarkers for the early stages of neurodegenerative diseases.

[0090] ③ Knowledge rule extraction and multimodal analogical reasoning based on association graphs. Based on the establishment of a multimodal association graph for early neurodegenerative diseases, algorithms such as graph matching and subgraph mining are used to extract potential knowledge rules and perform multimodal analogical reasoning. This provides effective knowledge embedding for the multimodal auxiliary diagnosis framework and an auxiliary basis for the subsequent hierarchical screening process, so as to better identify early biomarkers of neurodegenerative diseases and improve the accuracy of early screening. Relational mapping theory points out that the mapping of relational structures is more important than the similarity between targets in the reasoning process. Inspired by this, the relaxation loss is used to make the module pay more attention to the similarity of relational structures:

[0091] Using multimodal analogical reasoning of association graphs, we can directly predict analogy targets without explicitly providing relationships. Specifically, we use multimodal association graphs G and analogical reasoning to predict connections, thereby determining new relationships and rule networks. This task can be formalized as (e h ,e t ):(e q ,? ), where e h 、e t or e q With different modal information, they represent high-level features of neurodegenerative diseases in different modalities. (e h ,e t ) is a given analogy example that characterizes a specific regular relationship between two entities, and then given entity e q , predicting entity pairs similar to rules in the association graph (e q ,? ), obtain new rule pairs or rule networks similar to the given analogy examples through analogy reasoning, and They represent the hidden features corresponding to different modal entities after unified coding, and the high-order features of node features and connection features in the multimodal association graph are uniformly coded to obtain hidden features and calculate their similarity. sim(·) is the cosine similarity. ε and They are two mutually exclusive areas delineated in the features, which are unified in an end-to-end learning manner and adjusted in a timely manner. The features corresponding to the nodes in the graph structure represent entities, and the relationships between them represent relations, that is, [R] in the symbol. Close relations represent close relationships, which refer to the connection features with close connections between nodes. Alienated entities represent alienated entities, which refer to the entity features without close connections between nodes. The relaxation loss consists of pulling closer and pulling away, corresponding to the close relationship and alienated entity items respectively, which can constrain the model's attention to the transfer of relational structures and implicitly realize the structural mapping process. Among them, L relrepresents the Relaxation loss value, |S| is the total number of training sets S, is the hidden feature of the output analogy example, is another hidden feature of the analogy example.

[0092] Furthermore, the multimodal feature generation and completion module 21 is further used to perform multimodal feature generation and completion on the historical multimodal data to obtain historical complete modal features;

[0093] The multimodal association graph training module 31 is further used to construct the multimodal association graph optimized by the expert knowledge according to the historical complete modal features.

[0094] (2) Multimodal association graph integrating inference rules and expert knowledge: as shown in Figure 6:

[0095] ① Fuzzy Intelligent Module Based on Text Information: Case questionnaires and scales have a small text space and are difficult to quantify. In natural language processing, conventional text embedding methods struggle to effectively describe questionnaire and scale information. To effectively represent questionnaire and scale information, this invention proposes a text vectorization method based on fuzzy embedding. Since text language is typically fuzzy, fuzzy sets can encode the ambiguity of language. The encoding process is as follows: For each question in the document, a corresponding fuzzy set and membership function are designed to encode the question. Using the same fuzzy encoding method, the entire text is vectorized and embedded. After the text is fuzzy encoded, a fuzzy inference system must be established. First, a fuzzy rule base is constructed based on expert knowledge. Then, a suitable fuzzy inference mechanism is designed to form input / output fuzzy inference logic, ultimately achieving the expression and digitization of text language knowledge. Regarding the implementation of the fuzzy intelligent module for text information, this invention proposes the following module: Fuzzy rules are combined with neural networks to construct a fuzzy neural network. Each node in the neural network is composed of rules, and the node weight describes the role of the rule in the next decision. Compared with traditional neural networks, fuzzy neural networks feature a network structure designed with fuzzy logic and expert knowledge guidance, as well as interpretable node weights. Therefore, this text fuzzification AI module is a fuzzy neural network module with expert knowledge and interpretability.

[0096] ② Establish an expert knowledge base: Expert knowledge is a crucial component of an expert system. Its function is to store and transform expert knowledge so that computers can apply it for reasoning and prediction. Expert knowledge comes from specialized knowledge found in literature and the practical experience of experts. Effective and rational utilization of expert knowledge can increase the accuracy and interpretability of modules, providing theoretical support for module integration and diagnostic decision-making.

[0097] Collecting an expert knowledge base. Through field surveys and expert interviews, we capture the experts' experience in actual diagnostic processes. For example, we explore how doctors utilize multimodal data to diagnose patients, how they couple these modal data, and how, after determining the severity of a condition, they analyze the cause and provide treatment plans. This collected expert knowledge is organized, classified, and graded. This knowledge is then formalized into a rule base.

[0098] ③ Multi-module fusion mechanism and optimization based on fuzzy rules: Multimodal sub-module fusion is an important component of the multimodal module. A reasonable and effective fusion strategy can make full use of the private features of different modalities and the coupled common features to obtain accurate diagnostic results. The present invention proposes a modal fusion strategy based on expert knowledge, designs coupling rules between different modalities, and thus obtains a fusion strategy for multimodal sub-modules, such as same-level parallel fusion, supervised hierarchical fusion, etc. Module fusion optimization will be achieved through knowledge-based supervised learning. Based on the fuzzy fusion strategy, a fuzzy neural network framework is first constructed; then the expert knowledge base and the diagnostic database are input into the module for training; the key to training is to design a suitable optimization strategy, decompose the complex problem, and gradually optimize and solve it through hierarchical / layered optimization to obtain the parameters of the module.

[0099] ④ Knowledge-based result generation under intelligent diagnosis: To address the problem that traditional diagnostic models only output the condition status but lack analysis and explanation, this invention proposes an expert knowledge-driven diagnostic reasoning system designed to analyze and explain the condition. This diagnostic reasoning module uses a fuzzy diagnosis module to perform fuzzy reasoning on the corresponding condition based on the condition level and probability. It can analyze the diagnosis results, reasons, precautions, treatment plans, and recommendations for the condition. This diagnostic reasoning system can make the diagnostic model results more explainable and understandable, more like a real doctor.

[0100] The knowledge-based intelligent screening unit 3 of the embodiment of this specification solves the problem that expert knowledge is difficult to be effectively used in the early detection of neurodegenerative diseases. Due to the problem of individual differences in the early manifestations of neurodegenerative diseases, the universality of existing auxiliary diagnosis models is poor. There are complex associations and high heterogeneity between multimodal data features, and existing expert knowledge is difficult to effectively embed and integrate into the multimodal auxiliary diagnosis framework. The present invention adopts knowledge rule extraction and multimodal analogy reasoning based on the "scale-behavior-EEG-brain imaging" association map to provide effective knowledge embedding for the multimodal auxiliary diagnosis framework and improve the universality and robustness of the model.

[0101] The knowledge-based intelligent screening unit 3 of the embodiment of this specification also solves the challenge that purely data-driven medical artificial intelligence methods cannot universally learn multimodal representations and provide biological explanations. The present invention constructs a fuzzy neural network based on inference rules and expert knowledge base, including designing effective knowledge fusion strategies, knowledge frameworks and knowledge transformation, establishing knowledge models, constructing inference engines and designing fuzzy rules to achieve knowledge-based result generation under intelligent diagnosis.

[0102] In the embodiment of the present specification, the historical multimodal data includes labeled historical multimodal data and unlabeled historical multimodal data, wherein the labeled historical multimodal data represents historical multimodal data with corresponding clinical diagnosis results, and the unlabeled historical multimodal data represents historical multimodal data without corresponding clinical diagnosis results;

[0103] Specifically, as shown in FIG7 , the multimodal data acquisition specific scheme of the multimodal data acquisition unit 1 implemented in this specification is as follows:

[0104] ① Questionnaire and scale collection: The experiment uses the most comprehensive and professional cognitive impairment assessment scale used in early screening for neurodegenerative diseases to assess cognitive function in all recruited subjects. During the data collection process, standardized questionnaire and scale completion instructions are provided to ensure data accuracy and consistency.

[0105] ②Behavioral Data Collection: Use video equipment to record the subject's behavior during specific tasks or free behavior to ensure consistency and standardization of data collection. By recording and saving this video data, quantitatively evaluate the subject's behavioral characteristics and establish a foundation for behavioral pattern analysis.

[0106] ③ EEG data acquisition: The present invention uses the currently advanced multi-channel acquisition equipment Neuroelectrics to collect EEG brain electrophysiological signals. The device is based on a high common mode rejection ratio design and uses a fully differential front-end amplifier circuit to enhance the overall anti-interference ability of the circuit. A 10-10 electrode placement system is used, and the number of monopolar channels is selected as 64, with a sampling frequency of 500Hz. A EEG stimulation paradigm with visual flicker stimulation as input is designed, which includes three visual stimuli: black and white stimulation, color stimulation, and photo stimulation. Each set of stimulation will appear on the display and last for 2 seconds, with a flicker frequency of 15Hz, and then disappear for 1 second, and repeat. EEG signals are an important physiological signal that can reflect the cognitive function and neural activity of the subjects.

[0107] ④ Image data acquisition: The present invention uses MR-PET integrated scanning scientific research equipment to complete data acquisition. The equipment has both PET and MRI inspection functions, with the industry's highest 2.8mm spatial resolution and 32cm ultra-large axial field of view. At the same time, the equipment is equipped with a 3.0T superconducting magnet and a high-performance gradient coil system with a gradient field strength of 50mT / m and a climbing rate of 220mT / m / ms, combined with a 48-channel radio frequency receiving system to meet clinical and scientific research needs. The collected MRI includes T1-weighted sequence, T2-weighted sequence, and FLAIR sequence images, with a resolution of 1×1×1mm 3 , TR time of 2000ms, TE time of 20ms, and whole-brain 3D imaging were selected. MRI technology can display detailed brain structure and is of great value for the early diagnosis of neurodegenerative diseases. Positron emission tomography (PET) technology can obtain information on biomolecular activity and metabolism. The combination of brain MRI / PET imaging will provide more comprehensive and multidimensional information, helping to accurately assess the early pathological processes of neurodegenerative diseases.

[0108] The specific data preprocessing process of the multimodal data preprocessing module 11 in the embodiment of this specification is as follows:

[0109] ① Questionnaire-scale preprocessing: The questionnaire-scale is converted into a vectorized feature representation for further analysis. First, data cleaning and normalization are performed to remove irrelevant information and outliers, and missing values ​​are filled in. Then, all clinical questionnaires-scales are encoded using word embedding to obtain a vectorized representation.

[0110] ② Behavioral data preprocessing: For video behavioral data, computer vision technology and behavioral analysis algorithms are used for preprocessing to extract the action features of interest. Motion tracking and posture estimation algorithms are then used to calibrate and correct the accuracy of the behavioral data.

[0111] ③ EEG data preprocessing: The acquired signals are first bandpass filtered from 0.05 to 100 Hz and notched at 50 Hz. The EEG signals from 200 ms before to 800 ms after the stimulus are then separated as the visual evoked potential (VEP) signal of interest. Finally, the acquired data undergoes post-processing, including noise removal, filtering, and artifact removal. This filtering process helps remove irrelevant frequency components and noise interference, facilitating further analysis and modeling.

[0112] ④ Image Data Preprocessing: MRI / PET data preprocessing includes skull removal, outlier and artifact removal, motion correction, spatial normalization, and noise reduction. First, the structural brain images are segmented to remove the skull and ventricles. Then, motion correction is performed using a head motion correction algorithm to correct for artifacts caused by subject movement. Spatial normalization is used to eliminate individual and scanning parameter differences. Finally, Gaussian denoising is performed on the images to address the effects of noise in MRI caused by varying density and signal intensity of gray matter, white matter, and cerebrospinal fluid.

[0113] During the data collection process, imaging data collection can be time-consuming and expensive, resulting in some subjects not receiving data from all modalities. Due to negligence by data collectors and omissions by physicians, a small number of subjects lack complete and accurate assessment reports and corresponding labels. Furthermore, inconsistent reports from subjects can hinder professional physicians' ability to accurately classify patients. Consequently, the raw data suffers from issues such as missing modalities and lacking labels. To address this issue, unlabeled data is utilized to improve the performance of supervised learning models.

[0114] Specifically, as shown in FIG2 , the multimodal data acquisition unit 1 further includes:

[0115] a pseudo-label generation module 12 for generating pseudo-labels for the patient's unlabeled historical multimodal data based on the correlation between the labeled historical multimodal data of the same patient and a trained self-supervised learning model, and labeling the unlabeled historical multimodal data with the pseudo-labels;

[0116] The multimodal feature generation, completion and fusion calculation training module 23 is further used to train the multimodal feature generation and completion module and the multimodal data feature-level fusion module using the clinical diagnosis results corresponding to the labeled historical multimodal data and the pseudo labels corresponding to the unlabeled historical multimodal data;

[0117] The multimodal association graph training module 31 is further used to train the graph neural network model based on the clinical diagnosis results corresponding to the labeled historical multimodal data and the pseudo labels corresponding to the unlabeled historical multimodal data.

[0118] The process of generating pseudo labels by the pseudo label generating module 12 can be shown in FIG8 :

[0119] The first step is to combine the correlations between multimodal data from the same subject and generate pseudo labels using a self-supervised learning module. The self-supervised learning algorithm includes a self-supervised generator and an adaptive guidance network. It uses labeled data to learn the sample distribution, making rational use of low-quality unlabeled modality data. Labeled data is directly optimized based on Ls. Unlabeled data is divided into reliable and unreliable parts based on the prediction results, and optimized based on Lu and Lc respectively, as shown in the following formula:

[0120] The optimization is for self-supervised pseudo-label generation. The goal is to generate pseudo-labels for large amounts of unlabeled data. The three formulas, Ls, Lu, and Lc, are loss functions. Ls and Lu represent the supervised and unsupervised losses applied to labeled and unlabeled samples, respectively. Lc is the contrastive loss that fully utilizes unreliable pseudo-labels. λu and λc are weights.

[0121] Ls and Lu are both cross entropy (CE) losses. The model consists of a CNN-based encoder h, a decoder f with a segmentation head, and a representation head g. l and y l is labeled data, belonging to the labeled set Dl, x u and y u It is unlabeled data, belonging to the unlabeled set Du, represents the generated pseudo-label, Denotes the cross entropy loss. In each training step, B labeled data B are sampled equally l and B unlabeled data B u θ represents the weight. For Lc, where M is the total number of feature anchors, C represents the total number of categories, z ci Represents the representation of the i-th anchor point of class c. Each anchor point image is followed by a positive sample and N negative samples, which are represented as follows: and This is the output of the representation head. <*, *> is the cosine similarity between two different features, which is limited to between -1 and 1, so the parameter τ is needed to control it.

[0122] The second step is to use the generated pseudo labels to label the unlabeled multimodal brain data: a self-supervised algorithm is used to learn the correlation features between multimodal data, generate pseudo labels, and label the missing label data to expand the training dataset. Labels refer to information used to represent the state or category of neurodegenerative diseases. These labels can be used to indicate the specific disease stage of the patient, the type of disease, or other relevant diagnostic information. For neurodegenerative diseases such as Alzheimer's disease (AD), labels can include different categories or states, for example: Normal Control (NC): This label applies to individuals who have no signs or symptoms of neurodegenerative diseases.

[0123] Early Mild Cognitive Impairment (EMCI): This label applies to individuals who experience early, mild decline in memory and cognitive function, which may be an early sign of a neurodegenerative disease.

[0124] Late Mild Cognitive Impairment (LMCI): This label applies to individuals who experience mild cognitive decline in the intermediate term, which may be a sign of progression of a neurodegenerative disease.

[0125] Alzheimer's Disease (AD): This label applies to individuals diagnosed with Alzheimer's disease, with significant cognitive decline and neurodegenerative changes.

[0126] These labels can be determined based on a doctor's diagnosis, clinical assessment, neuroimaging, or other relevant test results. By associating multimodal data with corresponding labels, they can be used to train and evaluate models to diagnose and predict neurodegenerative disease states.

[0127] The third step is to combine pseudo-labels with existing annotated data for supervised training. The model building and training process for this stage will be detailed later. By combining supervised and self-supervised learning algorithms, more accurate and reliable models can be built to assist in the early diagnosis of neurodegenerative diseases.

[0128] In supervised training, the model is trained using an existing labeled dataset to learn the relationship between the input data and the corresponding labels. In the embodiments of this specification, supervised training can produce the following results:

[0129] Trained model: Through supervised training, an optimized model can be obtained, which can map the input data to the corresponding labels or categories.

[0130] Model parameters and weights: During supervised training, the model’s parameters and weights are adjusted and optimized to minimize the error between the input data and the labels.

[0131] Pseudo-labels: During the training process, the trained model can be used to make predictions on unlabeled data, and these predictions are used as pseudo-labels. These pseudo-labels can be used to support the training of related models in subsequent steps.

[0132] Supervised training can obtain a trained model, model parameters and weights, and possible pseudo labels. These results can be used to support the training and prediction tasks of related models in subsequent steps.

[0133] It can be understood that the data collection, preprocessing, and pseudo-label generation methods in multimodal data collection solve the following problems:

[0134] Technical Problem 1: Missing Category Labels. During the data collection phase, it is impossible to guarantee that all subjects will receive an accurate diagnosis from a professional physician, and the original data often has the problem of missing labels. Existing methods to solve missing labels often use unsupervised methods, which cannot fully mine the effective information of multimodal data. The self-supervised pseudo-label generation module proposed in the present invention can utilize unreliable samples to make reasonable use of low-quality unlabeled modal data, and ultimately obtain accurate pseudo-labels based on the shared features between multimodal data.

[0135] Technical Problem 2: Imbalance between labels and modalities. Due to psychological and experimental conditions, it is difficult to collect the same number of patient groups, and the number of healthy control groups is often far greater than that of the patient group, resulting in label imbalance problems. In addition, due to factors such as the acquisition time and equipment cost of different modalities, there will be more questionnaires-scales, behavioral data and EEG data in the actual acquisition process, while there will be relatively few high-cost and time-consuming brain imaging data. The use of data with unbalanced categories will lead to model bias and seriously affect the generalization of the model. Modal imbalance can easily cause the model to be biased towards the dominant modality. The present invention considers the differences in gradient propagation from different modalities and defines quantitative indicators to describe the difference ratio between modalities. According to the adaptive adjustment strategy of dynamically adjusting the coefficients of different modalities, the model gradually converges to the global optimal solution.

[0136] In the embodiments of this specification, to address the high cost, low efficiency, and large error of existing early screening methods for neurodegenerative diseases, an intelligent hierarchical screening system is developed and optimized online through dynamic incremental learning, achieving low-cost, high-efficiency, and high-precision screening for patients with early neurodegenerative diseases. The main contents include the following: (1) deployment of the intelligent hierarchical screening system, and (2) optimization of the intelligent hierarchical screening system through dynamic incremental learning.

[0137] (1) Deployment of intelligent hierarchical screening system:

[0138] Step 1: Integrate and deploy the multimodal data acquisition unit (1), the multimodal feature generation, completion, and fusion calculation unit (2), the knowledge-based intelligent screening unit (3), and the intelligent tiered screening unit (4) to build a personalized adaptive detection framework that is compatible with and capable of processing multiple input modalities. This framework can quickly and accurately predict the early onset probability of neurodegenerative diseases using data from any modality, providing technical support for tiered screening.

[0139] The second step is to develop an intelligent hierarchical screening system based on the personalized adaptive detection framework. By optimizing the hierarchical screening process, the cost of early screening for neurodegenerative diseases can be reduced while improving screening efficiency. The hierarchical screening process is optimized by comprehensively considering the cost and difficulty of collecting multimodal data and the psychological acceptance of the subjects. By performing progressive hierarchical screening on the subjects, the screening results of the subjects are finally obtained. The intelligent hierarchical screening system can be optimized online through dynamic incremental learning. It has good robustness and flexibility and can adapt to the data requirements of different scenarios. With the help of the intelligent hierarchical screening system, clinicians can quickly and accurately identify patients with early-stage neurodegenerative diseases, improving diagnostic efficiency and accuracy.

[0140] As shown in FIG9 , the intelligent hierarchical screening system may include primary screening, secondary screening, and tertiary screening.

[0141] ① Level 1 screening

[0142] Primary screening utilizes bimodal data and is primarily targeted at large-scale social groups. Considering the cost and difficulty of data collection in large-scale social screening, as well as the psychological receptiveness of subjects, it is recommended that questionnaires, scales, and behavioral data be prioritized for primary screening. These are also commonly used auxiliary diagnostic methods in clinical practice. Patients with early-stage neurodegenerative diseases experience symptoms such as memory loss, decreased language expression and comprehension, and decline in spatial cognition. These changes can be captured using questionnaires and scales. Behavioral data reflects a subject's motor skills, coordination, and self-care abilities, and can capture abnormal performance during specific tasks or free movement. Combined analysis of questionnaires, scales, and behavioral data can improve diagnostic accuracy and effectiveness. By integrating questionnaires, scales, and behavioral data into a personalized adaptive testing framework, the intelligent tiered screening system provides reliable screening results, helping physicians to preliminarily assess a subject's disease risk and identify individuals at high risk for early-stage neurodegenerative diseases.

[0143] ② Secondary screening

[0144] Secondary screening uses trimodal data and is primarily targeted at individuals at high risk for early-stage neurodegenerative diseases. Given the lower acquisition cost and complexity of EEG data compared to imaging data, it is recommended that EEG data be added to tertiary screening. Questionnaires, scales, and behavioral data capture the outward manifestations of early neurodegenerative diseases and are susceptible to interference from other conditions. For example, neurological conditions such as dementia and Parkinson's disease can mimic early symptoms of neurodegenerative diseases, making screening using questionnaires, scales, and behavioral data more challenging. In contrast, EEG data provides a more direct and physiologically consistent measure for diagnosing early-stage neurodegenerative diseases. EEG data records the spatiotemporal characteristics of brain activity, helping to analyze discharge activity and synchronization across different regions, ultimately understanding the subject's cognitive status. With the addition of EEG data, the intelligent tiered screening system will provide more accurate screening results, allowing clinicians to further assess the subject's risk and eliminate a large number of false positives.

[0145] ③Three-level screening

[0146] The tertiary screening uses full-modality data and is primarily targeted at patients suspected of early-stage neurodegenerative diseases. Considering the high cost of imaging data acquisition, the need for specialized equipment and technical support, and the need for specialized doctors or technicians to operate and interpret it, it is recommended that imaging data be added to the tertiary screening. Structural and functional disorders of the brain are the root causes of early-stage neurodegenerative diseases. Although EEG data can provide the spatiotemporal characteristics of brain activity, their observation scale is large and cannot provide more detailed information. In contrast, imaging data provides a large amount of fine-grained information, allowing for observation of details of brain structure and function, as well as interactions between different regions. Imaging data can help better understand the pathogenesis of neurodegenerative diseases. After adding imaging data, the intelligent tiered screening system provides the final screening results, helping clinicians make accurate diagnoses for patients suspected of early-stage neurodegenerative diseases.

[0147] (2) Dynamic incremental learning:

[0148] According to one embodiment of the present specification, as shown in FIG2 , the system further includes a dynamic incremental learning unit 5 for optimizing the early intelligent graded screening system for brain diseases based on multimodal computing using the final screening results of the target patient and the clinical diagnosis results of the target patient.

[0149] In the process of dynamic incremental learning of the embodiments of the specification, all the neurodegenerative disease early diagnosis related models involved in the previous steps are optimized using new data based on the actual diagnosis results of the doctor. According to the doctor's diagnosis results, new data samples can be introduced to optimize the model. This means that the model will use these new data samples to update weights and parameters to improve its performance and accuracy. By comparing with the doctor's diagnosis results, the performance and error of the model can be evaluated. Based on these comparison results, the model can be optimized accordingly. This includes adjusting the parameters of the model, improving the feature extraction method, or adjusting the structure of the classifier to make the model more accurate and reliable.

[0150] FIG10 is a schematic diagram of the structure of a computer device according to an embodiment of the present specification. The apparatus in the embodiment of the present specification may be a computer device according to the embodiment, executing the method according to the embodiment of the present specification. Computer device 1002 may include one or more processing devices 1004, such as one or more central processing units (CPUs), each of which may implement one or more hardware threads. Computer device 1002 may also include any storage resources 1006 for storing any type of information, such as code, settings, data, etc. For example, and without limitation, storage resources 1006 may include any one or more combinations of the following: any type of RAM, any type of ROM, flash memory devices, hard disks, optical disks, etc. More generally, any storage resource may use any technology to store information. Furthermore, any storage resource may provide volatile or non-volatile retention of information. Furthermore, any storage resource may represent a fixed or removable component of computer device 1002. In one embodiment, when processing device 1004 executes associated instructions stored in any storage resource or combination of storage resources, computer device 1002 may perform any operation of the associated instructions. The computer device 1002 also includes one or more drive mechanisms 1008 for interacting with any storage resources, such as a hard disk drive mechanism, an optical disk drive mechanism, and the like.

[0151] The computer device 1002 may also include an input / output module 1010 (I / O) for receiving various inputs (via input devices 1012) and for providing various outputs (via output devices 1014). A specific output mechanism may include a presentation device 1016 and an associated graphical user interface (GUI) 1018. In other embodiments, the input / output module 1010 (I / O), input devices 1012, and output devices 1014 may not be included, and the computer device 1002 may simply be a computer device in a network. The computer device 1002 may also include one or more network interfaces 1020 for exchanging data with other devices via one or more communication links 1022. One or more communication buses 1024 couple the components described above together.

[0152] The communication link 1022 may be implemented in any manner, for example, via a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 1022 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.

[0153] The embodiments of this specification also provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above method when executed by a processor.

[0154] The embodiments of this specification also provide a computer-readable instruction, wherein when a processor executes the instruction, the program therein causes the processor to execute the above method.

[0155] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

Claims

1. An intelligent early-stage brain disease screening system based on multimodal computing, characterized by: The system comprises: a multimodal data acquisition unit, configured to acquire multimodal data of a target patient at a specified screening level to form a screening data set; the screening level corresponds to the number of modal features in the multimodal data, and the multimodal data includes data of at least one modal feature from a predetermined plurality of modal features; a multimodal feature generation, completion, and fusion calculation unit, configured to generate and complete feature data of each modality feature in the screening data set to obtain a complete modality feature, extract pathological features from the complete modality feature for calculation, and obtain a first screening result; a knowledge-based intelligent screening unit, configured to encode the feature data of each modal feature in the screening data set into corresponding graph structure data features, match the graph structure data features using a multimodal association graph optimized by expert knowledge to obtain knowledge-based association features, and obtain a second screening result based on the knowledge-based association features; The intelligent hierarchical screening unit is used to perform weighted calculation on the first screening result and the second screening result to obtain the final screening result of the target patient.

2. The system according to claim 1, wherein: The multimodal feature generation, completion and fusion calculation unit further includes: a multimodal feature generation and completion module, configured to extract modal features of the multimodal data in the screening dataset using a variational autoencoder, complete the missing modal features in the screening dataset using a generative artificial intelligence algorithm, and align the completed modal features with multiple modal features in the screening dataset using a binding model to obtain complete modal features; A multimodal data feature-level fusion module is used to extract specific pathological features based on the complete modal features, where the specific pathological features are valid pathological features in the complete modal features. The specific pathological features are spliced ​​to obtain invariant pathological features, and the specific pathological features and invariant pathological features are input into a classifier through a modal fusion encoder with a cross-modal multi-head attention mechanism for calculation to obtain the first screening result.

3. The system according to claim 2, characterized in that The multimodal feature generation, completion and fusion calculation unit further includes: A multimodal feature generation, completion and fusion calculation training module is used to train the multimodal feature generation and completion module and the multimodal data feature-level fusion module using the historical multimodal data of multiple patients collected by the multimodal data acquisition unit and the clinical diagnosis results of brain diseases corresponding to the historical multimodal data.

4. The system according to claim 3, characterized in that The knowledgeable intelligent screening unit includes: The multimodal association graph training module is used to train a graph neural network model using the historical multimodal data and the clinical diagnosis results corresponding to the historical multimodal data, and to constrain the training process of the graph neural network model based on expert knowledge rules to obtain the multimodal association graph after expert knowledge optimization.

5. The system according to claim 4, characterized in that The multimodal association graph training module further includes: A multimodal association map construction submodule is used to quantify the correlation between each modality in the historical multimodal data of multiple patients, mine the common features between the features of each modality based on the correlation, and construct the association map based on the common features; The expert knowledge rule extraction submodule is used to regularize the expert knowledge of diagnosing brain diseases and obtain a knowledge rule base; The graph neural network model training submodule is used to perform graph structure learning on different modal features in the association graph, aggregate the information within each modality through multiple graph representation learning and graph convolution layer channels, and then use the knowledge rule base to interact and aggregate information between different modalities, extract effective graph features, and send them to the classification model to obtain prediction results. The loss value is calculated based on the prediction results and the corresponding clinical diagnosis results, and iterative training is performed based on the loss value to obtain the multimodal association graph optimized by the expert knowledge.

6. The system according to claim 4, characterized in that The multimodal feature generation and completion module is further used to perform multimodal feature generation and completion on the historical multimodal data to obtain historical complete modal features; The multimodal association graph training module is further used to construct the multimodal association graph optimized by expert knowledge based on the historical complete modal features.

7. The system according to claim 4, wherein: The historical multimodal data includes labeled historical multimodal data and unlabeled historical multimodal data, wherein the labeled historical multimodal data represents historical multimodal data with corresponding clinical diagnosis results, and the unlabeled historical multimodal data represents historical multimodal data without corresponding clinical diagnosis results; The multimodal data acquisition unit further comprises: a pseudo-label generation module, configured to generate pseudo-labels for the patient's unlabeled historical multimodal data based on the correlation between the labeled historical multimodal data of the same patient and based on a trained self-supervised learning model, and label the unlabeled historical multimodal data using the pseudo-labels; The multimodal feature generation, completion and fusion calculation training module is further used to train the multimodal feature generation and completion module and the multimodal data feature-level fusion module using the clinical diagnosis results corresponding to the labeled historical multimodal data and the pseudo labels corresponding to the unlabeled historical multimodal data; The multimodal association graph training module is further used to train the graph neural network model based on the clinical diagnosis results corresponding to the labeled historical multimodal data and the pseudo labels corresponding to the unlabeled historical multimodal data.

8. The system according to claim 1, wherein: The system also includes a dynamic incremental learning unit for optimizing the early intelligent graded screening system for brain diseases based on multimodal computing by utilizing the final screening results of the target patient and the clinical diagnosis results of the target patient.

9. The system according to claim 1, wherein: The multimodal data acquisition unit further comprises: The multimodal data preprocessing module is used to preprocess the multimodal data and form the screening data set based on the preprocessed multimodal data.

10. The system according to claim 9, characterized in that The multiple modality features include questionnaire features, behavioral features, EEG features, and imaging features; The multimodal data preprocessing module further includes: The questionnaire feature data preprocessing submodule is used to perform data cleaning and data missing processing on the questionnaire feature data, and vectorize the questionnaire feature data after data cleaning and data missing processing to obtain word vectors; A behavior feature data preprocessing submodule, used to extract action features from the behavior feature data; The EEG and image feature data preprocessing submodule is used to denoise, eliminate artifacts, and standardize the EEG feature data and the image feature data.

Citation Information

Patent Citations

  • Multi-modal disease auxiliary reasoning system and method, terminal and storage medium

    CN117012370A

  • Auxiliary diagnosis system for screening and monitoring breast cancer images

    CN117373653A

  • Adaptive pattern recognition for psychosis risk modelling

    US20160192889A1

Cited By

  • Criminal investigation case three-dimensional scene reconstruction and investigation deduction virtual reality platform

    CN120931841A

  • A virtual reality platform for 3D scene reconstruction and investigation simulation in criminal cases

    CN120931841B

  • Physical examination management platform

    CN121034636A

  • Metasystem for medical multi-modal data knowledge graph development

    CN121144531A

  • Multi-modal data integrated analysis system for screening children with autism

    CN121256711A