Neuropsychiatric disease early-stage auxiliary diagnosis method and system based on brain-like multi-mode large model
By constructing a brain-like multimodal hybrid expert diagnostic model and combining knowledge distillation and compression techniques, the problems of low efficiency and low accuracy in traditional neuropsychiatric disease diagnosis have been solved, achieving efficient and accurate early diagnosis of neuropsychiatric diseases.
Patent Information
- Application Number
- CN202511557806.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-02-10
AI Technical Summary
Traditional diagnostic methods for neuropsychiatric diseases are inefficient and inaccurate, failing to meet the need for rapid and precise diagnosis. Multimodal model-assisted diagnosis has low accuracy and insufficient generalization ability.
We construct a large-scale expert diagnostic model based on brain-like multimodal hybrid technology. We extract and fuse features through brain-like mechanisms, combine knowledge distillation and compression techniques to build a lightweight brain-like diagnostic model, and use reinforcement learning and incremental learning to optimize the model, thereby achieving efficient fusion and diagnosis of multimodal data.
It improves the accuracy of early diagnosis of neuropsychiatric diseases, enhances the diagnostic logic and interpretability of the model, reduces computational complexity, adapts to individual differences and disease evolution, and has efficient and accurate diagnostic capabilities.
Smart Images

Figure CN121506440A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal assisted diagnostic technology, and in particular to a method and system for early assisted diagnosis of neuropsychiatric diseases based on a brain-like multimodal large model. Background Technology
[0002] With the rapid development of modern society and the accelerated pace of life, the incidence of neuropsychiatric disorders continues to rise globally. According to the World Health Organization (WHO), neuropsychiatric disorders such as depression, anxiety disorders, and schizophrenia have become a significant part of the global disease burden, and it is projected that by 2030, depression will become the second leading cause of disability worldwide. Neuropsychiatric disorders not only severely impact the physical and mental health of patients but also impose a huge economic and emotional burden on families and society, significantly reducing patients' quality of life and increasing the demand for medical resources. Therefore, early and accurate diagnosis is crucial for effective intervention and treatment.
[0003] However, traditional diagnostic methods for neuropsychiatric disorders have significant limitations in real-world clinical practice. Some diagnoses often rely on the physician's clinical experience and single medical examination methods, such as medical history taking, scale assessments, and neuroimaging. While these methods can provide some diagnostic evidence, they often suffer from low efficiency and accuracy when processing complex patient data. For example, medical history taking and scale assessments depend on patient self-reporting and subjective judgment, making them susceptible to the influence of the patient's emotions and expressive abilities. While neuroimaging can provide structured data, it is not always sufficiently sensitive in the early diagnosis of functional disorders. Furthermore, traditional methods often fall short when faced with large amounts of multidimensional data, failing to meet the clinical need for rapid and accurate diagnosis.
[0004] To address these shortcomings, the medical community is extensively exploring new diagnostic pathways that combine multimodal data with artificial intelligence. Multimodal data encompasses not only neuroimaging data such as electroencephalograms (EEGs), magnetic resonance imaging (MRI), and functional magnetic resonance imaging (fMRI), but also multiple dimensions including behavioral performance, voice, language, and genetic information. This data can reflect a patient's physiological and psychological state from multiple perspectives, providing doctors with more comprehensive and detailed diagnostic evidence. For example, audio data can effectively reveal emotional and psychological fluctuations in a patient's speech, while EEG signals reflect dynamic changes in brain function. By integrating these heterogeneous data, early signs and subtle symptoms of neuropsychiatric disorders can be captured more accurately.
[0005] In recent years, artificial intelligence technology has made groundbreaking progress, especially the widespread application of large-scale multimodal models and brain-inspired multimodal diagnostic models based on Mixture of Experts (MoE), which are bringing revolutionary changes to the auxiliary diagnosis of neuropsychiatric diseases. Large-scale multimodal models have powerful cross-modal learning capabilities, and can automatically mine deep correlation features between different modalities from massive heterogeneous data. However, the accuracy of auxiliary diagnosis is relatively low, and the model's generalization ability and adaptability are insufficient. Summary of the Invention
[0006] To address the issue of low accuracy in multimodal model-assisted diagnosis, this invention provides an early auxiliary diagnostic method and system for neuropsychiatric diseases based on a brain-like multimodal large model. By processing and fusing multimodal data to obtain a fused feature vector, multidimensional information can be readily obtained. Subsequently, a brain-like multimodal hybrid expert diagnostic large model is constructed and processed to obtain a lightweight brain-like diagnostic model for diagnosing neuropsychiatric diseases, thereby improving the accuracy of multimodal model-assisted diagnosis.
[0007] To achieve the above objectives, the technical solution of the present invention is as follows:
[0008] The first aspect of this invention proposes an early auxiliary diagnostic method for neuropsychiatric diseases based on a brain-like multimodal large model, comprising:
[0009] Step 1: Collect and preprocess the patient's clinical history data to obtain multimodal data, laying the foundation for subsequent analysis;
[0010] Step 2: Use a brain-like mechanism to extract and fuse features from the preprocessed data to obtain a fused feature vector;
[0011] Step 3: Construct a large-scale brain-like multimodal hybrid expert diagnostic model, and process it using knowledge distillation and compression to obtain a lightweight brain-like diagnostic model, which facilitates the reduction of computational redundancy and improves diagnostic accuracy.
[0012] Step 4: Input the fused feature vector into the lightweight brain-like diagnostic model to obtain the predicted diagnostic results;
[0013] Step 5: Retrain the lightweight brain-like diagnostic model based on the predicted diagnostic results to obtain the optimal lightweight brain-like diagnostic model, which will help improve the diagnostic accuracy.
[0014] Step Six: After preprocessing the patient's clinical data, feature extraction and fusion are performed using a brain-like mechanism to obtain the target fusion feature vector. The target fusion feature vector is then input into the optimal lightweight brain-like diagnostic model to obtain the final diagnostic result.
[0015] Furthermore, step two specifically includes:
[0016] Preliminary feature extraction is performed on the preprocessed data to obtain preliminary feature vectors for multiple modalities; wherein the multiple modalities include at least two modalities from image data, audio data, or text data, to facilitate the acquisition of multimodal data;
[0017] The preliminary feature vectors of multiple modalities are input into the leaked integral firing neuron model to obtain the time step feature vectors of the corresponding modalities. All time step feature vectors are directly concatenated in chronological order to obtain the historical global temporal feature vector of each modality, which facilitates cross-time step information fusion and captures long-range temporal dependencies in the data.
[0018] The historical global temporal feature vector of each modality is input into the multi-channel attention module to obtain the weighted features of each modality, which helps to enhance the model's attention to local features;
[0019] By fusing the weighted features of all modalities, a fused feature is obtained, which facilitates the alignment, weighting, and fusion of deep semantic features.
[0020] Furthermore, the multi-channel attention module includes a main branch and a sub-branch; the main branch includes an interpolation sub-module and a scaling sub-module; and the sub-branch includes a linear transformation sub-module and a Softmax function.
[0021] The interpolation submodule is used to adjust the size of the input historical global temporal feature vector;
[0022] The linear transformation submodule is used to linearly transform the output of the interpolation submodule;
[0023] The Softmax function is used to process the output of the linearly changing submodule;
[0024] The scaling submodule is used to scale based on the outputs of the Softmax function and the interpolation submodule to obtain weighted features.
[0025] Furthermore, the brain-like multimodal hybrid expert diagnostic big model includes a knowledge graph, a main routing module, multiple expert sub-networks, a MoE output fusion module, and a diagnostic prediction layer;
[0026] The knowledge graph is used to represent the relationships between diseases, symptoms, and examination methods;
[0027] The main routing module is used to initially determine the major disease types based on input features and knowledge graphs, and select the corresponding expert subnetworks based on the major disease types; wherein, the main routing module includes a spiking neural network;
[0028] Each of the expert subnetworks is used to determine the specific disease type and the weights of the expert subnetwork based on the input features; wherein each of the expert subnetworks includes a feedforward neural network and a gating network;
[0029] The MoE output fusion module is used to obtain hybrid features based on the output of the expert subnetwork;
[0030] The diagnostic prediction layer is used to process the mixed features to obtain the prediction result.
[0031] Furthermore, the MoE output fusion module is represented by the following formula:
[0032]
[0033] Among them, f MoE For mixed features, N is the number of expert subnetworks involved in the computation. Let e be the weight of the j-th expert subnetwork. j This is the output of the j-th expert subnetwork;
[0034] The diagnostic prediction layer is represented by the following formula:
[0035] y=σ(W o f tr +b o )
[0036] f tr =ψ(f MoE )
[0037] Where y is the prediction result, σ is the activation function, and W o and b o These are the weights and biases, f, respectively. tr For encoding features, ψ(·)Transformer encoder.
[0038] Furthermore, the process of using knowledge distillation and compression to process the large-scale brain-like multimodal hybrid expert diagnostic model yields a lightweight brain-like diagnostic model, specifically including:
[0039] Knowledge distillation was performed on a large-scale brain-inspired multimodal hybrid expert diagnostic model to obtain a lightweight clinical brain-inspired diagnostic model.
[0040] During training, a quantization function is introduced to map the weights in the lightweight clinical brain-like diagnostic model from 32-bit floating-point numbers to 8-bit integer representations.
[0041] Singular value decomposition was performed on key layers in a lightweight clinical brain-like diagnostic model, and low-rank matrices were selected to be retained.
[0042] By jointly optimizing all losses and updating the parameters of the lightweight clinical brain-like diagnostic model until convergence, the final lightweight brain-like diagnostic model is obtained.
[0043] Furthermore, step five specifically includes:
[0044] A reward signal is constructed based on the predicted diagnostic results and the clinical diagnostic results. The reward signal is expressed by the following formula: Where sim(·) represents the similarity function. R is the reward amplification factor, R is the reward value, and Y is the predicted diagnostic result. clin This is a clinical diagnostic result;
[0045] Using the reward signal as feedback, the parameters of the lightweight brain-inspired diagnostic model are updated using a Monte Carlo-based policy gradient reinforcement learning algorithm. The update formula is expressed as follows:
[0046]
[0047] Where θ is a parameter. For learning rate, Let P(Y|x;θ) be the gradient operator for P(Y|x;θ), where P(Y|x;θ) is the probability of the predicted result under parameter θ.
[0048] By updating the parameters of a lightweight brain-inspired diagnostic model using a Monte Carlo-based policy gradient reinforcement learning algorithm, and simultaneously performing joint optimization using supervised learning loss, the optimal lightweight brain-inspired diagnostic model is obtained.
[0049] The second aspect of this invention proposes an early auxiliary diagnostic system for neuropsychiatric diseases based on a brain-like multimodal large model, comprising:
[0050] The data collection unit is used to collect and preprocess patients' clinical history data to obtain multimodal data, laying the foundation for subsequent analysis.
[0051] The fusion unit is used to extract and fuse features from preprocessed data using a brain-like mechanism to obtain a fused feature vector.
[0052] The model building unit is used to construct a large-scale brain-like multimodal hybrid expert diagnostic model. Knowledge distillation and compression are used to process the large-scale brain-like multimodal hybrid expert diagnostic model to obtain a lightweight brain-like diagnostic model, which facilitates the reduction of computational redundancy and improves diagnostic accuracy.
[0053] The prediction unit is used to input the fused feature vector into the lightweight brain-like diagnostic model to obtain the prediction and diagnosis results.
[0054] The training unit is used to retrain the lightweight brain-like diagnostic model based on the predicted diagnostic results to obtain the optimal lightweight brain-like diagnostic model, which facilitates the improvement of diagnostic accuracy.
[0055] The diagnostic unit is used to preprocess the patient's clinical data and then use a brain-like mechanism to extract and fuse features to obtain a target fused feature vector. The target fused feature vector is then input into the optimal lightweight brain-like diagnostic model to obtain the final diagnostic result.
[0056] A third aspect of the present invention provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the method for early auxiliary diagnosis of neuropsychiatric diseases based on a brain-like multimodal large model as described in the first aspect above.
[0057] A fourth aspect of the present invention provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device on which the storage medium is located to perform the early auxiliary diagnostic method for neuropsychiatric diseases based on a brain-like multimodal large model as described in the first aspect above.
[0058] The beneficial effects of this invention are:
[0059] This invention provides a method and system for early auxiliary diagnosis of neuropsychiatric diseases based on a brain-like multimodal large-scale model. It effectively integrates heterogeneous data from multiple sources, including speech, images, and text, and combines a unified preprocessing workflow to achieve collaborative modeling of information, thereby improving the comprehensiveness and accuracy of diagnosis. By introducing a brain-like spiking neural network (SNN), it simulates the sparse activation and temporal response mechanism of biological neurons, enhancing the model's ability to model long-range dependent features and reducing energy consumption during training and inference, making it suitable for deployment scenarios with low-power devices. Furthermore, this invention integrates a neuropsychiatric medical knowledge graph to construct a brain-like multimodal hybrid expert diagnostic large-scale model (MME-LMD) with semantic understanding capabilities, realizing a brain-like decision-making mechanism that combines "data-driven and knowledge-based reasoning," enhancing the model's diagnostic logic and interpretability. By guiding expert routing and multimodal feature fusion, the model can simulate the human brain's regional response patterns, achieving efficient sparse computation. To meet the needs of practical clinical applications, this invention also introduces knowledge distillation and model compression techniques to optimize the hybrid expert model and construct a lightweight, high-performance neuromorphic diagnostic model (NMDM). This significantly reduces model complexity and improves reasoning efficiency while maintaining diagnostic accuracy. Furthermore, to address the challenges posed by individual patient differences and disease evolution, this invention designs a continuous optimization mechanism based on reinforcement learning and incremental learning. This mechanism dynamically updates the medical knowledge graph and model parameters using newly added case information, continuously improving the model's generalization ability and adaptability. Overall, this invention possesses advantages such as high diagnostic accuracy, efficient and deployable model, interpretable structure, and continuous evolution. Attached Figure Description
[0060] Figure 1 The flowchart illustrates an early auxiliary diagnostic method for neuropsychiatric diseases based on a brain-like multimodal large model, as provided in an embodiment of the present invention.
[0061] Figure 2 This is one of the schematic diagrams of an early auxiliary diagnostic method for neuropsychiatric diseases based on a brain-like multimodal large model provided in an embodiment of the present invention.
[0062] Figure 3 This is a schematic diagram of spiking neural network feature processing provided in an embodiment of the present invention.
[0063] Figure 4 This is a schematic diagram of a multi-channel attention mechanism module provided in an embodiment of the present invention.
[0064] Figure 5 This is the second schematic diagram of an early auxiliary diagnostic method for neuropsychiatric diseases based on a brain-like multimodal large model provided in an embodiment of the present invention.
[0065] Figure 6 This is a schematic diagram of extracting Mel frequency cepstral coefficients according to an embodiment of the present invention.
[0066] Figure 7 This is a schematic diagram of an early auxiliary diagnostic method for neuropsychiatric diseases based on a brain-like multimodal large model, provided in an embodiment of the present invention.
[0067] Figure 8 This is an architecture diagram of an early auxiliary diagnostic system for neuropsychiatric diseases based on a brain-like multimodal large model, provided in an embodiment of the present invention. Detailed Implementation
[0068] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0069] Example 1
[0070] like Figure 1 As shown, early auxiliary diagnostic methods for neuropsychiatric diseases based on brain-like multimodal large models include:
[0071] S101: Collect and preprocess the patient's clinical history data.
[0072] S102: Use a brain-like mechanism to extract and fuse features from the preprocessed data to obtain a fused feature vector.
[0073] S103: Construct a large-scale brain-like multimodal hybrid expert diagnostic model, and process the large-scale brain-like multimodal hybrid expert diagnostic model using knowledge distillation and compression to obtain a lightweight brain-like diagnostic model.
[0074] Specifically, the brain-like multimodal hybrid expert diagnostic big model (MME-LMD) includes a knowledge graph, a main routing module, multiple expert sub-networks, a MoE output fusion module, and a diagnostic prediction layer.
[0075] The knowledge graph is used to represent the relationships between diseases, symptoms, and examination methods. The main routing module is used to initially determine the major disease types based on input features and the knowledge graph, and select the corresponding expert sub-networks based on these disease types; the main routing module includes a spiking neural network. Each expert sub-network is used to determine the specific disease type and the weights of the expert sub-network based on the input features; each expert sub-network includes a feedforward neural network and a gated network. The MoE output fusion module is used to obtain mixed features based on the outputs of the expert sub-networks. The diagnostic prediction layer is used to process the mixed features to obtain the prediction results.
[0076] S104: Input the fused feature vector into the lightweight brain-like diagnostic model to obtain the predicted diagnostic results.
[0077] S105: Retrain the lightweight brain-like diagnostic model based on the predicted diagnostic results to obtain the optimal lightweight brain-like diagnostic model.
[0078] S106: After preprocessing the patient's clinical data, feature extraction and fusion are performed using a brain-like mechanism to obtain a target fusion feature vector. The target fusion feature vector is then input into the optimal lightweight brain-like diagnostic model to obtain the final diagnostic result.
[0079] This invention uses the MoE (Mixture of Experts) structure as its core, leveraging its dynamic expert routing mechanism. Based on the characteristic distribution of the input data, it automatically activates suitable expert sub-networks for diagnostic calculations, thereby achieving efficient resource scheduling and model expansion capabilities. Furthermore, by employing a brain-like spiking neural network, it simulates the event-driven and sparse activation mechanisms of human brain neurons, improving the modeling ability for time-dependent features, enhancing the processing efficiency of complex time-series modalities such as audio signals and EEG data, and improving the accuracy of identifying early symptoms of neuropsychiatric diseases.
[0080] This invention constructs a large-scale brain-like multimodal hybrid expert diagnostic model, integrating the brain-like computational mechanism of spiking neural networks with the dynamic expert routing capabilities of the MoE structure, to achieve efficient modeling and feature fusion of complex time-series data. This method incorporates medical prior knowledge using knowledge graphs, achieving a fusion of data-driven and knowledge-driven approaches, enhancing the model's diagnostic accuracy, interpretability, and scalability. It provides an innovative path for intelligent and precise clinical auxiliary diagnosis and is suitable for constructing intelligent auxiliary diagnostic systems for neuropsychiatric diseases.
[0081] Example 2
[0082] Based on the above embodiments, such as Figure 2 As shown, this invention proposes a specific process for an early auxiliary diagnostic method for neuropsychiatric diseases based on a brain-like multimodal large model, including:
[0083] S201: Collect and preprocess the patient's clinical history data.
[0084] Specifically, at least two of the following data types are collected from the same or different data sources from the same patient: audio data (including voice recordings and ambient sounds); image data (including brain MRI images and functional magnetic resonance imaging (fMRI)); and text data (including electronic medical records and patient self-report text). The collected clinical data undergoes denoising, standardization, alignment, and formatting to ensure consistency in the temporal and numerical distribution of the multimodal data, constructing a multimodal big data dataset to lay the foundation for subsequent analysis.
[0085] Preferably, the following operations are also required for image data, audio data, and text data:
[0086] Image data: Resizing: Standardize resolution to the same size; Denoising: Use Gaussian or median filtering; Normalization: Standardize pixel values to 0 and 1; Data augmentation: Random rotation, cropping, and flipping. Audio data: Standardize length: Truncate or fill with silence to duration T; Denoising: Eliminate ambient noise using spectral subtraction; Standardize sampling rate: Resample to f. s Feature extraction: Calculate Mel-frequency cepstral coefficients (MFCC). Text data: Text cleaning: Remove stop words and special symbols; Lexical analysis: Segmentation to generate a vocabulary set W = {w1, w2, ..., w...} n Semantic embedding: Generate sentence vectors using the BERT model.
[0087] S202: Use a brain-like mechanism to extract and fuse features from the preprocessed data to obtain a fused feature vector.
[0088] Specifically, by leveraging the brain-like spiking mechanism and utilizing the discrete impulse response characteristics of brain-like spiking neural networks (SNNs), we can efficiently extract and fuse multimodal features and model the temporal dependencies of data.
[0089] First, feature extraction is performed on the preprocessed data to obtain preliminary feature vectors for multiple modalities. For each modality i (where i∈{image, audio, text}), where image is an image, audio is speech, and text is text, the input data X... i Features are extracted using a dedicated sub-network:
[0090]
[0091] Among them, f i (·) represents the feature extractor for the corresponding mode, F i Let be the initial feature vector of the i-th mode. This represents the i-th modal feature after preprocessing.
[0092] For feature extractors, the image modality employs multi-scale convolution and dynamic attention mechanisms to extract spatial features. The audio modality combines Mel-frequency cepstral coefficients (MFCC) and temporal attention mechanisms to extract speech features. The text modality utilizes pre-trained language models (such as BERT) and semantic attention mechanisms to extract text features.
[0093] The preliminary feature vectors of multiple modalities are input into the leaked integral firing neuron model to obtain the time step feature vectors of the corresponding modalities. All time step feature vectors are directly concatenated in chronological order to obtain the historical global temporal feature vector of each modality.
[0094] Specifically, a Leaky Integrate-and-Fire (LIF) neuron model is used for dynamic calculation of membrane potential. The output pulse of each neuron is determined by its membrane potential V(t) and a preset threshold Vt. th The comparison results determine whether the following conditions are met:
[0095]
[0096] Where S(t) is the output pulse exponent, S(t) = 1 generates an output pulse, and S(t) = 0 generates no output pulse; V(t) is the membrane potential at time step t. th This is a preset threshold.
[0097] When S(t) = 1, the membrane potential V(t) is reset to the resting potential V. th And in the next time step, update according to the attenuation coefficient λ as follows:
[0098] V(t+1)=V(t)+I(t)-λV(t)
[0099] Where I(t) is the weighted cumulative value of the input pulse signal, V(t+1) is the membrane potential at time step t+1, and λ is the attenuation coefficient.
[0100] The time-step feature vectors for the corresponding modality are obtained through a leaky integral firing neuron model. By accumulating weights from historical pulse signals and using a membrane potential update mechanism, long-range temporal dependencies in the data are captured. Unlike recursive updates or weighted accumulation methods, this invention directly concatenates the data generated at each time step. Let the total number of time steps be T, then the global temporal feature vector obtained by concatenating the feature vectors from all time steps is expressed as:
[0101]
[0102] in, Let F(T) be the global temporal feature vector, and F(T) be the feature vector at time step T. [·] represents the concatenation of feature vectors from each time step along the time axis. This concatenation method ensures that detailed information at each time step is preserved, thus fully describing the temporal evolution of the data. This allows the model to fully capture long-range temporal dependencies, providing rich and continuous temporal features for subsequent classification tasks, while maintaining the low power consumption and event-driven advantages of spiking neural networks. A schematic diagram of cross-time step information fusion is shown below. Figure 3 As shown.
[0103] The feature vector at each time step of the global temporal feature vector is input into a multi-channel attention mechanism (Linear-Softmax-Scale, LSS) module to obtain weighted features. Specifically, in this invention, as shown... Figure 4 As shown, by referring to the following Figure 4 Based on the idea of the SE module (a), a multi-channel attention mechanism module is proposed, the structure of which is as follows: Figure 4 As shown in (b), the LSS module first obtains a suitable feature vector through interpolation, then maps the input features to a fixed dimension (Linear) using a linear transformation, and finally calculates the channel-level attention weights using softmax. The specific construction principle is as follows:
[0104] e t =tanh(W h x t +b h )
[0105]
[0106] Among them, e t x is the output after linear transformation at time step t. t W is the input data at time step t. h For the linear layer weights, b h Here, α is the bias term, tanh is the activation function, and α is the alpha term. t Let V be the feature weights at time step t, and V be the weighted features.
[0107] Multimodal weighted features utilize a multi-channel attention mechanism to dynamically weight attention enhancement features across different modalities. Let the attention score for each modality be a. i The normalized weights are calculated using the softmax function:
[0108]
[0109] Among them, w i For the weighted feature weights, a i The attention score is given for each modality, where j represents the modality category, and image, audio, and text represent image, audio, and text, respectively.
[0110] The final fused feature vector is obtained by weighted summation of each modality:
[0111]
[0112] Among them, F fusion To fuse feature vectors, V i is the weighted feature of mode i.
[0113] S203: Construct a large-scale brain-like multimodal hybrid expert diagnostic model, and process the large-scale brain-like multimodal hybrid expert diagnostic model using knowledge distillation and compression to obtain a lightweight brain-like diagnostic model.
[0114] Specifically, by integrating brain-like partitioning sparse processing mechanisms and neuropsychiatric medical knowledge graphs, a brain-like multimodal hybrid expert diagnostic model (MME-LMD) is constructed. Through dynamic gating mechanisms, expert sub-networks are selectively activated to achieve sparse routing computation, simulate the specialized response of human brain regions, and reduce computational redundancy.
[0115] The brain-like multimodal hybrid expert diagnosis big model includes a knowledge graph, a main routing module, multiple expert sub-networks, a MoE output fusion module, and a diagnosis prediction layer.
[0116] The medical knowledge graph is embedded in the model reasoning process in the form of a graph structure. Nodes represent entities such as diseases, symptoms, and examination methods, and edges represent semantic relationships between entities (such as causal relationships and diagnostic associations).
[0117]
[0118] Where G represents the knowledge graph. Let E be the set of nodes and E be the set of edges.
[0119] By fusing knowledge graphs with multimodal features through Graph Attention Network (GAT), the diagnostic decision-making logic is optimized to meet the following requirements:
[0120] P(d|F fusion ,G)=softmax(f diagnosis (F fusion ,v d ,v s ,f c ))
[0121] Wherein P(d|F fusion G) represents the probability of predicting the disease, f diagnosis Let F be the decision function. fusion To fuse feature vectors, v d v s and v c These are the embedding vectors of disease, symptom, and examination method nodes in the knowledge graph.
[0122]
[0123] Among them, Y model For the output of the brain-like multimodal hybrid expert diagnostic large model, f fusion To integrate multimodal data with information from the knowledge graph layer, For including v d v s and v c The knowledge graph.
[0124] Main routing classification module (first-level decision unit):
[0125] Based on multimodal input data, the disease is first preliminarily classified using the main routing classification module (spiking neural network), dividing the input samples into broad categories, such as mental illnesses and neurological diseases. This step utilizes information on basic disease classifications in the knowledge graph to provide prior semantic guidance for subsequent expert selection.
[0126] Expert subnetwork (second-layer gating):
[0127] Based on the initial classification results from the main routing module, disease types are further subdivided through expert sub-networks. Independent expert models are designed for specific diseases such as depression and Parkinson's disease. Each expert sub-network uses deep learning to study the feature patterns corresponding to its disease, avoiding the "general but not precise" shortcomings of a single model and providing more professional and detailed diagnosis for specific diseases. Each expert sub-network includes a feedforward neural network and a gating network. Each expert sub-network can focus on different tasks or modal features, such as expert sub-networks specifically handling depression, Parkinson's disease, or Alzheimer's disease. Internally, each expert sub-network can employ an architecture optimized for its respective domain, and each expert sub-network produces different outputs for input f.
[0128] e j =φ j (f), j = 1, 2, ..., N
[0129] Where, φ j Let f represent the j-th expert network, f be the input features, N be the number of expert subnetworks participating in the computation, and e be the number of expert subnetworks participating in the computation. j This is the output of the j-th expert subnetwork.
[0130] Simultaneously, a gating network is used to dynamically assign expert weights based on the fusion features:
[0131]
[0132] in, W represents the activation weights of each expert subnetwork, where softmax is the activation function. g and b g For weight and bias.
[0133] Preferably, in order to compensate for the judgment bias that a single expert subnetwork may have in boundary samples (where symptoms are atypical or multiple diseases overlap), the present invention introduces a shared expert subnetwork, which is responsible for learning common features across diseases and providing supplementary judgments among multiple disease expert subnetworks.
[0134] The MoE output fusion module is used to obtain hybrid features based on the output of the expert subnetwork, as expressed by the following formula:
[0135]
[0136] Among them, f MoE It is a mixed feature. Let be the weight of the j-th expert subnetwork.
[0137] The diagnostic prediction layer further integrates information based on the MoE output using a fusion layer. After processing by a Transformer layer, the specific formula is as follows:
[0138] F tr =ψ(f MoE )
[0139] Among them, f tr The ψ(·)Transformer encoder is used to capture complex relationships between modes, representing the encoding features.
[0140] Finally, the diagnostic prediction layer provides disease classification, risk prediction, or treatment recommendations:
[0141] y=σ(W o f tr +b o )
[0142] Where y is the prediction result, σ is the activation function, and W o and b o These are the weights and the biases, respectively.
[0143] Preferably, when constructing and training a large brain-like multimodal hybrid expert diagnostic model, single-modal features are used first for model construction and training to obtain the model parameters under single-modal features. Then, multimodal features are used for training. After obtaining the trained large brain-like multimodal hybrid expert diagnostic model, it is optimized into a lightweight brain-like diagnostic model.
[0144] We optimize MME-LMD by employing knowledge distillation and model compression techniques. By distilling the semantic information, soft labels, and structural sparsity features of the intermediate layer, and combining pruning and quantization operations, we generate a lightweight and efficient neuropsychiatric clinical brain-like diagnostic model (NMDM), thereby improving actual reasoning efficiency.
[0145] Specifically, firstly, a large-scale brain-inspired multimodal expert diagnostic model is trained, fusing data from different modalities (such as images, audio, and text) to diagnose neuropsychiatric disorders. Then, the knowledge is transferred to a smaller clinical model. This clinical brain-inspired diagnostic model learns from the output and intermediate features of the large-scale model, maintaining high diagnostic accuracy while reducing computational complexity. Through this process, a highly efficient, lightweight brain-inspired diagnostic model is obtained, capable of processing collected multimodal patient data in real time, providing rapid and accurate assisted diagnosis, and supporting clinicians in making decisions. The diagnostic logic of knowledge distillation can be represented as follows:
[0146] NMD model (D new =SNN_Diagnosis(D new )
[0147] Among them, NMD model SNN_Diagnosis is a lightweight brain-inspired diagnostic model, while D is a large-scale brain-inspired multimodal hybrid expert diagnostic model. new For training data.
[0148] This method effectively improves the diagnostic efficiency of the model, while achieving the goal of accurate diagnosis by preserving key information in the multimodal large model.
[0149] Based on the lightweight clinical brain-like diagnostic model obtained through knowledge distillation, further compression is achieved using pruning, quantization, and low-rank decomposition. The process is as follows:
[0150] The weights of the clinical brain-like diagnostic model are pruned by setting the weight matrix to zero if it is less than a preset threshold. After pruning, the model is fine-tuned to minimize the loss.
[0151] During training, a quantization function Q(W) is introduced to map the weights W from 32-bit floating-point numbers to 8-bit integer representations. Singular value decomposition (SVD) is performed on key layers, selecting and retaining low-rank matrices. All losses are jointly optimized, and the parameters of the clinical brain-inspired diagnostic model are updated until convergence. The resulting lightweight model achieves diagnostic accuracy close to that of large multimodal diagnostic models while maintaining lower parameter count and computational cost.
[0152] The patient's new data is input into the distilled, lightweight clinical brain-like diagnostic model, which outputs diagnostic results:
[0153] Y = NMD model (D new )
[0154] Where Y represents the predicted diagnostic result, i.e., the output of the lightweight clinical brain-like diagnostic model.
[0155] Described by the above procedural formula, starting with knowledge distillation, the soft labels and intermediate features of the large model are transferred to the clinical brain-like diagnostic model. Combined with auxiliary compression techniques such as pruning, quantization, and low-rank decomposition, the brain-like diagnostic model can be systematically lightweighted. The entire process ensures that diagnostic performance is preserved as much as possible while compressing the model, and reduces energy consumption and improves inference efficiency during actual deployment.
[0156] S204: Input the fused feature vector into the lightweight brain-like diagnostic model to obtain the predicted diagnostic results.
[0157] S205: Retrain the lightweight brain-like diagnostic model based on the predicted diagnostic results to obtain the optimal lightweight brain-like diagnostic model.
[0158] Specifically, the medical knowledge graph is updated using new case data, and MME-LMD is fine-tuned through reinforcement learning and incremental learning mechanisms. The concepts and relationships in the knowledge graph are expanded, and expert routing or fusion weights are adjusted with the help of policy reward signals. At the same time, parameters are incrementally fine-tuned to avoid catastrophic forgetting, continuously optimizing the diagnostic performance of NMDM and adapting it to new clinical conditions and population differences.
[0159] The incremental learning process compares the model's output with the actual diagnostic results provided by clinical experts, constructs reward feedback, and then uses reinforcement learning algorithms to update the model parameters. Clinical experts serve as feedback, while the model continuously improves its diagnostic accuracy through algorithmic learning. This process mainly includes the following stages:
[0160] Model Diagnosis: For a given input data x including data of each modality, the current model outputs a diagnostic result Y.
[0161] Clinical expert feedback: Clinical experts independently assess the diagnosis of the same patient and provide a clinical diagnostic result. clin The system combines the diagnostic results from the model with feedback based on expert experience. Feedback may include an assessment of diagnostic accuracy (correct / incorrect), a confidence score, and suggestions for improvement or questions.
[0162] Reward signal construction: Calculate rewards based on the consistency between model output and clinical outcomes.
[0163]
[0164] Where sim(·) represents the similarity function. R is the reward amplification factor, and R is the reward value. If the model predictions perfectly match clinical outcomes, the reward value is higher; otherwise, the reward value is lower.
[0165] Reinforcement learning update stage: The reward signal is used as feedback to update the model parameters using reinforcement learning. The classic REINFORCE algorithm (a Monte Carlo-based policy gradient reinforcement learning algorithm) is employed to update the model parameters. The specific update formula is as follows:
[0166]
[0167] Where θ is a parameter. For learning rate, Let P(Y|x;θ) be the gradient operator for P(Y|x;θ), where P(Y|x;θ) is the probability of the predicted result under parameter θ.
[0168] Meanwhile, during the update process, traditional supervised learning loss is combined for joint optimization to balance the integration of new feedback and historical knowledge.
[0169] Model update: The updated model is put back into the diagnostic task. Through multiple rounds of incremental learning, the model's performance on new data is continuously corrected and improved.
[0170] S206: After preprocessing the patient's clinical data, feature extraction and fusion are performed using a brain-like mechanism to obtain a target fusion feature vector. The target fusion feature vector is then input into the optimal lightweight brain-like diagnostic model to obtain the final diagnostic result.
[0171] Example 3
[0172] Based on the above embodiments, such as Figure 5 As shown, this invention proposes a process for an early auxiliary diagnostic method for neuropsychiatric diseases based on a brain-like multimodal large model, specifically including:
[0173] Collect clinical data to construct multimodal big data and perform classification preprocessing:
[0174] Multimodal clinical data from patients is acquired using specialized data acquisition equipment. The data collected includes:
[0175] Medical history: The doctor inquires in detail about the patient's clinical symptoms, emotional state, cognitive ability, motor performance, sleep status, etc., and records the patient's self-reported specific symptom descriptions, such as "recent difficulty concentrating, significant mood swings, slow movement or sluggish reaction", etc.
[0176] Audio data: Voice data is collected using specialized recording equipment during doctor-patient interactions or when the patient is alone. The audio may contain information such as speech rate, pitch, and volume variations, which are important for capturing subtle changes in language expression.
[0177] Video or sensor data: Video or sensor data such as the patient's movement status, facial expressions, gait and hand movements are collected through cameras or wearable devices to provide objective evidence for motor dysfunction and behavioral abnormalities.
[0178] Preprocessing of the collected modal data:
[0179] The audio data undergoes a unified sampling rate, frame segmentation, and noise removal.
[0180] Clean and standardize the text data;
[0181] Format conversion and time alignment of video or sensor data are performed to ensure consistency in temporal sequence and numerical distribution of data across modalities, laying the foundation for subsequent feature extraction.
[0182] Utilizing brain-like mechanisms for feature extraction and deep data fusion:
[0183] After preprocessing the data for each modality, a neuromorphic spiking neural network was used for feature extraction. The specific processing is as follows:
[0184] Audio feature extraction: Mel frequency cepstral coefficients (MFCC) were used to process the patient's speech and extract features representing sound characteristics (such as volume, pitch, and rhythm). The MFCC preprocessing process includes pre-emphasis, framing, windowing, Fast Fourier Transform (FFT), Mel filter bank, and Discrete Cosine Transform (DCT).
[0185] Video / Motion Data Feature Extraction: For video or sensor data, a brain-like spiking neural network is used to model motion trajectories, gait changes, and hand tremors, capturing subtle temporal changes and spatial dynamic features.
[0186] Text data vectorization: Using a pre-trained large model (such as BERT) to vectorize the consultation records, converting the doctor's description and the patient's self-report into high-dimensional semantic features.
[0187] Subsequently, through the cross-time-step information fusion mechanism of the spiking neural network, the feature vectors generated at each time step are directly concatenated in chronological order to form a global temporal feature vector:
[0188] F fusion = [F(1); F(2); ...; F(T)]
[0189] Finally, the historical global temporal feature vector of each modality is input into the multi-channel attention module to obtain the weighted features of each modality. The weighted features of all modalities are then fused to obtain the fused features.
[0190] Construct and train a large-scale brain-like multimodal hybrid expert diagnostic model based on knowledge graphs:
[0191] Based on a medical knowledge graph of neuropsychiatric diseases, a brain-like MoE multimodal brain-like diagnostic model was constructed. The specific process includes:
[0192] Knowledge graph construction: Information such as disease-related symptoms, signs, behavioral indicators, and examination results are embedded into the knowledge graph in a graph structure to provide structured medical knowledge for the model.
[0193] Main routing module: Utilizes fused multimodal features to identify anomalies in the currently collected information, and determines the most likely disease category based on knowledge graph information, thereby activating the corresponding expert sub-network.
[0194] Expert Subnetwork: Within the activated expert subnetwork, specific disease categories are further determined through a gating mechanism. The expert subnetwork undergoes specialized training tailored to different disease characteristics.
[0195] MoE Output Fusion: The outputs of each expert are fused using a weighted summation method, expressed as follows:
[0196]
[0197] Among them, f MoE It is a mixed feature. Let be the weight of the j-th expert subnetwork.
[0198] Subsequently, the Transformer encoder is used to capture the complex relationships between modalities, obtaining cross-modal fused features:
[0199] F tr =ψ(f MoE )
[0200] Among them, f tr The ψ(·)Transformer encoder is used to capture complex relationships between modes, representing the encoding features.
[0201] Finally, the diagnostic prediction layer provides disease classification, risk prediction, or treatment recommendations:
[0202] y=σ(W o f tr +b o )
[0203] Where y is the prediction result, σ is the activation function, and W o and b o These are the weights and the biases, respectively.
[0204] A lightweight neuromorphic diagnostic model (NMDM) is obtained using knowledge distillation and model compression. To adapt to real-time clinical applications, knowledge distillation is employed to transfer the deep semantic information and graph logic of the large multimodal model to the lightweight model.
[0205] The new model is trained using intermediate layer semantic representations and expert outputs as teacher signals.
[0206] By combining techniques such as pruning, quantization, and low-rank decomposition, the model weights are optimized and compressed, unimportant parameters are set to zero, floating-point numbers are mapped to low-order integers, and SVD decomposition is used for key layers to reduce computational load.
[0207] After the above knowledge distillation and compression, the resulting lightweight neuropsychiatric clinical brain-like diagnostic model (NMDM) can rapidly process multimodal data in resource-constrained environments and provide accurate diagnostic suggestions.
[0208] Y = NMD model (D new )
[0209] The model is incrementally learned and continuously updated based on feedback from new cases.
[0210] To maintain the model's adaptability to new clinical data, an incremental learning and update mechanism based on feedback rewards was designed:
[0211] Data feedback and knowledge graph update: During the actual diagnosis process, the model output results and the independent diagnosis results of clinical experts are collected, and the new clinical data is analyzed and updated into the medical knowledge graph.
[0212] Reward signal calculation: Compare the similarity between the model prediction and the clinical diagnosis to calculate the reward signal.
[0213]
[0214] Reinforcement learning update: The REINFORCE algorithm is used to feed back the reward signal to update the model parameters. The update formula is as follows:
[0215]
[0216] Joint optimization: By combining traditional supervised loss in the reinforcement learning process, new data and historical knowledge are integrated. Through multiple iterations, the model continuously improves the accuracy and generalization ability of the task.
[0217] Example 4
[0218] Based on the above embodiments, the present invention provides a process for an early auxiliary diagnostic method for depression based on a brain-like multimodal large model, specifically including:
[0219] Collect clinical data to construct multimodal big data, and perform classification and preprocessing on it;
[0220] In this embodiment, the patient's multimodal clinical data is first collected using a mobile terminal and a dedicated recording device, mainly including:
[0221] During the consultation, the doctor meticulously recorded the patient's emotional state, sleep patterns, changes in appetite, daily activities and interests, concentration level, and other symptoms, as well as the patient's subjective descriptions, such as "long-term low mood, loss of interest in anything, frequent awakenings at night and difficulty falling back asleep, and loss of appetite."
[0222] Audio data: Using specialized recording equipment, audio recordings of patients' voices were collected during interactions with doctors and when they were alone in the relaxation room. Analysis revealed that patients with depression spoke in a low, slow, and monotonous voice, lacking emotional variation and vitality.
[0223] Furthermore, preprocessing of audio and text modalities includes unifying the audio sampling rate, performing audio frame segmentation, and removing noise.
[0224] Perform text cleaning and text standardization on the text.
[0225] Utilizing brain-like mechanisms for feature extraction modules and deep data fusion;
[0226] Specifically, according to the knowledge graph of depression, audio MFCC (Melbourne Frequency Cepstral Coefficients) features are an important part of objectively reflecting the physiological state of patients with depression, and text is an important way to reflect the psychological characteristics of depression. Therefore, in feature processing, the audio signal is converted into MFCC feature representation. The MFCC preprocessing process includes steps such as pre-emphasis, framing, windowing, FFT, Mel filter bank, and DCT.
[0227] Similarly, pre-trained large models can be used to vectorize text data, making subsequent processing and analysis easier.
[0228] After converting audio and text into features, features are extracted and fused using the characteristics of the impulse mechanism.
[0229] The pulse time step is weighted by the accumulated weight W of historical pulse signals. i Unlike recursive updates or weighted accumulation methods, this invention uses a membrane potential update mechanism to capture long-range temporal dependencies in the data. Instead, it directly concatenates the data generated at each time step. Let the total number of time steps be T, then the global temporal feature vector obtained by concatenating the feature vectors of all time steps is expressed as:
[0230] F fusion = [F(1); F(2); ...; F(T)]
[0231] Here, "[·]" indicates concatenation of feature vectors from each time step along the time axis. This concatenation method ensures that detailed information at each time step is preserved, thus fully describing the temporal evolution of the data. This allows the model to fully capture long-range temporal dependencies, providing rich and continuous temporal features for subsequent classification tasks, while maintaining the low power consumption and event-driven advantages of spiking neural networks. A schematic diagram of cross-time step information fusion is shown below. Figure 3 As shown.
[0232] Finally, the historical global temporal feature vector of each modality is input into the multi-channel attention module to obtain the weighted features of each modality. The weighted features of all modalities are then fused to obtain the fused features.
[0233] Using brain-like partitioning sparsity processing mechanisms and a medical knowledge graph of neuropsychiatric diseases, a brain-like multimodal hybrid expert diagnostic model (MME-LMD) is constructed and trained.
[0234] To address the diagnostic needs of depression, this embodiment introduces a medical knowledge graph related to depression. This graph contains structured descriptions and clinical indicators of symptoms such as depressed mood, sleep disturbances, loss of interest, and poor concentration. The MME-LMD model is constructed and trained using this graph. The model construction process includes:
[0235] Main routing module: Based on the fused multimodal features, the main routing module determines that there are emotional and behavioral abnormalities in the currently collected information and determines that the possibility of mental illness is high, thereby activating the corresponding expert sub-network.
[0236] Expert subnetwork activation: Among the activated experts, the gating mechanism further determines whether they meet the clinical characteristics of depression. If they do, a specialized expert network for depression is activated for training.
[0237] Expert output fusion: Multimodal features determine that the current collected information is more likely to be a mental illness through the main route, and activate the mental illness classification expert. The mental illness classification expert determines that it belongs to the depression category through the gating mechanism, and activate the depression expert for training.
[0238] The MoE output fusion module is used to obtain mixed features based on the output of the expert subnetwork.
[0239] Ultimately, the diagnostic prediction layer provides disease classification, risk prediction, or treatment recommendations.
[0240] By optimizing MME-LMD using knowledge distillation and model compression, a highly efficient and lightweight clinical brain-like diagnostic model (NMDM) for neuropsychiatric diseases is obtained.
[0241] After model training, knowledge distillation is used to transfer the deep knowledge (such as intermediate layer semantic representations and graph logical structures) of the multimodal fusion model to a lightweight neuromorphic diagnostic model (NMDM). NMDM can process collected patient multimodal data in real time, providing rapid and accurate assisted diagnosis to support clinicians in making decisions. The diagnostic logic of knowledge distillation can be represented as follows:
[0242] NMDM model (D new =SNN_Diagnosis(D new )
[0243] Among them, SNN_Diagnosis(D new This is a large-scale brain-like multimodal hybrid expert diagnostic model.
[0244] Furthermore, based on the lightweight clinical brain-like diagnostic model obtained through knowledge distillation, pruning, quantization, and low-rank decomposition are used for auxiliary compression. The final lightweight model is obtained, which achieves diagnostic accuracy close to that of a large multimodal diagnostic model while having lower parameter count and computational cost.
[0245] In clinical diagnosis, patient data is input into the distilled NMDM, and diagnostic results are output:
[0246] Y = NMD model (D new )
[0247] We updated the medical knowledge graph of depression using newly added depression cases, and fine-tuned the MME-LMD using reinforcement learning and incremental learning to improve the diagnostic effectiveness of NMDM.
[0248] In the clinical diagnosis of patients with depression, the output of a clinical brain-like diagnostic model is compared with the actual diagnostic results given by clinical experts to construct a reward feedback mechanism. Then, a reinforcement learning algorithm is used to update the model parameters. Clinical experts serve as feedback, while the model continuously adapts to new cases through incremental learning.
[0249] Data feedback and knowledge graph updates: During the actual diagnosis process, the model prediction results and independent diagnosis results of clinical experts are collected, and the detailed information of new cases is updated into the medical knowledge graph of depression.
[0250] Reward signal calculation: The reward is calculated based on the consistency between the model output and clinical outcomes.
[0251]
[0252] Reinforcement learning update stage: The reward signal is used as feedback to update the model parameters using reinforcement learning. The classic REINFORCE algorithm is employed to update the model parameters. The specific update formula is as follows:
[0253]
[0254] Meanwhile, during the update process, traditional supervised learning loss is combined for joint optimization to balance the integration of new feedback and historical knowledge.
[0255] Model update: The updated model is put back into the diagnostic task. Through multiple rounds of incremental learning, the model's performance on new data is continuously corrected and improved.
[0256] Example 5
[0257] Based on the above embodiments, this invention proposes a process for an early auxiliary diagnostic method for Parkinson's disease based on a brain-like multimodal large model, specifically including:
[0258] Collect clinical data to construct multimodal big data and perform classification preprocessing:
[0259] In this Parkinson's disease diagnostic example, multimodal clinical data of the patient is acquired via mobile phone and dedicated data acquisition equipment. The data acquisition content includes:
[0260] Medical history: The doctor inquired in detail about the patient's motor disorders, such as hand tremors, slow movements (bradykinesia), muscle stiffness, and gait instability. The doctor also recorded the patient's self-report, such as specific descriptions like "significant hand tremors, shortened stride, forward leaning, slow and stiff movements when walking."
[0261] Audio data: Voice data of patients during consultations and daily conversations were collected using specialized recording equipment. Analysis showed that Parkinson's disease patients may exhibit characteristics such as a weak voice, slow speech rate, and flat tone.
[0262] Video or sensor data: Motion videos or sensor data such as acceleration and gyroscopes are collected from patients' daily activities via mobile phones or wearable devices to capture motion abnormalities (such as tremors or unstable gait).
[0263] The collected modal data are preprocessed as follows: audio data is subjected to uniform sampling rate, frame segmentation and noise removal; text data is cleaned and standardized; video or sensor data is format converted and time aligned to ensure consistency of temporal sequence and numerical distribution of each modal data, laying the foundation for subsequent feature extraction.
[0264] Utilizing brain-like mechanisms for feature extraction and deep data fusion:
[0265] Based on the descriptions of movement disorders and voice changes in the Parkinson's disease diagnostic knowledge graph, the system focuses on extracting key information from audio, video, and motion data. The specific processing is as follows:
[0266] Audio feature extraction: Mel frequency cepstral coefficients (MFCC) were used to process the patient's speech and extract features representing sound characteristics (such as volume, pitch, and rhythm). The MFCC preprocessing process includes pre-emphasis, framing, windowing, Fast Fourier Transform (FFT), Mel filter bank, and Discrete Cosine Transform (DCT).
[0267] Video / Motion Data Feature Extraction: For video or sensor data, a brain-like spiking neural network is used to model motion trajectories, gait changes, and hand tremors, capturing subtle temporal changes and spatial dynamic features.
[0268] Text data vectorization: A pre-trained large model is used to vectorize the consultation records, converting the doctor's description and the patient's self-report into high-dimensional semantic features.
[0269] Subsequently, through the cross-time-step information fusion mechanism of the spiking neural network, the feature vectors generated at each time step are directly concatenated in chronological order to form a global temporal feature vector:
[0270] F fusion = [F(1); F(2); ...; F(T)]
[0271] This approach preserves the detailed information of each time step and fully captures long-range temporal dependencies, ensuring that subsequent classification tasks can utilize continuous and rich temporal features while maintaining the advantages of low power consumption and event-driven computation.
[0272] Finally, the historical global temporal feature vector of each modality is input into the multi-channel attention module to obtain the weighted features of each modality. The weighted features of all modalities are then fused to obtain the fused features.
[0273] Construct and train a knowledge graph-based MoE multimodal diagnostic large-scale model:
[0274] In the diagnosis of Parkinson's disease, a knowledge graph of Parkinson's disease was constructed, with related clinical symptoms such as bradykinesia, tremor, rigidity, and gait instability embedded in its graph structure. The model training process includes:
[0275] Main routing module: Based on the multimodal fusion features, it determines that there are characteristics such as movement disorders and abnormal sounds in the currently collected information, and determines that the possibility of psychomotor diseases is high, thereby activating the corresponding movement disorder expert network.
[0276] Expert subnetwork activation: In the expert subnetwork, the gating mechanism further identifies specific symptoms of Parkinson's disease and activates experts specifically trained for Parkinson's disease.
[0277] The MoE output fusion module is used to obtain hybrid features based on the output of the expert subnetwork.
[0278] Finally, the diagnostic prediction layer provides disease classification, risk prediction, or treatment recommendations:
[0279] y=σ(W o f tr +b o )
[0280] Where y is the prediction result, σ is the activation function, and W o and b o These are the weights and the biases, respectively.
[0281] Lightweight neuromorphic diagnostic models (NMDMs) are obtained by using knowledge distillation and model compression.
[0282] After MME-LMD training is completed, knowledge distillation techniques are used to transfer the deep semantic information and graph logic structure contained in the multimodal large model to a lightweight brain-like diagnostic model. Specifically, this includes:
[0283] Knowledge distillation is used to transfer intermediate layer features, graph association information, and expert output logic to the new model:
[0284] By combining pruning, quantization, and low-rank decomposition techniques to compress model weights, such as setting the weights below the threshold to zero, mapping 32-bit floating-point numbers to 8-bit integer representations, and performing SVD decomposition on key layers, the number of model parameters and computational cost can be reduced.
[0285] After optimization through the above steps, the final clinical lightweight neuromorphic diagnostic model (NMDM) can efficiently output auxiliary diagnostic results for Parkinson's disease based on real-time acquired multimodal data.
[0286] Y = NMD model (D new )
[0287] Incremental learning and continuous updating of the model using newly added Parkinson's disease cases:
[0288] In practical clinical applications, the system continuously collects diagnostic data from newly diagnosed Parkinson's disease patients and performs knowledge graph updates, calculates feedback rewards, updates through reinforcement learning, and performs joint optimization. The specific process includes:
[0289] Example 6
[0290] Based on the above embodiments, this invention presents the validation results of an early auxiliary diagnostic method for neuropsychiatric diseases based on a brain-like multimodal large model, specifically including:
[0291] To verify the effectiveness of the cross-time-step information fusion method provided in this invention, the E-DAIC dataset was used to verify the impact of various pulse time-step information fusion strategies, including multi-pulse time-step information addition, step-by-step concatenation, and time-step attention mechanisms. The experimental results are shown in Table 1.
[0292] Table 1 Comparison of Text-Based Pulse Time Step Information Fusion Methods
[0293] Advantages of stepwise concatenation: Stepwise concatenation achieves the highest accuracy (87.5%) with 9 time steps. This approach effectively captures complex patterns in time-series data by preserving the independent features of each time step and allowing the model to combine these features in subsequent layers.
[0294] The potential of time-step attention mechanisms: Although the time-step attention mechanism achieved lower accuracy (80.3%) with 10 time steps, this demonstrates its potential application value in SNNs. By introducing an attention mechanism, the model can dynamically adjust the degree of attention given to features at different time steps, potentially improving performance on complex tasks. However, the specific implementation and optimization of the attention mechanism in this experiment may still require further exploration.
[0295] Limitations of multi-time-step information addition: The multi-time-step information addition method achieved low accuracy (75.6%) at 9 time steps. This may be because the simple addition operation failed to fully capture the complex relationships between features at different time steps. Therefore, employing a more complex fusion strategy in SNNs may be more effective.
[0296] In time-step-based information fusion, the setting of the time step is crucial for information acquisition. Therefore, we explored the specific effects of different time step settings on the experimental results within the step-by-step splicing framework. Table 2 details this impact, summarized below:
[0297] Table 2: The impact of time step on accuracy
[0298] As the number of time steps increased from 1 to 9, the model's accuracy gradually improved, indicating that increasing the number of time steps helps the model better capture dependencies in time-series data. This reflects the ability of SNNs to accumulate information over time, enabling the model to establish connections over a longer time span.
[0299] The model achieved its highest accuracy (87.5%) when the time step reached 9. However, the accuracy decreased slightly (86.3%) as the time step was further increased to 10. This may be because too many time steps introduce noise or redundant information, making it difficult for the model to effectively distinguish between relevant and irrelevant features.
[0300] To verify the effectiveness of the brain-like intelligence-based early auxiliary diagnostic method for neuropsychiatric diseases based on multimodal fusion provided in this invention, the E-DAIC dataset was used to verify the accuracy and energy consumption of the multimodal depression auxiliary diagnostic method. The test results are shown in Tables 1 and 2, respectively. The expert model design for depression auxiliary diagnosis is as follows: Figure 7 As shown.
[0301] The experiment used accuracy (Acc), mean average precision (mAP), and F1-Score as evaluation metrics. Table 3 shows the comparison results of this invention with other methods on the EDAIC dataset.
[0302] Table 3 shows the comparison results with other methods on the EDAIC dataset.
[0303]
[0304] Example 7
[0305] Based on the above embodiments, such as Figure 8 As shown, this invention proposes an early auxiliary diagnostic system for neuropsychiatric diseases based on a brain-like multimodal large model, comprising:
[0306] The data collection unit is used to collect and preprocess patients' clinical history data.
[0307] The fusion unit is used to extract and fuse features from preprocessed data using a brain-like mechanism to obtain a fused feature vector.
[0308] The model building unit is used to construct a large-scale brain-like multimodal hybrid expert diagnostic model, and to process the large-scale brain-like multimodal hybrid expert diagnostic model using knowledge distillation and compression to obtain a lightweight brain-like diagnostic model.
[0309] The prediction unit is used to input the fused feature vector into the lightweight brain-like diagnostic model to obtain the prediction and diagnosis results.
[0310] The training unit is used to retrain the lightweight brain-like diagnostic model based on the predicted diagnostic results to obtain the optimal lightweight brain-like diagnostic model.
[0311] The diagnostic unit is used to preprocess the patient's clinical data and then use brain-like mechanisms to extract and fuse features to obtain a target fused feature vector. The target fused feature vector is then input into the optimal lightweight brain-like diagnostic model to obtain the final diagnostic result.
[0312] It should be noted that the early auxiliary diagnostic system for neuropsychiatric diseases based on a brain-like multimodal large model provided in this embodiment of the invention is for realizing the above-mentioned early auxiliary diagnostic method for neuropsychiatric diseases based on a brain-like multimodal large model. Its specific functions can be referred to in the above-mentioned method embodiments, and will not be repeated here.
[0313] Example 8
[0314] Based on the above embodiments, the present invention proposes an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the early auxiliary diagnosis method for neuropsychiatric diseases based on a brain-like multimodal large model as described in Embodiment 1 above.
[0315] Example 9
[0316] Based on the above embodiments, the present invention proposes a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the storage medium is located to perform the early auxiliary diagnosis method for neuropsychiatric diseases based on a brain-like multimodal large model as described in Embodiment 1 above.
[0317] In summary, this invention provides an early auxiliary diagnostic method and system for neuropsychiatric diseases based on a brain-like multimodal large-scale model. It effectively integrates heterogeneous data from multiple sources, including speech, images, and text, and combines a unified preprocessing workflow to achieve collaborative modeling of information, thereby improving the comprehensiveness and accuracy of diagnosis. By introducing a brain-like spiking neural network (SNN), it simulates the sparse activation and temporal response mechanism of biological neurons, enhancing the model's ability to model long-range dependent features and reducing energy consumption during training and inference, making it suitable for deployment scenarios with low-power devices. Furthermore, this invention integrates a neuropsychiatric medical knowledge graph to construct a brain-like multimodal hybrid expert diagnostic large-scale model (MME-LMD) with semantic understanding capabilities, realizing a brain-like decision-making mechanism that combines "data-driven and knowledge-based reasoning," enhancing the model's diagnostic logic and interpretability. By guiding expert routing and multimodal feature fusion, the model can simulate the human brain's regional response patterns, achieving efficient sparse computation. To meet the needs of practical clinical applications, this invention also introduces knowledge distillation and model compression techniques to optimize the hybrid expert model and construct a lightweight, high-performance neuromorphic diagnostic model (NMDM). This significantly reduces model complexity and improves reasoning efficiency while maintaining diagnostic accuracy. Furthermore, to address the challenges posed by individual patient differences and disease evolution, this invention designs a continuous optimization mechanism based on reinforcement learning and incremental learning. This mechanism dynamically updates the medical knowledge graph and model parameters using newly added case information, continuously improving the model's generalization ability and adaptability. Overall, this invention possesses advantages such as high diagnostic accuracy, efficient and deployable model, interpretable structure, and continuous evolution.
[0318] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An early auxiliary diagnostic method for neuropsychiatric diseases based on a brain-like multimodal large model, characterized in that, include: Step 1: Collect and preprocess the patient's clinical history data; Step 2: Use a brain-like mechanism to extract and fuse features from the preprocessed data to obtain a fused feature vector; Step 3: Construct a large-scale brain-like multimodal hybrid expert diagnostic model, and process the large-scale brain-like multimodal hybrid expert diagnostic model using knowledge distillation and compression to obtain a lightweight brain-like diagnostic model; Step 4: Input the fused feature vector into the lightweight brain-like diagnostic model to obtain the predicted diagnostic results; Step 5: Retrain the lightweight brain-like diagnostic model based on the predicted diagnostic results to obtain the optimal lightweight brain-like diagnostic model. Step Six: After preprocessing the patient's clinical data, feature extraction and fusion are performed using a brain-like mechanism to obtain the target fusion feature vector. The target fusion feature vector is then input into the optimal lightweight brain-like diagnostic model to obtain the final diagnostic result.
2. The method for early auxiliary diagnosis of neuropsychiatric diseases based on a brain-like multimodal large model according to claim 1, characterized in that, Step two specifically includes: Preliminary feature extraction is performed on the preprocessed data to obtain preliminary feature vectors for multiple modalities; wherein the multiple modalities include at least two modalities from image data, audio data, or text data; The preliminary feature vectors of multiple modalities are input into the leaked integral firing neuron model to obtain the time step feature vectors of the corresponding modalities. All time step feature vectors are directly concatenated in chronological order to obtain the historical global temporal feature vector of each modality. The historical global temporal feature vector of each modality is input into the multi-channel attention module to obtain the weighted features of each modality; The weighted features of all modalities are fused to obtain the fused features.
3. The method for early auxiliary diagnosis of neuropsychiatric diseases based on a brain-like multimodal large model according to claim 2, characterized in that, The multi-channel attention module includes a main branch and a sub-branch. The main branch includes an interpolation sub-module and a scaling sub-module. The sub-branch includes a linear transformation sub-module and a Softmax function. The interpolation submodule is used to adjust the size of the input historical global temporal feature vector; The linear transformation submodule is used to linearly transform the output of the interpolation submodule; The Softmax function is used to process the output of the linearly changing submodule; The scaling submodule is used to scale based on the outputs of the Softmax function and the interpolation submodule to obtain weighted features.
4. The method for early auxiliary diagnosis of neuropsychiatric diseases based on a brain-like multimodal large model according to claim 1, characterized in that, The brain-like multimodal hybrid expert diagnostic big model includes a knowledge graph, a main routing module, multiple expert sub-networks, a MoE output fusion module, and a diagnostic prediction layer. The knowledge graph is used to represent the relationships between diseases, symptoms, and examination methods; The main routing module is used to initially determine the major disease types based on input features and knowledge graphs, and select the corresponding expert subnetworks based on the major disease types; wherein, the main routing module includes a spiking neural network; Each of the expert subnetworks is used to determine the specific disease type and the weights of the expert subnetwork based on the input features; wherein each of the expert subnetworks includes a feedforward neural network and a gating network; The MoE output fusion module is used to obtain hybrid features based on the output of the expert subnetwork; The diagnostic prediction layer is used to process the mixed features to obtain the prediction result.
5. The method for early auxiliary diagnosis of neuropsychiatric diseases based on a brain-like multimodal large model according to claim 4, characterized in that, The MoE output fusion module is represented by the following formula: Among them, f MoE For mixed features, N is the number of expert subnetworks involved in the computation. Let e be the weight of the j-th expert subnetwork. j This is the output of the j-th expert subnetwork; The diagnostic prediction layer is represented by the following formula: y=σ(W p f tr +b o ) f tr =ψ(f MoE ) Where y is the prediction result, σ is the activation function, and W o and b o These are the weights and biases, f, respectively. tr For encoding features, ψ(·)Transformer encoder.
6. The method for early auxiliary diagnosis of neuropsychiatric diseases based on a brain-like multimodal large model according to claim 1, characterized in that, The process of using knowledge distillation and compression to process the large-scale brain-like multimodal hybrid expert diagnostic model yields a lightweight brain-like diagnostic model, specifically including: Knowledge distillation was performed on a large-scale brain-inspired multimodal hybrid expert diagnostic model to obtain a lightweight clinical brain-inspired diagnostic model. During training, a quantization function is introduced to map the weights in the lightweight clinical brain-like diagnostic model from 32-bit floating-point numbers to 8-bit integer representations. Singular value decomposition was performed on key layers in a lightweight clinical brain-like diagnostic model, and low-rank matrices were selected to be retained. By jointly optimizing all losses and updating the parameters of the lightweight clinical brain-like diagnostic model until convergence, the final lightweight brain-like diagnostic model is obtained.
7. The method for early auxiliary diagnosis of neuropsychiatric diseases based on a brain-like multimodal large model according to claim 1, characterized in that, Step five specifically includes: A reward signal is constructed based on the predicted diagnostic results and the clinical diagnostic results. The reward signal is expressed by the following formula: Where sim(·) represents the similarity function. R is the reward amplification factor, R is the reward value, and Y is the predicted diagnostic result. clin This is a clinical diagnostic result; Using the reward signal as feedback, the parameters of the lightweight brain-inspired diagnostic model are updated using a Monte Carlo-based policy gradient reinforcement learning algorithm. The update formula is expressed as follows: Where θ is a parameter. For learning rate, Let P(Y|x;θ) be the gradient operator for P(Y|x;θ), where P(Y|x;θ) is the probability of the predicted result under parameter θ. By updating the parameters of a lightweight brain-inspired diagnostic model using a Monte Carlo-based policy gradient reinforcement learning algorithm, and simultaneously performing joint optimization using supervised learning loss, the optimal lightweight brain-inspired diagnostic model is obtained.
8. An early auxiliary diagnostic system for neuropsychiatric diseases based on a brain-like multimodal large model, characterized in that, include: The data collection unit is used to collect and preprocess patients' clinical history data. The fusion unit is used to extract and fuse features from preprocessed data using a brain-like mechanism to obtain a fused feature vector. The model building unit is used to construct a large-scale brain-like multimodal hybrid expert diagnostic model, and to process the large-scale brain-like multimodal hybrid expert diagnostic model using knowledge distillation and compression to obtain a lightweight brain-like diagnostic model. The prediction unit is used to input the fused feature vector into the lightweight brain-like diagnostic model to obtain the prediction and diagnosis results. The training unit is used to retrain the lightweight brain-like diagnostic model based on the predicted diagnostic results to obtain the optimal lightweight brain-like diagnostic model. The diagnostic unit is used to preprocess the patient's clinical data and then use a brain-like mechanism to extract and fuse features to obtain a target fused feature vector. The target fused feature vector is then input into the optimal lightweight brain-like diagnostic model to obtain the final diagnostic result.
9. An electronic device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the method for early auxiliary diagnosis of neuropsychiatric diseases based on a brain-like multimodal large model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device where the storage medium is located to perform the early auxiliary diagnosis method for neuropsychiatric diseases based on a brain-like multimodal large model as described in any one of claims 1 to 7.
Citation Information
Cited By
A mental illness recognition method combining electronic medical records and electroencephalogram signals
CN122314351A