Knowledge and data fusion driven interpretable primary headache auxiliary identification method
By combining fuzzy logic and KAN feature fusion network, the problems of rigid knowledge expression and insufficient interpretability in the diagnosis of primary headache are solved, realizing accurate and interpretable auxiliary identification of primary headache, and improving the adaptability and credibility of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF SHANGHAI FOR SCI & TECH
- Filing Date
- 2026-01-15
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies for diagnosing primary headaches suffer from rigid knowledge representation and poor adaptability. Furthermore, deep learning models lack interpretability, making them difficult to apply in high-risk clinical scenarios and unable to effectively handle complex scenarios with ambiguous and overlapping symptoms.
We employ a knowledge and data fusion-driven approach, constructing an interpretable mapping of the ICHD-3 diagnostic criteria through fuzzy logic, and combining it with a KAN feature fusion network to enhance the model's ability to identify complex symptom patterns, outputting interpretable headache types and their confidence levels.
It enables accurate and interpretable auxiliary identification of primary headache, shortens the diagnosis time, improves the robustness and interpretability of the model, and can be widely applied in clinical settings.
Smart Images

Figure CN122135924A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart healthcare technology, and in particular to an interpretable primary headache auxiliary identification method driven by knowledge and data fusion. Background Technology
[0002] Primary headache is one of the most common chief complaints in neurology. The International Classification of Headache Disorders, 3rd edition (ICHD-3) provides a structured and systematic standard for the diagnosis of primary headache, detailing the diagnostic criteria for each type of headache, covering multiple dimensions such as frequency of headache attacks, duration, nature of pain, location and distribution, accompanying symptoms, triggering factors, and relieving factors. However, these diagnostic criteria are mostly expressed in qualitative and descriptive language, such as "moderate to severe pain," "often accompanied by nausea and / or vomiting," and "headache aggravated by daily activities." This ambiguity and subjectivity can easily lead to interpretation bias in actual diagnosis, especially when symptoms are atypical or have overlapping characteristics, easily resulting in misdiagnosis or missed diagnosis.
[0003] Currently, existing auxiliary diagnostic models and systems for primary headaches can be mainly divided into two categories technically: one is expert systems developed based on ICHD3. These systems rely on rule-based knowledge representation, which suffers from rigid knowledge expression and cannot adapt to complex and changing clinical conditions. The other is data-driven methods. With the continuous development of artificial intelligence technology in the medical field, data-driven auxiliary diagnostic models have become an important direction for intelligent headache identification. Traditional machine learning methods (such as decision trees and support vector machines) have improved diagnostic efficiency to some extent, but their performance is often limited because their model structures cannot clearly express flexible and complex clinical knowledge, especially when faced with problems such as ambiguous symptom boundaries and unclear semantic descriptions. In recent years, the rapidly developing deep learning models have achieved remarkable results in tasks such as medical image recognition and pathological classification, and they have powerful capabilities in multi-layer nonlinear feature extraction. However, deep neural network models generally suffer from the "black box" problem, making it difficult to explain the logical basis of their predictions to doctors, which limits their practical application value in high-risk clinical scenarios.
[0004] Chinese patent application CN120954735A discloses a method, system, device, and medium for predicting migraine attacks in advance. It establishes a migraine attack prediction model based on the Gradient Boosting algorithm. The input of the prediction model consists of key features, and the output is the probability prediction of a migraine attack. The interpretability analysis of the prediction model's output is performed using the SHAP method. The core objective of this invention is to predict the probability of a migraine attack in the next 48 hours. This is a single-disease time-dimensional prediction, only applicable to patients with confirmed or suspected migraines, and cannot cover the broad clinical needs of primary headaches. Furthermore, the single data-driven architecture of Gradient Boosting + SHAP cannot handle the complex scenarios of ambiguous and overlapping symptoms in primary headaches. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of existing technologies by providing a knowledge- and data-driven, interpretable auxiliary identification method for primary headache. On one hand, it constructs an interpretable mapping of the ICHD-3 diagnostic criteria using fuzzy logic, enabling the model to embed clinical knowledge. On the other hand, it extracts high-order nonlinear relationships from symptom data using KAN, improving the model's ability to identify complex symptom patterns. The combination of these two approaches will provide a solution for intelligent headache diagnosis that combines accuracy, robustness, and interpretability, potentially promoting the feasibility and reliability of auxiliary diagnostic systems in real-world clinical settings.
[0006] The objective of this invention can be achieved through the following technical solutions: A knowledge and data fusion-driven, interpretable method for the auxiliary identification of primary headache, the method comprising: Information on headache patients is collected and preprocessed to obtain multidimensional feature vectors. These multidimensional feature vectors are then input into a pre-trained interpretable primary headache auxiliary recognition model to obtain a probability distribution of the headache recognition results. The headache type with the highest probability in the probability distribution and its recognition confidence are output, and a risk warning is generated based on the headache type with the highest probability. The interpretable primary headache auxiliary identification model includes a hierarchical fuzzy logic layer, a KAN feature fusion network, and a classifier. The hierarchical fuzzy logic layer calculates the fuzzy membership degree of each feature when it belongs to different headache types based on a multi-dimensional feature vector and a differentiated membership function for different headache types, thereby obtaining fuzzy membership degree features. The KAN feature fusion network fuses the fuzzy membership features with the original multi-dimensional feature vector to obtain classification features. The classifier performs linear transformation and Softmax function transformation on the classification features to output the probability distribution of headache types.
[0007] Furthermore, the headache patient information includes demographic information, headache attack characteristics, triggering / relieving factors, accompanying symptoms, autonomic nervous system-related symptoms, aura and lifestyle, family history, prodromal symptoms, and education level.
[0008] Furthermore, the demographic information includes gender and age; The characteristics of the headache attacks include location, nature, severity, duration, and frequency; The triggering / relieving factors include whether daily activities worsen symptoms and whether there are clear ways to relieve them; The accompanying symptoms include nausea, vomiting, photophobia, and phonophobia; The autonomic symptoms include nasal congestion and runny nose, conjunctival congestion or tearing, facial sweating, eyelid edema / ptosis, and miosis; The warning signs and lifestyle factors include whether one smokes, drinks alcohol, and exercises regularly.
[0009] Furthermore, the preprocessing includes: Real-time verification of data integrity; if information is missing beyond a preset threshold, prompt for additional data collection and end the current recognition process. If the data is complete, remove data on primary headaches that are not the target type and exclude follow-up cases; The qualitative description of headache patient information is mapped to VAS scores, the graded description of headache patient information is encoded into corresponding values according to preset coding rules, and outlier identification and noise removal are performed on the mapped or encoded data to obtain a multidimensional feature vector consistent with the input dimension of the interpretable primary headache auxiliary identification model.
[0010] Furthermore, the hierarchical fuzzy logic layer includes a fine-grained fuzzy layer, a network layer, and a KAN linear layer; The fine-grained fuzzy layer is initialized with differences based on key indicators of different headache types, and the fuzzy membership degree of each feature when it belongs to different headache types is calculated by the differential membership function. The network layer calculates fuzzy membership degrees based on fine-grained fuzzy layers, performs fine-grained weighting on multi-dimensional feature vectors, and performs feature mapping through KAN linear layers to output fuzzy membership degree features.
[0011] Furthermore, the headache types include migraine, tension headache, and cluster headache; The differential membership function sets a steep rising interval for migraine features. When the migraine feature enters the steep rising interval, the fuzzy membership degree is raised from a low level to a high level. When the migraine feature does not enter the steep rising interval, the fuzzy membership degree is kept at a low level. The differential membership function is designed for tension headaches and uses a symmetric trapezoidal function to calculate fuzzy membership. The differential membership function sets a sensitive response interval for the cluster headache feature of cluster headache. The trigger threshold of the sensitive response interval is small. When the cluster headache feature enters the sensitive response interval, the membership degree rises rapidly.
[0012] Furthermore, the migraine features include nausea, vomiting, photophobia, and phonophobia, and the cluster headache features include nasal congestion and runny nose, conjunctival congestion or tearing, facial sweating, eyelid edema / ptosis, and pupillary constriction.
[0013] Furthermore, the KAN feature fusion network fuses fuzzy membership features with the original multidimensional feature vector to obtain classification features. The specific process includes: Using the GroupFeatures function, the membership degree of fuzzy membership features is grouped and mapped back to the feature space through KAN linear transformation, and the fuzzy membership features are converted into feature representations with the same dimension as the original multidimensional feature vectors. The feature representations are then multiplied element-wise with the original multidimensional feature vectors to obtain enhanced features. The nonlinear combination relationship between the enhanced features is extracted through KAN network mining to obtain classification features.
[0014] Furthermore, after outputting the headache type with the highest probability in the probability distribution and its recognition confidence, the system also outputs the fuzzy membership distribution of the multidimensional feature vector corresponding to the headache patient information under the headache type with the highest probability, based on the fuzzy membership degree of each feature output by the hierarchical fuzzy logic layer when it belongs to different headache types.
[0015] Furthermore, after training, the performance of the explained primary headache auxiliary identification model is measured and tested using sensitivity, specificity, positive predictive value, negative predictive value, Cohen's Kappa value, and Youden index.
[0016] Compared with the prior art, the beneficial effects of the present invention include: 1. This invention utilizes an interpretable primary headache auxiliary identification model to achieve auxiliary identification of interpretable primary headaches, and can output the probability of headache type based on the headache patient's information. In this invention, a dedicated trapezoidal membership function is customized for each clinical feature of primary headache, which can achieve accurate quantitative mapping of qualitative diagnostic rules, quickly output diagnostic suggestions and confidence levels, and help users shorten the identification and qualitative time of headache diseases. The model formed by the combination of multiple modules achieves synergistic gain of knowledge and data, taking into account both clinical priors and data patterns. It not only solves the problems of rigid knowledge expression and poor adaptability of traditional expert systems, but also makes up for the shortcomings of pure data-driven models that lack clinical logic constraints and are prone to de-clinical fitting.
[0017] 2. The trapezoidal membership function of this invention incorporates the contribution of features to diagnostic results into the design of fuzzy rules, making the rules both in line with clinical standards and adapted to data patterns, thereby improving the rationality and pertinence of the rules. Through the weighted calculation of the original features and fuzzy membership, fine-grained fuzzy features are generated, laying the foundation for subsequent feature extraction and solving the problem of adapting clinical knowledge to the original data.
[0018] 3. This invention combines hierarchical fuzzy logic with a KAN network. The fuzzy layer provides clinical prior constraints, while the KAN network mines complex correlations in the data. The two complement each other, and the nonlinear transformation of the KAN network provides both strong feature extraction capabilities and explicit analysis of feature response curves, solving the problem of the trade-off between high performance and interpretability. By mapping fuzzy features to the original feature space through the GroupFeatures function, the element-wise multiplication of two features is enhanced, improving feature discriminativeness and capturing more key information than models with single feature inputs.
[0019] 4. The output of this invention includes interpretable information such as membership degree visualization, enabling users to clearly understand the recognition logic and intuitively know the recognition basis. Attached Figure Description
[0020] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a distribution diagram of some attribute values of the present invention; Figure 3 This is a diagram illustrating the overall network architecture of the primary headache auxiliary identification model of this invention. Figure 4 This invention provides an explanation for the loss curve of the primary headache auxiliary identification model. Figure 5 This invention provides a membership distribution diagram of key factors for predicting migraine using an auxiliary identification model for primary headache. Figure 6 This invention provides a membership distribution diagram of key factors for predicting cluster headaches using an auxiliary identification model for primary headaches. Figure 7 This invention provides a membership distribution diagram of key factors for predicting tension headaches using an auxiliary identification model for primary headaches. Figure 8 This invention explains the top 10 features with the highest responsivity in the KAN network, an auxiliary identification model for primary headaches. Figure 9 This invention explains the top 10 characteristics of the response peaks in the KAN network, an auxiliary identification model for primary headaches. Figure 10This invention provides an interpretable primary headache auxiliary identification model that predicts the distribution of key factors in different headache SHAPs. Figure 11 This is a line graph summarizing the ablation experiment results of the model of this invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0022] Example 1 This embodiment discloses a knowledge and data fusion-driven, interpretable primary headache auxiliary identification method, the method as follows: Figure 1 As shown, steps S1-S3 are included, and the specific details of each step are as follows: Step S1: Collect information on headache patients and obtain multidimensional feature vectors after preprocessing.
[0023] Data collection methods include questionnaires and face-to-face consultations.
[0024] This method constructs a structured questionnaire covering 31 variables based on ICHD3 standards and the experience of headache experts.
[0025] Information on headache patients includes demographic information, headache attack characteristics, triggering / relieving factors, accompanying symptoms, autonomic symptoms, aura and lifestyle, family history, prodromal symptoms and education level.
[0026] Demographic information includes sex and age; Headache attack characteristics include location, nature, intensity, duration, and frequency; Triggering / relieving factors include whether daily activities worsen symptoms and whether there are clear ways to relieve them; Accompanying symptoms include nausea, vomiting, photophobia, and phonophobia; Autonomic symptoms include nasal congestion and runny nose, conjunctival congestion or tearing, facial sweating, eyelid edema / ptosis, and miosis; Warning signs and lifestyle factors include whether one smokes, drinks alcohol, and exercises regularly.
[0027] Primary headache types specifically include migraine, tension headache, and cluster headache.
[0028] Diagnostic labels strictly follow ICHD-3 and are uniformly classified into three categories: tension headache (2.1–2.4), migraine (1.1, 1.2, 1.3, 1.5), and cluster headache (3.1, 3.5.1). Variable coding combines structured discretization and numerical methods: for example, gender (1 female / 2 males), headache location (0 blanks / 1 unilateral / 2 bilateral), headache nature (0 blanks / multi-level categories), and headache severity are coded using VAS 1–10 points, with some items using binary coding (no: 1 / yes: 2). Education level and diagnostic labels are separately coded in a hierarchical manner (e.g., Diagnosis: TTH=0, Migraine=1, CH=2).
[0029] Preprocessing includes: Real-time verification of data integrity; if information is missing beyond a preset threshold, prompt for additional data collection and end the current recognition process. If the data is complete, remove data on primary headaches that are not the target type and exclude follow-up cases; The qualitative description of headache patient information is mapped to VAS scores, the graded description of headache patient information is encoded into corresponding values according to preset coding rules, and outlier identification and noise removal are performed on the mapped or encoded data to obtain a multidimensional feature vector consistent with the input dimension of the interpretable primary headache auxiliary identification model.
[0030] Specifically, after removing data on primary headaches that are not the target type, the rules for excluding follow-up cases are as follows: Intracranial neuropathies and other facial pain, secondary headaches, non-cluster headaches, trigeminal autonomic headaches and other primary headaches; for multiple diagnoses, only the "most troubling headache" is retained; follow-up cases are deleted and only the first visit record is retained, while for "drug overuse headache combined with chronic headache", only the chronic headache diagnosis is retained according to the rules.
[0031] To visually examine the differences in characteristics among the three types of headache, key variables such as age, disease duration, and headache severity were further visualized, such as... Figure 2 As shown in the figure. The results show that there are considerable differences in the distribution patterns of the above variables among the three types of headaches, providing intuitive clues and feature selection criteria for subsequent modeling.
[0032] Step S2: Input the multidimensional feature vector into the pre-trained interpretable primary headache auxiliary recognition model (Fuzzy-KAN model) to obtain the probability distribution of the headache recognition results.
[0033] The interpretable primary headache auxiliary identification model includes a hierarchical fuzzy logic layer, a KAN feature fusion network, and a classifier. Its network structure is as follows: Figure 3 As shown.
[0034] Figure 3The flowchart on the right shows the overall structure of the interpretable primary headache auxiliary identification model. SHAP&ICHD Guided Fuzzy Layer / SHAP&ICHD Guided Fuzzy Encoding are hierarchical fuzzy logic layers, their specific structures corresponding to the parts within the dashed boxes. Input is the input layer, Norm is the normalization layer of the interpretable primary headache auxiliary identification model, KAN is the KAN linear layer in the KAN feature fusion network of the interpretable primary headache auxiliary identification model, mining the high-order nonlinear relationship between fuzzy features and original features, SiLU is the activation unit between KAN layers, introducing nonlinearity to enhance feature expressiveness and avoid gradient vanishing, and Softmax is the classifier component, mapping the KAN layer features to the probability distribution of three types of headaches. This is the final output of the model, namely the classification and diagnosis result of primary headache.
[0035] Figure 3 Within the dashed box, Trapezoidal Membership is the trapezoidal membership function, KAN-Linear is the KAN linear layer, Sigmoid is the Sigmoid activation function, and ICHD Guided-A / B / C / D Fuzzy Block is the ICHD-guided A / B / C / D fuzzy block, corresponding to different types of primary headaches. Each fuzzy block has its own fuzzy rules designed based on the ICHD standard. SHAP_weights: 0-1 is the SHAP weight (value 0-1), which determines the weight allocation of input features to each ICHD fuzzy block. The combination of clinical knowledge and data patterns is strengthened based on the contribution of SHAP features.
[0036] The hierarchical fuzzy logic layer is based on multi-dimensional feature vectors. It calculates the fuzzy membership degree of each feature when it belongs to different headache types according to the differential membership degree function of different headache types, and obtains fuzzy membership degree features.
[0037] The KAN feature fusion network fuses fuzzy membership features with the original multidimensional feature vector to obtain classification features.
[0038] The classifier performs linear transformation and Softmax function transformation on the classification features, and outputs the probability distribution of headache type.
[0039] The hierarchical fuzzy logic layer is the foundation of the entire model. Its main function is to transform clinical knowledge from the ICHD-3 standard into computable fuzzy rules. It incorporates the normalized contribution values (SHAP - SHapley Additive exPlanations) of various attributes obtained from data modeling during the prediction process as rule guides, providing a basis for subsequent feature processing and classification. Specifically, this layer uses a trapezoidal membership function to construct a fine-grained fuzzy layer, performing hierarchical differentiated initialization of key indicators for different headache types.
[0040] The hierarchical fuzzy logic layer includes a fine-grained fuzzy layer, a network layer, and a KAN linear layer.
[0041] The fine-grained fuzzy layer performs differential initialization based on key indicators of different headache types, and calculates the fuzzy membership degree of each feature when it belongs to different headache types through the differential membership function.
[0042] The network layer calculates fuzzy membership degrees based on fine-grained fuzzy layers, performs fine-grained weighting on multi-dimensional feature vectors, and performs feature mapping through KAN linear layers to output fuzzy membership degree features.
[0043] The differential membership function sets a steep rising interval for migraine features. When a migraine feature enters the steep rising interval, the fuzzy membership is raised from a low level to a high level. When a migraine feature does not enter the steep rising interval, the fuzzy membership is kept at a low level. Migraine features include nausea, vomiting, photophobia, and phonophobia.
[0044] A prominent characteristic of migraines is that they are often accompanied by strong symptoms such as nausea and photophobia, which, once they appear, tend to be severe and develop rapidly. By setting a steep ascending range, the membership degree of indicators such as nausea and photophobia will rise rapidly when they reach a certain level, accurately capturing these typical migraine symptoms. For example, when a patient experiences significant nausea, the membership degree of the nausea indicator will quickly rise from a low value to a high value, thus highlighting the characteristics of migraines.
[0045] For tension headaches, the differential membership function uses a symmetric trapezoidal function to calculate fuzzy membership.
[0046] Tension-type headaches typically exhibit relatively stable and persistent symptoms, with a relatively mild intensity and lack of dramatic fluctuations. Therefore, using a symmetrical trapezoidal function allows for a smoother change in membership degree, effectively reflecting this characteristic of tension-type headaches. For example, regarding headache intensity, when the intensity remains within a certain range, the membership degree changes slowly with increasing or decreasing intensity, without sudden jumps, which aligns with the actual symptom presentation of tension-type headaches.
[0047] The differential membership function sets a sensitive response interval for cluster headache features. The sensitive response interval has a low trigger threshold, and the membership degree increases rapidly when a cluster headache feature enters the sensitive response interval. Cluster headache features include nasal congestion and runny nose, conjunctival congestion or tearing, facial sweating, eyelid edema / ptosis, and pupillary constriction.
[0048] Cluster headaches are accompanied by significant autonomic symptoms, such as nasal congestion and tearing. These symptoms appear suddenly and are quite pronounced, so setting a sensitive response interval can make the model more sensitive to these symptoms. When indicators such as nasal congestion and tearing change slightly, the membership degree can react quickly, thus accurately identifying cluster headache attacks.
[0049] In the specific implementation process, for each feature index x, its trapezoidal membership function can be expressed as: in, , , , These are parameters adjusted according to the ICHD-3 criteria and the characteristics of different headache types. For example, for the nausea index in migraines, and The value of will cause the membership degree to quickly reach 1 when the nausea symptoms are more obvious, so as to reflect the characteristics of its steep rising range.
[0050] The network layer performs fine-grained weighting on the original data and obtains the network layer output through feature mapping via a KAN linear layer. The specific expression is as follows: in, Represents the first data entry. The feature values of each feature Represents the first data entry. A fuzzy membership value, This represents the fusion of fuzzy, fine-grained features for each data point. , These represent the SiLU activation values and the spline basis functions after B-spline feature projection, respectively, during the data processing through the KAN linear layer. This is the output of the fuzzy feature extraction layer.
[0051] The main task of the KAN feature fusion network is to effectively fuse the fuzzy membership features output by the hierarchical fuzzy logic layer with the original features to enhance the model's ability to express data features. This network is designed based on the Kolmogorov-Arnold representation theorem and can efficiently handle multivariate nonlinear relationships.
[0052] The KAN feature fusion network fuses fuzzy membership features with the original multidimensional feature vector to obtain classification features. The specific process includes: Using the GroupFeatures function, the membership degree groups of fuzzy membership features are mapped back to the feature space through KAN linear transformation, and the fuzzy membership features are converted into feature representations with the same dimension as the original multidimensional feature vectors. The feature representations are then multiplied element-wise with the original multidimensional feature vectors to obtain enhanced features. The nonlinear combination relationship between the enhanced features is extracted through KAN network mining to obtain classification features.
[0053] The KAN linear transformation maps group-level membership back to the feature space. The specific calculation expression is as follows: The GroupFeatures function can be represented as: here, It is a weight matrix. It is the bias vector. Through this linear transformation, the model converts the fuzzy membership features into feature representations with the same dimensions as the original features, and then performs element-wise multiplication with the original features to enhance the features.
[0054] The KAN network enhances features through its unique structure. The KAN network processes data to learn complex nonlinear relationships between different features, further uncovering potential information within the data. After processing by the KAN network, features are more effectively refined and transformed, providing more discriminative feature representations for subsequent classification tasks. The results of the KAN network calculation are shown below: KAN is composed of multiple KAN-Linear layers. These are the classification features extracted by the KAN network, which are then used by the classifier for classification.
[0055] The classifier is the final output module of the model. Its main function is to classify headache types based on the feature representation output by the KAN feature fusion network. This model constructs a three-class neural network, which includes a Softmax output layer.
[0056] The classifier receives features H_2 from the KAN feature fusion network as input, maps them to a three-dimensional output space through a linear transformation, and then transforms the output into a probability distribution through a softmax function, ultimately outputting the probability distribution of headache type. Specifically, assuming the weight matrix of the linear transformation is W and the bias vector is b, the output P of the classifier can be expressed as: in, (i = 1, 2, 3) represents tension headache, migraine, and cluster headache, respectively. This represents the characteristics of the input. The function is defined as follows: in, This is the j-th output value after linear transformation. Using the Softmax function, the model converts the linearly transformed output into a probability distribution, ensuring that the probability of each headache type is between 0 and 1, and the sum of the probabilities of all headache types is 1. In this way, the model can determine the headache type of the input sample based on the magnitude of the probability values; the headache type with the highest probability value is the model's classification result.
[0057] After training, the performance of the interpretable primary headache auxiliary identification model was measured and tested using sensitivity, specificity, positive predictive value, negative predictive value, Cohen's Kappa value, and Youden index.
[0058] In this embodiment, during the training of the interpretable primary headache auxiliary recognition model, the Adam optimizer is used, with the number of learning epochs set to 100, the learning rate to 1e-3, and the batch size to 32. At the same time, an early stopping mechanism is introduced to prevent the model from overfitting.
[0059] The dataset used to train the interpretable primary headache auxiliary identification model came from 12,532 consecutively recruited headache patients; the data was organized by doctors based on questionnaire case data and cleaned simultaneously.
[0060] The cleaning rules mainly included: removing cases where key information was missing, leading to an inability to make a diagnosis (missing information rate >30%), intracranial neuropathies and other facial pain, secondary headaches, trigeminal autonomic headaches (non-cluster headaches), and other primary headaches; for multiple diagnoses, only the "most troubling headache" was retained; follow-up cases were deleted, and only the initial visit record was retained; for "medication overuse headaches combined with chronic headaches," only the chronic headache diagnosis was retained according to the rules. After cleaning, a total of 1731 follow-up cases (focusing only on the initial visit), 847 cases with missing key diagnostic items, 13 cases of cranial neuralgia, 58 cases of secondary headaches, 101 cases of other headache disorders, and 643 cases of trigeminal autonomic headaches other than cluster headaches were excluded, resulting in a final sample of 9139 cases, including 1527 cases of tension headaches, 7011 cases of migraines, and 601 cases of cluster headaches.
[0061] Diagnostic tests can be used to determine whether an individual suspected of having a disease actually has one. In both the retrospective and prospective experiments, diagnostic tests were performed by comparing fuzzy logic-based diagnostic results with those of headache experts for each headache case. Statistical analysis was performed using SPSS for Windows, and the following metrics were used to measure the diagnostic performance of the proposed method: sensitivity, specificity, positive predictive value, negative predictive value, Cohen's Kappa index, and Youden index. The specific calculation formulas are shown below: in, This indicates one of the primary headache diseases that can be diagnosed by the model; TP, FN, TN, and FP represent the number of true positives, false negatives, true negatives, and false positives, respectively. Sensitivity refers to the proportion of individuals diagnosed with the disease under the "gold standard" criteria who are actually ill, and the proportion of individuals diagnosed with the disease under the "gold standard" criteria who are actually not ill, and the proportion of individuals diagnosed with the disease under the "gold standard" criteria who are actually negative, and the probability that an individual diagnosed as positive actually has the disease. The probability that an individual diagnosed as negative actually does not have the disease is also considered positive. The overall concordance rate (π) represents the degree of agreement between the assessed diagnostic method and the gold standard diagnostic method. The Youden index reflects the ability of the diagnostic method to distinguish between patients and non-patients. The false diagnosis rate (α), also known as the false positive rate, is the probability that a healthy individual is misdiagnosed as a patient by the assessed diagnostic method. The false negative rate (β), also known as the false negative rate, is the probability that a patient is misdiagnosed as a healthy individual by the assessed diagnostic method. In addition, a consistency test was performed, and Cohen's kappa value was calculated to assess the consistency between diagnoses. The consistency test not only indicates whether the two methods are consistent, but also assesses the degree of consistency by calculating the Cohen's kappa value. The kappa value is calculated as follows: in, It is the actual observed consistency rate. This represents the expected consistency rate. Different intervals are defined for the Kappa value: if kappa ≥ 0.85, the consistency is considered excellent; if 0.6 ≤ kappa < 0.85, the consistency is good; if 0.45 ≤ kappa < 0.6, the consistency is moderate; and if kappa < 0.45, the consistency is poor. This example uses a 5% significance level and a 95% confidence interval (CI) to assess the range of fluctuations in the observed values.
[0062] The constructed network was trained on the dataset and obtained the following results: Figure 4 The results shown are as follows: First, there is the loss curve obtained from model training, such as... Figure 4As shown in the loss curve, the training loss decreases rapidly from its initial value, dropping to around 0.25 by the 5th epoch. This indicates that the model can quickly capture feature information from the data and has a good fitting ability. As training progresses, the test loss gradually stabilizes in the 0.15-0.18 range after the 10th epoch. This stable test loss range demonstrates the model's good generalization ability, maintaining good performance on unseen data. The early stopping mechanism is triggered when the validation set smoothing loss fails to improve for five consecutive epochs, effectively preventing overfitting and ensuring the model's stability and reliability.
[0063] The learning rate decay analysis shows that using the default β parameters (0.9, 0.999) of the Adam optimizer ensures the stability of gradient estimation.
[0064] In terms of overall performance, the model is excellent. The accuracy reached 96.26% (95% CI: 0.9447–0.9628), indicating high accuracy in headache type classification. The Kappa coefficient was 89.84% (95% CI: 0.8458–0.8988), indicating high consistency between the classification results and the true labels.
[0065] From the perspective of category specificity indicators, cluster headaches achieved 100% correct classification (all samples were correctly classified within the corresponding sample size), with sensitivity, specificity, PPV, and NPV all at 1.0 (95% CI: 1.0–1.0). This indicates that the model has a very strong ability to identify cluster headaches and can accurately classify cluster headache samples correctly. The sensitivity for migraines was 97.88% (95% CI: 0.9704–0.9857), showing high recognition ability and accurately identifying most migraine samples. The specificity for tension headaches was 98.02% (95% CI: 0.9725–0.9868), indicating that the model can effectively exclude other types of headaches and correctly classify non-tension headache samples.
[0066] Step S3: Output the headache type with the highest probability in the probability distribution and its identification confidence level, and generate a risk warning based on the headache type with the highest probability.
[0067] For specific examples of risk warnings, see below: If it is a migraine: it suggests that "there is a small possibility of misdiagnosis as tension headache. It is recommended to further confirm whether the patient has a family history of migraines." If it is a tension headache: it suggests "there is a risk of missed diagnosis, and atypical migraines (such as migraines without obvious nausea symptoms) need to be ruled out"; If it is cluster headache: "The diagnosis has a very high confidence level. It is recommended to further verify the diagnosis by considering the frequency of attacks (e.g., daily attacks during the cluster phase)."
[0068] After outputting the headache type with the highest probability in the probability distribution and its recognition confidence, the system also outputs the fuzzy membership distribution of the multidimensional feature vector corresponding to the headache patient information under the headache type with the highest probability, based on the fuzzy membership degree of each feature output by the hierarchical fuzzy logic layer when it belongs to different headache types.
[0069] Examples of key feature distributions visualized through model fuzzy membership degrees include: Figures 5-7 As shown.
[0070] As shown in the distribution in the figure, in the Migraine group, 90% of the samples had a membership degree of nausea above the 0.8 threshold. This is consistent with the clinical characteristic that migraine is often accompanied by nausea, indicating that the model can accurately capture this typical feature of migraine. In the CH (cluster headache) group, the membership degree distribution of conjunctival congestion above the 0.7 threshold showed a significant peak, indicating that conjunctival congestion symptoms are more pronounced during cluster headache attacks, and the model has high sensitivity to this feature. In the TTH (tension headache) group, the membership degree distribution of headache intensity showed a gentle trapezoidal feature, consistent with the relatively stable pain intensity of tension headache.
[0071] The KAN network itself is interpretable. Each layer of this neural network uses a set of B-spline functions to perform nonlinear transformations on the input features, and then generates the network output through linear combination. This makes the response curve of each feature explicitly interpretable, and the trained model can directly show the impact of each feature on the model output. Therefore, in another embodiment, after outputting the headache type with the highest probability in the probability distribution and its recognition confidence, the top ten features with the highest response rates input to the KAN network layers are statistically analyzed, as shown in the example diagram. Figure 8 and Figure 9 As shown.
[0072] Figure 8 These are the top 10 features with the highest responsivity in the KAN network model. Figure 9 The top 10 features with peak response times in the KAN network model are shown in the figure. As can be seen, when predicting headaches, the KAN network responds quickly in the first layer to factors such as headache severity, nausea, vomiting, phonophobia, and whether the headache worsens with daily activities. These features highly align with clinical experience provided by headache experts and the key elements for diagnosing primary headaches in the ICHD-3 criteria. This indicates that the interpretable primary headache auxiliary identification model constructed using this method can quickly capture key features for headache diagnosis, and further validates the model's high clinical interpretability.
[0073] Next, the SHAP technique was used to perform interpretability analysis on the trained model, and the results are as follows: Figure 10 As shown in the results, the top five most important features in the model's predictions were, in order, headache duration, nausea, vomiting, headache severity, and photophobia. This ranking of importance aligns with clinical experience and the ICHD-3 criteria, further validating the model's interpretability—that is, the model prioritizes key features closely related to headache type during the classification process.
[0074] Example 2 This embodiment, based on Embodiment 1 above, discloses an example of the specific ablation experiment process of a knowledge and data fusion-driven interpretable primary headache auxiliary identification method, demonstrating the significant superiority of the proposed interpretable primary headache auxiliary identification model in assisting headache symptom identification. The interpretable primary headache auxiliary identification model proposed in this method will be referred to as a Fuzzy+MLP network below.
[0075] The ablation experiments were conducted using MLP networks, KAN networks, Fuzzy+MLP networks, randomly initialized RandFuzzy+KAN networks, and Fuzzy+KAN networks for comparison. The experimental results, trained on the same dataset, are shown in Table 2. Figure 11 As shown.
[0076] Table 1. Average results of 10-fold cross-validation for each ablation experiment model The ablation experiment results showed that the Fuzzy+KAN network exhibited the best performance, with an accuracy of 0.9584, significantly higher than that of the MLP (0.9442), indicating that the model has a stronger overall ability to distinguish between the three types of headaches. Sensitivity (0.9513) and Specificity (0.9618) were both the highest, demonstrating that the model can effectively identify different types of headaches, which is crucial for clinical diagnostic support. The Youden index comprehensively reflects the balance between sensitivity and specificity of the model, further validating its high efficiency in distinguishing headache types. The Kappa value indicates a high degree of consistency between the model's predictions and actual clinical diagnoses, demonstrating outstanding reliability.
[0077] Compared to other models, the MLP network performs relatively poorly across various metrics, highlighting the limitations of basic neural networks in handling complex headache classification tasks. While the KAN and Fuzzy+MLP networks show some improvement, they still lag behind the Fuzzy+KAN network in key metrics. The randomly initialized Fuzzy+KAN network shows slightly lower metrics in some areas, indicating that proper initialization has a significant impact on model performance.
[0078] Figure 11The line graph visually presents the differences between the models on different metrics. The Fuzzy+KAN network ranks highly in metrics such as Accuracy and Specificity, especially showing a significant advantage in Specificity, which is of great importance in reducing misdiagnosis (such as misclassifying other headaches as cluster headaches).
[0079] In summary, the Fuzzy+KAN network, by integrating the characteristics of fuzzy logic and KAN structure, achieves optimization across multiple dimensions, effectively improving the accuracy and reliability of auxiliary diagnosis of primary headache. It also verifies the model's superiority in complex clinical classification tasks and provides strong support for the application of interpretable artificial intelligence in headache diagnosis.
[0080] Example 3 Based on Embodiment 1, this embodiment provides an electronic device, including: one or more processors and a memory, wherein the memory stores one or more programs, the one or more programs including instructions for executing the aforementioned knowledge and data fusion-driven interpretable primary headache auxiliary identification method.
[0081] At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to implement the aforementioned knowledge and data fusion-driven interpretable primary headache auxiliary identification method. Of course, in addition to software implementation, this invention does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0082] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0083] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0084] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A knowledge and data fusion-driven, interpretable primary headache auxiliary identification method, characterized in that, The method includes: Information on headache patients is collected and preprocessed to obtain multidimensional feature vectors. These multidimensional feature vectors are then input into a pre-trained interpretable primary headache auxiliary recognition model to obtain a probability distribution of the headache recognition results. The headache type with the highest probability in the probability distribution and its recognition confidence are output, and a risk warning is generated based on the headache type with the highest probability. The interpretable primary headache auxiliary identification model includes a hierarchical fuzzy logic layer, a KAN feature fusion network, and a classifier. The hierarchical fuzzy logic layer calculates the fuzzy membership degree of each feature when it belongs to different headache types based on a multi-dimensional feature vector and a differentiated membership function for different headache types, thereby obtaining fuzzy membership degree features. The KAN feature fusion network fuses the fuzzy membership features with the original multi-dimensional feature vector to obtain classification features. The classifier performs linear transformation and Softmax function transformation on the classification features to output the probability distribution of headache types.
2. The knowledge and data fusion-driven method for auxiliary identification of interpretable primary headache according to claim 1, characterized in that, The information on headache patients includes demographic information, headache attack characteristics, triggering / relieving factors, accompanying symptoms, autonomic symptoms, aura and lifestyle, family history, prodromal symptoms and education level.
3. The knowledge and data fusion-driven method for auxiliary identification of interpretable primary headache according to claim 2, characterized in that, The demographic information includes gender and age; The characteristics of the headache attacks include location, nature, severity, duration, and frequency; The triggering / relieving factors include whether daily activities worsen symptoms and whether there are clear ways to relieve them; The accompanying symptoms include nausea, vomiting, photophobia, and phonophobia; The autonomic nervous system-related symptoms include nasal congestion and runny nose, conjunctival congestion or tearing, facial sweating, eyelid edema / ptosis, and miosis; The warning signs and lifestyle factors include whether one smokes, drinks alcohol, and exercises regularly.
4. The knowledge and data fusion-driven method for auxiliary identification of interpretable primary headache according to claim 1, characterized in that, The preprocessing includes: Real-time verification of data integrity; if information is missing beyond a preset threshold, prompt for additional data collection and end the current recognition process. If the data is complete, remove data on primary headaches that are not the target type and exclude follow-up cases; The qualitative description of headache patient information is mapped to VAS scores, the graded description of headache patient information is encoded into corresponding values according to preset coding rules, and outlier identification and noise removal are performed on the mapped or encoded data to obtain a multidimensional feature vector consistent with the input dimension of the interpretable primary headache auxiliary identification model.
5. The knowledge and data fusion-driven method for auxiliary identification of interpretable primary headache according to claim 1, characterized in that, The hierarchical fuzzy logic layer includes a fine-grained fuzzy layer, a network layer, and a KAN linear layer; The fine-grained fuzzy layer is initialized with differences based on key indicators of different headache types, and the fuzzy membership degree of each feature when it belongs to different headache types is calculated by the differential membership function. The network layer calculates fuzzy membership degrees based on fine-grained fuzzy layers, performs fine-grained weighting on multi-dimensional feature vectors, and performs feature mapping through KAN linear layers to output fuzzy membership degree features.
6. The knowledge and data fusion-driven method for auxiliary identification of interpretable primary headache according to claim 5, characterized in that, The headache types include migraine, tension headache, and cluster headache; The differential membership function sets a steep rising interval for migraine features. When the migraine feature enters the steep rising interval, the fuzzy membership degree is raised from a low level to a high level. When the migraine feature does not enter the steep rising interval, the fuzzy membership degree is kept at a low level. The differential membership function is designed for tension headaches and uses a symmetric trapezoidal function to calculate fuzzy membership. The differential membership function sets a sensitive response interval for the cluster headache feature of cluster headache. The trigger threshold of the sensitive response interval is small. When the cluster headache feature enters the sensitive response interval, the membership degree rises rapidly.
7. The knowledge and data fusion-driven method for auxiliary identification of interpretable primary headache according to claim 6, characterized in that, The migraine features include nausea, vomiting, photophobia, and phonophobia, while the cluster headache features include nasal congestion and runny nose, conjunctival congestion or tearing, facial sweating, eyelid edema / ptosis, and miosis.
8. The knowledge and data fusion-driven method for auxiliary identification of interpretable primary headache according to claim 1, characterized in that, The KAN feature fusion network fuses fuzzy membership features with the original multidimensional feature vector to obtain classification features. The specific process includes: Using the GroupFeatures function, the membership degree of fuzzy membership features is grouped and mapped back to the feature space through KAN linear transformation, and the fuzzy membership features are converted into feature representations with the same dimension as the original multidimensional feature vectors. The feature representations are then multiplied element-wise with the original multidimensional feature vectors to obtain enhanced features. The nonlinear combination relationship between the enhanced features is extracted through KAN network mining to obtain classification features.
9. The knowledge and data fusion-driven method for auxiliary identification of interpretable primary headache according to claim 1, characterized in that, After outputting the headache type with the highest probability in the probability distribution and its recognition confidence, the system also outputs the fuzzy membership distribution of the multidimensional feature vector corresponding to the headache patient information under the headache type with the highest probability, based on the fuzzy membership degree of each feature output by the hierarchical fuzzy logic layer when it belongs to different headache types.
10. The knowledge and data fusion-driven method for auxiliary identification of interpretable primary headache according to claim 1, characterized in that, After training, the performance of the explained primary headache auxiliary identification model was measured and tested using sensitivity, specificity, positive predictive value, negative predictive value, Cohen's Kappa value, and Youden index.
Citation Information
Patent Citations
Migraine attack advanced prediction method, system and device and medium
CN120954735A