Method and system for classifying attention deficit hyperactivity disorder and computer equipment
By constructing a functional connectivity matrix based on resting-state functional magnetic resonance images and multiple voting decisions, the problem of insufficient individual differences in ADHD classification was solved, and higher accuracy and reliable classification results were achieved.
Patent Information
- Application Number
- CN202510725828.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-19
AI Technical Summary
Existing attention deficit hyperactivity disorder (ADHD) classification methods fail to fully consider individual differences, resulting in insufficient classification accuracy and susceptibility to noise.
By constructing a functional connectivity matrix based on individual resting-state functional magnetic resonance images, obtaining similarity feature vectors, using a support vector machine classifier to train a classification model, and adopting multiple voting decisions and difference classification mechanisms for classification, the impact of individual noise is reduced and the classification accuracy is improved.
It improves the individual difference expression ability of ADHD classification and the anti-interference ability of the classification model, reduces the risk of misjudgment of a single dimension, and enhances the clinical practical value and reliability.
Smart Images

Figure CN120674062A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular to a classification method, system and computer equipment for attention deficit hyperactivity disorder. Background Art
[0002] Attention Deficit Hyperactivity Disorder (ADHD) is a common neurodevelopmental disorder characterized by inattention, hyperactivity, and impulsive behaviors that are disproportionate to age and developmental level. As research into the disease continues to deepen, accurate early classification is crucial for developing personalized treatment plans, predicting prognosis, and evaluating the effectiveness of interventions. However, the symptoms are complex and diverse, and there are individual differences. Different patients exhibit varying degrees and combinations of attention deficit, hyperactivity, and impulsivity, making accurate classification a challenge.
[0003] Currently, some classification methods use resting-state functional magnetic resonance imaging (rs-fMRI) to construct a whole-brain functional connectivity matrix, and achieve classification prediction of diseased individuals by directly using connection strength or extracting topological network features. However, classification methods based on functional connectivity are usually based on group-level data analysis, which makes it difficult to fully consider the heterogeneity between individuals. Summary of the Invention
[0004] In order to solve the problem that the existing classification method for attention deficit hyperactivity disorder cannot provide more sensitive and accurate classification accuracy for individual differences, the present invention provides a classification method, system and computer device for attention deficit hyperactivity disorder.
[0005] To solve the above technical problems, the present invention provides the following technical solutions: a method for classifying attention deficit hyperactivity disorder, the method comprising the following steps: obtaining a training set and a similarity feature vector of the training set, the training set comprising a normal group and a diseased group, the similarity feature vector of the training set comprising a similarity feature vector of the normal group and a similarity feature vector of the diseased group; Acquire an individual to be classified, match the individual to be classified with a normal individual in the normal group to form a normal pair to be classified, and match the individual to be classified with a diseased individual in the diseased group to form a diseased pair to be classified; acquire a normal pair similarity feature vector of the normal pair to be classified and a diseased pair similarity feature vector of the diseased pair to be classified; Constructing a classification model with a preset classifier, using the similarity feature vector of the normal group and the similarity feature vector of the diseased group as training data to train the classification model, and stopping the training when the preset classification accuracy is met to obtain a trained classification model; Inputting the normal pairs to be classified and the diseased pairs to be classified into the classification model, the classification model outputting prediction results for the normal pairs to be classified and the diseased pairs to be classified based on the similarity feature vectors of the normal pairs to be classified, the similarity feature vectors of the diseased pairs to be classified and the trained classification model; The output of the prediction results of the normal to-be-classified pairs and the diseased to-be-classified pairs is carried out through multiple voting decisions, and the multiple voting decisions include: voting on the prediction results of the individuals to be classified to obtain the classification results of the individuals to be classified; and when there is a tie, the classification results of the individuals to be classified are classified as diseased based on the difference classification mechanism.
[0006] Preferably, before obtaining the training set and the similarity feature vector of the training set, the following steps are included: obtaining a preset number of initial individuals, wherein the initial individuals include initial normal individuals and initial diseased individuals; matching the initial normal individuals and the initial diseased individuals based on the propensity score matching method; and grouping the matched pairs of the initial normal individuals and the initial diseased individuals into the normal group and the diseased group as the training set, and treating the unmatched initial individuals as independent individuals to be classified.
[0007] Preferably, obtaining a training set and a similarity feature vector of the training set includes the following steps: obtaining resting-state magnetic resonance images of the normal individuals in the normal group and the diseased individuals in the diseased group, and performing standardized preprocessing on the resting-state magnetic resonance images; performing brain region segmentation on the preprocessed resting-state magnetic resonance images using a preset brain map template, and generating functional connection matrices corresponding to different brain region pairs for the resting-state magnetic resonance images of the normal individuals and the diseased individuals; comparing the functional connection matrices of any two normal individuals in the normal group, obtaining the similarity feature vectors between the two normal individuals, and matching them to form a normal training pair; comparing the functional connection matrices of any two diseased individuals in the diseased group, obtaining the similarity feature vectors between the two diseased individuals, and matching them to form a diseased training pair.
[0008] Preferably, the similarity feature vector includes preset dimensions based on the preset brain map template; a classification model is constructed with a preset classifier, and the similarity feature vector of the normal group and the similarity feature vector of the diseased group are used as training data to train the classification model. When the training reaches a preset classification accuracy, the training is stopped to obtain a trained classification model, which includes the following steps: training a support vector machine classifier based on each of the preset dimensions of the similarity feature vector, calculating the overall classification accuracy through cross-validation, and stopping the training when the preset classification accuracy is reached; the similarity feature vector is represented by a feature space in the support vector machine classifier, and the support vector machine classifier is used to distribute the similarity feature vectors of the normal training pair and the similarity feature vectors of the diseased training pair at intervals in the feature space.
[0009] Preferably, the normal pairs to be classified and the diseased pairs to be classified are input into the classification model, and the classification model outputs prediction results for the normal pairs to be classified and the diseased pairs to be classified based on the similarity feature vectors of the normal pairs to be classified, the similarity feature vectors of the diseased pairs to be classified, and the trained classification model, including the following steps: The normal pairs to be classified and the diseased pairs to be classified are input into the classifier, and classification decisions are made based on the similarity feature vectors of the normal pairs to be classified and the similarity feature vectors of the diseased pairs to be classified with the preset classification accuracy, and prediction results for the normal pairs to be classified and the diseased pairs to be classified are output.
[0010] Preferably, the multiple voting decisions specifically include the following steps: if the normal pairs to be classified are distributed in the distribution area of the normal training pairs, the prediction result of the normal pairs to be classified is correct, and a normal vote is cast for the classification result of the individual to be classified; if the normal pairs to be classified are distributed in the distribution area of the diseased training pairs, the prediction result of the normal pairs to be classified is wrong, and no vote is cast for the classification result of the individual to be classified; If the diseased pair to be classified is distributed in the distribution area of the normal training pair, the prediction result of the diseased pair to be classified is wrong, and no vote is cast for the classification result of the individual to be classified; If the diseased pair to be classified is distributed in the distribution area of the diseased training pair, the prediction result of the diseased pair to be classified is correct, and a vote of diseased is cast for the classification result of the individual to be classified; based on the majority voting decision result, the votes are calculated to obtain the classification result of the individual to be classified.
[0011] Preferably, the propensity score matching method is caliper matching without replacement, and matching the initial normal individuals and the initial diseased individuals based on the propensity score matching method includes the following steps: calculating the propensity scores for the initial normal individuals and the initial diseased individuals based on preset confounding variables; taking the initial diseased individuals as the benchmark, matching the initial normal individuals with a preset caliper threshold without replacement; wherein, before each matching, the initial normal individuals are randomly arranged.
[0012] Preferably, the similarity feature vector between the two normal individuals or the similarity feature vector between the two diseased individuals is the Spearman correlation coefficient between the functional connectivity matrices.
[0013] To solve the above technical problems, the present invention provides another technical solution as follows: a classification system for attention deficit hyperactivity disorder, applied to the above-mentioned classification method for attention deficit hyperactivity disorder, the system comprising: a training set acquisition module, configured to construct a classification model using a preset classifier, train the classification model using the similarity feature vectors of the normal group and the similarity feature vectors of the diseased group as training data, and obtain a classification boundary separating the similarity feature vectors of the normal group and the similarity feature vectors of the diseased group; The module for obtaining individuals to be classified is used to match the individuals to be classified with normal individuals in the normal group to form normal pairs to be classified, and to match the individuals to be classified with diseased individuals in the diseased group to form diseased pairs to be classified; and to obtain similarity feature vectors of the normal pairs to be classified and similarity feature vectors of the diseased pairs to be classified; a training module, configured to construct a classification model using a preset classifier, train the classification model using the normal group similarity feature vector and the diseased group similarity feature vector as training data, and obtain a classification boundary separating the normal group similarity feature vector and the diseased group similarity feature vector; A classification prediction module is used to input the normal pairs to be classified and the diseased pairs to be classified into the classification model, and the classification model outputs prediction results for the normal pairs to be classified and the diseased pairs to be classified based on the similarity feature vectors of the normal pairs to be classified, the similarity feature vectors of the diseased pairs to be classified and the trained classification model; wherein, the output prediction results for the normal pairs to be classified and the diseased pairs to be classified are decided by multiple voting, and the multiple voting decisions include: voting on the prediction results of the individuals to be classified to obtain the classification results of the individuals to be classified; wherein, when the votes are tied, the classification results of the individuals to be classified are classified as diseased based on the difference classification mechanism.
[0014] In order to solve the above technical problems, the present invention provides another technical solution as follows: a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the various steps of the aforementioned method for classifying attention deficit hyperactivity disorder.
[0015] Compared with the prior art, the method, system, and computer device for classifying attention deficit hyperactivity disorder provided by the present invention have the following beneficial effects: 1. A classification method for attention deficit hyperactivity disorder provided by an embodiment of the present invention aims to solve the problem that traditional classification methods have large individual differences and are easily affected by noise. By establishing group similarity connections, that is, based on the similarity feature vectors of normal individuals in the normal group of the training set and diseased individuals in the diseased group, they are matched with the individuals to be classified respectively, shifting from the individual level to the group similarity level, and replacing the original individual features with the similarity features of the matching pair dimensions, reducing computational complexity and individual noise, and capturing the typical brain connection pattern between each diseased individual or each normal individual through the similarity feature vector, making the similarity feature more group discriminative; using the similarity feature vector of the training set to train the classifier, find the preset classification accuracy that meets the requirements for dividing diseased and normal individuals, and the classification accuracy under the setting avoids data overtraining. Fitting, the voting mechanism for the subjects to be classified adopts a double prediction voting mechanism, which constructs pairs of subjects to be classified by respectively comparing the normal individuals and diseased individuals in the training set with the subjects to be classified, that is, the input data, and then compares the similarity feature vectors of the normal pairs of subjects to be classified and the diseased pairs of subjects to be classified with the trained model, providing a reference benchmark of "normal mode" and "disease mode" for the subjects to be classified, so that the similarity calculation has a clear biological significance and reduces the risk of misjudgment in a single dimension; multiple voting judgments and tie-vote judgments improve the anti-interference ability of the classification model, and introduce a classification mechanism based on difference in the case of tie-votes, which has a stronger theoretical basis and practical rationality in uncertain judgment, improves the ability to identify critical individuals, meets clinical safety, and provides further basis for medical evaluation, thereby improving the clinical practical value of the model.
[0016] 2. In the embodiment of the present invention, not all of the initial individuals obtained can be used for classifier training. Propensity score matching constructs a logistic regression model to calculate the propensity score of each individual on the confounding variables, measure the probability of the individual belonging to the diseased group, and then pair individuals with similar propensity scores in different groups, and pair individuals with similar propensity scores in the diseased group and the control group. The matched pairs of initial normal individuals and initial diseased individuals are grouped into normal and diseased groups as training sets to train the classifier. Propensity score matching balances the influence of confounding variables between groups, improves the comparability between the two groups of data, and reduces selection bias. The remaining samples that do not participate in the matching are retained as an independent test set to verify the performance of the classifier training degree.
[0017] 3. In the embodiment of the present invention, the resting-state magnetic resonance image reflects the spontaneous neural activity of the brain in the resting state, and its functional connectivity matrix can reveal the coordination pattern between brain regions, providing objective physiological indicators for the study of brain mechanisms of neurodevelopmental disorders in diseased individuals. Standardization preprocessing includes time layer correction, head motion correction, spatial normalization, denoising and other steps to ensure the comparability of data from different subjects, reduce the interference of data acquisition and processing errors on subsequent analysis, ensure the objectivity and stability of the data, and ensure comparability between them; the resting-state magnetic resonance image after preprocessing is subjected to brain region segmentation to generate a functional connectivity matrix, that is, the brain is divided into multiple regions of interest based on a preset brain map, the time series correlation between different brain regions is calculated, and the functional connectivity matrix is generated, providing a structured brain The partitioning framework makes the calculation of functional connectivity biologically meaningful and interpretable. The functional connectivity matrix converts the brain functional network into computable numerical features, providing a basis for subsequent similarity analysis; the similarity features between any two normal individuals in the normal group and any two diseased individuals in the diseased group are calculated, and the corresponding matching forms normal training pairs and diseased training pairs. By comparing the functional connectivity matrices, similarity feature vectors representing the similarity of functional connectivity patterns between individuals are generated, which is convenient for capturing common patterns within the group, helping the classification model to learn the statistical laws within the two groups of people, and reducing the noise interference of single data points; compared with the individual functional connectivity matrix, the similarity features between training pairs pay more attention to relative differences, revealing the pathological mechanism at the group level, and improving the sensitivity of the classification model.
[0018] 4. In an embodiment of the present invention, the overall classification accuracy is calculated through cross-validation to ensure that the trained support vector machine classifier has the most discriminative power for classification. The preset classification accuracy reduces the complexity of model training, reduces redundancy between features, reduces the risk of overfitting, and makes the classification model more robust for predicting the individuals to be classified. The classifier uses a support vector machine. The support vector machine maximizes the interval between the two categories (normal training pairs and diseased training pairs) and the plane by finding a hyperplane. The similarity feature vectors of the normal training pairs and the diseased training pairs are input into the support vector machine. The model can be trained to learn the distribution pattern of the two categories in the feature space. The normal training pairs and the diseased training pairs are clearly classified as normal or diseased based on the position of the similarity feature vectors in the feature space. The similarity feature vectors are represented in the classifier as having a feature space. Each dimension of the similarity feature vector corresponds to the coordinate in the feature space, so that the relationship between the feature vectors can be visualized and analyzed in the geometric space. The interval distribution reflects the degree of separation of samples of different categories in the feature space, which helps to understand how the classification model distinguishes the similarity feature vectors of normal training pairs and diseased training pairs.
[0019] 5. In an embodiment of the present invention, the similarity feature vectors of the normal pairs to be classified and the diseased pairs to be classified are used as input data of the classifier. The classifier can make accurate classification decisions for the normal pairs to be classified and the diseased pairs to be classified, and use the preset classification accuracy as a standard to ensure the reliability of the classification results, providing a valuable reference basis for clinical diagnosis or research; the classification model captures the relative differences between them, and the similarity feature vectors reflect the differences in brain network patterns between the normal pairs to be classified and the diseased pairs to be classified, while the support vector machine classifier maps such differences to the decision space of the classifier, so that the abstract neuroimaging features are converted into quantifiable classification basis, and the individual differences are converted into the distribution degree of the pairs to be classified in the feature space, thereby realizing inclusive classification of heterogeneity; by quantifying the classification accuracy and capturing group differences and individual heterogeneity, interpretable inference from neuroimaging data to disease classification is realized.
[0020] 6. In this embodiment of the present invention, a dual matching voting mechanism is implemented for normal pairs and diseased pairs. By matching with individuals with known disease conditions within the two groups, individual characteristics are converted into the degree of deviation from the normal training pairs and the degree of fit with the diseased training pairs. This avoids the limitations of a single classification boundary. The connection pattern of a single diseased individual may not represent the group characteristics. Using matching pairs can capture group commonalities and improve the classifier's tolerance for heterogeneous individuals. Voting is performed to judge the correctness of the prediction results. Only when the prediction result is consistent with the category trained by the classifier, that is, the normal pair to be classified falls within the distribution area of its corresponding normal training pair or the diseased pair to be classified falls within the distribution area of its corresponding diseased training pair, is the vote counted as valid. Otherwise, it is considered invalid and no vote is cast. This filters out unreliable predictions, avoids noise interference in the final decision, strengthens group pattern consistency, ensures that the classification result is based on matching the prediction pattern, and not on accidental fluctuations, and reduces the impact of accidental errors and quantifies classification uncertainty. At the same time, multiple voting results support dynamic evaluation, monitoring individual changes over time, such as before and after treatment, and evaluating the effectiveness of intervention by comparing changes in voting proportions.
[0021] 7. In an embodiment of the present invention, the propensity score matching method uses caliper matching without replacement to balance confounding variables such as age, gender, head motion parameters, and other non-brain function differences between the initial normal individuals and the initial diseased individuals, reduce inter-group bias, and make the initial individuals more statistically comparable. Matching without replacement ensures that each initial normal individual participates in only one matching, avoiding repeated use that may lead to redundancy or bias in the training set data. Caliper matching further limits the similarity of matching by setting a caliper threshold, thereby improving matching quality. Before each matching, the initial normal individuals are randomly arranged into normal groups to eliminate order-dependent bias and enhance the stability of the results.
[0022] 8. In an embodiment of the present invention, the Spearman correlation coefficient is a non-parametric statistical method used to measure the monotonic relationship between two variables, that is, the consistency of the trends between the variables, rather than a strict linear relationship. By calculating the Spearman correlation coefficient of the functional connectivity matrices of the brain regions of two individuals, the consistency of the brain network connection patterns of the two individuals is characterized. The Spearman correlation coefficient is performed by rank transformation on the data, that is, converting the original data into a sorted rank, eliminating the influence of outliers and distribution patterns, and characterizing the trend consistency of the brain connection pattern.
[0023] 9. An embodiment of the present invention further provides an attention deficit hyperactivity disorder classification system, which is applied to the above-mentioned attention deficit hyperactivity disorder classification method. Therefore, it also has the same beneficial effects as the above-mentioned attention deficit hyperactivity disorder classification method, and will not be described in detail here.
[0024] 10. An embodiment of the present invention further provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the computer device implements the various steps of the above-mentioned method for classifying attention deficit hyperactivity disorder. Therefore, the computer device also has the same beneficial effects as the above-mentioned method for classifying attention deficit hyperactivity disorder, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is the process of the classification method for attention deficit hyperactivity disorder provided by the first embodiment of the present invention Figure 1 .
[0026] Figure 2 This is the process of the classification method of attention deficit hyperactivity disorder provided by the first embodiment of the present invention. Figure 2 .
[0027] Figure 3 This is a flowchart of model training and classification prediction of the attention deficit hyperactivity disorder classification method provided by the second embodiment of the present invention.
[0028] Figure 4 4 is a block diagram of a classification system for attention deficit hyperactivity disorder provided by a third embodiment of the present invention.
[0029] Figure 5 It is a block diagram of a computer device provided by a fourth embodiment of the present invention.
[0030] Figure 6 is the prediction accuracy of the classification result of the attention deficit hyperactivity disorder classification method provided by the first embodiment of the present invention Description of the accompanying drawings: 100. Attention Deficit Hyperactivity Disorder Classification System; 101. Training Set Acquisition Module; 102. To-Be-Classified Individual Acquisition Module; 103. Training Module; 104. Classification Prediction Module; 200. Computer device; 201. Memory; 202. Processor; 203. Computer program. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and implementation examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0032] In the embodiments provided herein, it should be understood that "B corresponding to A" means that B is associated with A and B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information.
[0033] It should be understood that references to "one embodiment" or "an embodiment" throughout this specification mean that specific features, structures, or characteristics associated with the embodiment are included in at least one embodiment of the present invention. Therefore, the appearance of "in one embodiment" or "in an embodiment" throughout this specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. Those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required for the present invention.
[0034] In various embodiments of the present invention, it should be understood that the size of the serial numbers of the above-mentioned processes does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0035] The flow charts and block diagrams in the accompanying drawings of the present invention illustrate the possible implementation architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementation schemes, the functions marked in the box can also occur in a different order than those marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which is determined based on the functions involved. It should be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0036] See also Figure 1 and Figure 3 The present invention provides a classification method for attention deficit hyperactivity disorder, comprising the following steps: Step S1, obtaining a training set and a similarity feature vector of the training set, wherein the training set includes a normal group and a diseased group, and the similarity feature vector of the training set includes a similarity feature vector of the normal group and a similarity feature vector of the diseased group; Step S2: obtaining individuals to be classified, matching the individuals to be classified with normal individuals in the normal group to form normal pairs to be classified, and matching the individuals to be classified with diseased individuals in the diseased group to form diseased pairs to be classified; obtaining similarity feature vectors of the normal pairs to be classified and similarity feature vectors of the diseased pairs to be classified; Step S3: constructing a classification model with a preset classifier, using the similarity feature vectors of the normal group and the similarity feature vectors of the diseased group as training data to train the classification model. When the training reaches a preset classification accuracy, the training is stopped to obtain a trained classification model; Step S4: input the normal pairs to be classified and the diseased pairs to be classified into the classification model, and the classification model outputs prediction results for the normal pairs to be classified and the diseased pairs to be classified based on the similarity feature vectors of the normal pairs to be classified, the similarity feature vectors of the diseased pairs to be classified, and the trained classification model; Among them, the output of the prediction results of the normal to be classified pairs and the diseased to be classified pairs adopts multiple voting decisions, and the multiple voting decisions include: voting on the prediction results of the individuals to be classified to obtain the classification results of the individuals to be classified; among them, when the votes are tied, the classification results of the individuals to be classified are classified as diseased based on the difference classification mechanism.
[0037] As a significant neurodevelopmental disorder affecting childhood, the diagnosis of Attention Deficit Hyperactivity Disorder (ADHD) primarily relies on clinical behavioral assessment scales and interview tools. However, these methods are subject to certain subjective assessments and rely heavily on the assessor's clinical experience, limiting their objectivity and reproducibility. Studies have shown that the onset of ADHD is closely associated with the default mode network (DMN), executive control network (ECN), and attention-related networks. Therefore, in ADHD classification research, existing methods have constructed whole-brain functional connectivity matrices based on resting-state functional magnetic resonance imaging (rs-fMRI) and predicted the classification of individuals with ADHD by directly using connection strength or extracting topological network features.
[0038] The present invention provides a classification method for attention deficit hyperactivity disorder (ADHD), based on a functional connectivity similarity modeling approach between individual resting-state functional magnetic resonance images. This method addresses technical issues such as insufficient sensitivity to individual differences and poor model generalization in existing ADHD classification methods. By constructing similarity feature vectors based on connectivity similarities between pairs of classifications, rather than traditionally representing connectivity features based on a single individual, the method uses similarity feature vectors between training pairs within a training set to generate high-dimensional features that reflect consistent patterns within different groups, thereby enhancing the ability to express heterogeneity between individuals.
[0039] It can be understood that the training set includes a data set for training the classification model, which contains samples with known labels such as normal group or disease group. The normal group represents a collection of normal individual samples, and the disease group represents a collection of diseased individual samples, providing learning training samples for the classification model, so that the classification model can distinguish normal and diseased individuals through the difference between the normal group similarity feature vector and the diseased group similarity feature vector; the similarity feature vector is a high-dimensional feature vector that describes the similarity of brain region functional connections between individual matching pairs, and each dimension corresponds to the functional connection strength of a specific brain region pair. By comparing the distribution differences of the similarity feature vectors of the normal group and the diseased group in each dimension, such as the different connection strengths of specific brain region pairs, the classifier achieves classification by learning these differences; the individual to be classified has an unknown label, and needs to be judged by the classification model as a normal or diseased individual. The individual to be classified is paired with the normal group individual to form a normal pair to be classified, and paired with the diseased group individual to form a diseased pair to be classified. In the classification model, the similarity feature vectors of the individual to be classified and the known category individuals are compared to output the prediction results.
[0040] Specifically, pre-selected classifiers such as support vector machines, logistic regression, and neural networks are used to construct classification models. The similarity feature vectors of the normal group and the diseased group are input as training data, and the model parameters are iteratively adjusted to stop training when the classification accuracy reaches the preset classification accuracy. Voting decision-making is an integrated learning strategy that votes on the prediction results of multiple sample pairs, and the majority vote determines the final classification category of the individual to be classified.
[0041] Among them, when making voting decisions, if the number of normal votes and sick votes is equal, the individual to be classified will be classified as sick through the differential classification mechanism, which can avoid missed diagnosis. It has a stronger theoretical basis and practical rationality in the judgment of uncertain samples and improves the ability to identify critical individuals.
[0042] See also Figure 2 Furthermore, before step S1, obtaining the training set and the similarity feature vector of the training set, the following steps are included: Step S01, obtaining a preset number of initial individuals, the initial individuals including initial normal individuals and initial diseased individuals; Step S02, matching the initial normal individuals and the initial diseased individuals based on the propensity score matching method; Step S03: matching the paired initial normal individuals and the initial diseased individuals into a normal group and a diseased group as training sets, and the unmatched initial individuals are used as independent individuals to be classified.
[0043] Understandably, at the beginning of the study, a certain number of initial individuals were obtained from the public database to provide sufficient sample size for subsequent analysis to avoid overfitting of the model or insufficient statistical power due to insufficient samples. The initial samples included initial normal individuals and initial diseased individuals of different ages, genders, and regions. The similarity feature vectors between individuals may be affected by factors such as age, gender, and region. Therefore, the propensity score matching method uses "simulated randomized controlled experiments" to make the normal individuals and diseased individuals in the initial individuals comparable in non-disease variables, thereby attributing the similarity feature differences between normal individuals and diseased individuals to the disease itself. In addition, the number of initial individuals obtained from the database is not balanced. The propensity matching method can balance the matching correspondence between diseased individuals and normal individuals, so that each diseased individual can be effectively utilized, thereby more accurately estimating the treatment and using it as a training set to train the classification model.
[0044] Furthermore, the propensity score matching method is caliper matching without replacement. In step 02, matching the initial normal individuals and the initial diseased individuals based on the propensity score matching method includes the following steps: Step 021, calculating the propensity score for the initial normal individuals and the initial diseased individuals based on the preset confounding variables; Step 022, taking the initial diseased individuals as the benchmark, the initial normal individuals are matched without replacement using a preset caliper threshold; wherein, before each matching, the initial normal individuals are randomly arranged.
[0045] Propensity score matching calculates a propensity score for each initial individual, which represents the probability that the initial individual has ADHD given the covariates. The initial normal individuals and the initial affected individuals are then matched based on the propensity score, making the two matched groups as similar as possible in terms of the covariate distribution.
[0046] Specifically, in the propensity score of the classification method for attention deficit hyperactivity disorder of the present invention, the defined covariates are age and gender. Taking the initial diseased individual as the reference benchmark, the two covariates of age and gender that may have a significant impact on the research results are comprehensively considered, and a logistic regression model is constructed to estimate the propensity score of each initial individual. The propensity score aims to give the conditional probability that the initial individual is assigned to the diseased group under the conditions of given age and gender. Based on the calculated propensity score, the present invention adopts a caliper matching method without replacement to match the initial normal individuals, that is, for each initial diseased individual, the initial normal individuals whose propensity score differences are within the preset caliper threshold range are searched among the initial normal individuals for matching, and once the initial normal individuals are matched during the matching process, they will no longer participate in subsequent matching to ensure the uniqueness and validity of the matching.
[0047] Optionally, the preset caliper threshold range in the present invention is set to 0.02.
[0048] Furthermore, the matching order of the initial normal individuals may affect the matching results, so the matching process was repeated 1000 times, and the order of the initial normal individuals was randomly disrupted before each matching, and finally the group with the best matching results was selected to construct the training set.
[0049] Among them, the remaining samples that did not participate in the matching are retained as an independent test set to verify the performance of the classifier training degree.
[0050] Furthermore, in step S1, obtaining the training set and the similarity feature vector of the training set includes the following steps: Step S11: obtaining resting-state magnetic resonance images of normal individuals in the normal group and diseased individuals in the diseased group, and performing standardization preprocessing on the resting-state magnetic resonance images. Step S12, performing brain region segmentation on the pre-processed resting-state MRI images using a preset brain atlas template, and generating functional connectivity matrices corresponding to different brain region pairs for the resting-state MRI images of normal individuals and diseased individuals; Step S13: comparing the functional connectivity matrices of any two normal individuals in the normal group, obtaining similarity feature vectors between the two normal individuals, and matching them to form a normal training pair; Step S14: compare the functional connectivity matrices of any two diseased individuals in the diseased group, obtain the similarity feature vectors between the two diseased individuals, and match them to form a diseased training pair.
[0051] Specifically, resting-state magnetic resonance images refer to brain images obtained through magnetic resonance imaging (MRI) technology when an individual is in a resting state, that is, not performing a specific task. They reflect the brain structure and functional state of the individual in a resting state and are often used to study brain functional connections. A series of preprocessing steps, such as denoising, alignment, and normalization, are performed on the acquired resting-state magnetic resonance images to eliminate differences caused by different individuals and different scanning parameters, so that the image data are comparable, thereby improving the accuracy and reliability of subsequent data analysis.
[0052] Furthermore, after obtaining the similarity feature vectors between two normal individuals or the similarity feature vectors between two diseased individuals, classification labels are assigned to these similarity feature vectors respectively.
[0053] For example, the normal training pair is marked as 1, and the diseased training pair is marked as 2.
[0054] The preset brain map template is a predefined standard template for brain region division, which is used to segment brain images into different brain regions. According to the brain map template, the preprocessed resting-state magnetic resonance images are segmented into different brain regions, and the functional connectivity relationship between different brain regions is analyzed to form a functional connectivity matrix; the functional connectivity matrix is a matrix that describes the functional connection strength between different brain regions. Each element in the matrix represents the functional connectivity value between two brain regions. The functional connectivity matrix quantitatively analyzes the functional connectivity relationship between brain regions.
[0055] Optionally, in an embodiment of the present invention, the Shen268 brain atlas is used to define 268 brain regions of interest (ROIs) in the whole brain. For each brain region of interest, the average value of the pre-processed blood oxygenation level dependent signal (BOLD signal) of all voxels in the brain region (in magnetic resonance imaging, a voxel represents a small area in the brain) is averaged to obtain a time series; the time series reflects the activity changes of the brain region during the scanning process, that is, the sequence of the pre-processed BOLD signals of all voxels in each brain region of interest changing with time. By averaging the time series of all voxels in the brain region, the overall activity pattern of the brain region can be extracted.
[0056] Based on the time series, the functional connectivity strength between the 268 brain regions of interest in the whole brain can be calculated, that is, the Pearson correlation coefficient of the time series between the 268 pairs of brain regions of interest in the whole brain can be calculated, thereby creating a 268X268 functional connectivity matrix. The 268X268 functional connectivity matrix describes the functional connectivity relationship between different brain regions and is the basis for subsequent analysis and comparison of the functional connectivity differences between normal training pairs and diseased training pairs. Each element (i, j) in the matrix represents the functional connectivity strength between the i-th brain region and the j-th brain region.
[0057] The Pearson correlation coefficient is a statistical indicator that measures the "degree of linear correlation" between two variables. In this embodiment, the Pearson correlation coefficient measures the functional connection strength between two brain regions. It calculates the linear correlation between the time series of the two brain regions, providing an objective quantitative indicator to facilitate statistical analysis.
[0058] By calculating the Pearson correlation coefficient, the result is a 268X268 functional connectivity matrix, in which the rows and columns of the matrix represent the 268 brain regions of interest (ROIs) defined in the Shen268 brain atlas. The values in the matrix represent the functional connectivity strength between the two brain regions, ranging from [-1,1]. A positive value indicates a positive correlation between the two brain regions, that is, their activity patterns are similar, and a negative value indicates a negative correlation between the two brain regions, that is, their activity patterns are opposite. The larger the absolute value of the value, the stronger the correlation, that is, the tighter the functional connection.
[0059] Furthermore, statistical transformation is performed on the values in the functional connectivity matrix.
[0060] It can be understood that in this embodiment, the Fisher Z transform is used to convert the Pearson correlation coefficient into a Z value that is approximately normally distributed. After the Fisher Z transform, the functional connectivity strength value originally ranging from -1 to 1 is converted into a Z value between negative infinity and positive infinity. The distribution of the Z value will be closer to the normal distribution, thereby improving the statistical properties of the functional connectivity strength value, making it more suitable for subsequent parameter testing and model construction.
[0061] Furthermore, the similarity feature vectors between the two normal individuals in the normal training pair or the similarity feature vectors between the two diseased individuals in the diseased training pair are the Spearman correlation coefficients between the functional connectivity matrices.
[0062] Specifically, the similarity feature vectors in the normal training pairs are obtained: based on the functional connection strength between a brain region of a normal individual in the normal training pair and the remaining 267 brain regions, all connection strength values related to the brain region are extracted from the functional connection matrix to form a connection vector, which reflects the functional connection pattern between the brain region and all other 267 brain regions; similarly, the connection vectors of the same brain region and the remaining 267 brain regions of another normal individual in the normal training pair are obtained; the Spearman correlation coefficient of the two connection vectors of this brain region of the two normal individuals is calculated, and similarly, the Spearman correlation coefficient of the connection vectors of the remaining 267 brain regions of the two normal individuals is calculated to obtain a 268-dimensional similarity feature vector containing 268 Spearman correlation coefficients, thereby converting the original individual-based connection representation into a similarity representation based on individual pairs.
[0063] Specifically, similarly, the similarity feature vectors of the normal pairs to be classified and the diseased pairs to be classified can be obtained, which are also the Spearman correlation coefficients between the functional connectivity matrices.
[0064] The Spearman correlation coefficient is an indicator for measuring "trend consistency", which measures the correlation between the ranking orders of two variables and reflects the trend consistency of the two variables, such as synchronous increase, synchronous decrease, or opposite trends. In this embodiment, the Spearman correlation coefficient measures the consistency of the hierarchical order between two functional connection vectors for the same brain region between individual pairs, and is applicable to data with nonlinear relationships or non-normal distributions.
[0065] As a non-limiting specific implementation scheme, an example of obtaining similarity feature vectors in a normal training pair is as follows: assuming that there is a normal individual A and a normal individual B in a normal training pair, a functional connectivity matrix a corresponding to 268 different brain region pairs is generated for the resting-state magnetic resonance image of normal individual A. The rows and columns of the functional connectivity matrix a represent brain regions (brain region 1...brain region 268), respectively. Each element (i, j) in the matrix represents the functional connection strength between the i-th brain region and the j-th brain region. Each element in the matrix a is calculated by the Pearson correlation coefficient of the time series between each pair of brain regions.
[0066] Similarly, a functional connectivity matrix b corresponding to 268 different brain region pairs is generated for the resting-state magnetic resonance image of a normal individual B, and each element in the matrix b is obtained by calculating the Pearson correlation coefficient of the time series between each pair of brain regions.
[0067] Extract the first column element a1 in the functional connectivity matrix a, which represents the functional connectivity vector between brain area 1 of student A and the remaining 267 brain areas, that is, the sequence of functional connectivity strengths between brain area 1 and the other 267 brain areas (brain area 2-brain area 268). Extract the first column element b1 in the functional connectivity matrix b, which represents the functional connectivity vector between brain area 1 of student B and the remaining 267 brain areas. Calculate the Spearman correlation coefficient between the first column element a1 in the functional connectivity matrix a and the first column element b1 in the functional connectivity matrix b, and obtain the functional connectivity similarity of brain area 1 of student A and student B as the first value in the similarity feature vector.
[0068] Similarly, the Spearman correlation coefficients are calculated for the remaining column elements in the functional connectivity matrix a and the corresponding elements in the columns of the functional connectivity matrix b as the remaining elements in the similarity feature vector, and a similarity feature vector containing 268 values is obtained. Each value represents the functional connectivity similarity between student A and student B for the same brain region.
[0069] The functional connectivity vector between student A's brain area 1 and the rest of the brain areas contains a sequence of functional connectivity strengths between brain area 1 and the other 267 brain areas (brain area 2-brain area 268). If the functional connectivity vector between student A's brain area 1 and the rest of the brain areas is more consistent with the functional connectivity vector between student B's brain area 1 and the rest of the brain areas in the order of connectivity strength, the closer the Spearman correlation coefficient is to 1, indicating that the functional connectivity patterns of the two brain areas are more similar. For example, the functional connectivity vector between student A's brain area 1 and the rest of the brain areas A1=[10,20,5,15], and the functional connectivity vector between student B's brain area 1 and the rest of the brain areas B1=[ 12, 22, 8, 18]. The values of A1 and B1 are sorted from smallest to largest and assigned a rank. The rank of the functional connectivity vector A1 corresponding to the original order is obtained as Rank(A1) = [2, 4, 1, 3], and the rank of the functional connectivity vector B1 corresponding to the original order is obtained as Rank(B1) = [2, 4, 1, 3]. Comparing each corresponding element in the rank vectors of A1 and B1, the ranks of A1 and B1 increase or decrease completely synchronously, indicating that their trends are highly consistent. Therefore, the Spearman correlation coefficient is +1, indicating that the functional connectivity pattern of brain area 1 is highly consistent across normal individuals. Conversely, if the ranks of A1 and B1 have opposite trends, it indicates poor consistency, with a coefficient of -1. If there is no pattern, the coefficient is close to 0.
[0070] Repeat the above calculation process and calculate the Spearman correlation coefficient for the functional connectivity matrices of all possible student pairs in the normal group (such as student C-student B, student C-student A) to obtain the similarity feature vectors between different student pairs.
[0071] Furthermore, the similarity feature vector includes preset dimensions based on a preset brain map template.
[0072] As a non-limiting embodiment, in this embodiment, the similarity feature vector has 268 dimensions based on the Shen268 brain map, and each dimension represents a numerical value in the similarity feature vector. Based on the above description, it can be seen that each numerical value in the similarity feature vector is the functional connectivity similarity between the corresponding individual pairs relative to the functional connectivity pattern of a brain region.
[0073] In step S3, a classification model is constructed using a preset classifier, and the similarity feature vectors of the normal group and the similarity feature vectors of the diseased group are used as training data to train the classification model. When the training reaches a preset classification accuracy, the training is stopped to obtain a trained classification model, which includes the following steps: Step S31: training the support vector machine classifier based on each preset dimension of the similarity feature vector, calculating the overall classification accuracy through cross-validation, and stopping the training when the preset classification accuracy is reached; In step S32, the similarity feature vectors are represented by a feature space in a support vector machine classifier, and the support vector machine classifier is used to distribute the similarity feature vectors of the normal training pair and the similarity feature vectors of the diseased training pair in intervals in the feature space.
[0074] Understandably, the overall classification accuracy is calculated through cross-validation to ensure that the trained support vector machine classifier has the most discriminative power for classification. The preset classification accuracy reduces the complexity of model training, reduces redundancy between features, reduces the risk of overfitting, and makes the classification model more robust in predicting the classified individuals. The classifier adopts a support vector machine, which maximizes the interval between the two categories (normal training pairs and diseased training pairs) to the plane by finding a hyperplane, and inputs the similarity feature vectors of the normal training pairs and the diseased training pairs into the support vector machine, which can train the model to learn the distribution pattern of the two categories in the feature space.
[0075] Specifically, the hyperplane is a "classification surface" in the feature space represented by the similarity feature vector in the support vector machine classifier. The function of the hyperplane is to separate data of different categories as much as possible. The values in the similarity feature vector are three-dimensional, which is the functional connectivity similarity of the functional connectivity pattern of a brain region containing many features. There are 268 dimensions for 268 brain regions.
[0076] Exemplarily, each dimension of the similarity feature vector represents the functional connectivity similarity of each pair of brain regions, corresponding to a coordinate axis in the feature space. In this embodiment, the similarity feature vector has 268 dimensions, and the data points are located in the 268-dimensional feature space. The similarity feature vectors of normal training pairs and diseased training pairs are two types of points in this high-dimensional space. In the high-dimensional space, the hyperplane is an abstract "dividing surface". The goal of the support vector machine classifier is to find the hyperplane, divide the normal similarity feature vector and the diseased similarity feature vector into two sides, and maximize the distance from the point closest to the hyperplane in the two types of points to the hyperplane.
[0077] After the new data enters the feature space, it is determined whether it belongs to the normal group or the diseased group based on which side of the hyperplane it is located on.
[0078] Furthermore, in step S31 , training the support vector machine based on the similarity feature vector and calculating the overall classification accuracy includes training data preparation, initializing support vector machine parameters, model training and cross-validation, and stopping condition judgment.
[0079] Among them, the training data preparation is based on the support vector machine classifier, which maps the similarity feature vector into a feature space. Each similarity feature vector corresponds to a data point. The similarity feature vectors of all normal training pairs constitute a normal class sample point set, and the similarity feature vectors of diseased training pairs constitute a diseased class sample point set. The labels corresponding to the normal class sample point set and the diseased class sample point set are obtained.
[0080] Specifically, the normal training pair labels and the diseased training pair labels are assigned in the step of obtaining the training set in step S1.
[0081] For example, the label of the normal sample point set is 1, and the label of the diseased sample point set is 2.
[0082] Initializing the support vector machine parameters includes setting the key hyperparameters of the support vector machine and optimizing them through the grid search method. The key hyperparameters include the penalty parameter C and the kernel function parameter γ, and the number of cross-validation folds is determined.
[0083] Specifically, as a non-restrictive implementation scheme, in this embodiment, the setting range of the hyperparameters is optimized by the grid search method: different values of the penalty parameter C and the kernel function parameter γ are combined to form a parameter grid. For each parameter combination, the support vector machine model is trained using the training set, and performance evaluation is performed on the validation set, such as classification accuracy, F1 value and other indicators.
[0084] Optionally, the penalty parameter C is set at 2 -5 to 2 15 The range is set with an exponential step size. The larger the penalty parameter C value is, the stricter the classification is, and the greater the penalty for misclassified samples is. The kernel function parameter γ is set in 2 -15 to 2 3 The γ value affects the distribution of samples in the feature space. The larger the γ is, the fewer support vectors there are and the higher the model complexity is.
[0085] After traversing all parameter combinations, based on the performance evaluation results of the validation set, the parameter combination with the best classification effect in the test set is selected. In this embodiment, C=2 is used. 9 ,γ=2 -¹³ As the optimal parameter combination, the model is retrained using this parameter combination to obtain the final support vector machine classifier.
[0086] Furthermore, the minimum redundancy maximum relevance (mRMR) feature selection algorithm can be further integrated to screen out the functional connectivity similarity features that are most relevant to the normal training pair or diseased training pair labels and have low redundancy.
[0087] The minimum redundancy maximum correlation algorithm selects the most valuable features for classification by calculating the correlation between features and category labels and the redundancy between features. Features with higher correlation and lower redundancy can effectively improve classifier performance.
[0088] Specifically, for each functional connectivity similarity feature, the mutual information between it and the category label (normal or diseased) is calculated to measure the correlation between the feature and the category; at the same time, the mutual information between features is calculated to evaluate the degree of redundancy between features.
[0089] It can be understood that according to the minimum redundancy and maximum relevance criterion, feature subsets with high comprehensive scores are iteratively screened out, redundant and irrelevant features are removed, data dimensions are reduced, feature dimensions are reduced, computational complexity is reduced, and model training efficiency is improved; redundant features are removed, information duplication is avoided, noise interference is reduced, and the generalization ability of the model is enhanced; key features are focused on so that the model is more focused on functional connection patterns related to the disease and classification accuracy is improved.
[0090] Specifically, if Figure 6 The analysis chart shown is a classification accuracy curve chart under different numbers of features screened out after integrating the mRMR feature selection algorithm.
[0091] Model training and cross-validation involve randomly dividing the training set (normal training pairs and diseased training pairs) into k mutually non-overlapping subsets, of which k-1 subsets are used to train the support vector machine classification model, and the remaining 1 subset is used for validation.
[0092] For each division, use k-1 subsets to train the support vector machine classification model. By solving the optimization problem, the weight vector, bias term and slack variable of the support vector machine classification model are determined. The trained model is used to predict the validation subset, and the classification accuracy is calculated. The above division, training and verification process is repeated k times so that each subset has an opportunity to be used as a verification set. Finally, the average classification accuracy of the k verification results is calculated as the overall classification accuracy of the model under this parameter combination, that is, the preset classification accuracy.
[0093] Furthermore, in step S4, the normal pairs to be classified and the diseased pairs to be classified are input into the classification model, and the classification model outputs prediction results for the normal pairs to be classified and the diseased pairs to be classified based on the similarity feature vectors of the normal pairs to be classified, the similarity feature vectors of the diseased pairs to be classified, and the trained classification model, including the following steps: In step S41, the normal pairs to be classified and the diseased pairs to be classified are input into the classifier, and classification decisions are made based on the similarity feature vectors of the normal pairs to be classified and the similarity feature vectors of the diseased pairs to be classified with a preset classification accuracy, and prediction results for the normal pairs to be classified and the diseased pairs to be classified are output.
[0094] Specifically, the values of the similarity feature vectors of the normal pairs to be classified and the diseased pairs to be classified reflect the differences in functional connection patterns between normal individuals and diseased individuals. The similarity feature vectors of the normal pairs to be classified and the diseased pairs to be classified are input into a trained support vector machine (SVM) classifier. During the training phase, the classifier has learned an optimal hyperplane through the similarity feature vectors of the normal group and the diseased group. The hyperplane divides the feature space into two regions: one side corresponds to the normal group and the other side corresponds to the diseased group. The feature vector of each pair to be classified will be regarded as a point in the feature space, and the value of each dimension corresponds to a coordinate axis in the feature space. In this embodiment, the feature vector has 268 dimensions, and the values of the similarity feature vectors of each pair to be classified are located in the 268-dimensional space. The value of each dimension determines its specific position in the space.
[0095] The classifier determines the predicted category based on which side of the hyperplane the point to be classified is located. If the similarity feature vector of the pair to be classified is on the normal side of the hyperplane: the prediction result of the normal pair to be classified is output, that is, the normal pair to be classified is judged to belong to the normal group; if the similarity feature vector of the pair to be classified is on the diseased group side of the hyperplane: the prediction result of the diseased pair to be classified is output, that is, the diseased pair to be classified is judged to belong to the diseased group.
[0096] In step S4, the multiple voting decision further includes the following steps: Step S42: If the normal pairs to be classified are distributed in the distribution area of the normal training pairs, the prediction result of the normal pairs to be classified is correct, and a normal vote is cast for the classification result of the individuals to be classified; If the normal pairs to be classified are distributed in the distribution area of the diseased training pairs, the prediction result of the normal pairs to be classified is wrong, and the classification result of the individuals to be classified will not be voted; If the diseased pairs to be classified are distributed in the distribution area of the normal training pairs, the prediction result of the diseased pairs to be classified is wrong, and the classification result of the individuals to be classified will not be voted; If the diseased pairs to be classified are distributed in the distribution area of the diseased training pairs, the prediction result of the diseased pairs to be classified is correct, and a vote for diseased is cast for the classification result of the individual to be classified; Based on the majority voting decision results, the voting results are calculated to obtain the classification results of the individuals to be classified.
[0097] Understandably, through multiple comparisons and voting, the influence of accidental errors in single judgments is effectively reduced, avoiding overall classification errors caused by misjudgment of individual pairs to be classified, and improving the model's ability to resist interference from noise and abnormal data. The double comparison mechanism, that is, simultaneously evaluating normal pairs to be classified and diseased pairs to be classified, verifies the similarity between the individuals to be classified and groups of different categories from two dimensions, makes the classification results more convincing, and reduces the risk of misjudgment that may occur from a single perspective. For groups with high heterogeneity such as attention deficit hyperactivity disorder, the characteristics between individuals vary significantly. The voting mechanism can integrate multiple similarity judgment results, better adapt to the diversity within the group, and improve the classification accuracy of complex samples. The rule of classifying as diseased by default when the vote is tied is in line with the conservative principle of misjudgment rather than missed diagnosis in medical diagnosis, minimizes the risk of missed diagnosis, and ensures the safety and reliability of diagnosis.
[0098] In step S42, the category of the individual to be classified is determined by a double comparison and majority voting mechanism, and the overall specific process of outputting the prediction results for the normal pair to be classified and the diseased pair to be classified is as follows: First, the similarity feature vector of the normal pair to be classified is spatially compared with the distribution area of the similarity feature vector of the normal training pair in the training set. If the similarity feature vector of the pair to be classified falls into the distribution area of the normal training pair, the prediction result of the normal pair to be classified is determined to be correct, and a vote of "normal" is cast for the classification result of the individual to be classified; if it falls into the distribution area of the diseased training pair, it indicates that the prediction is wrong and no voting is performed.
[0099] Similarly, the diseased pairs are evaluated. If the similarity feature vectors of the diseased pairs fall within the distribution of the diseased training pairs, the prediction is considered correct and a vote is cast for "disease." If they fall within the distribution of the normal training pairs, the individual is not voted on. Finally, based on the majority voting decision rule, all valid votes are tallied, and the category with the most votes is determined as the final classification result for the individual to be classified. In the event of a tie, the default classification is determined as diseased according to the pre-set differential classification mechanism.
[0100] The distribution of voting results, such as the proportion of votes, can intuitively reflect the degree of uncertainty in the classification results and provide a reference for whether further medical evaluation is needed.
[0101] A second embodiment of the present invention provides a method for classifying attention deficit hyperactivity disorder. The difference between this embodiment and the method for classifying attention deficit hyperactivity disorder provided by the first embodiment is that step S2 further includes the following steps: Step S21, based on the similarity feature vectors of the initial normal individuals and the initial diseased individuals, obtaining the average mean individual of the normal individuals in the normal group and the average mean individual of the diseased individuals in the diseased group; Step S22: Match the individuals to be classified with the average mean individuals of normal individuals to form a normal pair to be classified, and match them with the average mean individuals of diseased individuals to form a diseased pair to be classified, and obtain the normal pair similarity feature vectors and the diseased pair similarity feature vectors of the normal pair to be classified and the diseased pair to be classified.
[0102] It can be understood that by establishing a group mean sample, that is, the functional connection matrix of individuals based on the training set, the mean individual of the normal group and the mean individual of the diseased group are obtained, which are used as the typical individuals of the normal group and the typical individuals of the diseased group, respectively, and matched with the individuals to be classified. The typical brain connection patterns of diseased individuals and normal individuals are captured through the mean individuals, making the similarity features more group-discriminative.
[0103] See also Figure 4 The remaining method steps of the second embodiment are the same as the classification method for attention deficit hyperactivity disorder provided in the first embodiment, and are not repeated here.
[0104] A third embodiment of the present invention provides an attention deficit hyperactivity disorder classification system 100. The attention deficit hyperactivity disorder classification system 100 is applied to the attention deficit hyperactivity disorder classification method provided in the first embodiment of the present invention. The system 100 includes: The training set acquisition module 101 is used to acquire a training set and a similarity feature vector of the training set, wherein the training set includes a normal group and a diseased group, and the similarity feature vector of the training set includes a similarity feature vector of the normal group and a similarity feature vector of the diseased group; The to-be-classified individual acquisition module 102 is configured to match the to-be-classified individual with the normal individual in the normal group to form a normal to-be-classified pair, and match the to-be-classified individual with the diseased individual in the diseased group to form a diseased to-be-classified pair, and obtain a normal to-be-classified pair similarity feature vector and a diseased to-be-classified pair similarity feature vector of the normal to-be-classified pair and the diseased to-be-classified pair; Training module 103; used to build a classification model with a preset classifier, use the similarity feature vectors of the normal group and the similarity feature vectors of the diseased group as training data to train the classification model, and obtain the classification boundary separating the similarity feature vectors of the normal group and the similarity feature vectors of the diseased group; The classification prediction module 104 is used to input the normal pairs to be classified and the diseased pairs to be classified into the classifier. The classifier outputs the prediction results for the normal pairs to be classified and the diseased pairs to be classified based on the classification boundary and the similarity feature vector of the pairs to be classified. Multiple voting decisions are made to vote on the prediction results of the individuals to be classified to obtain the classification results of the individuals to be classified. Among them, when the votes are tied, the classification results of the individuals to be classified are classified as diseased based on the difference classification mechanism.
[0105] As a specific implementation scheme, it has the same beneficial effects as the classification method for attention deficit hyperactivity disorder provided in the first embodiment, and will not be described in detail here.
[0106] See also Figure 5 The fourth embodiment of the present invention provides a computer device 200, which includes a memory 201, a processor 202, and a computer program 203 stored in the memory 201 and executable on the processor 202. When the computer program 203 is executed by the processor, the various steps of the method for classifying attention deficit hyperactivity disorder provided in the first embodiment of the present invention are implemented.
[0107] Compared with the prior art, the method, system, and computer device for classifying attention deficit hyperactivity disorder provided by the present invention have the following beneficial effects: 1. A classification method for attention deficit hyperactivity disorder provided by an embodiment of the present invention aims to solve the problem that traditional classification methods have large individual differences and are easily affected by noise. By establishing group similarity connections, that is, based on the similarity feature vectors of normal individuals in the normal group of the training set and diseased individuals in the diseased group, they are matched with the individuals to be classified respectively, shifting from the individual level to the group similarity level, and replacing the original individual features with the similarity features of the matching pair dimensions, reducing computational complexity and individual noise, and capturing the typical brain connection pattern between each diseased individual or each normal individual through the similarity feature vector, making the similarity feature more group discriminative; using the similarity feature vector of the training set to train the classifier, find the preset classification accuracy that meets the requirements for dividing diseased and normal individuals, and the classification accuracy under the setting avoids data overtraining. Fitting, the voting mechanism for the subjects to be classified adopts a double prediction voting mechanism. By constructing pairs of subjects to be classified with normal individuals and diseased individuals in the training set, that is, input data, the similarity feature vectors of the normal pairs of subjects to be classified and the diseased pairs of subjects to be classified are compared with the trained model, providing a reference benchmark of "normal mode" and "disease mode" for the subjects to be classified, so that the similarity calculation has a clear biological significance and reduces the risk of misjudgment in a single dimension; the discrimination decision and tie-vote rules of multiple votes improve the anti-interference ability of the classification model, and introduce a classification mechanism based on difference in the case of tie-votes, which has a stronger theoretical basis and practical rationality in uncertain discrimination, improves the ability to identify critical individuals, meets clinical safety, and provides further basis for medical evaluation, thereby improving the clinical practical value of the model.
[0108] 2. In the embodiment of the present invention, not all of the initial individuals obtained can be used for classifier training. Propensity score matching constructs a logistic regression model to calculate the propensity score of each individual on the confounding variables, measure the probability of the individual belonging to the diseased group, and then pair individuals with similar propensity scores in different groups, and pair individuals with similar propensity scores in the diseased group and the control group. The matched pairs of initial normal individuals and initial diseased individuals are grouped into normal and diseased groups as training sets to train the classifier. Propensity score matching balances the influence of confounding variables between groups, improves the comparability between the two groups of data, and reduces selection bias. The remaining samples that do not participate in the matching are retained as an independent test set to verify the performance of the classifier training degree.
[0109] 3. In the embodiment of the present invention, the resting-state magnetic resonance image reflects the spontaneous neural activity of the brain in the resting state, and its functional connectivity matrix can reveal the coordination pattern between brain regions, providing objective physiological indicators for the study of brain mechanisms of neurodevelopmental disorders in diseased individuals. Standardization preprocessing includes time layer correction, head motion correction, spatial normalization, denoising and other steps to ensure the comparability of data from different subjects, reduce the interference of data acquisition and processing errors on subsequent analysis, ensure the objectivity and stability of the data, and ensure comparability between them; the resting-state magnetic resonance image after preprocessing is subjected to brain region segmentation to generate a functional connectivity matrix, that is, the brain is divided into multiple regions of interest based on a preset brain map, the time series correlation between different brain regions is calculated, and the functional connectivity matrix is generated, providing a structured brain The partitioning framework makes the calculation of functional connectivity biologically meaningful and interpretable. The functional connectivity matrix converts the brain functional network into computable numerical features, providing a basis for subsequent similarity analysis; the similarity features between any two normal individuals in the normal group and any two diseased individuals in the diseased group are calculated, and the corresponding matching forms normal training pairs and diseased training pairs. By comparing the functional connectivity matrices, similarity feature vectors representing the similarity of functional connectivity patterns between individuals are generated, which is convenient for capturing common patterns within the group, helping the classification model to learn the statistical laws within the two groups of people, and reducing the noise interference of single data points; compared with the individual functional connectivity matrix, the similarity features between training pairs pay more attention to relative differences, revealing the pathological mechanism at the group level, and improving the sensitivity of the classification model.
[0110] 4. In an embodiment of the present invention, the overall classification accuracy is calculated through cross-validation to ensure that the trained support vector machine classifier has the most discriminative power for classification. The preset classification accuracy reduces the complexity of model training, reduces redundancy between features, reduces the risk of overfitting, and makes the classification model more robust for predicting the individuals to be classified. The classifier uses a support vector machine. The support vector machine maximizes the interval between the two categories (normal training pairs and diseased training pairs) and the plane by finding a hyperplane. The similarity feature vectors of the normal training pairs and the diseased training pairs are input into the support vector machine. The model can be trained to learn the distribution pattern of the two categories in the feature space. The normal training pairs and the diseased training pairs are clearly classified as normal or diseased based on the position of the similarity feature vectors in the feature space. The similarity feature vectors are represented in the classifier as having a feature space. Each dimension of the similarity feature vector corresponds to the coordinate in the feature space, so that the relationship between the feature vectors can be visualized and analyzed in the geometric space. The interval distribution reflects the degree of separation of samples of different categories in the feature space, which helps to understand how the classification model distinguishes the similarity feature vectors of normal training pairs and diseased training pairs.
[0111] 5. In an embodiment of the present invention, the similarity feature vectors of the normal pairs to be classified and the diseased pairs to be classified are used as input data of the classifier. The classifier can make accurate classification decisions for the normal pairs to be classified and the diseased pairs to be classified, and use the preset classification accuracy as a standard to ensure the reliability of the classification results, providing a valuable reference basis for clinical diagnosis or research; the classification model captures the relative differences between them, and the similarity feature vectors reflect the differences in brain network patterns between the normal pairs to be classified and the diseased pairs to be classified, while the support vector machine classifier maps such differences to the decision space of the classifier, so that the abstract neuroimaging features are converted into quantifiable classification basis, and the individual differences are converted into the distribution degree of the pairs to be classified in the feature space, thereby realizing inclusive classification of heterogeneity; by quantifying the classification accuracy and capturing group differences and individual heterogeneity, interpretable inference from neuroimaging data to disease classification is realized.
[0112] 6. In this embodiment of the present invention, a dual matching voting mechanism is implemented for normal pairs and diseased pairs. By matching with individuals with known disease conditions within the two groups, individual characteristics are converted into the degree of deviation from the normal training pairs and the degree of fit with the diseased training pairs. This avoids the limitations of a single classification boundary. The connection pattern of a single diseased individual may not represent the group characteristics. Using matching pairs can capture group commonalities and improve the classifier's tolerance for heterogeneous individuals. Voting is performed to judge the correctness of the prediction results. Only when the prediction result is consistent with the category trained by the classifier, that is, the normal pair to be classified falls within the distribution area of its corresponding normal training pair or the diseased pair to be classified falls within the distribution area of its corresponding diseased training pair, is the vote counted as valid. Otherwise, it is considered invalid and no vote is cast. This filters out unreliable predictions, avoids noise interference in the final decision, strengthens group pattern consistency, ensures that the classification result is based on matching the prediction pattern, and not on accidental fluctuations, and reduces the impact of accidental errors and quantifies classification uncertainty. At the same time, multiple voting results support dynamic evaluation, monitoring individual changes over time, such as before and after treatment, and evaluating the effectiveness of intervention by comparing changes in voting proportions.
[0113] 7. In an embodiment of the present invention, the propensity score matching method uses caliper matching without replacement to balance confounding variables such as age, gender, head motion parameters, and other non-brain function differences between the initial normal individuals and the initial diseased individuals, reduce inter-group bias, and make the initial individuals more statistically comparable. Matching without replacement ensures that each initial normal individual participates in only one matching, avoiding repeated use that may lead to redundancy or bias in the training set data. Caliper matching further limits the similarity of matching by setting a caliper threshold, thereby improving matching quality. Before each matching, the initial normal individuals are randomly arranged into normal groups to eliminate order-dependent bias and enhance the stability of the results.
[0114] 8. In an embodiment of the present invention, the Spearman correlation coefficient is a non-parametric statistical method used to measure the monotonic relationship between two variables, that is, the consistency of the trends between the variables, rather than a strict linear relationship. By calculating the Spearman correlation coefficient of the functional connectivity matrices of the brain regions of two individuals, the consistency of the brain network connection patterns of the two individuals is characterized. The Spearman correlation coefficient is performed by rank transformation on the data, that is, converting the original data into a sorted rank, eliminating the influence of outliers and distribution patterns, and characterizing the trend consistency of the brain connection pattern.
[0115] 9. An embodiment of the present invention further provides an attention deficit hyperactivity disorder classification system, which is applied to the above-mentioned attention deficit hyperactivity disorder classification method. Therefore, it also has the same beneficial effects as the above-mentioned attention deficit hyperactivity disorder classification method, and will not be described in detail here.
[0116] 10. An embodiment of the present invention further provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the computer device implements the various steps of the above-mentioned method for classifying attention deficit hyperactivity disorder. Therefore, the computer device also has the same beneficial effects as the above-mentioned method for classifying attention deficit hyperactivity disorder, which will not be elaborated here.
[0117] The above is a detailed introduction to the attention deficit hyperactivity disorder classification method and attention deficit hyperactivity disorder classification system disclosed in the embodiments of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention. Any modifications, equivalent replacements and improvements made within the principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A classification method for attention deficit hyperactivity disorder, characterized by: The method comprises the following steps: Obtaining a training set and a similarity feature vector of the training set, wherein the training set includes a normal group and a diseased group, and the similarity feature vector of the training set includes a normal group similarity feature vector and a diseased group similarity feature vector; Acquire an individual to be classified, match the individual to be classified with a normal individual in the normal group to form a normal pair to be classified, and match the individual to be classified with a diseased individual in the diseased group to form a diseased pair to be classified; acquire a normal pair similarity feature vector of the normal pair to be classified and a diseased pair similarity feature vector of the diseased pair to be classified; Constructing a classification model with a preset classifier, using the similarity feature vector of the normal group and the similarity feature vector of the diseased group as training data to train the classification model, and stopping the training when the preset classification accuracy is met to obtain a trained classification model; Inputting the normal pairs to be classified and the diseased pairs to be classified into the classification model, the classification model outputting prediction results for the normal pairs to be classified and the diseased pairs to be classified based on the similarity feature vectors of the normal pairs to be classified, the similarity feature vectors of the diseased pairs to be classified and the trained classification model; The output of the prediction results of the normal to-be-classified pairs and the diseased to-be-classified pairs is carried out through multiple voting decisions, and the multiple voting decisions include: voting on the prediction results of the individuals to be classified to obtain the classification results of the individuals to be classified; and when there is a tie, the classification results of the individuals to be classified are classified as diseased based on the difference classification mechanism.
2. The method for classifying attention deficit hyperactivity disorder according to claim 1, wherein: Before obtaining the training set and the similarity feature vector of the training set, the following steps are included: Acquiring a preset number of initial individuals, wherein the initial individuals include initial normal individuals and initial diseased individuals; Matching the initial normal individuals and the initial diseased individuals based on a propensity score matching method; The matched pairs of the initial normal individuals and the initial diseased individuals are correspondingly grouped into the normal group and the diseased group as the training set, and the unmatched initial individuals are used as independent individuals to be classified.
3. The method for classifying attention deficit hyperactivity disorder according to claim 2, wherein: Obtaining a training set and a similarity feature vector of the training set includes the following steps: Acquiring resting-state magnetic resonance images of the normal individuals in the normal group and the diseased individuals in the diseased group, and performing standardized preprocessing on the resting-state magnetic resonance images. performing brain region segmentation on the preprocessed resting-state magnetic resonance images using a preset brain atlas template, and generating functional connectivity matrices corresponding to different brain region pairs for the resting-state magnetic resonance images of the normal individual and the diseased individual; Comparing the functional connectivity matrices of any two normal individuals in the normal group, obtaining similarity feature vectors between the two normal individuals, and matching them to form a normal training pair; The functional connectivity matrices of any two diseased individuals in the diseased group are compared to obtain similarity feature vectors between the two diseased individuals, and the two diseased individuals are matched to form a diseased training pair.
4. The method for classifying attention deficit hyperactivity disorder according to claim 3, wherein: The similarity feature vector includes a preset dimension based on the preset brain map template; a classification model is constructed using a preset classifier, and the similarity feature vector of the normal group and the similarity feature vector of the diseased group are used as training data to train the classification model. When the training reaches a preset classification accuracy, the training is stopped to obtain a trained classification model, comprising the following steps: Training a support vector machine classifier based on each of the preset dimensions of the similarity feature vector, calculating the overall classification accuracy through cross-validation, and stopping the training when the preset classification accuracy is reached; The similarity feature vector represents a feature space in the support vector machine classifier, and the support vector machine classifier is used to distribute the similarity feature vectors of the normal training pair and the similarity feature vectors of the diseased training pair at intervals in the feature space.
5. The method for classifying attention deficit hyperactivity disorder according to claim 4, wherein: Inputting the normal pairs to be classified and the diseased pairs to be classified into the classification model, the classification model outputting prediction results for the normal pairs to be classified and the diseased pairs to be classified based on the similarity feature vectors of the normal pairs to be classified, the similarity feature vectors of the diseased pairs to be classified, and the trained classification model, comprising the following steps: The normal pairs to be classified and the diseased pairs to be classified are input into the classifier, and classification decisions are made based on the similarity feature vectors of the normal pairs to be classified and the similarity feature vectors of the diseased pairs to be classified with the preset classification accuracy, and prediction results for the normal pairs to be classified and the diseased pairs to be classified are output.
6. The method for classifying attention deficit hyperactivity disorder according to claim 5, wherein: The multiple voting decision specifically includes the following steps: If the normal pairs to be classified are distributed in the distribution area of the normal training pairs, the prediction result of the normal pairs to be classified is correct, and a normal vote is cast for the classification result of the individual to be classified; If the normal pairs to be classified are distributed in the distribution area of the diseased training pairs, the prediction result of the normal pairs to be classified is wrong, and no vote is cast for the classification result of the individual to be classified; If the diseased pair to be classified is distributed in the distribution area of the normal training pair, the prediction result of the diseased pair to be classified is wrong, and no vote is cast for the classification result of the individual to be classified; If the diseased pair to be classified is distributed in the distribution area of the diseased training pair, the prediction result of the diseased pair to be classified is correct, and a vote of diseased is cast for the classification result of the individual to be classified; Based on the majority voting decision result, the votes are calculated to obtain the classification result of the individual to be classified.
7. The method for classifying attention deficit hyperactivity disorder according to claim 2, wherein: The propensity score matching method is caliper matching without replacement. Matching the initial normal individuals and the initial diseased individuals based on the propensity score matching method includes the following steps: Calculating a propensity score for the initial normal individual and the initial diseased individual based on a preset confounding variable; The initial diseased individuals are used as a benchmark, and the initial normal individuals are matched without replacement using a preset caliper threshold; wherein, before each matching, the initial normal individuals are randomly arranged.
8. The method for classifying attention deficit hyperactivity disorder according to claim 3, wherein: The similarity feature vector between the two normal individuals or the similarity feature vector between the two diseased individuals is the Spearman correlation coefficient between the functional connectivity matrices.
9. A classification system for attention deficit hyperactivity disorder, characterized in that The method for classifying attention deficit hyperactivity disorder according to any one of claims 1 to 8, wherein the system comprises: A training set acquisition module, configured to acquire a training set and a similarity feature vector of the training set, wherein the training set includes a normal group and a diseased group, and the similarity feature vector of the training set includes a normal group similarity feature vector and a diseased group similarity feature vector; The module for obtaining individuals to be classified is used to match the individuals to be classified with normal individuals in the normal group to form normal pairs to be classified, and to match the individuals to be classified with diseased individuals in the diseased group to form diseased pairs to be classified; and to obtain similarity feature vectors of the normal pairs to be classified and similarity feature vectors of the diseased pairs to be classified; a training module, configured to construct a classification model using a preset classifier, train the classification model using the normal group similarity feature vector and the diseased group similarity feature vector as training data, and obtain a classification boundary separating the normal group similarity feature vector and the diseased group similarity feature vector; A classification prediction module is used to input the normal pairs to be classified and the diseased pairs to be classified into the classification model, and the classification model outputs prediction results for the normal pairs to be classified and the diseased pairs to be classified based on the similarity feature vectors of the normal pairs to be classified, the similarity feature vectors of the diseased pairs to be classified and the trained classification model; wherein, the output prediction results for the normal pairs to be classified and the diseased pairs to be classified are decided by multiple voting, and the multiple voting decisions include: voting on the prediction results of the individuals to be classified to obtain the classification results of the individuals to be classified; wherein, when the votes are tied, the classification results of the individuals to be classified are classified as diseased based on the difference classification mechanism.
10. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, each step of the method for classifying attention deficit hyperactivity disorder as described in any one of claims 1 to 8 is implemented.