A class-sharing semi-supervised feature selection method for depression brain images

By introducing a class-shared semi-supervised feature selection method in the brain image data processing of depression, fuzzy similarity relationships are constructed using fuzzy convex hemispheres and fuzzy information measurements, the problems of inaccurate calculation of high-dimensional data and scarcity of labeled data are solved, and more efficient feature selection and diagnostic performance are achieved.

CN118470373BActive Publication Date: 2025-06-06NANTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410394369.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-02
Publication Date
2025-06-06
Estimated Expiration
2044-04-02

AI Technical Summary

Technical Problem

In the prior art, when processing brain image data for high-dimensional depression, Euclidean distance calculation is inaccurate and complex, and the supervised and unsupervised feature selection performance of some labeled data is poor, especially when the data volume is large and labeled data is scarce.

Method used

A method of selecting feature of the class-shared semi-supervised, which forms a fuzzy similarity relationship by constructing a fuzzy convex hemisphere, defines fuzzy information measurement, redefines fuzzy correlation and redundancy, and uses the principles of fuzzy correlation maximization and redundancy to search for feature subsets, and combines tag-specific feature selection methods and dynamic optimization strategies.

Benefits of technology

It improves the prediction performance of brain image data for depression, reduces the cost of data collection and labeling, enhances the generalization ability of the model, and significantly improves the accuracy and efficiency of intelligent assisted diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118470373B_ABST
    Figure CN118470373B_ABST
Patent Text Reader

Abstract

The present invention provides a category-sharing semi-supervised feature selection method for depression brain images, which belongs to the technical field of medical image label selection. The technical problem of the presence of unlabeled data and excessive redundant pathological attributes in the image samples of the brain region of patients with depression is solved. The technical scheme is as follows: comprising the following steps: S10, reading the depression brain image data, preprocessing and dividing it, and finally constructing a four-tuple decision information system; S20, for the labeled data and the unlabeled data, respectively, constructing fuzzy information granules according to the distance metric to form a fuzzy similarity relationship; S30, characterizing the importance of depression data features according to the maximum relevance minimum redundancy strategy; S40, combining the label-specific feature selection method and the dynamic optimization strategy to select important brain regions for predicting depression. The beneficial effects of the present invention are: it is helpful for the prediction of depression and improves the diagnosis and treatment of depression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image label selection, and in particular to a category-sharing semi-supervised feature selection method for depression brain images. Background Art

[0002] Depression is a common mental disorder characterized by a significant and persistent low mood, accompanied by a loss of interest and pleasure, which often affects an individual's work, study and social function. Depression may be caused by a variety of factors, including genetic, biological, psychosocial and environmental factors.

[0003] At present, Wang Yue et al. proposed to select the classification feature set by combining hybrid feature algorithm with genetic algorithm in the study of feature selection related to brain function of depression in "Depression classification method based on hybrid feature selection algorithm". The brain function connection matrix of two groups of subjects in five frequency bands was constructed by phase locking between signals, and the connection values ​​with significant differences were used as features according to the results of t-test. In the face of high-dimensional features, it is proposed to use quadratic programming feature selection based on mutual information and Fisher score to sort all features separately, and package the first 100 features of the two by intersection or union. The optimal subset is further selected by genetic algorithm for classification. However, it does not consider the situation of a large amount of unlabeled data in practical applications and has certain limitations. S.Kavi Priya et al. introduced a new embedded feature selection method based on statistical correlation class frequency (SRCF-WOA) of whale optimization algorithm in "An embedded feature selection approach for depression classification using short text sequences". This method extracts unit features and composite features to capture semantic and structural information. There is no feature selection for specific labels, which lacks precision to a certain extent.

[0004] The high incidence of depression has attracted great attention from the World Health Organization. Due to the large number of patients and the huge amount of data that needs to be analyzed, there is an urgent need to explore a new method to effectively extract brain area information related to depression from a large amount of data. Through feature selection, this method can effectively help doctors identify patients at high risk of depression, thus providing a new research direction in the field of medical image analysis. In addition, this method is expected to improve the diagnosis and treatment of depression, provide doctors and researchers with more precise tools to better understand the neurobiological basis of depression, and provide more effective medical services for patients with depression. Summary of the invention

[0005] The traditional method of using Euclidean distance to calculate the distance between two samples to construct similarity relations cannot process high-dimensional data. In the case of high-dimensional data, Euclidean distance will be affected by the dimension, the calculation of similarity may be inaccurate, and the calculation complexity is relatively high. At the same time, in practical applications, the collected data is always partially labeled, especially the number of unlabeled samples is much larger than the number of labeled samples. However, supervised and unsupervised feature selection of partially labeled data cannot achieve good performance. In order to make up for the above defects, the present invention proposes a category-sharing semi-supervised feature selection method for depression brain images. First, fuzzy convex hemispheres are constructed for labeled brain data and unlabeled brain data to form fuzzy similarity relations; second, several appropriate fuzzy information measures are defined to quantify inconsistency or uncertainty. Third, fuzzy correlation considering feature-decision relations and fuzzy redundancy considering feature-feature relations are redefined to collaboratively evaluate the quality of candidate features. Accordingly, the principle of maximizing fuzzy correlation and minimizing fuzzy redundancy is used to search for feature subsets. Finally, the label-specific feature selection method and the dynamic optimization strategy are combined, and the proposed feature selection method is used to select brain regions to improve the prediction performance. It has strong application value in intelligent auxiliary diagnosis of depression.

[0006] In order to achieve the above objectives, the present invention adopts a technical solution: a category-sharing semi-supervised feature selection method for depression brain images, comprising the following steps:

[0007] S10, read the depression brain image data, preprocess and divide it, including labeled and unlabeled data. Finally, build a four-tuple decision information system;

[0008] S20, constructing fuzzy convex hemispheres for labeled data and unlabeled data according to distance metrics to form fuzzy similarity relationships;

[0009] S30, characterize the importance of depression data features according to the maximum relevance and minimum redundancy strategy;

[0010] S40. Combine label-specific feature selection method and dynamic optimization strategy to select important brain regions for predicting depression.

[0011] Further optimizing the scheme, the step S10 includes the following steps:

[0012] S11, read the data set of depression brain region image data, determine its attribute set and decision class, the decision information system S = <U, C ∪ D, V, f >, where U = {x 1 ,..,x j ,...,x n} represents the sample set of depression brain region image data, n represents the number of brain region image data sets, xj represents the jth sample, x n represents the nth sample, C = {a 1 ,a 2 ,...,a m} represents the conditional attributes in the depression brain region image data, m represents the number of attributes in the depression brain region image data, a 1 Indicates the first attribute, a m represents the mth attribute, D={d 1 ,d 2 ,...d i} represents a non-empty finite set of decision attributes of depression brain region image data, called decision attributes, i represents the number of decision categories in depression brain region image data, d 1 represents the first decision category, d i represents the i-th decision category, and Since it is very expensive to collect data under full supervision, there are often a large number of unlabeled samples. Accordingly, this kind of decision system is called a partial decision system, which usually satisfies U = U L ∪U U , where U L and U U Respectively represent the labeled depression brain region image data and the unlabeled depression brain region image data, usually |U L |<<|U U |, where |·| represents the cardinality of the set. x represents a non-empty finite set of attributes of depression brain region image data. In particular, d(x) for each labeled sample x∈U L is known, and d(x) for each unlabeled sample x∈U U is unknown. V = ∪ a∈C∪D V a , V a is a possible case of attribute a of the depression brain region image data; f:U×(C∪D)→V is an information function that assigns an information value to each brain region image, that is, x∈U,f(x,a)∈V a .

[0013] S12, according to the number of different information values ​​of the decision feature D in the data set, the depression brain region image data set U is divided into two parts: the marked part U L and the unmarked part U U , there are n data subsets in total, satisfying U=U L ∪U U .

[0014] Further optimizing the solution, step S20 includes the following steps:

[0015] S21, there is labeled data U for depression brain region images L ={x 1 ,x 2 ,...,x n}, each decision i, i = 1, 2, ..., k, k represents the decision category in the labeled data of depression, select the brain area image data of each category of depression, and take the data center point of this category

[0016] S22. Calculation of depression brain samples and other samples The Euclidean distance between two pairs dis(x j ,x p ).

[0017] S23. Calculation of depression brain samples and all other data centers Euclidean distance between two

[0018] S24. According to the Euclidean distance, for Construct a sample x j is the center of the circle, x j Go to each category center sample in turn The distance between them is the radius of the fuzzy information particle, which is defined as follows:

[0019]

[0020] Where i represents the sample x j The corresponding decision class, Indicates that the sample x j is the center of the circle, The set of neighbors selected for the fuzzy ball of radius.

[0021] S25. According to the Euclidean distance, for Then construct the current category center sample is the center of the circle, is the fuzzy information granule with radius, which is defined as follows:

[0022]

[0023] Where i represents the sample x j The corresponding decision class, Indicates the sample centered on the current category is the center of the circle, The set of neighbors selected for the fuzzy ball of radius.

[0024] S26, select the intersection of two fuzzy balls to construct a shared fuzzy information particle Depression brain samples can be obtained p With x j The shared fuzzy information particles form a new fuzzy similarity relationship R for each category. cir (x p ,x j ).

[0025]

[0026] S27, for depression brain region image unlabeled dataset U U ={y 1 ,y 2 ,...,y n},y n represents the nth sample, Classify the labeled samples of depression brain area images into categories. For all categories D = {d 1 ,d 2 ,...,d i}Take the sample center point Take the current unlabeled samples y one by one j , construct the current unlabeled sample y j is the center of the circle, is the fuzzy information granule with radius, which is defined as follows:

[0027]

[0028] in, Indicates that the current unlabeled sample y j is the center of the circle, The set of neighbors selected for the fuzzy ball of radius.

[0029] S28, reconstruct the sample center point is the center of the circle, is the fuzzy information granule with radius, which is defined as follows:

[0030]

[0031] in, Indicates the center point is the center of the circle, The set of neighbors selected for the fuzzy ball of radius.

[0032] S29, select the intersection of the two constructed fuzzy balls, that is, the shared fuzzy information granule, add an unlabeled sample of the depression brain area in turn, update the neighborhood set δ, the samples in the neighborhood and the fuzzy similarity relationship R cir .

[0033] Further optimizing the solution, step S30 includes the following steps:

[0034] S31. Due to the lack of label information, the feature evaluation and selection in the partial decision system are hindered. In order to solve the problem of scarce label information, it is intended to generate fuzzy decisions for a given partial decision system, which is expressed as Marking process.

[0035] S32. Calculate each labeled depression brain sample, then sample x j and fuzzy decision making The fuzzy membership of Its definition is as follows:

[0036]

[0037] where [x j ] c is relative to the fuzzy similarity relation R cir The fuzzy similarity class of It is a depression brain sample x j The distance from the center of each class.

[0038] S33. For unlabeled samples in the brain region of depression, this will inevitably lead to uncertainty in the labeling process. Therefore, we use class membership to further induce fuzzy decisions for unlabeled samples. Specifically, for x t ∈U U , its membership can be defined as:

[0039]

[0040] where x t represents unlabeled samples of depression brain regions, [x t ] c is the fuzzy similarity relation R after update cir The fuzzy similarity class of Represents an updated fuzzy decision.

[0041] S34. Calculate the fuzzy decision corresponding to each decision category of the unlabeled data of the depression brain region, denoted as The specific expressions are as follows:

[0042]

[0043] Where T is a parameter that controls the sharpening strength, as T→0, the fuzzy decision will be close to a one-hot decision, i.e. similar to a hard label. The fuzzy decision is generated by using a sharpening function.

[0044] S35. Evaluate the uncertainty of the model based on entropy and calculate the candidate features c∈C of the depression brain region and the fuzzy decision Fuzzy correlation J rel .

[0045] S36, based on fuzzy decision Calculate the fuzzy redundancy J between the candidate feature c∈C of the depression brain region and another candidate feature c'∈C of the depression brain region red .

[0046] S37. Based on the principle of maximizing fuzzy relevance and minimizing fuzzy redundancy, a new semi-supervised criterion combining fuzzy relevance and fuzzy redundancy of candidate feature c is defined:

[0047]

[0048] Among them, C s is a subset of features that have been selected, Represents fuzzy correlation J rel , Represents the fuzzy redundancy J red .

[0049] Further optimizing the solution, step S40 includes the following steps:

[0050] S41. Combining label-specific feature selection methods with dynamic optimization strategies to calculate candidate features of brain regions in depression for each category The corresponding fuzzy correlation J rel , get the correlation value table J corresponding to each feature rel , select the correlation value table J rel The feature c with the largest median value 1 , the feature c 1 Delete it from attribute set C and set the correlation value table J rel This feature c 1 The correlation value of feature c is deleted. 1 Add to the feature sorting set Sort and mark this feature c 1 If only one feature needs to be selected at this time, the current feature with the largest correlation c is output 1 .

[0051] S42. Calculate candidate features of brain regions in depression (c' represents the feature subset C without fuzzy correlation J rel The fuzzy redundancy J corresponding to the feature subset after the maximum feature red , and obtain the fuzzy redundancy value table corresponding to each feature.

[0052] S43, Calculation J rel -J red The maximum value of the corresponding feature c 2 Index idx and update the current category feature sorting set Sort.

[0053] S44, feature c 2 Delete from attribute set C, and continue to delete correlation value table J rel Corresponding to the feature of index idx. Update the fuzzy correlation value table and fuzzy redundancy value table.

[0054] S45, after looping through all attributes in C in turn, output the reduced set red of each category data set of the depression brain region to obtain the selected brain region.

[0055] Compared with the prior art, the present invention has the following beneficial effects:

[0056] (1) The class-sharing semi-supervised feature selection method for depression brain images provided by the present invention is more suitable for the common situation of limited data and scarce labeled data in many practical applications. It can improve model performance, reduce data collection and labeling costs, and improve the generalization ability of the model.

[0057] (2) For labeled samples, fuzzy similarity relationships are formed by constructing a fuzzy convex hemisphere. For unlabeled samples, constructing a fuzzy convex hemisphere can greatly reduce the number of other samples in the convex hemisphere, greatly reducing the computational cost.

[0058] (3) By sorting the importance of features according to the maximum relevance and minimum redundancy strategy, a more streamlined feature subset can be obtained, which is more efficient and has higher application value. This method has significantly improved the feature selection of depression data, thereby improving the accuracy and efficiency of intelligent auxiliary diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The accompanying drawings are used to further explain the specific steps of the present invention and constitute a part of the specification.

[0060] Figure 1 The figure is an overall flow chart of the category-sharing semi-supervised feature selection method for depression brain images of the present invention.

[0061] Figure 2 This is a diagram of the overall data processing framework of the category-sharing semi-supervised feature selection method for depression brain images of the present invention. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. Of course, the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.

[0063] Example 1

[0064] See also Figure 1 to Figure 2 The technical solution provided in this embodiment is a category-sharing semi-supervised feature selection method for depression brain images, comprising the following steps:

[0065] S10, read the depression brain image data, preprocess and divide it, including labeled and unlabeled data. Finally, build a four-tuple decision information system;

[0066] S20, constructing fuzzy convex hemispheres for labeled data and unlabeled data according to distance metrics to form fuzzy similarity relationships;

[0067] S30, characterize the importance of depression data features according to the maximum relevance and minimum redundancy strategy;

[0068] S40. Combine label-specific feature selection method and dynamic optimization strategy to select important brain regions for predicting depression.

[0069] The step S10 comprises the following steps:

[0070] S11, read the data set of depression brain region image data, determine its attribute set and decision class, the decision information system S = <U, C ∪ D, V, f >, where U = {x 1 ,..,x j ,...,x n} represents the sample set of depression brain region image data, n represents the number of brain region image data sets, x j represents the jth sample, x n represents the nth sample; C = {a 1 ,a 2 ,...,a m} represents the conditional attributes in the depression brain region image data, m represents the number of attributes in the depression brain region image data, a 1 Indicates the first attribute, a m represents the mth attribute, D={d 1 ,d 2 ,...d i} represents a non-empty finite set of decision attributes of depression brain region image data, called decision attributes, i represents the number of decision categories in depression brain region image data, d 1 represents the first decision category, di represents the i-th decision category, and

[0071] S12, according to the number of different information values ​​of the decision D in the data set, the depression brain area image data set U is divided into two parts: the marked part U L and the unmarked part U U , there are n data subsets in total, satisfying U=U L ∪U U .

[0072] The original data set is converted into a four-tuple decision information system as follows:

[0073]

[0074] Set the depression brain region image dataset to a labeled dataset U L and the unlabeled dataset U U .

[0075]

[0076]

[0077] Specifically, step S20 includes the following steps:

[0078] S21, there is labeled data U for depression brain region images L ={x 1 ,x 2 ,...,x n}, each decision i, i = 1, 2, ..., k, k represents the decision category in the labeled data of depression, select the brain area image data of each category of depression, and take the data center point of this category

[0079] In this embodiment, the number of decision classes is 4, and the corresponding central samples are:

[0080]

[0081] S22. Calculation of depression brain samples and other samples The Euclidean distance between two pairs dis(x j ,x p ).

[0082]

[0083] S23. Calculation of depression brain samples and all other data centers Euclidean distance between two

[0084]

[0085] S24. According to the Euclidean distance, for Construct a sample x j is the center of the circle, x j Go to each category center sample in turn The distance between them is the radius of the fuzzy information particle, which is defined as follows:

[0086]

[0087] Where i represents the sample x j The corresponding decision class, Indicates that the sample x j is the center of the circle, The set of neighbors selected for the fuzzy ball of radius.

[0088] For ease of understanding, the sample x in the embodiment is 1 To expand the description, for sample x 1 For example, i = 1, so the neighborhood set selected by the fuzzy ball is

[0089] S25. According to the Euclidean distance, for Then construct the current category center sample is the center of the circle, is the fuzzy information granule with radius, which is defined as follows:

[0090]

[0091] Where i represents the sample x j The corresponding decision class, Indicates the sample centered on the current category is the center of the circle, The set of neighbors selected for the fuzzy ball of radius.

[0092] In this embodiment, sample 1 is also used to describe the selected neighborhood set:

[0093] S26, select the intersection of two fuzzy balls to construct a shared fuzzy information particle Depression brain samples can be obtained p With x j The shared fuzzy information particles form a new fuzzy similarity relationship R for each category. cir (x p ,x j ).

[0094]

[0095] This example takes category 1 as an example, calculates the distance of each sample, determines whether it satisfies the neighborhood relationship, and obtains the fuzzy similarity relationship between samples:

[0096]

[0097] S27, for depression brain region image unlabeled dataset U U ={y 1 ,y 2 ,...,y n},y n represents the nth sample, Classify the labeled samples of depression brain area images into categories. For all categories D = {d 1 ,d 2 ,...,d i}Take the sample center point Take the current unlabeled samples y one by one j , construct the current unlabeled sample y j is the center of the circle, is a fuzzy ball with a radius of , which is defined as follows:

[0098]

[0099] in, Indicates that the current unlabeled sample y j is the center of the circle, The set of neighbors selected for the fuzzy ball of radius.

[0100] S28, reconstruct the sample center point is the center of the circle, is the fuzzy information granule with radius, which is defined as follows:

[0101]

[0102] in, Indicates the center point is the center of the circle, The set of neighbors selected for the fuzzy ball of radius.

[0103] S29, select the intersection of the two constructed fuzzy balls, that is, the shared fuzzy information granule, add an unlabeled sample of the depression brain area in turn, update the neighborhood set δ, the samples in the neighborhood and the fuzzy similarity relationship R cir .

[0104] For ease of understanding, this example assumes that the first unlabeled sample of depression brain is x 25 , then we get

[0105] δ={x 3 ,x 9 ,x 12 ,x 25}, D i ={1,3}.

[0106] The step S30 comprises the following steps:

[0107] S31. Due to the lack of label information, the feature evaluation and selection in the partial decision system are hindered. In order to solve the problem of scarce label information, it is intended to generate fuzzy decisions for a given partial decision system, which is expressed as Marking process.

[0108] S32. Calculate each labeled depression brain sample, then sample x j and fuzzy decision making The fuzzy membership of Its definition is as follows:

[0109]

[0110] where [x j ] c is relative to the fuzzy similarity relation R cir The fuzzy similarity class of It is a depression brain sample x j The distance from the center of each class.

[0111] Taking the example in the article as an example, there are four categories, then the depression brain has fuzzy decision corresponding to the labeled data for:

[0112]

[0113]

[0114] S33. For unlabeled samples in the brain region of depression, this will inevitably lead to uncertainty in the labeling process. Therefore, we use class membership to further induce fuzzy decisions for unlabeled samples. Specifically, for x t ∈U U , its membership can be defined as:

[0115]

[0116] where x t represents unlabeled samples of depression brain regions, [x t ] c is the fuzzy similarity relation R after update cir The fuzzy similarity class of Represents an updated fuzzy decision.

[0117] S34. Calculate the fuzzy decision corresponding to each decision category of the unlabeled data of the depression brain region, denoted as The specific expressions are as follows:

[0118]

[0119] Where T is a parameter that controls the sharpening strength, as T→0, the fuzzy decision will be close to a one-hot decision, i.e. similar to a hard label. The fuzzy decision is generated by using a sharpening function.

[0120] This example corresponds to unlabeled brain data of depression, fuzzy decision for:

[0121]

[0122]

[0123] S35. Evaluate the uncertainty of the model based on entropy and calculate the candidate features c∈C of the depression brain region and the fuzzy decision Fuzzy correlation J rel .

[0124] S36, based on fuzzy decision Calculate the fuzzy redundancy J between the candidate feature c∈C of the depression brain region and another candidate feature c'∈C of the depression brain region red .

[0125] S37. Based on the principle of maximizing fuzzy relevance and minimizing fuzzy redundancy, a new semi-supervised criterion combining fuzzy relevance and fuzzy redundancy of candidate feature c is defined:

[0126]

[0127] Among them, C s is a subset of features that have been selected, Represents fuzzy correlation J rel , Represents the fuzzy redundancy J red .

[0128] In this embodiment, the corresponding J can be calculated rel =[0.57850.36570.31490.49740.54200.42720.57800.22590.5922], the corresponding J red =[0.9294 0.6907 0.6765 0.8445 0.80470.74920.8234 0.8245].

[0129] The step S40 comprises the following steps:

[0130] S41. Combining label-specific feature selection methods with dynamic optimization strategies to calculate candidate features of brain regions in depression for each category The corresponding fuzzy correlation J rel , get the correlation value table J corresponding to each feature rel , select the correlation value table J rel The feature c with the largest median value 1 , the feature c 1 Delete it from attribute set C and set the correlation value table J rel This feature c 1 The correlation value of feature c is deleted. 1 Add to the feature sorting set Sort and mark this feature c 1 If only one feature needs to be selected at this time, the current feature with the largest correlation c is output 1 .

[0131] In this embodiment, taking category 1 as an example, according to the J calculated above, rel =[0.5785 0.3657 0.31490.49740.5420 0.4272 0.5780 0.2259 0.5922], we can see that the feature with the greatest correlation is a 9 , then feature a 9 Delete feature a from attribute set C. 9 Add to the feature sorting set Sort. If you only need to select one feature output at this time, that is, output a 9 .

[0132] S42. Calculate candidate features of brain regions in depression (c' represents the feature subset C without fuzzy correlation J rel The fuzzy redundancy J corresponding to the feature subset after the maximum feature red , and obtain the fuzzy redundancy value table corresponding to each feature.

[0133] S43, Calculation J rel -J red The maximum value of the corresponding feature c 2 Index idx and update the current category feature sorting set Sort.

[0134] In this embodiment, according to J rel =[0.5785 0.3657 0.3149 0.4974 0.54200.42720.57800.2259] and Jred =[0.9294 0.6907 0.6765 0.8445 0.8047 0.7492 0.82340.8245], calculate J rel -J red The maximum value of , the corresponding feature index is 7, and the feature sorting set Sort = [9 7].

[0135] S44, feature c 2 Delete from attribute set C, and continue to delete correlation value table J rel Corresponding to the feature of index idx. Update the fuzzy correlation value table and fuzzy redundancy value table.

[0136] S45, after looping through all attributes in C in turn, output the reduced set red of the depression brain region data set to obtain the selected brain region.

[0137] The final output of this example is Sort = [9 7 5 1 4 2 6 3 8].

[0138] Example 2

[0139] See also Figure 1 to Figure 2 The technical solution provided in this embodiment is a category-sharing semi-supervised feature selection method for depression brain images, comprising the following steps:

[0140] S10, read the depression brain image data, preprocess and divide it, including labeled and unlabeled data. Finally, build a four-tuple decision information system;

[0141] S20, constructing fuzzy convex hemispheres for labeled data and unlabeled data according to distance metrics to form fuzzy similarity relationships;

[0142] S30, characterize the importance of depression data features according to the maximum relevance and minimum redundancy strategy;

[0143] S40. Combine label-specific feature selection method and dynamic optimization strategy to select important brain regions for predicting depression.

[0144] The step S10 comprises the following steps:

[0145] S11, read the data set of depression brain region image data, determine its attribute set and decision class, the decision information system S = <U, C ∪ D, V, f >, where U = {x 1 ,..,x j ,...,x n} represents the sample set of depression brain region image data, n represents the number of brain region image data sets, x jrepresents the jth sample, x n represents the nth sample; C = {a 1 ,a 2 ,...,a m} represents the conditional attributes in the depression brain region image data, m represents the number of attributes in the depression brain region image data, a 1 Indicates the first attribute, a m represents the mth attribute, D={d 1 ,d 2 ,...d i} represents a non-empty finite set of decision attributes of depression brain region image data, called decision attributes, i represents the number of decision categories in depression brain region image data, d 1 represents the first decision category, d i represents the i-th decision category, and

[0146] S12, according to the number of different information values ​​of the decision D in the data set, the depression brain area image data set U is divided into two parts: the marked part U L and the unmarked part U U , there are n data subsets in total, satisfying U=U L ∪U U .

[0147] The original data set is converted into a four-tuple decision information system as follows

[0148]

[0149] Set the depression brain region image dataset to a labeled dataset U L and the unlabeled dataset U U .

[0150]

[0151] Specifically, step S20 includes the following steps:

[0152] S21, there is labeled data U for depression brain region images L ={x 1 ,x 2 ,...,x n}, each decision i, i = 1, 2, ..., k, k represents the decision category in the labeled data of depression, select the brain area image data of each category of depression, and take the data center point of this category

[0153] In this embodiment, the number of decision classes is 4, and the corresponding central samples are:

[0154]

[0155] S22. Calculation of depression brain samples and other samples The Euclidean distance between two pairs dis(x j ,x p ).

[0156]

[0157] S23. Calculation of depression brain samples and all other data centers Euclidean distance between two

[0158]

[0159] S24. According to the Euclidean distance, for Construct a sample x j is the center of the circle, x j Go to each category center sample in turn The distance between them is the radius of the fuzzy information particle, which is defined as follows:

[0160]

[0161] Where i represents the sample x j The corresponding decision class, Indicates that the sample x j is the center of the circle, The set of neighbors selected for the fuzzy ball of radius.

[0162] For ease of understanding, the sample x in the embodiment is 1 To expand the description, for sample x 1 For example, i = 1, so the neighborhood set selected by the fuzzy ball is

[0163] S25. According to the Euclidean distance, for Then construct the current category center sample is the center of the circle, is the fuzzy information granule with radius, which is defined as follows:

[0164]

[0165] Where i represents the sample x j The corresponding decision class, Indicates the sample centered on the current category is the center of the circle, The set of neighbors selected for the fuzzy ball of radius.

[0166] In this embodiment, sample 1 is also used to describe the selected neighborhood set:

[0167] S26, select the intersection of two fuzzy balls to construct a shared fuzzy information particle Depression brain samples can be obtained p With x j The shared fuzzy information particles form a new fuzzy similarity relationship R for each category. cir (x p ,x j ).

[0168]

[0169] This example takes category 1 as an example, calculates the distance of each sample, determines whether it satisfies the neighborhood relationship, and obtains the fuzzy similarity relationship between samples:

[0170]

[0171] S27, for depression brain region image unlabeled dataset U U ={y 1 ,y 2 ,...,y n},y n represents the nth sample,

[0172] Classify the labeled samples of depression brain area images into categories. For all categories D = {d 1 ,d 2 ,...,d i}Take the sample center point Take the current unlabeled samples y one by one j , construct the current unlabeled sample y j is the center of the circle, is a fuzzy ball with a radius of , which is defined as follows:

[0173]

[0174] in, Indicates that the current unlabeled sample y j is the center of the circle, The set of neighbors selected for the fuzzy ball of radius.

[0175] S28, reconstruct the sample center point is the center of the circle, is the fuzzy information granule with radius, which is defined as follows:

[0176]

[0177] in, Indicates the center point is the center of the circle, The set of neighbors selected for the fuzzy ball of radius.

[0178] S29, select the intersection of the two constructed fuzzy balls, that is, the shared fuzzy information granule, add an unlabeled sample of the depression brain area in turn, update the neighborhood set δ, the samples in the neighborhood and the fuzzy similarity relationship R cir .

[0179] For ease of understanding, this example assumes that the first unlabeled sample of depression brain is x 25 , then we get

[0180] δ={x 3 ,x 5 ,x 12 ,x 25}, D i ={1,3}.

[0181] The step S30 comprises the following steps:

[0182] S31. Due to the lack of label information, the feature evaluation and selection in the partial decision system are hindered. In order to solve the problem of scarce label information, it is intended to generate fuzzy decisions for a given partial decision system, which is expressed as Marking process.

[0183] S32. Calculate each labeled depression brain sample, then sample x j and fuzzy decision making The fuzzy membership of Its definition is as follows:

[0184]

[0185] where [x j ] c is relative to the fuzzy similarity relation R cir The fuzzy similarity class of It is a depression brain sample x j The distance from the center of each class.

[0186] Taking the example in the article as an example, there are four categories, then the depression brain has fuzzy decision corresponding to the labeled data for:

[0187]

[0188]

[0189] S33. For unlabeled samples in the brain region of depression, this will inevitably lead to uncertainty in the labeling process. Therefore, we use class membership to further induce fuzzy decisions for unlabeled samples. Specifically, for x t ∈U U , its membership can be defined as:

[0190]

[0191] where x t represents unlabeled samples of depression brain regions, [x t ] c is the fuzzy similarity relation R after update cir The fuzzy similarity class of Represents an updated fuzzy decision.

[0192] S34. Calculate the fuzzy decision corresponding to each decision category of the unlabeled data of the depression brain region, denoted as The specific expressions are as follows:

[0193]

[0194] Where T is a parameter that controls the sharpening strength, as T→0, the fuzzy decision will be close to a one-hot decision, i.e. similar to a hard label. The fuzzy decision is generated by using a sharpening function.

[0195] This example corresponds to unlabeled brain data of depression, fuzzy decision for:

[0196]

[0197] S35. Evaluate the uncertainty of the model based on entropy and calculate the candidate features c∈C of the depression brain region and the fuzzy decision Fuzzy correlation J rel .

[0198] S36, based on fuzzy decision Calculate the fuzzy redundancy J between the candidate feature c∈C of the depression brain region and another candidate feature c'∈C of the depression brain region red .

[0199] S37. Based on the principle of maximizing fuzzy relevance and minimizing fuzzy redundancy, a new semi-supervised criterion combining fuzzy relevance and fuzzy redundancy of candidate feature c is defined:

[0200]

[0201] Among them, C s is a subset of features that have been selected, Represents fuzzy correlation J rel , Represents the fuzzy redundancy J red .

[0202] In this embodiment, the corresponding J can be calculated rel =[0.3410 0.4636 0.3411 0.3399 0.4128 0.3647 0.3954 0.3139 0.3402], the corresponding J red =[0.9294 0.6907 0.6765 0.84450.80470.7492 0.8234 0.8245].

[0203] The step S40 comprises the following steps:

[0204] S41. Combining label-specific feature selection methods with dynamic optimization strategies to calculate candidate features of brain regions in depression for each category The corresponding fuzzy correlation J rel , get the correlation value table J corresponding to each feature rel , select the correlation value table J rel The feature c with the largest median value 1 , the feature c 1 Delete it from attribute set C and set the correlation value table J rel This feature c 1 The correlation value of feature c is deleted. 1 Add to the feature sorting set Sort and mark this feature c 1 If only one feature needs to be selected at this time, the current feature with the largest correlation c is output 1 .

[0205] In this embodiment, taking category 1 as an example, according to the J calculated above, rel =[0.3410 0.4636 0.3411 0.3399 0.4128 0.3647 0.3954 0.3139 0.3402], we can see that the feature with the greatest correlation is a 2 , then feature a 2 Delete feature a from attribute set C. 2 Add to the feature sorting set Sort. If you only need to select one feature output at this time, that is, output a 9 .

[0206] S42. Calculate candidate features of brain regions in depression (c' represents the feature subset C without fuzzy correlation J rel The fuzzy redundancy J corresponding to the feature subset after the maximum featurered , and obtain the fuzzy redundancy value table corresponding to each feature.

[0207] S43, Calculation J rel -J red The maximum value of the corresponding feature c 2 Index idx and update the current category feature sorting set Sort.

[0208] In this embodiment, according to J rel =[0.3410 0.3411 0.3399 0.4128 0.3647 0.39540.31390.3402] and J red =[0.6892 0.7289 0.7085 0.7262 0.7393 0.7379 0.69680.6869], calculate J rel -J red The maximum value of , the corresponding feature index is 7, and the feature sorting set Sort = [2 5].

[0209] S44, feature c 2 Delete from attribute set C, and continue to delete correlation value table J rel Corresponding to the feature of index idx. Update the fuzzy correlation value table and fuzzy redundancy value table.

[0210] S45, after looping through all attributes in C in turn, output the reduced set red of the depression brain region data set to obtain the selected brain region.

[0211] The final output of this example is Sort = [2 5 3 1 7 6 4 9 8].

[0212] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A class-sharing semi-supervised feature selection method for depression brain images, characterized in that: The steps include: S10, read the depression brain image data, preprocess and divide it, including labeled and unlabeled data, and construct a four-tuple decision information system S = <U, C ∪ D*, V, f>, where U represents the sample set of depression brain region image data, C represents the conditional attribute, D* represents the non-empty finite set of decision attributes, called decision attributes, and S20, constructing fuzzy information granules for labeled data and unlabeled data according to distance metrics to form fuzzy similarity relationships; S30, characterize the importance of depression data features according to the maximum relevance and minimum redundancy strategy; S40, combining label-specific feature selection method and dynamic optimization strategy to select important brain regions for predicting depression; The step S20 comprises the following steps: S21, for depression brain region images with labeled data to form fuzzy similarity relationship R cir ,; S27. For the unlabeled dataset of depression brain region images, take the current unlabeled samples y one by one j , Construct the current unlabeled sample y j is the center of the circle, The fuzzy information particle with radius x c i is the category center of category i; S28, reconstruct the sample center point is the center of the circle, The fuzzy information particle with radius S29, select two constructed fuzzy balls The intersection of the two is the shared fuzzy information granule. We add an unlabeled sample of the depression brain region in turn, update the neighborhood set δ, and the fuzzy similarity relationship R between the samples in the neighborhood and the unlabeled data cir ; The step S30 comprises the following steps: S31. Generate fuzzy decisions for a given partial decision system, expressed as D = {D1, D2, ..., D k }Marking process; S32. Calculate each labeled depression brain sample x j Fuzzy Membership Degree of Fuzzy Decision Making Its definition is as follows: where [x j ] c is the fuzzy similarity relation R relative to the labeled data cir Fuzzy similarity class of x c i For x j Category center of S33, for any sample x in the unlabeled samples of the depression brain region t , its membership is defined as: where [x t ] c is the fuzzy similarity relation R relative to the unlabeled data cir Fuzzy similarity class of ; S34. Calculate the fuzzy decision corresponding to each decision category of the unlabeled data of the depression brain region, denoted as D i , specifically expressed as follows: Where T is a parameter that controls the sharpening strength. As T→0, the fuzzy decision will be close to a hot decision, that is, similar to a hard label, the fuzzy decision is generated by using the sharpening function, and k represents the number of decision categories; S37. Based on the principle of maximizing fuzzy relevance and minimizing fuzzy redundancy, a new semi-supervised criterion combining fuzzy relevance and fuzzy redundancy of candidate feature c is defined: Among them, C s is a subset of features that have been selected, Represents fuzzy correlation J rel , Represents the fuzzy redundancy J red , candidate feature c∈C, another candidate feature c'∈C for depression brain region.

2. The class-sharing semi-supervised feature selection method for depression brain images according to claim 1, characterized in that: The step S10 comprises the following steps: S11, read the data set of depression brain area image data, determine its attribute set and decision class, the decision information system S = <U, C ∪ D *, V, f>, where U = {x1, .., x j ,...,x n } represents the sample set of depression brain region image data, n represents the number of samples in the brain region image data set, x j represents the jth sample, x n represents the nth sample; C = {a1, a2, ..., a m } represents the conditional attribute in the depression brain region image data, m represents the number of attributes in the depression brain region image data, a1 represents the first attribute, a m represents the mth attribute, D*={d1,d2,...d k } represents a non-empty finite set of decision attributes of depression brain region image data, is called the decision attribute, k represents the number of decision categories in the depression brain region image data, d1 represents the first decision category, and d k represents the kth decision category; S12, according to the number of different information values ​​of the decision D* in the data set, the depression brain region image data set U is divided into two parts: the marked part U L and the unmarked part U U , there are n3 data subsets in total, satisfying U=U L ∪U U .

3. The class-sharing semi-supervised feature selection method for depression brain images according to claim 1, characterized in that: The step S21 comprises the following steps: S211, for depression brain area image has labeled data U L ={x1,x2,...,x n1 }, for each decision i, i = 1, 2, ..., k, k represents the number of decision categories in the labeled data of depression, select the brain region image data of each category of depression, and take the data center point of the category; S221. Calculation of depression brain samples and other samples The Euclidean distance between two pairs dis(x j ,x p ); S231. Calculation of depression brain samples and its category data center points Euclidean distance between two S241. According to the Euclidean distance, for Construct a sample x j is the center of the circle, x j Go to the category center sample in turn The distance between them is the radius of the fuzzy information particle, which is defined as follows: Where i represents the sample x j The corresponding decision class, Indicates that the sample x j is the center of the circle, The neighborhood set selected for the fuzzy ball of radius; S251. According to the Euclidean distance, for Sequentially construct samples centered on the current category is the center of the circle, is the fuzzy information granule with radius, which is defined as follows: Where i represents the sample x j The corresponding decision class, Indicates the sample centered on the current category is the center of the circle, The neighborhood set selected for the fuzzy ball of radius; S261, select the intersection of two fuzzy balls to construct a shared fuzzy information particle Obtained depression brain sample x p With x j The shared fuzzy information particles form a new fuzzy similarity relationship R for each category. cir (x p ,x j ); In step S27 In step S28 .

4. The class-sharing semi-supervised feature selection method for depression brain images according to claim 1, characterized in that: The step S40 includes the following steps: S41. Combining label-specific feature selection methods with dynamic optimization strategies to calculate candidate features of brain regions in depression for each category The corresponding fuzzy correlation J rel , get the correlation value table J corresponding to each feature rel , select the correlation value table J rel The feature c1 with the largest median value is deleted from the attribute set C, and the correlation value table J is rel The correlation value of feature c1 is deleted, feature c1 is added to the feature sorting set Sort, and the index idx of feature c1 is marked. If a feature is selected at this time, the feature c1 with the maximum correlation is output; S42. Calculate candidate features of brain regions in depression c' represents the feature subset C without fuzzy correlation J rel The feature subset after the maximum feature, the corresponding fuzzy redundancy J red , get the fuzzy redundancy value table corresponding to each feature; S43, Calculation J rel -J red The maximum value of the feature c2 index idx is obtained, and the current category feature sorting set Sort is updated at the same time; S44, delete feature c2 from attribute set C, and continue to delete correlation value table J rel Corresponding to the features of index idx, update the fuzzy correlation value table and the fuzzy redundancy value table; S45. After looping through all the attributes in C in turn, the reduced set red of each category data set of the depression brain region is output to obtain the selected brain region.

Citation Information

Patent Citations

  • An unsupervised pedestrian re-recognition method based on fuzzy depth clustering

    CN109299707A

  • Brain structure feature selection method, mobile terminal and computer readable storage medium

    CN111062420A