Hierarchical Feature Selection Method and System Based on Tag Relevance and Instance Relevance
By introducing label correlation and instance correlation constraints in the hierarchical feature selection method, using the k-nearest neighbor mechanism optimized by hierarchical tree structure and Gaussian kernel function, the problem of failure to fully utilize hierarchical structure and instance correlation in the existing technology is solved, and more efficient feature selection and classification performance is achieved.
Patent Information
- Application Number
- CN202510093336.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-01-21
AI Technical Summary
When handling large-scale classification tasks, existing hierarchical feature selection methods fail to fully utilize the semantic information and instance correlation in the hierarchy, resulting in huge and inefficient feature subsets.
A hierarchical feature selection method (HFS-LCIC) based on label correlation and instance correlation is proposed. Through the k-nearest neighbor mechanism optimized by hierarchical tree structure and Gaussian kernel function, label correlation and instance correlation constraints are constructed to optimize feature selection.
Effectively reduce the number of features, improve classification performance, improve the accuracy and robustness of the model, and is suitable for feature selection and classification tasks of large-scale hierarchical data.
Smart Images

Figure CN119537898B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of machine learning and data processing, and particularly relates to a hierarchical feature selection method and system based on label correlation and instance correlation. Background Art
[0002] In the big data era, the number of categories in classification tasks is increasing rapidly, from a few or dozens of categories in the past to hundreds of thousands of categories today. These categories are organized into hierarchical structures, such as tree structures and directed acyclic graphs, according to the semantic information between categories to solve the classification learning problem with a huge number of categories. Compared with traditional classification learning that divides categories into flat sets, hierarchical classification learning can decompose the classification task into multiple subtasks. First, identify the abstract coarse-grained categories of the instances, and then refine them layer by layer to specific fine-grained categories. This method significantly reduces the difficulty of classification.
[0003] In recent years, hierarchical classification learning has been widely applied to multiple research fields, including image classification, text classification, disease diagnosis, and gene recognition. In hierarchical classification learning, as the number of categories increases sharply, the number of features describing the classification task also increases rapidly, which leads to the problem of dimensionality disaster. Feature selection, as an effective dimensionality reduction method, has received extensive attention. However, traditional feature selection methods usually select a general feature subset to distinguish all categories, resulting in the selected feature subset still being large, and some features may only be useful for distinguishing individual categories. These methods assume that categories are independent, but ignore the possible hierarchical structure between them.
[0004] Inspired by the divide-and-conquer strategy, researchers have proposed a new hierarchical feature selection method: select respective feature subsets for different subtasks to effectively reduce the number of features. In addition, between categories in the hierarchical structure, such as "mobile phone" and "tablet computer" under "electronic products", there is a high degree of correlation between them and they share many features. In medical image data, the same disease may exhibit multiple different features and lesion morphologies, which are divided into different clusters in the feature space. By considering instance correlation, a medical image classification system can more accurately classify image instances with similar features into the same category, thereby improving classification accuracy. Therefore, the correlation in the label space and the similarity in the instance space can be used as auxiliary information to make the hierarchical feature selection method achieve better results in improving classification performance.
[0005] Although current research on label and instance relevance has made some progress in hierarchical feature selection, there are still several significant deficiencies. In the label space, existing hierarchical feature selection methods mainly focus on the relationships between parent-child and sibling classes. Although these methods utilize the relevance between classes to a certain extent, they ignore the connections between broader classes and fail to fully exploit the rich semantic information embedded in the hierarchical structure to consider the relevance of each class to other classes to varying degrees. At the same time, it is particularly important to focus on the association between features and labels under local subtasks. On the one hand, this reduces model complexity, and on the other hand, it can highlight the differences between different classes within each subtask and more deeply explore the subtle but crucial features in instances for classification subtasks, thereby obtaining a compact and powerful classification model. However, existing hierarchical feature selection methods lack a full consideration of both label relevance and instance relevance. In addition, existing methods mainly use mechanisms such as cosine similarity to evaluate instance relevance, which cannot effectively handle the noisy data and redundant features existing in the instance space.
[0006] To address the above problems, we propose a hierarchical feature selection method (HFS-LCIC) and system based on label relevance and instance relevance; converting the information of the hierarchical structure and instance distribution into effective regularization terms to improve classification performance. Summary of the Invention
[0007] The object of the present invention is to propose a hierarchical feature selection method and system based on label relevance and instance relevance; under the global structure, a label relevance constraint is constructed using the path distance between internal classes in the hierarchical tree structure; under local subtasks, an instance relevance constraint is constructed through the k-nearest neighbor mechanism optimized by the Gaussian kernel function; not only considering the hierarchical relationship between classes but also utilizing the similarity between instances, thereby being able to more effectively select features and improve classification performance.
[0008] To achieve the above object, the technical solution of the present invention is as follows:
[0009] For the hierarchical feature selection method based on label relevance and instance relevance, the feature selection task is decomposed into several subtasks, and a feature subset is selected for each subtask to improve classification accuracy, specifically including the following steps:
[0010] S1. Input the hierarchical training data set, including an instance matrix containing n instances and m features , and a label matrix containing n instances and one label , and convert the single-label matrix into a multi-label matrix;
[0011] S2. Describe the semantic relationships between categories in the hierarchical training dataset using a hierarchical tree structure, and set the number of internal nodes in the hierarchical tree structure to N + 1; each internal node corresponds to a subtask;
[0012] S3. Construct a loss term for the internal nodes, and at the same time use sparse learning, label correlation, and instance correlation as regularization terms. Use the loss term and the regularization terms including sparse learning, label correlation, and instance correlation to construct the objective function of the subtask corresponding to each internal node. Optimize the weight matrix of each internal node according to the objective function, sort the weight matrices of each subtask in descending order, and select the features with large weight values to complete the hierarchical feature selection task.
[0013] Preferably, the conversion of the single-label matrix to the multi-label matrix is specifically that the single-label matrix is converted to a multi-label matrix through binary relevance , where represents the label of instance j , d represents the number of categories.
[0014] Preferably, the description of the semantic relationships between categories in the hierarchical training dataset using the hierarchical tree structure is specifically as follows: The instance matrix X is divided into according to the internal nodes, where represents the instance matrix corresponding to the internal node i , n i represents the number of instances corresponding to the internal node i ; X 0 is the instance matrix corresponding to the root node; represents 's label matrix, where is the label matrix corresponding to the internal node i , and the label i of the instance j of the internal node , d max represents the maximum value of the number of categories in the internal nodes; is used to represent the weight matrix of the i th internal node.
[0015] Preferably, the construction of the loss term for the internal nodes is specifically as follows: Use the least squares loss as the loss function for hierarchical feature selection:
[0016] (1)
[0017] Denotes the Frobenius norm of a matrix.
[0018] Preferably, the regularization term for sparse learning is selected norm.
[0019] Preferably, the i label correlation regularization term of the
[0020] th internal node is specifically:
[0021] In the formula, tr denotes the trace of a matrix, that is, the sum of the elements on the main diagonal of the matrix; denotes the internal node i and the internal node l the path distance between:
[0022] (3)
[0023] In the formula, 0 ≤ i ≤ N, 0 ≤ l ≤ N, i = 0 or l = 0 represents the root node in the hierarchical tree structure, denotes the i th internal node and the l th internal node's lowest common ancestor node.
[0024] Preferably, the i instance correlation regularization term of the
[0025] th internal node is specifically:
[0026] In the formula, and respectively denote the instance i under the p th sub-task and the instance q , denotes the instance correlation matrix, denotes 's Laplacian matrix, denotes the correlation between the instance i under the p th sub-task and the instance q :
[0027] (6)
[0028] In the formula, the parameter , denotes calculated by the Euclidean distance of k the k nearest neighbor instances.
[0029] Preferably, the objective function for constructing the sub-task corresponding to the internal node using the loss term and the regularization terms including sparse learning, label correlation, and instance correlation is specifically:
[0030] (8)
[0031] where λ, α, and β are all non-negative constants, controlling sparsity, label correlation, and instance correlation respectively.
[0032] Preferably, optimize the weight matrix of each internal node according to the objective function, specifically:
[0033] Calculate with respect to the derivative:
[0034] (9)
[0035] where is a diagonal matrix, and the r -th diagonal element is . Substitute equation (8) with equation (9) according to equation (9):
[0036] (10)
[0037] The derivative of equation (10) with respect to is:
[0038] (11)
[0039] Let equation (11) be equal to 0 to obtain the expression of as follows:
[0040] (12)
[0041] Initialize the weight d max according to the maximum value of the number of categories in the internal node, ; perform iterative calculation through equation (12) until convergence to obtain the optimized weight matrix.
[0042] A hierarchical feature selection system based on label correlation and instance correlation includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the above-mentioned hierarchical feature selection method based on label correlation and instance correlation.
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] The present invention realizes effective feature selection and classification for large-scale hierarchical data by optimizing an objective function that comprehensively considers multiple factors. The least squares loss ensures the accuracy of prediction, while sparse learning promotes the effective selection of features. At the same time, the label correlation and instance correlation regularization terms respectively utilize the class relationships in the hierarchical structure and the similarities between instances, further improving the classification accuracy and robustness. Thus, features can be selected more effectively, and the classification performance can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a model diagram of the hierarchical feature selection method based on label correlation and instance correlation of the present invention.
[0046] Figure 2 It is a schematic diagram of the algorithm of the hierarchical feature selection method based on label correlation and instance correlation of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The technical solution of the present invention will be specifically described below with reference to the accompanying drawings.
[0048] The present invention proposes a hierarchical feature selection method based on label correlation and instance correlation (HFS-LCIC). The feature selection task is decomposed into several sub-tasks, and a feature subset is selected for each sub-task to improve the classification accuracy. The specific steps are as follows:
[0049] S1. Input the hierarchical training data set, including an instance matrix containing n instances and m features , and a label matrix containing n instances and one label , and convert the single-label matrix into a multi-label matrix;
[0050] S2. Use a hierarchical tree structure to describe the semantic relationships between classes in the hierarchical training data set. The number of internal nodes (non-leaf nodes) in the hierarchical tree structure is set to N + 1; each internal node in the hierarchical tree structure is used to recursively model the problem, and each internal node corresponds to a sub-task;
[0051] S3. Construct a loss term for the internal node, and at the same time use sparse learning, label correlation, and instance correlation as regularization terms . Use the loss term and the regularization terms including sparse learning, label correlation, and instance correlation to construct the objective function of the sub-task corresponding to the internal node . Optimize the weight matrix of each internal node according to the objective function, sort the weight matrices of each sub-task in descending order, select the features with large weight values, and complete the hierarchical feature selection task.
[0052] The HFS-LCIC method of the present invention consists of four main components, as Figure 1 shown; (1) Least squares loss: This loss term aims to minimize the error between the predicted label and the actual label; (2) Sparse learning: The sparse regularization term constrains the feature coefficients to quickly select the optimal shared features; (3) Label correlation: This regularization term uses the rich semantic information in the hierarchical tree structure to analyze the degree of association between each category and other categories from a global perspective; (4) Instance correlation: This regularization term uses the instance distribution information to analyze the similarity between instances under subtasks to further distinguish fine-grained categories.
[0053] In this embodiment, the conversion of the single-label matrix into a multi-label matrix is specifically that the single-label matrix is converted into a multi-label matrix through binary relevance , where represents the label of instance j , and d represents the number of categories.
[0054] In this embodiment, the use of the hierarchical tree structure to describe the semantic relationship between categories in the hierarchical training dataset is specifically as follows: The instance matrix X is divided into according to the internal nodes, where represents the instance matrix corresponding to the internal node i , n i represents the number of instances corresponding to the internal node i ; X 0 is the instance matrix corresponding to the root node; represents 's label matrix, where is the label matrix corresponding to the internal node i , and the instance i of the internal node j has the label , d max represents the maximum value of the number of categories in the internal nodes; is used to represent the weight matrix of the i th internal node.
[0055] In this embodiment, the construction of the loss term for the internal node is specifically as follows: The least squares loss is used as the loss function for hierarchical feature selection:
[0056] (1)
[0057] represents the Frobenius norm of the matrix; the least squares loss uses the Euclidean distance as the similarity metric, which is simple and efficient to calculate.
[0058] In this embodiment, the regularization term of sparse learning is selected norm, As a convex function that meets the requirements of feature sparsity, the norm is easier to achieve global optimization. Then, the objective function of the basic hierarchical feature selection method is:
[0059] (2)
[0060] where λ represents a non-negative constant.
[0061] In this embodiment, the semantic relationships embedded in the hierarchical structure, as prior knowledge, help to capture the connections between categories, thus assisting in the design of the feature selection method; the smaller the path distance between internal nodes (categories) in the hierarchical tree structure, the closer the relationship between the categories, and the more similar their corresponding weight matrices are; the i th internal node and other internal nodes l the path distance between them is calculated as:
[0062] (3)
[0063] In the formula, 0 ≤ i ≤ N, 0 ≤ l ≤ N, i = 0 or l = 0 represents the root node in the hierarchical tree structure, represents the i th internal node and the l th internal node's lowest common ancestor node. According to the path distance between internal nodes, we can obtain the corresponding label correlation matrix P where . Therefore, the label correlation regularization term of the i th internal node is defined as:
[0064] (4)
[0065] After adding the regularization term, formula (2) can be rewritten as:
[0066] (5)
[0067] In the formula, λ and α are non-negative constants that control sparsity and label correlation, respectively.
[0068] In this embodiment, in the hierarchical tree structure, instances under the same inner class have higher similarity; the more similar the instances are, the more similar their corresponding labels are; instance correlation is a factor to be considered in the subtask that focuses on the association between features and labels in local instances; this increases the gap between different categories within each subtask and further differentiates fine-grained categories; this results in a more accurate selection of an effective feature subset for each subtask, thus generating a compact and robust classification model; however, there are still noise data and redundant features in the feature space of the subtasks, and the commonly used cosine similarity is not sufficient to accurately evaluate the similarity between instances; therefore, based on the k-nearest neighbor mechanism in the subtasks, the present invention uses a Gaussian kernel function to assign weights to instances at different distances to obtain an instance correlation matrix composed of the instance similarity between each pair of instances ; and denotes the correlation between instance i and instance p under the q th subtask:
[0069] (6)
[0070] In the formula, the parameter , denotes the nearest neighbor instances of k calculated by the Euclidean distance, and respectively denote instance i and instance p under the q th subtask; the instance correlation regularization term of the i th internal node is specifically:
[0071] (7)
[0072] In the formula, denotes the Laplacian matrix of ; combining this regularization term, the final objective function is rewritten as:
[0073] (8)
[0074] In the formula, λ, α, and β are all non-negative constants, which control sparsity, label correlation, and instance correlation respectively.
[0075] In this embodiment, the weight matrix of each internal node is optimized according to the objective function, specifically:
[0076] Calculate Regarding Derivative of:
[0077] (9)
[0078] Wherein, is a diagonal matrix, and the r th diagonal element is According to Equation (9), Equation (8) is replaced with:
[0079] (10)
[0080] The derivative of Equation (10) with respect to is:
[0081] (11)
[0082] Let Equation (11) be equal to 0 to obtain the expression of as follows:
[0083] (12)
[0084] The whole process of the HFS-LCIC method is as Figure 2 shown. Initialize the weight d max according to the maximum value of the number of categories in the internal nodes, ; perform iterative calculations through Equation (12) until convergence to obtain the optimized weight matrix. Then, sort the weight matrices of each subtask in descending order, and finally select the features with larger weight values to complete the hierarchical feature selection task.
[0085] The present invention also proposes a hierarchical feature selection system based on label correlation and instance correlation, including a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the above-mentioned hierarchical feature selection method based on label correlation and instance correlation.
[0086] In summary, first, the present invention introduces the least square loss and the sparse regularization term. Secondly, when considering the global hierarchical structure, the present invention considers the different degrees of association between categories. A correlation matrix is constructed by calculating the path distance between internal categories, the regularization term is used to limit the label correlation, and supervision information is provided for each subtask. Finally, under the local subtask, the present invention uses the k-nearest neighbor mechanism optimized by the Gaussian kernel function to construct a regularization term to constrain the instance correlation. This helps to widen the gap between different categories under each subtask. Thus, features can be selected more effectively, and the classification performance can be improved.
[0087] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, as long as the functions and effects produced do not exceed the scope of the technical solution of the present invention, fall within the protection scope of the present invention.
Claims
1. A hierarchical feature selection method based on label correlation and instance correlation, characterized in that: The feature selection task is decomposed into several subtasks, and a feature subset is selected for each subtask to improve classification accuracy. The classification application field is image classification. The specific steps include: S1. Input hierarchical training dataset, including an instance matrix X∈R containing n instances and m features n×m , and a single-label matrix D∈R containing n instances and one label n×1 , and convert the single-label matrix into a multi-label matrix; S2, use a hierarchical tree structure to describe the semantic relationship between categories in the hierarchical training dataset. The number of internal nodes in the hierarchical tree structure is set to N+1; each internal node corresponds to a subtask; S3. Construct loss terms for internal nodes, and use sparse learning, label relevance, and instance relevance as regularization terms. Use loss terms and regularization terms including sparse learning, label relevance, and instance relevance to construct the objective function of the subtask corresponding to the internal node. Optimize the weight matrix of each internal node according to the objective function, sort the weight matrix of each subtask in descending order, select features with large weight values, and complete the hierarchical feature selection task. The use of a hierarchical tree structure to describe the semantic relationship between categories in a hierarchical training data set is specifically as follows: the instance matrix X is divided into X0, X1, X2, ...X according to the internal nodes. i ...,X N , 0≤i≤N,n i ≤n,X i represents the instance matrix corresponding to internal node i, n i represents the number of instances corresponding to the internal node i, where X0 is the instance matrix corresponding to the root node; Y0,Y1,Y2,...Y i ...,Y N represents X0,X1,X2,...X i ...,X N A multi-label matrix, where is the multi-label matrix corresponding to internal node i, the label of instance j of internal node i 1≤j≤n i , d max Indicates the maximum number of categories in the internal node; The weight matrix used to represent the i-th internal node; The objective function of the subtask corresponding to the internal node constructed using the loss term and the regularization terms including sparse learning, label correlation and instance correlation is specifically: In the formula, The loss function selected for the hierarchical features, ||·|| F Represents the Frobenius norm of the matrix; Regularization term selection for sparse learning 2,1 norm, ||W i || 2,1 is the weight matrix W i l 2,1 norm; is the label correlation regularization term, tr represents the trace of the matrix, that is, the sum of the elements on the main diagonal of the matrix; P il =Path(i,l) represents the path distance between internal node i and internal node l: Path(i,l)=Path(i,0)+Path(l,0)-2×Path(LCA(i,l),0) (3) 0≤i≤N, 0≤l≤N, i=0 or l=0 represents the root node in the hierarchical tree structure, LCA(i,l) represents the lowest common ancestor node of the i-th internal node and the l-th internal node; 2tr((X i W i ) T L i (X i W i )) is the instance relevance regularization term of the i-th internal node, L i Represents C i n i ×n i The Laplace matrix, C i represents the instance correlation matrix, represents the correlation between instance p and instance q under the i-th subtask; λ, α, and β are all non-negative constants, controlling sparsity, label relevance, and instance relevance, respectively; According to the objective function, the weight matrix of each internal node is optimized. The weight matrix W i The expression is as follows: According to the maximum number of categories in the internal node d max Initialize weights W=[W0,W1,...,W N ]; Iterate the calculation through formula (12) until convergence to obtain the optimized weight matrix.
2. The hierarchical feature selection method based on label correlation and instance correlation according to claim 1, characterized in that: The conversion of the single-label matrix into a multi-label matrix is specifically as follows: the single-label matrix D is converted into a multi-label matrix Y=[y 1 ;y 2 ;...;y n ]∈R n×d , where y j ={0,1} d (1≤j≤n) represents the label of instance j, and d represents the number of categories.
3. The hierarchical feature selection method based on label correlation and instance correlation according to claim 1, characterized in that: The correlation between instance p and instance q under the i-th subtask is calculated as follows: In the formula, and They represent instance p and instance q under the i-th subtask respectively, and the parameters Represents the Euclidean distance The k nearest neighbor instances of .
4. A hierarchical feature selection system based on label correlation and instance correlation, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory. When the processor executes the computer program, the method specifically performs the steps in the hierarchical feature selection method based on label correlation and instance correlation as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Hierarchical multi-label categorization method suitable for legal identification
CN107577785A
A feature selection method based on a hierarchical deep network
CN109919177A