Multi-view feature selection method, model training method, equipment and program product
By determining the weight of the feature subset in the multi-view feature selection method and optimizing feature selection, the problem of poor feature selection effect under the category imbalance problem is solved, and a better feature selection effect is achieved.
Patent Information
- Application Number
- CN202411853509.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-06
AI Technical Summary
When facing the problem of category imbalance, the feature selection effect of existing multi-view feature selection is not ideal, resulting in the inability to obtain better performance for subsequent machine learning tasks.
By obtaining the multi-view dataset, the distribution difference of the eigenvalue in each feature subset is determined, the weight of the feature subset is determined based on the distribution difference, and the optimal feature subset is output through the objective function optimization iteration.
By paying attention to the difference level of the characteristic value distribution of original data of different categories, we can treat the original data of all categories equally, alleviate the impact of category imbalance problem, and thus improve the effect of feature selection.
Smart Images

Figure CN119939201A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and in particular, relates to a multi-view feature selection method, a model training method, a device and a program product. Background Art
[0002] Multi-view Feature Selection is a feature selection method for processing multi-view data. Multi-view data often comes from different sources or has different modalities (such as text, images, and sounds, etc.), and each view provides a different description of the same object. Multi-view feature selection aims to extract a set of the most representative and relevant feature subsets from these different views in order to obtain better performance (such as improving accuracy, efficiency, etc.) in subsequent machine learning tasks (such as classification tasks, clustering tasks, regression tasks, anomaly detection, time series prediction, information retrieval, image processing, dimensionality reduction, etc.).
[0003] However, in the current multi-view feature selection method, when faced with the problem of class imbalance (such as the amount of data in one category is much larger than the amount of data in another category), the feature selection effect is not ideal, resulting in the subsequent machine learning tasks unable to achieve better performance. Summary of the invention
[0004] The embodiments of the present application provide a multi-view feature selection method, a model training method, a device and a program product to solve the problem that the feature selection effect is not ideal when facing the problem of class imbalance in the current multi-view feature selection method.
[0005] In a first aspect, an embodiment of the present application provides a multi-view feature selection method, comprising:
[0006] Acquire a first multi-view data set; the first multi-view data set includes a plurality of first views, any of the first views represents a first feature set obtained by extracting features from an original data set in a description manner, the first feature set includes one or more first feature subsets, and the first feature subsets include first feature values corresponding to original data in the original data set; optionally, a description manner may be a modality, an angle, a source, or a feature extraction method;
[0007] Determining a distribution difference of first feature values in each of the first feature subsets;
[0008] determining a weight of the first feature subset according to a distribution difference of first feature values in the first feature subset;
[0009] An optimal first feature subset is determined from a plurality of first feature subsets according to the weights of the first feature subsets.
[0010] In a possible implementation manner of the first aspect, after determining the weight of the first feature subset according to the distribution difference of the first feature value in the first feature subset, the method further includes:
[0011] determining the first view weight according to the weight of the first feature subset in the first view;
[0012] adjusting the weight of the first view according to the view adjustment factor;
[0013] determining a comprehensive weight of the first feature subset according to the adjusted weight of the first view and the weight of the first feature subset in the first view;
[0014] The determining an optimal first feature subset from a plurality of first feature subsets according to the weight of the first feature subset comprises:
[0015] An optimal first feature subset is determined from a plurality of first feature subsets according to the comprehensive weights of the first feature subsets.
[0016] In a second aspect, an embodiment of the present application provides a multi-view feature selection model training method, comprising:
[0017] Acquire a second multi-view dataset; the second multi-view dataset includes a plurality of second views, any of the second views represents a second feature set obtained by extracting features from a sample set in a description manner, the second feature set includes one or more second feature subsets, the second feature subsets include second feature values of the samples, the sample set includes a plurality of category samples, and each of the second feature values is annotated with category information;
[0018] Inputting the second multi-view data set into an objective function of a multi-view feature selection model, wherein the objective function is determined according to a binary function, a weight factor, and a sparse regularization term; wherein the binary function is used to determine the distribution difference between the second feature values of samples of different categories in each second feature subset according to the category information, the weight factor is used to determine the weight of the second feature subset according to the distribution difference between the second feature values of samples of different categories, and the sparse regularization term is used to perform sparse learning on the second feature subset according to the weight of the second feature subset;
[0019] The objective function of the multi-view feature selection model is optimized iteratively until the objective function converges, and an optimal second feature subset is output; the optimal second feature subset includes a plurality of selected second feature subsets.
[0020] In a possible implementation manner of the second aspect, determining, according to the category information, a distribution difference between the second feature values of samples of different categories in each of the second feature subsets includes:
[0021] Determine at least one of the mean, variance, standard deviation, and range of the second feature values of samples of different categories in each of the second feature subsets;
[0022] Determine distribution differences between the second eigenvalues of samples of different categories in the second feature subset according to at least one of the mean, variance, standard deviation, and range of the second eigenvalues of samples of different categories in the second feature subset.
[0023] In a possible implementation manner of the second aspect, the weight factor is further used for:
[0024] determining a weight of the second view according to the weight of the second feature subset in the second view;
[0025] adjusting the weight of the second view according to the view adjustment factor;
[0026] determining a comprehensive weight of the second feature subset according to the adjusted weight of the second view and the weight of the second feature subset in the second view;
[0027] The sparse regularization term is used to perform sparse learning on the second feature subset according to the weight of the second feature subset, including:
[0028] The sparse regularization term is used to perform sparse learning on the second feature subset according to the comprehensive weight of the second feature subset.
[0029] In a possible implementation manner of the second aspect, the objective function is:
[0030]
[0031] Among them, α v is the vth second view weight factor, used to determine the weight of the second view; feature selection variable W v is the vth second view mapping matrix, in which each element represents the weight information of the second feature subset; represents the distribution difference between the i-th sample and the j-th sample of the second feature subset in the v-th second view; represents the distribution difference between any two sample categories of the second feature subset in the vth second view except the i-th sample and the j-th sample; p and λ are adjustable hyperparameters, where p is the view adjustment factor and λ is the regularization strength hyperparameter.
[0032] In a possible implementation manner of the second aspect, the iteratively optimizing the objective function of the multi-view feature selection model until the objective function converges and outputting an optimal second feature subset includes:
[0033] Dividing the second multi-view dataset into a training set and a test set based on multi-fold cross validation;
[0034] Initialize the feature selection variable W v and the second view weight factor α v ;
[0035] According to the training set and the test set, the feature selection variable W in the objective function is v and the second view weight factor α v Perform alternating iterative updates;
[0036] When each fold iteration satisfies the convergence condition, the next fold iteration is performed until each fold iteration satisfies the convergence condition, thereby obtaining the multi-view feature selection model that has been trained;
[0037] According to the trained multi-view feature selection model, an optimal second feature subset is output; the optimal second feature subset is the second feature subset whose weight is greater than a preset threshold.
[0038] In a possible implementation manner of the second aspect, the feature selection variable W in the objective function is selected according to the training set and the test set. v and the second view weight factor α v Perform alternating iterative updates, including:
[0039] Determine the distribution difference between the second feature values of samples of different categories in each of the second feature subsets by KL divergence;
[0040] Determine the feature selection variable W v The difference in the standardized KL divergence distribution under ;
[0041] According to the standardized distribution difference, optimize the feature selection variable W v ;
[0042] Select variable W based on the optimized features v And the view adjustment factor, optimize the second view weight factor α v .
[0043] In a third aspect, an embodiment of the present application provides a multi-view feature selection device, comprising:
[0044] an acquisition module, configured to acquire a first multi-view data set; the first multi-view data set includes a plurality of first views, any of the first views represents a first feature set obtained by extracting features from an original data set in a description manner, the first feature set includes one or more first feature subsets, and the first feature subsets include first feature values corresponding to original data in the original data set;
[0045] A first determination module, configured to determine a distribution difference of a first feature value in each of the first feature subsets;
[0046] A second determination module, configured to determine a weight of the first feature subset according to a distribution difference of first feature values in the first feature subset;
[0047] The third determination module is used to determine an optimal first feature subset from multiple first feature subsets according to the weights of the first feature subsets.
[0048] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method described in any one of the first aspect or the second aspect is implemented.
[0049] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of the first aspect or the second aspect is implemented.
[0050] In a sixth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device executes any one of the methods in the first or second aspect above.
[0051] It can be understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0052] Compared with the prior art, the embodiments of the present application have the following beneficial effects: the embodiments of the present application determine the weight of the first feature subset by the distribution difference of the first eigenvalue in the first feature subset, and then determine the optimal first feature subset according to the weight of the first feature subset, thereby achieving equal treatment of original data of all categories by focusing on the distribution difference of the first eigenvalue of original data of different categories, rather than focusing on the quantity difference of original data of different categories, so as to reduce the impact of the category imbalance problem, and thus achieve better effect of feature selection. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0054] Figure 1 It is a flowchart of a multi-view feature selection model training method provided in one embodiment of the present application;
[0055] Figure 2 It is a schematic diagram of the principle of a multi-view feature selection model training method provided in one embodiment of the present application;
[0056] Figure 3 is a dendrogram of a second multi-view data set provided by an embodiment of the present application;
[0057] Figure 4 is a sample set about leaves provided in an embodiment of the present application;
[0058] Figure 5 is a flowchart of a multi-view feature selection method provided by an embodiment of the present application;
[0059] Figure 6 It is a structural diagram of a multi-view feature selection device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0060] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0061] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0062] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0063] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0064] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0065] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0066] Glossary:
[0067] KL divergence (Kullback-Leibler divergence) is an information-theoretic measure for measuring the difference between two probability distributions. It can be used to compare the similarity of two distributions.
[0068] Sparse learning is a class of machine learning methods that aims to identify a small number of important features or variables from high-dimensional data while ignoring most irrelevant or redundant information. Sparse learning introduces sparsity constraints or regularization terms to force the model parameters (such as weights or coefficients in feature selection) to have only a few non-zero values, which helps to simplify the model, improve generalization ability, and enhance the interpretability of the model when processing high-dimensional data. Its core idea is "sparse representation", that is, using only a few features to represent data in a high-dimensional data space, thereby reducing dimensions and avoiding overfitting. Commonly used sparse learning techniques include Lasso regression (l1-regularization), sparse coding, and compressed sensing.
[0069] Feature selection is a technique in machine learning and data mining that aims to select the most useful features or variables for model prediction from the raw data while removing redundant, irrelevant or noise features. This method helps to simplify the model, reduce overfitting, increase the training speed of the model, and enhance the generalization ability and interpretability of the model. The model in this paragraph can be a classification model, a prediction model, a clustering model, etc.
[0070] The existing multi-view feature selection first extracts features from samples of each view to obtain one or more feature subsets, and then combines feature subsets from different views to obtain a multi-view dataset. Then, statistical indicators (such as mutual information, correlation coefficient) or machine learning models (such as decision trees, random forests) are used to evaluate the relationship between each feature subset and the class label. In this process, regularization terms (l1-norm, l 2,1 -norm) to control the sparsity of feature subset selection and prevent overfitting. Finally, an optimization algorithm (such as genetic algorithm, particle swarm optimization) is used to select the optimal feature subset to ensure that the representativeness of the features is maintained in multiple views.
[0071] However, the existing multi-view feature selection has the following disadvantages:
[0072] (1) When dealing with the problem of multi-view category imbalance, the feature selection effect is poor;
[0073] (2) Lack of dynamic adjustment of the importance of different views;
[0074] (3) When the distribution difference of eigenvalues in feature subsets is small, it is difficult to achieve sparse learning.
[0075] To this end, this embodiment provides a multi-view feature selection model training method, see Figure 1-Figure 2 , Figure 1 FIG. 4 is a flow chart showing a multi-view feature selection model training method provided by this embodiment. Figure 2 The schematic diagram of the principle of a multi-view feature selection model training method provided by this embodiment is shown. The method includes S110-S130:
[0076] S110: Acquire a second multi-view dataset; the second multi-view dataset includes multiple second views, any second view represents a second feature set obtained by extracting features from a sample set in a description manner, the second feature set includes one or more second feature subsets, the second feature subset includes second feature values of the samples, the sample set includes multiple category samples, and each second feature value is annotated with category information.
[0077] In this embodiment, the elements in the second multi-view dataset are feature values extracted from samples.
[0078] For details, see Figure 3 , Figure 3 A dendrogram of a second multi-view data set of an embodiment of the present application is shown, wherein the second multi-view data set includes multiple second views, each second view is a second feature set. Each second feature set includes one or more second feature subsets. Each second feature subset includes second feature values corresponding to the second feature of all samples.
[0079] For example, for the sample categories in the sample set: different animals, such as cats, dogs, cows, etc., belong to different categories; different objects, such as pots, bowls, ladles, basins, etc., also belong to different categories; the same leaves, but belonging to different species, such as maple leaves, poplar leaves, willow leaves, etc., also belong to different categories. The sample set includes multiple categories, and each sample is annotated with its own category information.
[0080] A description can be a modality, an angle, a source, or a feature extraction method, etc.
[0081] The modality of the sample can be image, audio, text, video, etc.
[0082] If the modality of some samples in the sample set is an image, then this part of the samples describes one or more categories through images, for example, the image can be a photo of a cat, a dog, a cow, etc. If this part of the samples includes RGB images, depth images, edge images, etc., then by extracting feature values from the RGB image, an RGB image view can be obtained; by extracting feature values from the depth image, a depth image view can be obtained; by extracting feature values from the edge image, an edge image view can be obtained. For the RGB image view, the following features can be extracted from the RGB image: R channel pixel value feature, G channel pixel value feature, B channel pixel value feature, image average brightness feature, image contrast feature, etc. For the depth image view, the following features can be extracted from the depth image: depth value (Z axis) feature, depth gradient feature, average depth feature, depth image variance feature, etc. For the edge image view, the following features can be extracted from the edge image: edge intensity feature, edge direction feature, edge quantity feature, average edge length feature, etc.
[0083] If the modality of some samples in the sample set is audio, then this part of the samples describes one or more categories through audio. For example, the audio can be the cry of a cat, dog, or cow, or a paragraph describing the shape of a cat, dog, or cow. Feature values can be extracted from this part of the audio samples according to different feature extraction methods and application scenarios to obtain a time domain view, a frequency domain view, a timbre view, and the like. For the time domain view, the following features can be extracted from the audio sample: energy features, amplitude features, waveform features, and the like. For the frequency domain view, the following features can be extracted from the audio sample: spectrum features, power spectral density (PSD) features, Mel-frequency cepstral coefficient (MFCC) features, and the like. For the timbre view, the following features can be extracted from the audio sample: loudness features, brightness features, roughness features, and the like.
[0084] Similarly, if the modality of the sample is text or video, the second feature value can also be extracted from the text sample or the video sample to obtain multiple second views (ie, multiple second feature sets).
[0085] The above-mentioned RGB image view, depth image view, edge image view, time domain view, frequency domain view, and timbre view are all second views described in this embodiment, and each second view provides information on different aspects of the sample. The above-mentioned R channel pixel value features, G channel pixel value features, B channel pixel value features, image average brightness features, image contrast features, energy features, amplitude features, waveform features, spectrum features, power spectral density (PSD) features, Mel frequency cepstral coefficient (MFCC) features, loudness features, brightness features, roughness features, etc. are all second features in this embodiment.
[0086] It is easy to understand that if the sample is an image, the second feature value cannot be extracted from the audio sample and the text sample, and the second feature (such as amplitude feature, loudness feature, etc.) of the audio sample and the text sample corresponding to the image sample can be recorded as 0. If the sample set includes 1000 samples, there is a corresponding second feature value for each second feature, thereby forming a second feature subset corresponding to each second feature, that is, each second feature subset has 1000 elements corresponding to 1000 second feature values.
[0087] Different second views can be obtained based on samples at different angles. For example, a piece of furniture may have one or more samples of its main view, rear view, top view, bottom view, left view, right view, axonometric view, and sectional view. Second feature values can be extracted from the samples at these angles to obtain multiple second views, namely: main view, rear view, top view, bottom view, left view, right view, axonometric view, and sectional view. Each second view may include one or more second features, and each second feature corresponds to its own second feature subset.
[0088] Different second views can be obtained based on samples from different sources. Different sources can be different data sources (such as image data, text data, medical data), different collection methods (sensor collection data, device collection data), different application scenario sources (medical data, financial data, etc.), different feature extraction methods (for example, for text data, different extraction methods such as bag-of-words model, TF-IDF, Word2Vec, etc.), and different data acquisition channels (such as experimental measurement, instrument collection, manual annotation, web crawling, etc.).
[0089] According to different feature extraction methods, features can be extracted from samples in the sample set to obtain different second views. For example, for text data, multiple second views can be obtained through different feature extraction methods such as bag-of-words model, TF-IDF, Word2Vec, etc.
[0090] S120: Inputting the second multi-view data set into the objective function of the multi-view feature selection model, the objective function is determined according to a binary function, a weight factor and a sparse regularization term; wherein the binary function is used to determine the distribution difference between the second eigenvalues of samples of different categories in each second feature subset according to the category information, the weight factor is used to determine the weight of the second feature subset according to the distribution difference between the second eigenvalues of samples of different categories, and the sparse regularization term is used to perform sparse learning on the second feature subset according to the weight of the second feature subset.
[0091] Current multi-view feature selection methods have poor feature selection effects when dealing with the multi-view class imbalance problem. The so-called class imbalance problem refers to the phenomenon that if the number of samples in one class is much larger than the number of samples in another class, this phenomenon will have a negative impact on the performance of the model because the model will tend to over-focus on the majority class and ignore the minority class.
[0092] This embodiment uses a binary function to calculate the distribution difference between the second eigenvalues of samples of different categories in each second feature subset, so as to solve the problem of class imbalance. This is because whether it is a sample of the majority category or a sample of the minority category, the weight of the second feature subset can be determined by determining the distribution of the second eigenvalues of samples of different categories, so as not to pay attention to the impact of the number of samples of different categories on the second feature subset, thereby solving the problem of class imbalance. In other words, this embodiment uses a binary function to make the multi-view feature selection model focus on the distribution difference level of the second eigenvalues of samples of different categories, rather than focusing on the quantitative difference level of samples of different categories, so as to treat samples of all categories equally to reduce the impact of the class imbalance problem.
[0093] The weight factor is used to determine the weight of the second feature subset according to the distribution difference between the second feature values of samples of different categories. Exemplarily, for a certain second feature, such as the image average brightness feature, in the second feature subset corresponding to the image average brightness feature, the distribution difference (such as normal distribution or discreteness, etc.) of the second feature values of different categories is determined. If the distribution difference between the image average brightness features of all or most of the category samples is very large, it means that the discrimination of the second feature is relatively large, and a larger weight can be given to the second feature subset. In this way, in subsequent tasks (such as classification tasks, prediction tasks, etc.), it is easy to separate samples of different categories through the second feature subset. On the contrary, for another second feature, such as the R channel pixel value feature, in the second feature subset corresponding to the R channel pixel value feature, if the distribution difference of the R channel pixel value features of all or most of the category samples is very small, it means that the discrimination of the second feature is relatively small, and a smaller weight can be given to the second feature subset. If the second feature is retained, it is difficult to separate samples of different categories through the second feature subset in subsequent tasks (such as classification tasks, prediction tasks, etc.). In the subsequent feature selection process, the second feature subsets with large weights will be retained, and the second feature subsets with small weights will be eliminated. This helps to simplify the model (referring to the model in subsequent tasks, such as prediction model, classification model or clustering model, etc.), reduce overfitting, increase the training speed of the model, and enhance the generalization ability and interpretability of the model.
[0094] The sparse regularization term is used to perform sparse learning on the second feature subset according to the weight of the second feature subset. Specifically, the sparse regularization term is used to learn the sparsity between the second feature subsets, which makes some second feature subsets with smaller weights as close to zero as possible, thereby achieving the effect of feature selection and model compression.
[0095] S130: Optimizing and iterating the objective function of the multi-view feature selection model until the objective function converges, and outputting an optimal second feature subset; the optimal second feature subset includes a plurality of selected second feature subsets.
[0096] After the objective function converges, an optimal second feature subset can be generated by retaining the second feature subsets with larger weights and discarding the second feature subsets with smaller or zero weights. The optimal second feature subset includes the second feature subsets with the greatest impact on the target variable in the second multi-view dataset, thereby significantly reducing the dimension of the sample data while retaining the key information of the sample data.
[0097] As an optional implementation manner, before S110: acquiring the second multi-view data set, the process further includes S101-S103.
[0098] S101: Obtain a sample set.
[0099] Specifically, the samples in the sample set can come from different modalities, different angles or different sources, so as to better describe objects of different categories. Each sample is annotated with category information.
[0100] S102: Fusing the sample sets.
[0101] The samples in the sample set may have information redundancy, modal differences, or even inconsistent dimensions, which requires sample data fusion. Sample data fusion involves feature extraction from the samples in the sample set, which can be achieved by directly concatenating features, transforming features, or by joint modeling. Fusion can integrate samples from different sources, different modalities, and different angles to obtain more comprehensive and accurate information.
[0102] S103: Standardize the fused sample set to obtain a second multi-view data set.
[0103] Since the dimensions and ranges of the second view data may be different, unstandardized features may cause some features to be over-emphasized or neglected during the optimization process. Therefore, standardization can be performed on the fused sample set (such as Z-score standardization with a mean of 0 and a variance of 1) to make each second feature at the same level and prevent the imbalance of weights between the second features from affecting the optimization results.
[0104] As an optional implementation, in S120, determining the distribution difference between the second feature values of samples of different categories in each second feature subset according to the category information includes S121-S122:
[0105] S121: Determine at least one of the mean, variance, standard deviation, and range of the second feature values of samples of different categories in each second feature subset.
[0106] S122: Determine the distribution difference between the second eigenvalues of the samples of different categories in the second feature subset according to at least one of the mean, variance, standard deviation, and range of the second eigenvalues of the samples of different categories in the second feature subset.
[0107] Specifically, any one of the mean, variance, standard deviation, and range of the second eigenvalue can reflect the distribution difference between the second eigenvalues of samples of different categories. This distribution difference can be reflected in the form of dispersion, normal distribution, histogram, density map, etc. It is easy to understand that if any one of the mean, variance, standard deviation, and range of the second eigenvalues of samples of two categories differs greatly, it means that the samples of the two categories are easier to distinguish, and the distribution difference between the two categories is greater.
[0108] As an alternative implementation, the weight factor in this embodiment not only focuses on the weights at the second feature subset level but also on the weights at the second view level. Specifically, the weight factor in this embodiment is also used for S123 - S125.
[0109] S123: Determine the weight of the second view according to the weights of the second feature subsets in the second view.
[0110] Since there is one or more second feature subsets in each second view, after knowing the weights of all the second feature subsets in the second view, the weight of the second view can be calculated.
[0111] S124: Adjust the weight of the second view according to the view adjustment factor.
[0112] The view adjustment factor is an adjustable hyperparameter, namely p, which is used to avoid obtaining a trivial solution. If p is 1, it is equivalent to not having this adjustable hyperparameter. Then it may lead to (when the quality gap between views is large) only selecting one second view, that is, there will be only one second view, and this view adjustment factor is 1. At this time, the other second views are completely useless.
[0113] When p > 1: It increases the imbalance of the weight distribution, making the more important second view obtain a higher weight, while the relatively unimportant second view is further weakened. This setting is suitable for emphasizing the dependence on the second view with high information content when the quality difference between second views is large.
[0114] When p = 1: The weight distribution is linear, and no additional amplification or suppression is performed on the importance of the second view. At this time, the adjustment of the weight is directly determined by the geometric description performance of the second view, which is suitable for the case when the quality difference between second views is small.
[0115] When 0 < p < 1: It enhances the balance of the weight distribution, that is, weakens the preference for the second view with high information content, making the weight distribution more uniform. This setting is suitable for avoiding over - dependence on certain second views when the multi - view information complementarity is strong.
[0116] S125: Determine the comprehensive weight of the second feature subset according to the adjusted weight of the second view and the weights of the second feature subsets in the second view.
[0117] Optionally, the comprehensive weight of the second feature subset can be the adjusted weight of the second view multiplied by the weights of the second feature subsets in the second view.
[0118] Correspondingly, in S120: The sparse regularization term is used to perform sparse learning on the second feature subset according to the weights of the second feature subset, including S126.
[0119] S126: The sparse regularization term is used to perform sparse learning on the second feature subset according to the comprehensive weight of the second feature subset.
[0120] This embodiment considers the second view level. If a feature of a second view is more advantageous in distinguishing different categories, its corresponding second view weight factor will become larger, resulting in the second view being given a higher influence in multi-view feature selection. The final comprehensive weight of the second feature subset can also better reflect the contribution of the second feature subset to category distinction.
[0121] As an optional implementation, the objective function of the multi-view feature selection model in the embodiment of the present application is based on a binary function, a second view weight factor (including a weight at the second view level and a weight at the second feature subset level) and a new sparse regularization term (i.e. -norm) is constructed. The objective function consists of the following parts:
[0122] Loss function: The KL divergence is used to measure the ability of the second feature subset to distinguish between categories, aiming to minimize the difference of the second features between the same categories and maximize the difference of the second features between different categories.
[0123] Sparse regularization term: using -norm regularization term is used to encourage sparse solutions so that only important second feature subsets are retained during feature selection, and redundant or irrelevant second feature subsets are discarded. The constructed objective function is:
[0124]
[0125] Among them, α v is the vth second view weight factor, used to determine the weight of the second view; feature selection variable W v is the vth second view mapping matrix, where each element represents the weight information of the second feature subset. represents the distribution difference between the i-th sample and the j-th sample of the second feature subset in the v-th second view; It represents the distribution difference of any two sample categories other than the i-th sample and the j-th sample in the second feature subset of the v-th second view. p and λ are adjustable hyperparameters, where p is the view adjustment factor and λ is the regularization strength hyperparameter. For α v restrictions. About W v It can make the selected second feature subsets independent of each other, thereby reducing redundant information and making the selected second feature subsets more representative.
[0126] The second view weight factor of this embodiment is equal to the weight of the second view. It has the following four functions:
[0127] (1) Adjusting the distribution sensitivity of view weights: The degree to which the weight distribution responds to the importance of the second view is controlled. As mentioned above, the difference in p determines the adjustment direction of the second view.
[0128] (2) Balancing the importance of global and local information: The introduction of actually adds a layer of nonlinear adjustment to the optimization of the second view weight, so that the importance of global and local information can be more flexibly balanced. By adjusting p, the sensitivity of the multi-view feature selection model to the extreme value of the second view weight factor can be changed, and the contribution ratio of different second views to the objective function can be controlled.
[0129] (3) Improving the robustness of the model to redundant views: A smaller p-value can prevent some second views from obtaining excessive weights due to accidental noise or redundant features, reduce the interference of redundant second views on the feature selection results, and thus improve the stability of the multi-view feature selection model.
[0130] (4) Hyperparameter tuning space: The adjustable hyperparameter p provides a hyperparameter tuning space for the multi-view feature selection model, making the objective function more adaptable to different tasks or datasets. For small sample sets or large differences between classes: increase p (such as p>1) to amplify the contribution of important views. For data with strong multimodal complementarity or more noise: reduce p (such as p<1) to enhance the balance and noise resistance of multi-views.
[0131] The objective function of this embodiment is obtained by introducing a new sparse regularization term ( -norm), compared with the existing l1-norm, l 2,1 -norm, etc. can perform feature learning more accurately. Therefore, when the distribution difference of the second eigenvalues in the second feature subset is small, accurate sparse learning can also be achieved.
[0132] In formula (1), the first term (i.e., the part before “+” in formula (1)) is mainly used to balance the differences between different categories, so that the samples in the majority class and the minority class are treated equally. In addition, it is also used to consider the different information of different second views, so as to fully explore the importance of different second views. The second term (i.e., the part after “+” in formula (1)) is a new sparse regularization term ( -norm), which can capture W more accurately vThe actual sparsity of the matrix. In summary, the proposed objective function (i.e., formula (1)) mainly focuses on the class imbalance problem, view importance, and sparse solution for multi-view feature selection. The construction of the objective function aims to find the optimal second feature subset by weighing the sparsity of feature selection, weight adjustment in view fusion, and balanced processing of multi-class samples.
[0133] This embodiment introduces a binary function to balance category differences, learns the importance of multiple views through the second view weight factor, and enhances the interpretability and computational efficiency of the model through the sparse regularization term. This method can maintain a stable and robust feature selection effect when facing massive data, category imbalance problems, and complex modal data. At the same time, its alternating iterative optimization strategy ensures the efficiency of the solution, enabling this method to process data with extremely high feature dimensions and has the potential to be widely used in actual big data scenarios.
[0134] As an optional implementation, in S130, the objective function of the multi-view feature selection model is optimized and iterated until the objective function converges, and the optimal second feature subset is output, including S131-S135.
[0135] S131: Divide the second multi-view dataset into a training set and a test set based on multi-fold cross validation.
[0136] In order to evaluate the robustness of the feature selection method and prevent the multi-view feature selection model from overfitting, the second multi-view dataset can be split using a ten-fold cross-validation strategy. Specifically, the entire second multi-view dataset is randomly and evenly divided into 10 parts, of which 9 are used as training sets and 1 is used as a test set. By using a different part as the test set and the remaining nine as the training set each time, the test is repeated (i.e., a different fold is selected as the test set each time, and the rest are used as the training set) until all data are tested. Finally, the ten verification results are averaged to obtain an overall performance evaluation of the multi-view feature selection model. Ten-fold cross-validation can make full use of sample data, improve the generalization performance of the multi-view feature selection model, and provide a reliable evaluation criterion for subsequent feature selection and classification verification.
[0137] S132: Initialize feature selection variable W v and the second view weight factor α v .
[0138] The feature selection variable W is a weight matrix, in which the elements represent the selection weights of the second feature subset. The initial value can be assigned randomly or preliminarily estimated based on the sample data distribution. The second view weight factor α vIt is used to measure the importance of each second view in feature selection. The initial weight of each second view is also randomly set, and then these weights are optimized through iterative updates.
[0139] S133: Select variable W in the objective function based on the training set and the test set v and the second view weight factor α v Perform alternating iterative updates.
[0140] In order to find the global optimal solution for feature selection, an alternating iterative algorithm can be used for the objective function based on the training set and the test set, that is, optimizing W v When α is fixed v , optimize α v When W is fixed v , alternating.
[0141] S134: When each fold iteration satisfies the convergence condition, the next fold iteration is performed until each fold iteration satisfies the convergence condition, thereby obtaining a trained multi-view feature selection model.
[0142] When each fold iteration meets the convergence condition, the next fold iteration means: when the current 9 training sets and 1 test set meet the convergence condition, one of the current 9 training sets is used as the test set, and the replaced training set is used as the test set, and the iteration continues.
[0143] S135: Outputting an optimal second feature subset according to the trained multi-view feature selection model; the optimal second feature subset is a second feature subset whose weight is greater than a preset threshold.
[0144] When the optimization process of the objective function meets the convergence condition, the iteration stops. The convergence condition is determined by the following criteria: The change in the objective function value tends to zero: If the change in the objective function tends to be stable after multiple iterations, it means that the optimization has approached the global optimal solution. At the end of the iteration, the final optimized feature selection variable W is output v , we can obtain the optimal second feature subset selected under the global optimal solution. These second feature subsets will be used for subsequent data processing tasks, such as classification, regression or cluster analysis.
[0145] The second feature subsets selected in the feature selection variable W are all second feature subsets whose weights are greater than a preset threshold. For example, by selecting a second feature subset whose weight is not 0, the optimal second feature subset is obtained. It is also possible to sort by weight and select a preset number of second feature subsets.
[0146] It should be noted that the multi-view feature selection model of the embodiment of the present application can be applied to the single-label case, that is, each sample in the sample set has only one label, and can also be applied to the multi-label case, that is, the samples in the sample set can have multiple labels. If the second multi-view dataset with a single label is input into the multi-view feature selection model, each second eigenvalue in the second feature subset in the feature selection variable is a vector. If the second multi-view dataset with multiple labels is input into the multi-view feature selection model, each second eigenvalue in the second feature subset in the feature selection variable is a matrix.
[0147] As an optional implementation, S133: select variable W in the objective function based on the training set and the test set. v and the second view weight factor α v Perform alternating iterative updates, including S136-S139.
[0148] S136: Determine the distribution difference between the second feature values of samples of different categories in each second feature subset by KL divergence.
[0149] Exemplarily, the distribution difference between the second eigenvalues of samples of different categories in the second feature subset is calculated by KL divergence. In multi-view feature selection, KL divergence can effectively capture the distribution difference of the second eigenvalues between categories. By measuring the distribution difference of the second eigenvalues of samples of different categories, it can help identify the second features that are discriminative for the classification task. The specific formula is as follows:
[0150]
[0151] Among them, X v represents the vth second view; N represents the normal distribution. and They respectively represent the normal distribution information of the second eigenvalues of the i-th class samples of the second feature subset in the v-th second view, and the normal distribution information of the second eigenvalues of the j-th class samples. and is the average of the second eigenvalue of the i-th sample and the second eigenvalue of the j-th sample of the second feature subset in the corresponding v-th second view. and is the covariance matrix of the second eigenvalue of the i-th class sample and the second eigenvalue of the j-th class sample under the v-th second view corresponding to the normal distribution information. and The covariance matrices are and The determinant of . is a symmetric matrix. tr is the trace of the matrix.
[0152] S137: Determine the feature selection variable W v The difference in the normalized KL divergence distribution under .
[0153] After calculating the KL divergence difference of samples of different categories, the difference is compared with the current feature selection variable W v Combined. Thus, W is used to reweight these differences to reflect the contribution of each second feature subset to category distinction under the current feature selection scheme. As shown in the following formula:
[0154]
[0155] Among them, W v is the mapping matrix of the v-th second view data, which retains the weight information of each second feature subset under the v-th second view. y=i and y=j represent the i-th and j-th samples respectively. P represents the probability, and the greater the probability, the greater the contribution. The distribution difference between different categories is calculated by KL divergence. In addition, weighted normalization is used to obtain the standardized difference value of the second eigenvalue of the second feature subset in each second view between different categories. This normalization process ensures that the contribution of the second feature subset matches its selection weight and balances the distribution difference of the second eigenvalues of different categories. As shown in the following formula:
[0156]
[0157] in, represents the distribution difference between the i-th sample and the j-th sample of the second feature subset in the v-th second view; r i and r j They represent the number of samples in the i-th and j-th categories of the second feature subset in the v-th second view respectively.
[0158] S138: Optimize feature selection variable W based on standardized distribution differences v .
[0159] Specifically, in each iteration, W is adjusted according to the following update formula: v The value of , so that it moves towards a better solution. The update formula is:
[0160]
[0161] in, Indicates derivative.
[0162] S139: Select variable W based on optimized features v And the view adjustment factor, optimize the second view weight factor α v .
[0163] The second view weight factor determines the importance of different second views in feature selection. Adaptively adjusting these second view weight factors helps balance the contribution of multi-view features. The second view weight factor is based on the contribution of each view to the overall goal. The formula is as follows:
[0164]
[0165] in
[0166] After completing one iteration (i.e., S136-S139), the number of iterations t is increased by 1, and the next iteration is entered. In each iteration, the feature selection variable W v and the second view weight α v It will be continuously updated to approach the global optimal solution. After each iteration, the current optimization result is evaluated to determine whether the convergence condition has been met.
[0167] As an optional implementation manner, after outputting the optimal second feature subset at S130, the method further includes S140.
[0168] S140: Perform classification verification on the optimal second feature subset.
[0169] In order to verify the effectiveness of the selected optimal second feature subset, the optimal second feature subset can be classified using a traditional classification algorithm such as support vector machine (SVM). SVM can effectively process high-dimensional data, and through kernel function mapping, it can process linearly inseparable feature sets. By classifying the optimal second feature subset, the actual effect of the feature selection method is evaluated, and the performance of the classification model is evaluated by the indicator accuracy. The results of the classification verification can help determine the adaptability and robustness of the selected optimal second feature subset in different tasks, thereby providing feedback for the improvement and application of the feature selection method.
[0170] The objective function of the embodiment of the present application uses a new binary function to calculate the distribution difference between the feature values of all category samples in the same second feature subset, so as to achieve equal treatment of sample data of all categories to alleviate the impact of the category imbalance problem. Subsequently, the second view weight factor is incorporated into the objective function to learn the importance of each second view. Finally, a novel sparse regularization term is introduced for the first time, namely -norm, to more accurately learn the feature selection variable W v The sparse representation of the matrix is used for efficient feature selection of large-scale data. For example: given a data set, the input is composed of a binary function, a second view weight factor and -norm sparse regularization term in the objective function. The objective function is optimized and solved by an alternating iterative strategy until it converges to the global optimal solution, and the final variable for feature selection (i.e., the optimal second feature subset) can be obtained, thereby completing the feature selection process. This method is particularly suitable for efficient feature selection tasks in large-scale data scenarios.
[0171] As an example, see Figure 4 , Figure 4 A sample set of leaves provided by an embodiment of the present application is shown. In the sample set, three leaves in each column are one leaf, which are from different angles and therefore belong to the same category. There are three leaf images of each type. Figure 4 Only some samples in the sample set are shown in FIG. Figure 4 Take the example to illustrate the steps of the multi-view feature selection model training method, as follows:
[0172] (1) The sample set and the initial values of the adjustable hyperparameters are input into the multi-view feature selection model, the initial value of the adjustable parameter λ is set to 1, and the initial value of p is set to 2. In addition, the adjustable hyperparameters can also be randomly generated by the multi-view feature selection model.
[0173] (2) Perform data fusion on the sample set to obtain a fused multi-view dataset. Since the leaves in the sample set come from three different angles, there are three views. After data fusion, a total of 18 features are extracted, each feature corresponds to a feature subset, and each view includes 6 feature subsets. The fused data (multi-view dataset) is as follows (only the feature values of the first 20 samples are shown):
[0174]
[0175] In the above table, each column is a feature, and each row is the 18 eigenvalues (in vector form) corresponding to a sample.
[0176] (3) Standardize the multi-view dataset. The standardized feature values are as follows:
[0177]
[0178] (4) Use ten-fold cross validation to divide the standardized multi-view dataset into training set and test set.
[0179] (5) Initialize the number of iterations to 0 and randomly initialize the feature selection variables Second View Weight The randomly initialized α = [0.6084; 0.5791; 0.0103], the feature selection variable W is as follows:
[0180]
[0181] The above table shows the initialization weights of 18 feature subsets.
[0182] (6) Use KL divergence to calculate the distribution difference between feature values of different categories in each feature subset. The calculated distribution difference between the first 20 categories of feature values is as follows:
[0183]
[0184] The table above shows the distribution differences between the first 20 feature values in a feature subset. If there are 100 categories, there are 100 rows and 100 columns. Since the distribution difference of the same category is 0, you can see that the diagonal in the table above is 0.
[0185] (7) Calculate the feature selection variable The standardized KL divergence distribution difference is as follows:
[0186]
[0187] (8) Update according to equation (5) The updated W is as follows:
[0188]
[0189] (9) Update according to equation (6) After updating, α=[0.0121; 0.0322; 0.8896].
[0190] (10)t=t+1.
[0191] (11) until convergence, output the final feature selection variable W. As shown below:
[0192]
[0193] (12) According to the feature selection variable W, the selected features can be obtained as the 1st, 2nd, 4th, 8th, 10th, 12th, 13th, and 14th features.
[0194] See also Figure 5 , Figure 5 A schematic flow chart of a multi-view feature selection method disclosed in an embodiment of the present application is shown, and the method includes S210 - S240 .
[0195] S210: Obtain a first multi-view dataset; the first multi-view dataset includes multiple first views, any first view represents a first feature set obtained by extracting features from an original dataset in a description manner, the first feature set includes one or more first feature subsets, and the first feature subset includes a first feature value corresponding to the original data in the original dataset.
[0196] Raw data refers to data that has not been preprocessed. For example, raw data can be image data, text data, audio data, video data, etc. collected from source devices (such as cameras, sensors, microphones, etc.). The raw data set often includes raw data of multiple categories, but it is worth noting that the raw data set here is different from the sample set in S110. The samples in the sample set are labeled / annotated data, while the raw data in the raw data set here are unlabeled / annotated data.
[0197] A description can be a modality, a perspective, a source, or a feature extraction method.
[0198] Acquiring the first multi-view dataset in this step may be understood as inputting the first multi-view dataset into a trained multi-view feature selection model.
[0199] S220: Determine the distribution difference of the first feature values in each first feature subset.
[0200] It is easy to understand that the first eigenvalues corresponding to the original data of the same category are often the same or similar. Therefore, the distribution of the first eigenvalues corresponding to the original data of the same category is more concentrated. If the distribution of the first eigenvalues corresponding to the original data of two categories is quite different, it means that the original data of the two categories are easier to distinguish.
[0201] This step can be understood as: determining the distribution difference of the first feature value in each first feature subset through the trained multi-view feature selection model.
[0202] S230: Determine the weight of the first feature subset according to the distribution difference of the first feature value in the first feature subset.
[0203] It is easy to understand that if it is found that there are multiple data groups with very different distributions in the first feature subset (the data in the data group is the first eigenvalue), then the same data group can be considered as a category, and the large distribution difference means that the categories are easier to distinguish, so a larger weight can be given to the first feature subset with large distribution differences. For a first feature subset, if all the first eigenvalues in it are distributed more concentratedly, then a smaller weight can be given to the first feature subset.
[0204] This step can be understood as: determining the weight of the first feature subset according to the distribution difference of the first feature value in the first feature subset through the trained multi-view feature selection model.
[0205] S240: Determine an optimal first feature subset from multiple first feature subsets according to the weight of the first feature subset.
[0206] The multiple first feature subsets mentioned above refer to the first feature subsets in all the first views.
[0207] Optionally, a first feature subset having a weight greater than a preset value is selected as the optimal first feature subset, or the first feature subsets are sorted according to the weights, and a preset number of first feature subsets are selected as the optimal first feature subset.
[0208] This step can be understood as: determining the optimal first feature subset according to the weight of the first feature subset through the trained multi-view feature selection model.
[0209] It can be understood that the multi-view feature selection method (S210-S240) disclosed in this embodiment is equivalent to the deployment stage of the multi-view feature selection model after training.
[0210] This embodiment determines the weight of the first feature subset by the distribution difference of the first eigenvalue in the first feature subset, and then determines the optimal first feature subset according to the weight of the first feature subset. In this way, by focusing on the distribution difference of the first eigenvalues of original data of different categories, rather than focusing on the quantity difference of original data of different categories, the original data of all categories are treated equally, so as to reduce the impact of the category imbalance problem, and thus the effect of feature selection is better.
[0211] As an optional implementation, S230: after determining the weight of the first feature subset according to the distribution difference of the first feature value in the first feature subset, the method further includes S250-S270.
[0212] S250: Determine a first view weight according to the weight of the first feature subset in the first view.
[0213] S260: Adjust the weight of the first view according to the view adjustment factor.
[0214] Corresponding to the trained multi-view feature selection model, the first view weight is equivalent to the first view weight factor.
[0215] S270: Determine a comprehensive weight of the first feature subset according to the adjusted weight of the first view and the weight of the first feature subset in the first view.
[0216] Correspondingly, S240: determining an optimal first feature subset from multiple first feature subsets according to the weight of the first feature subset, including S241.
[0217] S241: Determine an optimal first feature subset from multiple first feature subsets according to the comprehensive weight of the first feature subset.
[0218] This embodiment considers the first view level. If the distribution difference of the first feature value of the first feature subset of a first view is larger, it means that it has more advantages in category distinction, and its corresponding first view weight can be larger, thereby causing the first view to be given a higher influence in multi-view feature selection. The obtained comprehensive weight of the first feature subset can also better reflect the contribution of the first feature subset to the distinction between different categories.
[0219] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0220] Corresponding to the multi-view feature selection method described in the above embodiment, Figure 6 A structural block diagram of a multi-view feature selection device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0221] Reference Figure 6 , the device comprises:
[0222] The acquisition module 310 is used to acquire a first multi-view data set; the first multi-view data set includes multiple first views, and any first view represents a first feature set obtained by extracting features from the original data set in a description method. The first feature set includes one or more first feature subsets, and the first feature subset includes a first feature value corresponding to the original data in the original data set.
[0223] A first determination module 320, configured to determine a distribution difference of a first feature value in each first feature subset;
[0224] The second determination module 330 is used to determine the weight of the first feature subset according to the distribution difference of the first feature value in the first feature subset.
[0225] The third determination module 340 is configured to determine an optimal first feature subset from multiple first feature subsets according to the weights of the first feature subsets.
[0226] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0227] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0228] An embodiment of the present application also provides an electronic device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor implements the steps of any of the above-mentioned method embodiments when executing the computer program.
[0229] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0230] An embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0231] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device capable of carrying the computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0232] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0233] The computer program code for performing the operation of the embodiment of the application can be written with one or more programming languages or their combination, and the programming language includes object-oriented programming languages, such as python, Matlab, Java, Smalltalk, C++, and also includes conventional procedural programming languages, such as "C" language or similar programming languages.Program code can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on the remote computer, or executed completely on the remote computer or server.In the case of a remote computer, the remote computer can include a local area network (LAN) or a wide area network (WAN)--connected to the user's computer through any type of network, or, can be connected to an external computer (for example, utilizing an Internet service provider to connect through the Internet).
[0234] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0235] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0236] In the embodiments provided in the present application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0237] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0238] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A multi-view feature selection method, characterized in that: include: Acquire a first multi-view data set; the first multi-view data set includes a plurality of first views, and any of the first views represents one or more first feature subsets obtained by extracting features from an original data set in a description manner, wherein the first feature subsets include first feature values corresponding to original data in the original data set; Determining a distribution difference of first feature values in each of the first feature subsets; determining a weight of the first feature subset according to a distribution difference of first feature values in the first feature subset; An optimal first feature subset is determined from a plurality of first feature subsets according to the weights of the first feature subsets.
2. The multi-view feature selection method according to claim 1, characterized in that: After determining the weight of the first feature subset according to the distribution difference of the first feature value in the first feature subset, the method further includes: determining a weight of the first view according to the weight of the first feature subset in the first view; adjusting the weight of the first view according to the view adjustment factor; determining a comprehensive weight of the first feature subset according to the adjusted weight of the first view and the weight of the first feature subset in the first view; The determining an optimal first feature subset from a plurality of first feature subsets according to the weight of the first feature subset comprises: An optimal first feature subset is determined from a plurality of first feature subsets according to the comprehensive weights of the first feature subsets.
3. A multi-view feature selection model training method, characterized in that: include: Acquire a second multi-view dataset; the second multi-view dataset includes a plurality of second views, any of the second views represents a second feature set obtained by extracting features from a sample set in a description manner, the second feature set includes one or more second feature subsets, the second feature subsets include second feature values of the samples, the sample set includes a plurality of category samples, and each of the second feature values is annotated with category information; Inputting the second multi-view data set into an objective function of a multi-view feature selection model, wherein the objective function is determined according to a binary function, a weight factor, and a sparse regularization term; wherein the binary function is used to determine the distribution difference between the second feature values of samples of different categories in each second feature subset according to the category information, the weight factor is used to determine the weight of the second feature subset according to the distribution difference between the second feature values of samples of different categories, and the sparse regularization term is used to perform sparse learning on the second feature subset according to the weight of the second feature subset; The objective function of the multi-view feature selection model is optimized iteratively until the objective function converges, and an optimal second feature subset is output; the optimal second feature subset includes a plurality of selected second feature subsets.
4. The multi-view feature selection model training method according to claim 3, characterized in that: The determining, according to the category information, the distribution difference between the second feature values of samples of different categories in each of the second feature subsets includes: Determine at least one of the mean, variance, standard deviation, and range of the second feature values of samples of different categories in each of the second feature subsets; Determine the distribution difference between the second eigenvalues of samples of different categories in the second feature subset according to at least one of the mean, variance, standard deviation and range of the second eigenvalues of samples of different categories in the second feature subset.
5. The multi-view feature selection model training method according to claim 3, characterized in that: The weighting factors are also used for: determining a weight of the second view according to the weight of the second feature subset in the second view; adjusting the weight of the second view according to the view adjustment factor; determining a comprehensive weight of the second feature subset according to the adjusted weight of the second view and the weight of the second feature subset in the second view; The sparse regularization term is used to perform sparse learning on the second feature subset according to the weight of the second feature subset, including: The sparse regularization term is used to perform sparse learning on the second feature subset according to the comprehensive weight of the second feature subset.
6. The multi-view feature selection model training method according to claim 5, characterized in that: The objective function is: Among them, α v is the vth second view weight factor, used to determine the weight of the second view; feature selection variable W v is the vth second view mapping matrix, in which each element represents the weight information of the second feature subset; represents the distribution difference between the i-th sample and the j-th sample of the second feature subset in the v-th second view; represents the distribution difference of any two sample categories other than the i-th sample and the j-th sample in the second feature subset in the v-th second view; p and λ are adjustable hyperparameters, where p is the view adjustment factor and λ is the regularization strength hyperparameter.
7. The multi-view feature selection model training method according to claim 6, characterized in that: The optimizing and iterating the objective function of the multi-view feature selection model until the objective function converges and outputting an optimal second feature subset comprises: Dividing the second multi-view dataset into a training set and a test set based on multi-fold cross validation; Initialize the feature selection variable W v and the second view weight factor α v ; According to the training set and the test set, the feature selection variable W in the objective function is v and the second view weight factor α v Perform alternating iterative updates; When each fold iteration satisfies the convergence condition, the next fold iteration is performed until each fold iteration satisfies the convergence condition, thereby obtaining the multi-view feature selection model that has been trained; According to the trained multi-view feature selection model, an optimal second feature subset is output; the optimal second feature subset is the second feature subset whose weight is greater than a preset threshold.
8. The multi-view feature selection model training method according to claim 7, characterized in that: The feature selection variable W in the objective function is selected according to the training set and the test set. v and the second view weight factor α v Perform alternating iterative updates, including: Determine the distribution difference between the second feature values of samples of different categories in each of the second feature subsets by KL divergence; Determine the feature selection variable W v The difference in the standardized KL divergence distribution under ; According to the standardized distribution difference, optimize the feature selection variable W v ; Select variable W based on the optimized features v And the view adjustment factor, optimize the second view weight factor α v .
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1-2 or 3-8 is implemented.
10. A computer program product, characterized in that When the computer program product runs on an electronic device, the electronic device executes the method according to any one of claims 1 to 2 or 3 to 8.