Multi-mark feature set generation method and system

By calculating mark weights and constructing constrained nearest neighbor graphs on multi-labeled data sets, and generating target feature sorting, the problem that the existing technology cannot accurately generate feature sets is solved, efficient and accurate feature set generation is achieved, and work efficiency is improved.

CN120144997APending Publication Date: 2025-06-13EAST CHINA JIAOTONG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510136161.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing multi-marked feature selection method cannot accurately generate feature sets of known features, resulting in the inability to effectively generate target feature sets, which reduces work efficiency.

Method used

By obtaining multi-label data sets in the preset database, generating training samples and marker sets in real time, calculating the marker weights of each original marker, building a constraint nearest neighbor graph, creating a feature importance evaluation framework, generating target feature sorting, and outputting the model in real time through the feature set output model.

Benefits of technology

It realizes the rapid and efficient generation of target feature sets, removes redundant data, improves work efficiency, and ensures that the generated feature sets are unique.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144997A_ABST
    Figure CN120144997A_ABST
Patent Text Reader

Abstract

The invention provides a multi-label feature set generation method and system, and the method comprises the steps: generating a corresponding training sample in real time according to a multi-label data set, and detecting a label set corresponding to the multi-label data set in real time; calculating a mark weight corresponding to each original mark in real time according to a preset rule, and constructing a corresponding constrained nearest neighbor graph in real time according to the training sample and the mark weights in real time; generating a target feature sequence corresponding to the multi-label data set in real time through an importance evaluation framework; generating a corresponding training feature set according to the target feature sequence, and correspondingly inputting the mark set and the training feature set into a preset network to correspondingly train a feature set output model; and acquiring actual feature data input by the user in real time, and outputting a target feature set corresponding to the actual feature data in real time through the feature set output model. According to the method, the feature set can be quickly and effectively generated, and the working efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a method and system for generating a multi-label feature set. Background Art

[0002] With the progress of technology and the development of the times, artificial intelligence technology has also developed rapidly and has become increasingly mature. Specifically, people have developed task processing methods such as single-label learning and multi-label learning to improve the processing efficiency of tasks.

[0003] Among them, single-label learning specifically means that an object is associated with a single label. Correspondingly, multi-label learning allows an object to be associated with multiple labels at the same time. Specifically, during the learning process, it is necessary to detect the features corresponding to the task and output the corresponding feature set in real time to complete subsequent processing.

[0004] Furthermore, in the process of creating a feature set in existing single-label learning, it is necessary to divide features into strongly correlated features, weakly correlated features, and uncorrelated features, and then construct corresponding feature subsets according to the division results, and further calculate the importance of each feature in the current feature subset to generate a corresponding feature sorting list. However, most existing multi-label feature selection methods only obtain a feature sorting sequence and cannot obtain a feature set with a known number of features, resulting in the inability to accurately generate the required target feature set and correspondingly reducing work efficiency. Summary of the Invention

[0005] Based on this, the purpose of the present invention is to provide a method and system for generating a multi-label feature set to solve the problem that most existing multi-label feature selection methods only obtain a feature sorting sequence and cannot obtain a feature set with a known number of features.

[0006] The first aspect of the embodiment of the present invention proposes:

[0007] A method for generating a multi-label feature set, wherein the method includes:

[0008] Obtain a multi-label data set from a preset database, generate corresponding training samples in real time according to the multi-label data set, and detect a label set corresponding to the multi-label data set in real time, where the label set contains a number of original labels;

[0009] Calculate the label weight corresponding to each original label in real time according to a preset rule, and construct a corresponding constrained nearest neighbor graph in real time according to the training samples and the label weights;

[0010] Create a corresponding feature importance evaluation framework in real time according to the constrained nearest neighbor graph, and generate a target feature ranking corresponding to the multi-label data set in real time through the importance evaluation framework;

[0011] Generate a corresponding training feature set according to the target feature ranking, and input the label set and the training feature set into the interior of a preset network correspondingly to train a feature set output model correspondingly;

[0012] Obtain the actual feature data input by the user in real time, and output a target feature set corresponding to the actual feature data in real time through the feature set output model, and the target feature set is unique.

[0013] The beneficial effects of the present invention are: by obtaining a multi-label data set in real time, several original labels for subsequent processing can be further obtained. Based on this, in order to be able to detect the features contained in the current data set correspondingly, it is necessary to further detect the label weights corresponding to each current original label in the label set. Based on this, a target feature set for subsequent model training will be further created, so that a required feature output model can be further trained. Based on this, the actual data input by the user in real time can be directly processed by this model, and a target feature set required by the current user can be correspondingly generated. In this process, redundant data can be effectively removed, and at the same time, the target feature set can be obtained quickly and effectively, corresponding to improving the work efficiency.

[0014] Further, the step of calculating the label weight corresponding to each of the original labels in real time according to a preset rule includes:

[0015] When several of the original labels are obtained in real time, add corresponding target identifiers to each of the original labels in turn;

[0016] Based on the target identifiers, detect the mutual information generated between two adjacent original labels in turn, and detect the overall mutual information between each of the original labels and the label set one by one;

[0017] Calculate the label weight corresponding to each of the original labels according to the mutual information and the overall mutual information correspondingly.

[0018] Further, the step of calculating the label weight corresponding to each of the original labels according to the mutual information and the overall mutual information correspondingly includes:

[0019] When the mutual information and the overall mutual information are obtained respectively, perform parsing processing on the mutual information and the overall mutual information to detect the importance probability of the mutual information in the overall mutual information in real time;

[0020] Perform real-time conversion processing on the importance probability to convert the importance probability into a marker weight corresponding to the original marker in real time.

[0021] Furthermore, the expression of the algorithm for detecting in real time the importance probability of the mutual information in the overall mutual information is:

[0022]

[0023] where y i and y j represent two adjacent original markers, IG(y i ; y j ) represents the information gain between y i and y j , W(y i ) represents the importance of marker y i in the marker set Y, that is, the ratio of the sum of the information gain between marker y i and other markers to the information gain between all markers, and both m and k represent constants.

[0024] Furthermore, the steps of generating in real time a target feature ranking corresponding to the multi-marker data set through the importance evaluation framework include:

[0025] When the multi-marker data set is obtained in real time, a number of the original markers are detected correspondingly;

[0026] Generate in real time a target constraint algorithm adapted to a number of the original markers through the importance evaluation framework, and perform feature constraint processing on a number of the original markers through the target constraint algorithm to correspondingly generate the target feature ranking.

[0027] Furthermore, the steps of performing feature constraint processing on a number of the original markers through the target constraint algorithm to correspondingly generate the target feature ranking include:

[0028] When the target constraint algorithm is obtained in real time, calculate in real time a constraint similarity matrix generated corresponding to each other between two adjacent ones of the original markers through the target constraint algorithm;

[0029] Call out in correspondence a target constraint nearest neighbor graph adapted to the constraint similarity matrix through the importance evaluation framework, and calculate the target feature ranking in correspondence according to the target constraint nearest neighbor graph and the constraint similarity matrix based on a preset algorithm.

[0030] Furthermore, the steps of generating a corresponding training feature set according to the target feature ranking include:

[0031] When the target feature ranking is obtained in real time, all features corresponding to the multi-label data set are divided into strongly relevant features, weakly relevant features, and irrelevant features according to the target feature ranking;

[0032] Correspondingly, retain the strongly relevant features, remove the irrelevant features, and correspondingly screen the weakly relevant features to correspondingly generate the training feature set.

[0033] The second aspect of the embodiments of the present invention proposes:

[0034] A multi-label feature set generation system, wherein the system includes:

[0035] A detection module, configured to obtain a multi-label data set in a preset database, generate corresponding training samples in real time according to the multi-label data set, and detect in real time a label set corresponding to the multi-label data set, where the label set includes several original labels;

[0036] A calculation module, configured to calculate in real time the label weight corresponding to each original label according to a preset rule, and construct a corresponding constrained nearest neighbor graph in real time according to the training samples and the label weight;

[0037] A processing module, configured to create a corresponding feature importance evaluation framework in real time according to the constrained nearest neighbor graph, and generate a target feature ranking corresponding to the multi-label data set through the importance evaluation framework;

[0038] A training module, configured to generate a corresponding training feature set according to the target feature ranking, and input the label set and the training feature set into the interior of a preset network correspondingly to train a feature set output model;

[0039] An output module, configured to obtain actual feature data input by a user in real time, and output a target feature set corresponding to the actual feature data through the feature set output model, where the target feature set is unique.

[0040] Further, the calculation module is specifically configured to:

[0041] When several original labels are obtained in real time, add corresponding target identifiers to each original label in turn;

[0042] Based on the target identifiers, detect in turn the mutual information generated between adjacent two original labels, and detect one by one the overall mutual information between each original label and the label set;

[0043] Calculate the label weight corresponding to each original label according to the mutual information and the overall mutual information.

[0044] Further, the calculation module is specifically configured to:

[0045] When the mutual information and the overall mutual information are respectively obtained, parse and process the mutual information and the overall mutual information to detect in real time the importance probability of the mutual information in the overall mutual information;

[0046] Perform real-time conversion processing on the importance probability to convert the importance probability into a marker weight corresponding to the original marker in real time.

[0047] Further, the expression of the algorithm for detecting in real time the importance probability of the mutual information in the overall mutual information is:

[0048]

[0049] where y i and y j represent two adjacent original markers, IG(y i ; y j ) represents the information gain between y i and y j , W(y i ) represents the importance of the marker y i in the marker set Y, that is, the ratio of the sum of the information gain between the marker y i and other markers to the information gain between all markers, and both m and k represent constants.

[0050] Further, the processing module is specifically configured to:

[0051] When the multi-marker data set is obtained in real time, detect a number of the original markers correspondingly;

[0052] Generate a target constraint algorithm adapted to a number of the original markers in real time through the importance evaluation framework, and perform feature constraint processing on a number of the original markers through the target constraint algorithm to generate the target feature ranking correspondingly.

[0053] Further, the processing module is specifically configured to:

[0054] When the target constraint algorithm is obtained in real time, calculate in real time a constraint similarity matrix generated between two adjacent original markers correspondingly through the target constraint algorithm;

[0055] Call out a target constraint nearest neighbor graph adapted to the constraint similarity matrix correspondingly through the importance evaluation framework, and calculate the target feature ranking correspondingly based on a preset algorithm according to the target constraint nearest neighbor graph and the constraint similarity matrix.

[0056] Further, the processing module is specifically configured to:

[0057] When the target feature sorting is obtained in real time, all features corresponding to the multi-label data set are divided into strongly relevant features, weakly relevant features, and irrelevant features according to the target feature sorting;

[0058] Correspondingly, retain the strongly relevant features, remove the irrelevant features, and correspondingly screen the weakly relevant features to correspondingly generate the training feature set.

[0059] The third aspect of the embodiments of the present invention proposes:

[0060] A computer includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the computer program, the multi-label feature set generation method as described above is implemented.

[0061] The fourth aspect of the embodiments of the present invention proposes:

[0062] A readable storage medium stores a computer program. Wherein, when the program is executed by a processor, the multi-label feature set generation method as described above is implemented.

[0063] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0064] Figure 1 It is a flowchart of the multi-label feature set generation method provided by the first embodiment of the present invention;

[0065] Figure 2 It is a structural block diagram of the multi-label feature set generation system provided by the third embodiment of the present invention.

[0066] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. Specific Embodiments

[0067] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0068] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.

[0069] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0070] Please refer to Figure 1 , which shows the multi-label feature set generation method provided by the first embodiment of the present invention. The multi-label feature set generation method provided by this embodiment can quickly and effectively generate corresponding feature sets based on the multi-label data set, correspondingly improving work efficiency.

[0071] Specifically, this embodiment provides:

[0072] A multi-label feature set generation method, specifically including the following steps:

[0073] Step S10, obtaining a multi-label data set in a preset database, generating corresponding training samples in real time according to the multi-label data set, and detecting in real time a label set corresponding to the multi-label data set, where the label set contains several original labels;

[0074] Step S20, calculating in real time the label weight corresponding to each of the original labels according to a preset rule, and constructing a corresponding constrained nearest neighbor graph in real time according to the training samples and the label weights;

[0075] Step S30, creating a corresponding feature importance evaluation framework in real time according to the constrained nearest neighbor graph, and generating a target feature ranking corresponding to the multi-label data set through the importance evaluation framework;

[0076] Step S40, generating a corresponding training feature set according to the target feature ranking, and inputting the label set and the training feature set into the interior of a preset network correspondingly to train a feature set output model;

[0077] Step S50, obtain the actual feature data input by the user in real time, and output the target feature set corresponding to the actual feature data in real time through the feature set output model, where the target feature set is unique.

[0078] Further, the step of calculating the marking weight corresponding to each of the original marks in real time according to the preset rules includes:

[0079] When several of the original marks are obtained in real time, add the corresponding target identifiers to each of the original marks in sequence;

[0080] Based on the target identifiers, detect the mutual information generated between adjacent two original marks in sequence, and detect the overall mutual information between each of the original marks and the mark set one by one;

[0081] Calculate the marking weight corresponding to each of the original marks respectively according to the mutual information and the overall mutual information.

[0082] Further, the step of calculating the marking weight corresponding to each of the original marks respectively according to the mutual information and the overall mutual information includes:

[0083] When the mutual information and the overall mutual information are obtained respectively, perform parsing processing on the mutual information and the overall mutual information to detect the importance probability of the mutual information in the overall mutual information in real time;

[0084] Perform real-time conversion processing on the importance probability to convert the importance probability into the marking weight corresponding to the original mark in real time.

[0085] Further, the expression of the algorithm for detecting the importance probability of the mutual information in the overall mutual information in real time is:

[0086]

[0087] where y i and y j represent two adjacent original marks, IG(y i ; y j ) represents the information gain between y i and y j , W(y i ) represents the importance of the mark y i in the mark set Y, that is, the ratio of the sum of the information gain between the mark y i and other marks to the information gain between all marks, and both m and k represent constants.

[0088] Further, the step of generating the target feature ranking corresponding to the multi-label data set in real time through the importance evaluation framework includes:

[0089] When the multi-label data set is obtained in real time, a number of the original labels are detected correspondingly;

[0090] The target constraint algorithm adapted to a number of the original labels is generated in real time through the importance evaluation framework, and the feature constraint processing is performed on a number of the original labels through the target constraint algorithm to correspondingly generate the target feature ranking.

[0091] Further, the step of performing feature constraint processing on a number of the original labels through the target constraint algorithm to correspondingly generate the target feature ranking includes:

[0092] When the target constraint algorithm is obtained in real time, the constraint similarity matrix generated corresponding to each other between two adjacent original labels is calculated in real time through the target constraint algorithm;

[0093] The target constraint nearest neighbor graph adapted to the constraint similarity matrix is called out correspondingly through the importance evaluation framework, and the target feature ranking is calculated correspondingly based on a preset algorithm according to the target constraint nearest neighbor graph and the constraint similarity matrix.

[0094] Further, the step of generating the corresponding training feature set according to the target feature ranking includes:

[0095] When the target feature ranking is obtained in real time, all features corresponding to the multi-label data set are divided into strongly relevant features, weakly relevant features, and irrelevant features according to the target feature ranking;

[0096] The strongly relevant features are retained correspondingly, the irrelevant features are removed, and the weakly relevant features are screened correspondingly to correspondingly generate the training feature set.

[0097] In addition, in this embodiment, it should also be noted that:

[0098] Step 1: Obtain data from the multi-label database on the network, and divide the data into a training set and a test set. Remove the outliers in the multi-label data, and normalize the features of the multi-label data. The normalization formula is as follows:

[0099]

[0100] where f nor is the normalized feature. f max and f min are the maximum and minimum values of the feature f respectively. This formula normalizes the feature values to [0-1].

[0101] Step 2: In multi-label learning, the training sample set can be expressed as: D = {x i , y i | 1 ≤ i ≤ m}, where X i ∈ R d is the sample, and y i ∈ {1, -1} q is the corresponding logical label vector, where 1 and -1 represent relevant and irrelevant to the training sample respectively.

[0102]

[0103] Among them, IG(y i ; y j ) represents the information gain between y i and y j , and W(y i ) represents the importance of the label y i in the label set Y, that is, the ratio of the sum of the information gain between the label y i and other labels to the information gain between all labels: Obviously, 0 < W(y i ) < 1, and

[0104] Step 3: First, based on the weight matrix W generated by the sample and its neighbor samples, it is used to obtain label-specific features. The weight of each label is calculated from all its features. In order to find the most discriminative features for each label. The weights corresponding to any label y i (1 ≤ i ≤ q) are V i , and they are sorted in descending order. Record the first u feature indicators, and the values after u are zero.

[0105]

[0106] Among them, the dimension of V is q × d, where q is the number of labels and d is the number of features. The element V qd = IG(y q ; f d ), I i represents the index set of the specific features related to the label when the label is y i , and h represents the index of the feature. It can be seen from formula (2) that the set I i represents the indices of the first u features in the feature ranking of the label y i . The weights corresponding to any label y i are V i , and they are sorted in descending order. Record the first u feature indicators, and the values after u are zero. A new weight matrix is defined as follows:

[0107]

[0108] Furthermore, the weight matrix W generated based on samples and neighbor samples is used as prior knowledge to explore the marker group-specific features. First, the K-means algorithm is used to group similar markers into a cluster. During this period, the c columns of the marker matrix Y are divided into G groups, namely {Y 1 , Y 2 ,......Y g ......Y G}, satisfying According to the partition of the marker space, the partition of the weight matrix V can be divided into {Y 1 , Y 2 ,......Y g ......Y G}, where Y g is called the weight of the g-th group. For any marker group g (1 ≤ g ≤ G), by adding up each row of the weight matrix of each group, a d-dimensional row vector is obtained, and the first u column indices are recorded. The first u column indices are defined as follows:

[0109]

[0110] where, I g represents the set of marker group-specific feature indices corresponding to the group weight V g , represents summing all rows of V g to obtain a row vector, and h represents the index of the feature. In formula (4), the first u feature index sets of the marker group can be obtained, so as to find the shared features of similar markers. Next, according to the obtained I g , the weight values corresponding to the column indices not belonging to I g are set to 0, and a new group weight matrix is defined as follows:

[0111]

[0112] Finally, through combined into a set

[0113] According to the above discussion, the marker-exclusive feature weight matrix and the marker group-exclusive feature weight matrix can be obtained.

[0114]

[0115] Sort the new features according to this formula, so as to obtain the marker-exclusive features and the marker group-exclusive features.

[0116] Step 4: Logical labels cannot effectively show the relative importance of each label. To this end, we use manifold learning to map logical labels to numerical labels to mine richer semantic information. Therefore, we further calculate the label weights by combining manifold learning, mutual information, and constructing label-specific features and label group-specific features, and construct a topological structure in the feature space to generate a weight matrix between samples. The weight matrix is ​​optimized by minimizing the objective function.

[0117]

[0118] Where x is sorted according to the V matrix obtained by formula 1 to select samples with marker-specific features and marker group-specific features. is the weight matrix between the sample and its neighbor samples, unless x j is x i One of the K-nearest neighbors, w k is the tag weight matrix used to measure the contribution of each tag. Then, by using tag-specific features and tag weights constructed using manifold learning and mutual information, we can achieve the transfer of the topological structure of the feature space. This process is completed through minimization operations, effectively obtaining numerical tags.

[0119]

[0120] Where Z i ∈R q For numerical value mark.

[0121] Step 5: In multi-label data, since there is a certain correlation between the labels, when calculating the similarity between a sample and its neighboring samples, the correlation between the sample and the neighboring samples in terms of labels is also introduced. In this way, the definition of neighboring samples is not only based on sample features, but also combines the local correlation between labels, thereby constructing the sample x i and x j The similarity matrix of the local label correlation between , its specific form is as follows:

[0122]

[0123] The sequence Sort from small to large (m=1,2), where and represents neighboring samples, k is the number of nearest neighbors, x i and x j The similarity between them depends on the distance between them and the similarity of their label correlation. The latter can be considered as x i and x jThe connection strength affected by label correlation. In the similarity matrix, the higher the similarity between x i and x j , the corresponding is larger. Calculate the labeled constrained nearest neighbor graph

[0124]

[0125] where D = diag(S) represents the degree matrix, and S is the sample labeled constrained similarity matrix defined above.

[0126] Step 6: If two samples are the nearest neighbors, the features should make them closer to each other; on the contrary, if two samples are not the nearest neighbors, the features should make them farther away from each other. Based on this constraint, define the constrained similarity matrix between the sample and its neighbor samples as follows:

[0127]

[0128] Calculate the two feature constrained nearest neighbor graphs and

[0129]

[0130] where and represent the degree matrices corresponding to the two similarity matrices respectively, and are the two sample feature constrained similarity matrices defined above respectively.

[0131] Step 7: Calculate the feature importance ranking

[0132]

[0133] where and MCLS - SPE is judged according to the local preservation ability of the features. The stronger the local preservation ability of the features, the lower the corresponding feature score indicates that the feature is more important. Finally, a feature ranking is obtained according to formula (11).

[0134] Step 8: Divide the features. Given a multi - labeled decision table M = <U, X, L>, where U represents a non - empty finite sample set, x represents the set of conditional attributes, L = {L 1 , L 2 ,..., L q} represents a non - empty label set containing q labels, F is a subset of x,

[0135] Feature fn For the features that are strongly positively correlated with the label, the following conditions are satisfied:

[0136]

[0137] Feature f n For the features that are weakly positively correlated with the label, the following conditions are satisfied:

[0138]

[0139] Feature f n For the features that are uncorrelated with the label, the following conditions are satisfied:

[0140]

[0141] Step 9: For the strongly positively correlated features, all are retained. For the uncorrelated features, all are deleted. For the weakly positively correlated features, it is necessary to further determine whether they are satisfied. We selectively retain a part, and the selection criterion is the magnitude of the label weight that is strongly positively correlated with the feature. Here, we will use the label weight defined by Formula 1. Set the parameter a to control the selection of weakly positively correlated features, and a ∈ [0, 1]. The larger the value, the fewer features are selected. When a = 0, all weakly correlated features are selected. Finally, the feature set G is obtained.

[0142] G = {f 1 , f 2 ,..., f n}

[0143] where n is the number of features in the obtained feature set.

[0144] Please refer to Figure 2 , the third embodiment of the present invention provides:

[0145] A multi-label feature set generation system, wherein the system includes:

[0146] A detection module, configured to obtain a multi-label data set in a preset database, generate corresponding training samples in real time according to the multi-label data set, and detect in real time a label set corresponding to the multi-label data set, where the label set includes several original labels;

[0147] A calculation module, configured to calculate in real time the label weight corresponding to each of the original labels according to a preset rule, and construct a corresponding constrained nearest neighbor graph in real time according to the training samples and the label weights;

[0148] A processing module, configured to create a corresponding feature importance evaluation framework in real time according to the constrained nearest neighbor graph, and generate a target feature ranking corresponding to the multi-label data set in real time through the importance evaluation framework;

[0149] A training module, configured to generate a corresponding training feature set according to the target feature ranking, and input the label set and the training feature set into the interior of a preset network correspondingly to train a feature set output model correspondingly;

[0150] An output module, configured to obtain actual feature data input by a user in real time, and output a target feature set corresponding to the actual feature data in real time through the feature set output model, where the target feature set is unique.

[0151] Further, the calculation module is specifically configured to:

[0152] When obtaining a plurality of the original labels in real time, add corresponding target identifiers to each of the original labels in sequence;

[0153] Based on the target identifiers, sequentially detect the mutual information generated between two adjacent original labels, and detect the overall mutual information between each original label and the label set one by one;

[0154] Calculate the label weight corresponding to each original label according to the mutual information and the overall mutual information correspondingly.

[0155] Further, the calculation module is specifically configured to:

[0156] When obtaining the mutual information and the overall mutual information respectively, perform parsing processing on the mutual information and the overall mutual information to detect the importance probability of the mutual information in the overall mutual information in real time;

[0157] Perform real-time conversion processing on the importance probability to convert the importance probability into the label weight corresponding to the original label in real time.

[0158] Further, the expression of the algorithm for detecting the importance probability of the mutual information in the overall mutual information in real time is:

[0159]

[0160] where y i and y j represent two adjacent original labels, and IG(y i ; y j ) represents the information gain between y i and y j , and W(yi ) represents the tag y i The importance in the tag set Y, that is, the tag y i And the ratio of the sum of the information gains between the tag y and other tags to the information gain between all tags. Both m and k represent constants.

[0161] Furthermore, the processing module is specifically configured to:

[0162] When the multi-tag data set is obtained in real time, a number of the original tags are detected correspondingly;

[0163] The target constraint algorithm adapted to a number of the original tags is generated in real time through the importance evaluation framework, and the target constraint algorithm is used to perform feature constraint processing on a number of the original tags to correspondingly generate the target feature ranking.

[0164] Furthermore, the processing module is specifically configured to:

[0165] When the target constraint algorithm is obtained in real time, the constraint similarity matrix generated correspondingly between two adjacent original tags is calculated in real time through the target constraint algorithm;

[0166] The target constraint nearest neighbor graph adapted to the constraint similarity matrix is retrieved correspondingly through the importance evaluation framework, and the target feature ranking is calculated correspondingly based on a preset algorithm according to the target constraint nearest neighbor graph and the constraint similarity matrix.

[0167] Furthermore, the processing module is specifically configured to:

[0168] When the target feature ranking is obtained in real time, all features corresponding to the multi-tag data set are divided into strongly relevant features, weakly relevant features, and irrelevant features according to the target feature ranking;

[0169] The strongly relevant features are retained correspondingly, the irrelevant features are removed, and the weakly relevant features are screened correspondingly to correspondingly generate the training feature set.

[0170] The fourth embodiment of the present invention provides a computer, including a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the computer program, the multi-tag feature set generation method as described above is implemented.

[0171] The fifth embodiment of the present invention provides a readable storage medium, on which a computer program is stored. Wherein, when the program is executed by a processor, the multi-tag feature set generation method as described above is implemented.

[0172] In summary, the multi-label feature set generation method and system provided by the above embodiments of the present invention can quickly and effectively generate corresponding feature sets based on a multi-label data set, thereby correspondingly improving work efficiency.

[0173] It should be noted that the above-mentioned respective modules can be functional modules or program modules, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned respective modules can be located in the same processor; or the above-mentioned respective modules can also be located in different processors in any combined form.

[0174] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0175] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.

[0176] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0177] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0178] The above-described embodiments only represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several variations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. A method for generating a multi-label feature set, characterized in that: The method comprises: Acquire a multi-label data set in a preset database, generate corresponding training samples according to the multi-label data set in real time, and detect a label set corresponding to the multi-label data set in real time, wherein the label set includes a plurality of original labels; Calculating the label weight corresponding to each of the original labels in real time according to preset rules, and constructing the corresponding constrained nearest neighbor graph in real time according to the training samples and the label weights; Creating a corresponding feature importance evaluation framework in real time according to the constrained nearest neighbor graph, and generating a target feature ranking corresponding to the multi-label data set in real time through the importance evaluation framework; Generating a corresponding training feature set according to the target feature sorting, and inputting the tag set and the training feature set into a preset network to train a feature set output model; The actual feature data input by the user is acquired in real time, and a target feature set corresponding to the actual feature data is output in real time through the feature set output model, wherein the target feature set is unique.

2. The method for generating a multi-label feature set according to claim 1, wherein: The step of calculating the tag weight corresponding to each of the original tags in real time according to the preset rules comprises: When a plurality of the original marks are acquired in real time, a corresponding target identifier is added to each of the original marks in turn; Based on the target identifier, the mutual information generated between two adjacent original tags is detected in sequence, and the overall mutual information between each original tag and the tag set is detected one by one; The label weight corresponding to each of the original labels is calculated according to the mutual information and the overall mutual information.

3. The method for generating a multi-label feature set according to claim 2, wherein: The step of calculating the label weight corresponding to each of the original labels according to the mutual information and the overall mutual information comprises: When the mutual information and the overall mutual information are respectively obtained, the mutual information and the overall mutual information are analyzed and processed to detect the importance probability of the mutual information in the overall mutual information in real time; The importance probability is converted in real time to convert the importance probability into a tag weight corresponding to the original tag in real time.

4. The method for generating a multi-label feature set according to claim 3, wherein: The expression of the algorithm for detecting the importance probability of the mutual information in the overall mutual information in real time is: Among them, y i and j Represents two adjacent original marks, IG(y i ;y j ) represents y i and j The information gain between i ) indicates the mark y i The importance of label y in the label set Y, that is, label y i The ratio of the sum of the information gain of the other labels to the information gain between all labels, where m and k are both constants.

5. The method for generating a multi-label feature set according to claim 1, wherein: The step of generating a target feature ranking corresponding to the multi-label data set in real time through the importance evaluation framework comprises: When the multi-label data set is acquired in real time, a number of the original labels are detected accordingly; A target constraint algorithm adapted to the plurality of original tags is generated in real time through the importance evaluation framework, and feature constraint processing is performed on the plurality of original tags through the target constraint algorithm to generate the target feature ranking accordingly.

6. The method for generating a multi-label feature set according to claim 5, characterized in that: The step of performing feature constraint processing on the plurality of original marks by the target constraint algorithm to generate the target feature sorting accordingly comprises: When the target constraint algorithm is acquired in real time, a constraint similarity matrix generated by the correspondence between two adjacent original marks is calculated in real time by the target constraint algorithm; The target constraint nearest neighbor graph adapted to the constraint similarity matrix is ​​called out through the importance evaluation framework, and the target feature ranking is calculated based on the target constraint nearest neighbor graph and the constraint similarity matrix based on a preset algorithm.

7. The method for generating a multi-label feature set according to claim 6, characterized in that: The step of generating a corresponding training feature set according to the target feature sorting comprises: When the target feature ranking is acquired in real time, all features corresponding to the multi-label data set are divided into strongly correlated features, weakly correlated features, and irrelevant features according to the target feature ranking; The strongly correlated features are correspondingly retained, the irrelevant features are removed, and the weakly correlated features are correspondingly screened to correspondingly generate the training feature set.

8. A multi-label feature set generation system, characterized in that: The system comprises: A detection module, used to obtain a multi-label data set in a preset database, to generate corresponding training samples according to the multi-label data set in real time, and to detect a label set corresponding to the multi-label data set in real time, wherein the label set includes a plurality of original labels; A calculation module, used to calculate the label weight corresponding to each of the original labels in real time according to a preset rule, and to construct a corresponding constrained nearest neighbor graph in real time according to the training samples and the label weight; A processing module, used to create a corresponding feature importance evaluation framework in real time according to the constrained nearest neighbor graph, and generate a target feature ranking corresponding to the multi-label data set in real time through the importance evaluation framework; A training module, used to generate a corresponding training feature set according to the target feature sorting, and input the tag set and the training feature set into a preset network to train a feature set output model; The output module is used to obtain the actual feature data input by the user in real time, and output the target feature set corresponding to the actual feature data in real time through the feature set output model, and the target feature set is unique.

9. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for generating a multi-marker feature set as described in any one of claims 1 to 7 is implemented.

10. A readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for generating a multi-marker feature set as described in any one of claims 1 to 7 is implemented.