Non-convex imbalance multi-example multi-label learning method based on dual-granularity labeling

Through the non-convex imbalanced multi-example multi-label learning method of two-particle size annotation, key examples are identified and reweighted loss function and low rank constraints are used to solve the problem of label imbalance and correlation in multi-example multi-label learning, and the labeling efficiency and model performance are improved.

CN120277462AActive Publication Date: 2025-07-08NAT UNIV OF DEFENSE TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510341283.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The existing multi-example multi-label learning methods ignore the differences in the role of examples in objects, resulting in label imbalance problems and poor generalization capabilities of models, and fail to effectively deal with label correlation.

Method used

Using a non-convex unbalanced multi-example multi-label learning method based on two-particle size annotation, key examples are identified through the DL-MIML model, reweighted loss function and low rank constraints are used, and the model is optimized in combination with the block coordinate-Nesterov accelerated projection descent algorithm to solve the problem of label imbalance and correlation.

Benefits of technology

It improves the labeling efficiency and accuracy, enhances the model's attention to a few types of samples, alleviates the imbalance of label distribution, and improves the classification performance of multi-example multi-label learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277462A_ABST
    Figure CN120277462A_ABST
Patent Text Reader

Abstract

The invention provides a non-convex unbalanced multi-instance multi-label learning method based on dual-granularity labeling, and the method comprises the steps: obtaining a target multi-instance package set of a target multi-instance multi-label learning task, wherein each packet in the multi-instance packet set consists of a plurality of instances and corresponds to a plurality of tag classes, and the instances can be distinguished into key instances and irrelevant instances according to different tag classes to which the instances belong; and detecting key examples of a target multi-example packet in the to-be-detected multi-example packet set through the DL-MIML model to obtain classification results of two levels of examples and packets. The method has the beneficial effects that the labeling efficiency and the labeling accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computers and artificial intelligence, and particularly to a non-convex unbalanced multi-instance multi-label learning method based on dual-granularity annotation. Background Art

[0002] A large number of multi-semantic objects with complex structures drive the emergence of the multi-instance multi-label learning, a weakly supervised paradigm, which describes polysemous objects with complex structures using multiple instances and learns the mapping from the instance set to the class label set. Currently, there are many methods aiming to efficiently identify multi-semantic objects. However, in practical tasks such as image text annotation and drug screening, the tasks require domain experts to manually check a large number of images and texts, or to identify effective drug molecules from a huge drug library, which is time-consuming and laborious.

[0003] Therefore, automatic annotation of instance-level labels is crucial. Existing methods that can achieve instance-level label prediction usually ignore some problems:

[0004] (1) Each instance in the object does not play an equal role. It is the diversity of the features of different instances that endows the object with multi-semanticity. It is unreasonable for some methods to regard each instance in the object as the same.

[0005] (2) There is a serious label imbalance problem in the reconstructed dataset. The label distribution in applications is usually skewed, which inevitably brings the label imbalance problem. Some methods extract instances from the object, resulting in a sharp increase in the number of samples in the reconstructed dataset, changing the label distribution and exacerbating the imbalance problem.

[0006] (3) Some methods simply degrade the multi-instance multi-label learning task into a binary classification task, ignoring label correlation, resulting in a decline in the classification performance of the model and poor generalization ability. Summary of the Invention

[0007] The main purpose of the embodiments of the present invention is to propose a non-convex unbalanced multi-instance multi-label learning method based on dual-granularity annotation, which improves the annotation efficiency and annotation accuracy.

[0008] One aspect of the present invention provides a non-convex unbalanced multi-instance multi-label learning method based on dual-granularity annotation, characterized by comprising:

[0009] Obtain a target multi-instance multi-label learning task's target multi-instance bag set, where each bag in the multi-instance bag set consists of several instances and corresponds to several label classes, and the instances can be distinguished into key instances and irrelevant instances according to different label classes to which they belong;

[0010] Detect the key examples of the target multi-instance bag in the multi-instance bag set to be measured through the DL-MIML model, and obtain the classification results at both the example and bag levels;

[0011] The DL-MIML model is obtained through the following steps:

[0012] Use the example weight vector to determine the contribution value of the examples in the bag of the target multi-instance bag set to the labeled label, and determine the key examples. Determine the bag representation according to the contribution value of the label and the key examples, where the key examples are used to characterize that the examples belong to the labeled label;

[0013] Calculate the label scores of the key examples and irrelevant examples in the bag, and determine the loss of each bag according to the label scores using the reweighted loss function;

[0014] Process the objective function based on the training bag set with low-rank constraints to obtain the second objective function, where the first objective function is to minimize the loss;

[0015] Process the example weight vector with sparse constraints to obtain the DL-MIML model;

[0016] Obtain the variables to be learned in the DL-MIML model, and perform optimization and solution through the block coordinate-Nesterov accelerated projection descent algorithm.

[0017] According to the non-convex unbalanced multi-instance multi-label learning method based on dual-grained annotation, where the example weight vector u i,c Learn the contribution of n i examples to the i-th bag being labeled with the c-th label through the classification model, where the example weight vector is used to align features and labels at the bag level; where the target multi-instance bag set {X1, X2,..., X N}, N represents the total number of bags, and the bag X i consists of a set of examples n i is the total number of examples in the i-th bag, and there is a label vector Y i = [y i,1 , y i,2 ,..., y i,C , when the bag X i is labeled with the c-th label, y i,c = 1, otherwise 0, where C represents the dimension of the label space; at the same time, each bag includes at least one key example representing the label feature.

[0018] According to the non-convex unbalanced multi-instance multi-label learning method based on dual-grained annotation, where the label scores of the key examples and irrelevant examples in the bag are calculated, and the loss of each bag is determined according to the label scores using the reweighted loss function, including:

[0019] Train the classification model using a reweighted loss function, where the reweighted loss function is as follows:

[0020]

[0021] where is the score that the i-th bag belongs to the c-th label class calculated based on key examples, is the label score calculated based on the j-th irrelevant example, and T is the transpose operation;

[0022] The loss of each bag is as follows:

[0023]

[0024] is the reweighted weight, and ρ is a hyperparameter.

[0025] According to the non-convex imbalanced multi-instance multi-label learning method based on dual granularity annotation, where the objective function based on the training bag set is processed using a low-rank constraint to obtain a second objective function, where the first objective function is to minimize the loss, including:

[0026] The processing of the objective function based on the training bag set using a low-rank constraint is as follows:

[0027]

[0028] where ① is determined by low-rank decomposition, and ② is determined by and is determined, where λ is a regularization factor, W represents a matrix with a low rank of K, and W = AB T , where and rank(W) can be decomposed into rank(A) + rank(B), and rank(W) = ‖W‖ * , ||*||F represents the Frobenius norm, is the real number space, rank represents the rank, and d represents the matrix dimension.

[0029] According to the non-convex imbalanced multi-instance multi-label learning method based on dual granularity annotation, where a sparse constraint is used to process the example weight vector to obtain a DL-MIML model, including:

[0030] Obtain the bag features obtained by embedding the examples based on the weight vector, that is, the embedded bag features are a convex combination of key examples;

[0031] Add a sparse constraint to the weight vector according to the convex combination, where the sparse constraint is a non-convex sparse-inducing regularization term, and the regularization term includes the l0-norm, l1-norm, and non-negativity constraint;

[0032] The DL-MIML model is obtained according to the weight vector with the added sparse constraint and the second loss function as follows:

[0033]

[0034] where r is a hyperparameter regarding the number of key examples.

[0035] According to the non-convex unbalanced multi-instance multi-label learning method based on dual granularity annotation, where the variables to be learned in the DL-MIML model are obtained and optimized by the block coordinate-Nesterov accelerated projection descent algorithm, including:

[0036] The variable to be learned is denoted as:

[0037] w = (A, {bc}, {ui @ c})

[0038] where, in the constraint set add the constraint on u i,c i.e.,

[0039] Determine the optimization problem according to the variable to be learned as

[0040]

[0041] where the objective function

[0042] According to the non-convex unbalanced multi-instance multi-label learning method based on dual granularity annotation, where the optimization is solved by the block coordinate-Nesterov accelerated projection descent algorithm, including:

[0043] When the block coordinate descent algorithm fixes the non-target variable blocks at the latest updated values, it iteratively minimizes F for each variable block ω j to obtain which is the value of ω j after the k-th update, i.e.:

[0044]

[0045] where, N v is the total number of variable blocks;

[0046] In each iteration, the DL-MIML model minimizes the approximate linear surrogate function for the variable block ωj Update, i.e.:

[0047]

[0048] where α k is the step size;

[0049] The algorithm introduces an auxiliary variable j during the optimization process of ω Initialize k = 0, a0 = 1,

[0050] where the auxiliary variable z j satisfies and In the k-th iteration, DL-MIML first calculates the smallest i ≥ 0 such that

[0051]

[0052] Through the smallest i, we can obtain Update the variable to be learned as

[0053]

[0054] According to the updated variable to be learned, obtain the auxiliary variable for the next iteration, i.e.,

[0055]

[0056] Repeat the iteration until convergence.

[0057] According to the non-convex unbalanced multi-instance multi-label learning method based on dual-grained annotation, the optimization of the variable A to be learned includes:

[0058] Fix {b c} and {u i,c}, and determine that the objective function is equivalent to:

[0059]

[0060] The gradient of the objective function f with respect to A is

[0061]

[0062] Use the Nesterov-accelerated gradient algorithm to calculate the optimization result of the variable A to be learned.

[0063] According to the non-convex unbalanced multi-instance multi-label learning method based on dual-grained annotation, the optimization of the variable b c to be learned includes:

[0064] Fix A, {b m|m≠c}, and {u i,c}, it is determined that the objective function is equivalent to:

[0065]

[0066] The gradient of the objective function f with respect to b c is

[0067]

[0068] The Nesterov-accelerated gradient algorithm is used to calculate the optimization result of the variable b to be learned c .

[0069] According to the non-convex unbalanced multi-instance multi-label learning method based on dual granularity annotation, where the variable u to be learned i,c optimization includes:

[0070] Fix A, {b c}, and {u m,n |m≠i, n≠c}, the objective function is equivalent to

[0071]

[0072] The DL-MIML model is used to handle the convex optimization problem min f(u i,c ) with non-convex sparse-inducing regularization through the projected gradient descent algorithm, and the optimal solution is The sparse Euclidean projection on the set , that is

[0073]

[0074] where represents the projection operator, and I * is the index set that retains the first r largest positive elements in , while is the complement of I * , can be obtained through a gradient-based algorithm. The gradient of the objective function f with respect to u i,c is

[0075]

[0076] The beneficial effects of the present invention are as follows: Considering the problem of key example recognition in the multi-instance multi-label learning process, it can achieve two-level data classification and annotation of bags and examples based on bag labels and example features; at the same time, it considers dealing with the unbalanced label distribution and label correlation in the multi-instance multi-label learning task. By using a weighted loss function, it well balances the quantity gap between key examples and irrelevant examples, enhances the attention degree of the model to minority-class samples, and effectively alleviates the problem of unbalanced label distribution. It uses a kernel function for low-rank approximation of label correlation, which can be combined with the reweighted loss in a "convex" form, so that the method can train a multi-instance multi-label learning model that can handle both label imbalance and label correlation simultaneously; the block coordinate-Nesterov accelerated projection descent algorithm enables the model to converge to the optimal solution at a relatively fast speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:

[0078] Figure 1 is a schematic flowchart of the non-convex unbalanced multi-instance multi-label learning method based on dual granularity annotation according to an embodiment of the present invention.

[0079] Figure 2 is a schematic diagram of the DL-MIML model framework according to an embodiment of the present invention.

[0080] Figure 3 is a schematic flowchart of the block coordinate-Nesterov accelerated projection descent algorithm according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. In the following description, suffixes such as "module", "component", or "unit" used to denote elements are only for the convenience of describing the present invention and have no specific meaning in themselves. Therefore, "module", "component", or "unit" can be used interchangeably. "First", "second", etc. are only used to distinguish technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features. In the following description, the consecutive numbering of method steps is for the convenience of review and understanding. Combining the overall technical solution of the present invention and the logical relationship between each step, adjusting the implementation order between steps will not affect the technical effect achieved by the technical solution of the present invention. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0082] Figure 1 It is a schematic flow chart of the non-convex imbalanced multi-instance multi-label learning method based on dual granularity annotation according to an embodiment of the present invention, which includes but is not limited to steps S100 to S700, where S100 to S200 are non-convex imbalanced multi-instance multi-label learning methods, and S300 to S700 are schematic flow charts of the DL-MIML model construction process.

[0083] S100, obtain the target multi-instance multi-label learning task's target multi-instance bag set, where each bag in the multi-instance bag set consists of several instances and corresponds to several label classes, and the instances can be distinguished into key instances and irrelevant instances according to different label classes they belong to.

[0084] S200, detect the key instances of the target multi-instance bags in the multi-instance bag set to be tested through the DL-MIML model, and obtain classification results at both the instance and bag levels;

[0085] S300, use the instance weight vector to determine the contribution value of the instances in the bags of the target multi-instance bag set to the labeled labels, and determine the key instances, and determine the bag representation according to the contribution value of the labels and the key instances, where the key instances are used to characterize that the instances belong to the labeled labels.

[0086] In some embodiments, in the multi-instance multi-label learning task, if a bag carries a certain label, then at least one instance in it belongs to that label class, otherwise it means that all instances in the bag do not belong to this label class. The instances belonging to the corresponding label class are called "key instances" or "Witness". For the multi-instance bag set to be trained {X1, X2,..., X N}, the bag X i consists of a set of instances and has a label vector Y i = [y i,1 , y i,2 ,..., y i,C . When the bag X i is marked by the c-th label, y i,c = 1, otherwise it is 0.

[0087] In some embodiments, reference Figure 2 is a schematic diagram of the DL-MIML model framework according to an embodiment of the present invention, and the instance weight vector is used to determine the bag representation of the bag. For the training of a general classification model, denote y and as the true value and predicted value of the data point x respectively. The model is learned by minimizing the loss function where . However, due to the inconsistent levels of features and labels in the multi-instance multi-label learning task, y and x cannot be obtained simultaneously. DL-MIML introduces the instance weight vector ui,c , to measure n i contributions of the examples in the i-th bag to the c-th labeled tag. It should be noted that the example weight vector can not only detect key instances, but also eliminate the influence of irrelevant examples, which helps to enhance the discriminative ability of the model. Through the learned example weight vector u i,c , can be used as the new representation of the i-th bag to identify the c-th tag class, thus aligning features and tags at the bag level.

[0088] S400, calculate the label scores of key examples and irrelevant examples in the bag, and according to the label scores, use a reweighted loss function to determine the loss of each bag.

[0089] In some embodiments, the model needs to select an appropriate Cross-entropy loss, as a classic loss function in the field of machine learning, can be extended to the multi-label scenario to train the model. The cross-entropy loss function is where s is the score of the model output corresponding to the label class. Therefore, for the i-th bag, DL-MIML calculates s based on key examples as Calculate s based on the j-th irrelevant example as Since there are n i instances in the i-th bag, the loss function is expressed as:

[0090]

[0091] To balance the number of key examples and irrelevant examples, by introducing a weight where ρ is a hyperparameter, the loss of the i-th bag is

[0092] S500, process the objective function based on the training bag set with a low-rank constraint to obtain a second objective function, where the first objective function is to minimize the loss.

[0093] In some embodiments, the training bag set is a multi-instance multi-label data set for training.

[0094] Considering the label correlation in the multi-instance multi-label sample space, DL-MIML incorporates the low-rank constraint of the model into the objective function. Specifically, a matrix W with low rank K can be represented by low-rank factorization as W = AB T , where and where rank(W) can be decomposed into rank(A) + rank(B), and denote rank(W) = ‖W‖ * , then there is Therefore, the loss of the i-th bag can also be expressed as Based on the goal of minimizing the loss function, there are

[0095]

[0096] Among them, ① holds by low-rank decomposition, ② holds by and where λ is the regularization factor, C: the dimension of the label space (i.e., the classification problem of the dataset involves a total of C label classes); ||*|| F represents the Frobenius norm in mathematics; the hollow R represents the real number space (i.e., A is a d*K-dimensional matrix and B is a C*K-dimensional matrix), and d represents the matrix dimension. rank represents the rank, that is, the rank of matrix W can be decomposed into the rank of matrix A + the rank of matrix B.

[0097] S600, uses sparse constraints to process the example weight vectors to obtain the DL-MIML model.

[0098] In some embodiments, to meet the actual application requirements, it is necessary to constrain the example weight vectors. On the one hand, based on the standard assumption, the bag contains at least one key example representing the label feature, and the other examples are irrelevant examples. DL-MIML uses the l0-norm to constrain the number of selected key examples. On the other hand, by definition, the elements in the example weight vector represent the contribution of the example to the label. Therefore, all its elements must be non-negative. At the same time, the example weight vector plays an important role in embedding the original bag space into the discriminative feature space. The new feature representation of the mapped bag can be regarded as a convex combination of the selected key examples. Naturally, it is required that the sum of all elements in the example weight vector is 1, which can be achieved by adding an l1-norm constraint. Therefore, DL-MIML contains a non-convex sparse-induced regularization term that adds the l0-norm, l1-norm, and non-negativity constraints to the example weight vectors.

[0099] In summary, the DL-MIML is modeled as follows:

[0100]

[0101] where r is a hyperparameter regarding the number of key examples.

[0102] S700, obtains the variables to be learned in the DL-MIML model and optimizes and solves them through the block coordinate-Nesterov accelerated projection descent algorithm.

[0103] In some embodiments, there are three groups of variables to be learned in the model. Denote ω=(A, {b c}, {u i,c}) and then the optimization problem is

[0104]

[0105] Among them the objective function is defined as

[0106] Note that: variables A, {b c} and {u i,c} are coupled in this framework that contains both smooth and non-smooth functions; due to the use of non-convex sparse-inducing regularization to constrain u i,c , each block of ω is non-convex and non-smooth, which makes the objective function a non-trivial non-convex and non-smooth optimization problem. It is extremely challenging to directly use the general gradient descent algorithm to solve such a problem. Therefore, DL-MIML developed the BC-Nesterov-PD algorithm to handle this problem.

[0107] Specifically, DL-MIML uses the Block Coordinate Descent (BCD) algorithm to iteratively minimize F for each block of variables while fixing the other blocks of variables at their latest updated values. Denote as the value of ω j after the k-th update, and let In each iteration, DL-MIML updates the block of variables ω j by minimizing an approximate linear surrogate function, that is

[0108]

[0109] where α k is the step size.

[0110] DL-MIML speeds up the convergence rate of this minimization problem by adding a "momentum term" formed by the weighted accumulation of all past gradients on the basis of gradient descent, and performing a "look-ahead" step of gradient calculation in the direction of the momentum before adjusting the update. The algorithm introduces j during the optimization process of ω Initialize k = 0, a0 = 1, where z j satisfies and In the k-th iteration, DL-MIML first calculates the smallest i ≥ 0 such that

[0111]

[0112] The existence of i can be proven. Through the smallest i, we can obtain Therefore, the updated variable is

[0113]

[0114] Based on this, the auxiliary variables for the next iteration are obtained, i.e.,

[0115]

[0116] Repeat the above iteration until convergence. The specific process is as shown Figure 3 The schematic diagram of the block coordinate - Nesterov accelerated projection descent algorithm process shown.

[0117] In some embodiments, the variables to be optimized are specifically as follows:

[0118] (1) Optimize A

[0119] Fix {b c} and {u i,c}, and the objective function is equivalent to:

[0120]

[0121] Therefore, the problem min f(A) is a convex optimization problem and can be solved by a gradient - based method. The gradient of the objective function f with respect to A is

[0122]

[0123] Next, solve it according to the steps of the Nesterov - Accelerated Gradient (NAG) algorithm.

[0124] (2) Optimize b c

[0125] Fix A, {b m | m ≠ c}, and {u i,c}, and the objective function is equivalent to

[0126]

[0127] Therefore, the problem min f(b c ) is a convex optimization problem and can be solved by a gradient - based method. The gradient of the objective function f with respect to b c is

[0128]

[0129] Next, solve it according to the steps of the NAG algorithm.

[0130] (3) Optimize u i,c

[0131] Fix A, {b c} and {um,n |m≠m≠i,n≠c}, the objective function is equivalent to

[0132]

[0133] DL-MIML addresses the convex optimization problem with non-convex sparse-inducing regularization min f(u i,c ) using the Projected Gradient Descent (PGD) algorithm. Its optimal solution is the sparse Euclidean projection onto the set , that is

[0134]

[0135] where denotes the projection operator. I * is the index set that retains the first r largest positive elements in , while is the complement of I * . can be obtained by a gradient-based algorithm. The gradient of the objective function f with respect to u i,c is

[0136]

[0137] In some embodiments, the DL-MIML model according to the embodiments of the present invention processes data required for human activity recognition (HAR) tasks, such as detecting, identifying, classifying, and recording the daily movement and behavior patterns of an individual by different sensors, for example, movement at different locations and body activity states (walking, running, driving, or stationary); collecting these data through a smartphone with sensors and communication facilities, preprocessing them, and extracting relevant features (such as features in the time and frequency domains); constructing multi-instance bags with these relevant features to represent combinations of various human activities. It can be understood that human individual activities are variable. Therefore, through key instance recognition, handling of imbalance problems, and low-rank approximation of label correlations in the DL-MIML model according to the embodiments of the present invention, the accuracy and recognition performance of individual activity recognition are improved. For another example, in the drug-target prediction task based on molecular conformation using the DL-MIML model according to the embodiments of the present invention, a single drug can be regarded as a "bag", which can have multiple molecular conformations or multiple chemical fragments, which are regarded as "instances". These different instances may be associated with different protein targets (labels) respectively, such that a drug can bind to multiple protein targets. It can be understood that the drug molecule library data is huge. Therefore, through key instance recognition, handling of imbalance problems, and low-rank approximation of label correlations in the DL-MIML model according to the embodiments of the present invention, classification detection tasks for data at two levels of molecular conformation (instances) and drugs (bags) are achieved, greatly improving the efficiency.

[0138] Embodiments of the present invention further provide an electronic device, which includes a processor and a memory;

[0139] The memory stores a program;

[0140] The processor executes the program to execute the aforementioned non-convex imbalanced multi-instance multi-label learning method based on dual granularity annotation; this electronic device has the function of carrying and running the software system for glass flow control provided by the embodiments of the present invention. For example, a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or communicating with charged particle tools or other imaging devices, etc.

[0141] Embodiments of the present invention further provide a computer-readable storage medium, where the storage medium stores a program, and the program is executed by a processor to implement the non-convex imbalanced multi-instance multi-label learning method based on dual granularity annotation as described above.

[0142] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may in fact be executed substantially simultaneously or the blocks may sometimes be executed in the reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and in which sub-operations described as part of a larger operation are executed independently.

[0143] Embodiments of the present invention also disclose a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the foregoing non-convex unbalanced multi-instance multi-label learning method based on dual granularity annotation.

[0144] Furthermore, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features described may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0145] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes of various kinds.

[0146] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a predefined sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0147] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0148] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0149] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0150] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0151] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A non-convex imbalanced multi-instance multi-label learning method based on dual-granularity annotation, characterized in that Including: Obtain a target multi-instance multi-label learning task's target multi-instance bag set, where each bag in the multi-instance bag set consists of several instances and corresponds to several label classes, and the instances can be distinguished into key instances and irrelevant instances according to different label classes they belong to; Detect the key instances of the target multi-instance bags in the multi-instance bag set to be tested through a DL-MIML model, and obtain classification results at both the instance and bag levels; The DL-MIML model is obtained through the following steps: Use an instance weight vector to determine the contribution value of the instances in the bags of the target multi-instance bag set to the labeled labels, and determine the key instances. Determine the bag representation according to the contribution value of the labels and the key instances, where the key instances are used to characterize that the instances belong to the labeled labels; Calculate the label scores of the key instances and irrelevant instances in the bag, and determine the loss of each bag according to the label scores using a reweighted loss function; Process the objective function based on the training bag set with a low-rank constraint to obtain a second objective function, where the first objective function is to minimize the loss; Process the instance weight vector with a sparse constraint to obtain a DL-MIML model; Obtain the variables to be learned in the DL-MIML model, and perform optimization and solution through a block coordinate-Nesterov accelerated projected gradient descent algorithm.

2. The non-convex unbalanced multi-instance multi-label learning method based on dual-grained annotation according to claim 1, wherein The example weight vector u i,c learns the contribution of n i examples to the i-th packet being labeled with the c-th label through a classification model, where the example weight vector is used to align features and labels at the packet level; Among them, the target multi-instance bag set {X1, X2,..., X N}, N represents the total number of bags, and bag X i is composed of a group of instances . n i is the total number of instances in the i-th bag, and there is a label vector Y i = [y i,1 , y i,2 ,..., y i,C . When bag X i is marked with the c-th label, y i,c = 1, otherwise it is 0, where C represents the dimension of the label space; meanwhile, each bag contains at least one key instance representing the label feature.

3. The non-convex unbalanced multi-instance multi-label learning method based on dual-grain annotation according to claim 2, wherein The calculating the label scores of the key instances and irrelevant instances in the bag, and determining the loss of each bag according to the label scores using a reweighted loss function includes: Training the classification model by using a reweighted loss function, where the reweighted loss function is as follows: wherein, is the score that the i-th packet calculated based on the key example belongs to the c-th label class, is the label score calculated based on the j-th irrelevant example, and T is the transpose operation; Loss per packet is is the reweighted weight, and ρ is a hyperparameter.

4. The non-convex unbalanced multi-instance multi-label learning method based on dual granularity annotation according to claim 3, wherein The processing the objective function based on the training bag set with a low-rank constraint to obtain a second objective function, where the first objective function is to minimize the loss includes: The processing the objective function based on the training bag set with a low-rank constraint is: Among them, ① is determined by low-rank decomposition, where ② is determined by and determined, where λ is the regularization factor, W represents a matrix with low rank K, and W = AB T , where and rank(W) can be decomposed into rank(A)+rank(B), and rank(W) = ‖W‖ * , ||*‖ F represents the Frobenius norm, is the real number space, rank represents the rank, and d represents the matrix dimension.

5. The non-convex unbalanced multi-instance multi-label learning method based on dual granularity annotation according to claim 4, wherein The processing the instance weight vector with a sparse constraint to obtain a DL-MIML model includes: Obtain the bag features after embedding the instances based on the weight vector, that is, the embedded bag features are a convex combination of key instances; Add a sparse constraint to the weight vector according to the convex combination, where the sparse constraint is a non-convex sparse-inducing regularization term, and the regularization term includes l0-norm, l1-norm, and non-negativity constraint; The DL-MIML model obtained according to the weight vector with the added sparse constraint and the second loss function is: where r is a hyperparameter regarding the number of key instances.

6. The non-convex unbalanced multi-instance multi-label learning method based on dual granularity annotation according to claim 5, wherein The obtaining the variables to be learned in the DL-MIML model and performing optimization and solution through a block coordinate-Nesterov accelerated projected gradient descent algorithm includes: The variables to be learned are denoted as: ω=(A,{b c},{u i,c}) Among them, in the constraint set C i,c add the constraint on u i,c , that is Determine the optimization problem according to the variables to be learned as Among them Index function 7. The non-convex unbalanced multi-instance multi-label learning method based on dual granularity annotation according to claim 6, wherein The performing optimization and solution through a block coordinate-Nesterov accelerated projected gradient descent algorithm includes: The block coordinate descent algorithm fixes the non-target variable blocks at the most recently updated values and iteratively minimizes F for each variable block ω j to obtain the value of ω after the k-th update, i.e.: j That is: where N v is the total number of variable blocks; In each iteration, the DL-MIML model updates the variable block ω by minimizing an approximate linear surrogate function, i.e.: j Update, i.e.: where α k is the step size; The algorithm introduces auxiliary variables during the optimization process of ω j and initializes k = 0, a0 = 1 ​ where the auxiliary variable z j satisfies and In the k-th iteration, DL-MIML first calculates the smallest i ≥ 0 such that By the smallest i, we can obtain Update the variable to be learned as According to the updated variables to be learned, obtain the auxiliary variables for the next iteration, that is Repeat the iteration until convergence.

8. The non-convex unbalanced multi-instance multi-label learning method based on double-grained annotation according to claim 6, characterized in that The optimization of the variable to be learned A includes: Fix {b c} and {u i,c}, and the objective function is determined to be equivalent to: The gradient of the objective function f with respect to A is Use the Nesterov-accelerated gradient algorithm to calculate the optimization result of the variable to be learned A.

9. The non-convex unbalanced multi-instance multi-label learning method based on dual-grain annotation according to claim 6, wherein The variable b to be learned c is optimized as follows: Fix A, {b m | m ≠ c}, and {u i,c}, determine that the objective function is equivalent to: The gradient of the objective function f with respect to b c is The optimized result of the variable b to be learned is calculated using the Nesterov-accelerated gradient algorithm c is obtained 10. The non-convex unbalanced multi-instance multi-label learning method based on dual-grain annotation according to claim 6, characterized in that, Variable u to be learned i,c The optimization includes: Fix A, {b c}, and {u m,n | m≠i, n≠c}, the objective function is equivalent to The DL-MIML model is used to solve the convex optimization problem with non-convex sparse-inducing regularization, min f(u i,c ), by the projected gradient descent algorithm, and the optimal solution is the sparse Euclidean projection on the set C i,c , that is, Among them represents a projection operator, and I * is to retain the index set of the first r largest positive elements in and * is the complement of I can be obtained through a gradient-based algorithm. The gradient of the objective function f with respect to u i,c is

Citation Information

Patent Citations

  • Feature selection based multi-example multi-tag learning method and system

    CN105046284A

  • Joint low-rank constraint cross-view discrimination subspace learning method and device

    CN110619367A

  • Scene image labeling method based on coarse-fine granularity multi-image multi-label learning

    CN111461265A

  • Multi-example multi-label learning method based on combined error correction coding strategy

    CN117011647A