A Method and Device for Multi-Label Classification to Address Incomplete Label Annotations
The method uses a partially labeled validation set to improve multi-label classification accuracy by adjusting loss functions, reducing labor costs and enhancing task precision.
Patent Information
- Application Number
- CN202211457173.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-11-21
AI Technical Summary
The problems of incomplete labeling and missing labels in existing multi-label classification tasks lead to huge workload and low accuracy.
The complete verification set of a small number of annotated verification sets are used to guide the correction of multi-label model training. The unlabeled label loss function is calculated by combining positive labels, negative labels and unlabeled label loss function, combined with rebalancing parameters, to improve the model training effect.
Effectively save manual tagging costs and improve the accuracy of multi-label classification tasks.
Smart Images

Figure CN115908921B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of deep learning, and particularly relates to a method and device for multi-label classification in response to incomplete label annotation. Background Art
[0002] With the rapid development of mobile Internet technology, taking photos casually and sharing them on the network has become a new living habit. How to quickly classify network pictures is the key to the next step of network big data processing. However, in real life, it is common for a single picture to contain multiple targets and objects. Therefore, multi-label classification technology is required to label multiple tags that appear in a single picture. With the development of deep learning technology in recent years, using a deep network to complete this task has become a trend.
[0003] However, deep networks often require a large number of labeled samples. For multi-label classification tasks, a single picture needs to be labeled with multiple category tags, and the manual labeling workload is huge and missing tags are very likely to occur.
[0004] In view of this, the present invention proposes a method and device for multi-label classification in response to incomplete label annotation, which uses a small number of completely labeled validation sets to guide and correct the training of a large multi-label model, and can effectively address the problems of incomplete training set annotation and missing tags in multi-label classification tasks. Summary of the Invention
[0005] In order to solve the technical problems such as incomplete training set annotation and missing tags in existing multi-label classification tasks, the present application provides a method and device for multi-label classification in response to incomplete label annotation to solve the above technical defect problems.
[0006] According to one aspect of the present invention, a method for multi-label classification in response to incomplete label annotation is proposed, including the following steps:
[0007] S1. Obtain a validation set and a training set, and manually label each picture in the validation set with multiple category tags;
[0008] S2. Train a multi-label classification model based on the training set and a total loss function, where the total loss function includes a positive label loss function, a negative label loss function, and an unlabeled label loss function;
[0009] S3. Calculate the unlabeled label loss function based on the validation set and a rebalancing parameter; and
[0010] S4. Finally, obtain a trained multi-label classification model.
[0011] In a specific embodiment, in step S2, the expression of the total loss function is:
[0012]
[0013] Among them, represents the positive label loss function; represents the negative label loss function; represents the unlabeled label loss function; P X represents the positive label set; N X represents the negative label set; U X represents the unlabeled label set.
[0014] In a specific embodiment, in step S3, the unlabeled label loss function is calculated based on the validation set and the rebalancing parameter, and the specific expression is:
[0015]
[0016] Among them, L(f(x), topK({p c})) represents taking the top K of the current network prediction values; βc represents the rebalancing parameter, β c is obtained from the calculation results of the same loss function on the validation set.
[0017] In a specific embodiment, in step S3, the calculation expression of the rebalancing parameter is:
[0018]
[0019] Among them, D val represents the validation set; for the unlabeled label loss function to always have a positive promotion effect on the training of the neural network, the rebalancing parameter β c should not be less than 0, so finally
[0020] In a specific embodiment, in step S1, it further includes obtaining a multi-label data set and dividing the multi-label data set into a validation set and a training set according to a ratio.
[0021] In a second aspect, the present application proposes a device for multi-label classification to address incomplete label annotation, including the following modules:
[0022] An acquisition module, configured to acquire a validation set and a training set, and manually label each picture in the validation set with multiple category labels; and
[0023] A training module, which trains a multi-label classification model based on the training set and the total loss function, where the total loss function includes a positive label loss function, a negative label loss function, and an unlabeled label loss function; and
[0024] A computing module that calculates an unlabeled label loss function based on a validation set and rebalancing parameters; and
[0025] An output module that finally obtains a trained multi-label classification model.
[0026] In a specific embodiment, in the training module, the expression of the total loss function is:
[0027]
[0028] Wherein, represents the positive label loss function; represents the negative label loss function; represents the unlabeled label loss function; P X represents the positive label set; N X represents the negative label set; U X represents the unlabeled label set.
[0029] In a specific embodiment, in the computing module, the unlabeled label loss function is calculated based on the validation set and rebalancing parameters, and the specific expression is:
[0030]
[0031] Wherein, L(f(x), topK({p c})) represents taking the top K of the current network prediction values; β c represents the rebalancing parameter, and β c is obtained from the calculation results of the same loss function on the validation set.
[0032] In a specific embodiment, in the computing module, the calculation expression of the rebalancing parameter is:
[0033]
[0034] Wherein, D val represents the validation set; for the unlabeled label loss function to always have a positive impact on the training of the neural network, the rebalancing parameter β c should not be less than 0, so finally
[0035] Thirdly, the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method according to any one of the above is implemented.
[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0037] The present invention proposes a new method that can utilize a small part of the verification set with complete annotations to guide the calculation of the loss function for correcting the incomplete label annotations in a large multi-label training set, effectively saving the manual labeling cost and improving the accuracy of the multi-label classification task. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent:
[0039] Figure 1 is a flowchart of the method for multi-label classification for coping with incomplete label annotations according to the present application;
[0040] Figure 2 is a schematic diagram of the device for multi-label classification for coping with incomplete label annotations according to the present application;
[0041] Figure 3 is a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The present application will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. In addition, it should be noted that for the convenience of description, only the parts related to the invention are shown in the drawings.
[0043] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0044] For the convenience of understanding by those skilled in the art, the general idea of the method for multi-label classification for coping with incomplete label annotations in the embodiments of the present invention is introduced as follows.
[0045] For a specific image X in a multi-label dataset containing N label categories, its annotated label where y c ∈{-1, 0, 1}, represents whether the category c appears {'1'} in the image X, does not appear {'-1'}, or is unknown {'0'}.
[0046] Define the positive label set of the image X negative label set unlabeled label set The loss functions of each part are statistically analyzed item by item, so the loss function of the multi-label classifier can be expressed as follows:
[0047]
[0048] Train a neural network f(x) using the above loss function to complete the prediction of the label category to which the picture X belongs.
[0049] During the training process of a multi-label classification task, the calculation of the loss function for the unknown labels of the samples is divided into two methods. The first is to ignore, in which case The other is to treat it as a negative sample, in which case This application aims to guide and correct the training noise caused by incomplete label annotation in a large multi-label training set through a fully annotated validation set during the training process, and provides another calculation method.
[0050] Figure 1 The flowchart of the method for multi-label classification that addresses incomplete label annotation in this application is shown. Please refer to Figure 1 This method includes the following steps:
[0051] S1. Obtain a validation set and a training set, and manually label each picture in the validation set with multiple category labels.
[0052] In a specific embodiment, in step S1, it further includes obtaining a multi-label data set and dividing the multi-label data set into a validation set and a training set according to a certain ratio.
[0053] S2. Train a multi-label classification model based on the training set and the total loss function, where the total loss function includes a positive label loss function, a negative label loss function, and an unlabeled label loss function;
[0054] S3. Calculate the unlabeled label loss function based on the validation set and the rebalancing parameter; and
[0055] S4. Finally, obtain the trained multi-label classification model.
[0056] In a specific embodiment, in step S2, the expression of the total loss function is:
[0057]
[0058] Among them, represents the positive label loss function; represents the negative label loss function; represents the unlabeled label loss function; P X represents the positive label set; N X represents the negative label set; U X represents the unlabeled label set.
[0059] To utilize the useful information of the fully labeled validation set and address the unknown labels of samples in the training set, a rebalancing parameter is introduced during the calculation of the unlabeled label loss function The specific expression of the unlabeled label loss function is as follows:
[0060]
[0061] In the formula, L(f(x), topK({p c})) represents taking the top K of the current network prediction values.
[0062] β c represents the rebalancing parameter, and β c is obtained from the calculation results of the same loss function on the validation set D val The calculation expression of the rebalancing parameter β c is as follows:
[0063]
[0064] In the formula, D val represents the validation set; to always have a positive impact on the training of the neural network for the unlabeled label loss function the rebalancing parameter β c should be no less than 0. Therefore, finally
[0065] during the training process, the rebalancing parameter β c is dynamically adjusted with the help of the validation set, enabling the network to obtain more positive feedback and further improving the accuracy.
[0066] This application can utilize a small portion of the fully labeled validation set to guide and correct the calculation of the loss function for the incomplete label annotation part in a large multi-label training set, effectively saving the manual labeling cost and improving the accuracy of the multi-label classification task.
[0067] Further referring to Figure 2 as an implementation of the above method, this application provides an embodiment of a device for multi-label classification dealing with incomplete label annotation. This device embodiment corresponds to the Figure 1 method embodiment shown and can be specifically applied to various electronic devices. The device 200 includes the following modules:
[0068] An acquisition module 210, configured to acquire a validation set and a training set, and manually label each image in the validation set with multiple category labels; and
[0069] A training module 220 that trains a multi-label classification model based on a training set and a total loss function, where the total loss function includes a positive label loss function, a negative label loss function, and an unlabeled label loss function; and
[0070] A calculation module 230 that calculates and obtains an unlabeled label loss function based on a validation set and a rebalancing parameter; and
[0071] An output module 240 that finally obtains a trained multi-label classification model.
[0072] In a specific embodiment, in the training module 220, the expression of the total loss function is:
[0073]
[0074] Wherein, represents the positive label loss function; represents the negative label loss function; represents the unlabeled label loss function; P X represents the positive label set; N X represents the negative label set; U X represents the unlabeled label set.
[0075] In a specific embodiment, in the calculation module 230, the unlabeled label loss function is calculated and obtained based on the validation set and the rebalancing parameter, and the specific expression is:
[0076]
[0077] Wherein, L(f(x), topK({p c})) represents taking the top K of the current network prediction values; β c represents the rebalancing parameter, and β c is obtained from the calculation results of the same loss function on the validation set.
[0078] In a specific embodiment, in the calculation module 230, the calculation expression of the rebalancing parameter is:
[0079]
[0080] Wherein, D val represents the validation set; for the unlabeled label loss function to always have a positive promotion effect on the training of the neural network, the rebalancing parameter β c should not be less than 0, so finally
[0081] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the above is implemented.
[0082] Refer to the following Figure 3 , which shows a schematic structural diagram of a computer system 300 of a terminal device or a server suitable for implementing the embodiments of the present application. Figure 3 The shown terminal device or server is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0083] As Figure 3 shown, the computer system 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 302 or the programs loaded from the storage section 308 into the random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the system 300 are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0084] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as required. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as required so that a computer program read from it can be installed into the storage section 308 as required.
[0085] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, the above functions defined in the methods of the present application are executed. It should be noted that the computer-readable medium described in the present application can be a computer-readable signal medium or a computer-readable medium or any combination of the two. The computer-readable medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.
[0086] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).
[0087] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0088] The modules described in the embodiments of this application can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a receiving module, an obtaining module, a determining module, a calculating module, and a generating module. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the receiving unit can also be described as "a module that, in response to determining that the verification request information includes a username, a request time, a user signature code, and a client application code, obtains the configured information of the target user preset."
[0089] As another aspect, the present application also provides a computer-readable medium, which may be included in the server described in the above embodiments; or may exist alone without being assembled into the server. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the server, the server is caused to: receive verification request information sent by the client of the target user; in response to determining that the verification request information includes a user name, a request time, a user signature code, and a client application code, obtain the preset configuration information of the target user, where the configuration information includes the preset user password corresponding to the user name; determine whether the verification request information is valid according to the request time, and in response to determining that it is valid, determine whether the user signature code is included in the preset storage area; in response to determining that it is not included, store the user signature code in the preset storage area, and calculate a server application code based on the user password, the request time, and the user signature code; in response to determining that the server application code and the client application code match, generate verification success information for indicating that the verification request is a legitimate request.
[0090] In addition, the above computer-readable medium may be included in the terminal device described in the above embodiments; or may exist alone without being assembled into the terminal device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the terminal device, the terminal device is caused to: obtain user information input by the target user, where the user information includes a user name and a user password; generate a user signature code for indicating the target user based on the user information; determine the request time; calculate a client application code based on the user password, the request time, and the user signature code; generate verification request information including the user name, the request time, the user signature code, and the client application code, and send the verification request information to the server.
[0091] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principle. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present application.
Claims
1. A method for multi-label classification to address incomplete label annotation, characterized in that, It includes the following steps: S1. Obtain a validation set and a training set, and manually label each image in the validation set with multiple category labels; S2. Train a multi-label classification model based on the training set and a total loss function, where the total loss function includes a positive label loss function, a negative label loss function, and an unlabeled label loss function; the expression of the total loss function is: Among them, represents the positive label loss function; represents the negative label loss function; represents the unlabeled label loss function; P X represents the positive label set; N X represents the negative label set; U X represents the unlabeled label set; S3. Calculate the unlabeled label loss function based on the validation set and a rebalancing parameter, and the specific expression is: Among them, L(f(x), topK({p c})) represents taking the top K of the current network prediction values; β c represents the rebalancing parameter, and β c is calculated from the calculation results of the same loss function on the validation set. The calculation expression of the rebalancing parameter is: Among them, D val represents the validation set; for the unlabeled label loss function to always have a positive impact on the training of the neural network, the rebalancing parameter β c should not be less than 0, so finally and S4. Finally, obtain the trained multi-label classification model.
2. The method for multi-label classification for coping with incomplete label annotation according to claim 1, characterized in that, In step S1, it further includes obtaining a multi-label data set, and dividing the multi-label data set into the validation set and the training set according to a ratio.
3. A device for multi-label classification to address incomplete label annotation, characterized in that, It includes the following modules: An acquisition module, which is used to obtain a validation set and a training set, and manually label each image in the validation set with multiple category labels; And A training module, which trains a multi-label classification model based on the training set and a total loss function, where the total loss function includes a positive label loss function, a negative label loss function, and an unlabeled label loss function, and the expression of the total loss function is: Among them, represents the positive label loss function; represents the negative label loss function; represents the unlabeled label loss function; P X represents the positive label set; N X represents the negative label set; U X represents the unlabeled label set; and A calculation module, which calculates the unlabeled label loss function based on the validation set and a rebalancing parameter, and the specific expression is: Among them, L(f(x), topK({p c})) represents taking the top K of the current network prediction values; β c represents the rebalancing parameter, and β c is calculated from the calculation results of the same loss function on the validation set. The calculation expression of the rebalancing parameter is: Among them, D val represents the validation set; for the unlabeled label loss function to always have a positive impact on the training of the neural network, the rebalancing parameter β c should not be less than 0, so finally and An output module, which finally obtains the trained multi-label classification model.
4. A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of claims 1-2 is implemented.
Citation Information
Patent Citations
Lung X-ray image classification method based on K-means clustering and GAN
CN113222072A
Target detection network self-supervised training method and device based on event camera
CN114049483A