Image classification method based on multi-instance learning
Through the multi-example learning method, the mapping relationship between example tags and package tags is learned, and the problem of insufficient feature extraction in the image classification network in the prior art is not rich enough in complex scenarios and multi-label classification tasks, achieving higher error tolerance and robustness.
Patent Information
- Application Number
- CN202510198105.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
When handling complex scenes and multi-label classification tasks, existing image classification networks are prone to ignore global information or local details of the image, resulting in insufficient feature extraction or inaccurateness.
Using an image classification method based on multi-example learning, the complex mapping relationship between sample tags and package tags is learned through neural networks, feature semantic information related to sample image classification tasks in the package is extracted, and packet features are obtained through feature fusion modules.
The error tolerance rate of image package classification tasks is improved, the robustness of the model in real scenes is enhanced, and the adaptability of different types of image classification tasks is improved.
Smart Images

Figure CN120125895A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning image classification, and particularly to an image classification method based on multi-instance learning. Background Art
[0002] Image classification networks usually rely on traditional convolutional neural network (CNN) structures such as LeNet, AlexNet, VGG, etc. When extracting features, these networks may overlook the global information or local details of images, resulting in insufficiently rich or accurate extracted features, and having certain limitations in dealing with tasks such as complex scenes and multi-label classification. Therefore, it is necessary to design an image classification method based on multi-instance learning that can learn the complex mapping relationship between instance labels and bag labels through a neural network, thereby improving the error tolerance rate of image bag classification tasks and adapting to different types of image classification tasks in reality. Summary of the Invention
[0003] The object of the invention is to provide an image classification method based on multi-instance learning, which can learn the complex mapping relationship between instance labels and bag labels through a neural network, thereby improving the error tolerance rate of image bag classification tasks and adapting to different types of image classification tasks in reality.
[0004] Technical Solution: The image classification method based on multi-instance learning according to the present invention includes the following steps:
[0005] Step 1, perform a preprocessing operation on the input image to mask out the irrelevant regional images in the input image and obtain a preprocessed image;
[0006] Step 2, use an object detection algorithm to extract each region of interest in the preprocessed image, and merge the regions of interest in the preprocessed image into a bag;
[0007] Step 3, input the bag into an optimized multi-instance image classification network. The multi-instance image classification network includes a feature extraction network and a feature fusion module. The feature extraction network extracts the feature semantic information related to the instance image classification task in the bag as instance features, and then the feature fusion module performs feature fusion on the instance features to obtain bag features;
[0008] Step 4, map the bag features to the sample label space through a fully connected layer, and then output the feature prediction probability of the bag through a classifier, and take the classification result with the maximum prediction probability as the classification result of the input image.
[0009] Further, in Step 1, when performing the preprocessing operation on the input image:
[0010] The input image is cropped using the two-stage object detection algorithm R-CNN to remove the region images irrelevant to classification. The region images irrelevant to classification include the edge region of the input image and the region where the image noise is greater than a preset noise threshold.
[0011] Further, in step 2, the region of interest is the region image related to the classification task of the example image.
[0012] Further, in step 3, after the multi-instance image classification network is constructed, it needs to be optimized. The optimization steps are as follows:
[0013] First, extract each region of interest in the preprocessed example image and merge the regions of interest into an example bag;
[0014] Then, input the example bag into the multi-instance image classification network to obtain the example bag features;
[0015] Then, map the example bag features to the sample label space through a fully connected layer, and then output the feature prediction probability of the example bag through a classifier;
[0016] Finally, use the cross-entropy function to calculate the network prediction error based on the label of the example image and the feature prediction probability of the example bag, and then optimize the hyperparameters in the multi-instance image classification model according to the network prediction error.
[0017] Further, in step 3, when optimizing the hyperparameters in the multi-instance image classification model according to the network prediction error:
[0018] Use the Adam optimizer to update the hyperparameters in the multi-instance image classification model through the gradient descent algorithm according to the gradient information of the cross-entropy function, so as to achieve the optimization of the hyperparameters.
[0019] Further, in step 3, the feature extraction network is composed of multiple spatial state convolutional encoders stacked together, which is used to learn the feature semantic information of different scales of the example images in the bag.
[0020] Further, in step 3, the spatial state convolutional encoder includes an independent first extraction branch and a second extraction branch; the first extraction branch is an ordinary convolution branch, which is used for local feature extraction; the second extraction branch is a state space model branch, which is used to capture long-range dependencies of the features.
[0021] Further, in step 3, the feature fusion module includes a long short-term memory module, a multi-head self-attention module, a dynamic multi-instance normalization module, and a one-dimensional convolution module;
[0022] The long short-term memory module is used to receive example features, capture the associated feature information between each example feature, and send the captured associated feature information to both the multi-head self-attention module and the adder simultaneously;
[0023] The multi-head self-attention module is used to receive the associated feature information, learn the relationships between different associated feature information through the multi-head self-attention mechanism to obtain attention features, and then assign different weights to each attention feature;
[0024] The dynamic multi-instance normalization module is used to receive the attention features, and use the statistics to achieve the consistency of the feature representations at the bag level and the example level for the attention features, obtain the normalized features, and send the normalized features to the adder;
[0025] The adder is used to add the associated feature information and the features to obtain a feature vector;
[0026] The one-dimensional convolution module is used to perform one-dimensional convolution layer fusion on the feature vector obtained by the adder to obtain the bag-level bag features.
[0027] Furthermore, in step 4, when mapping the bag features to the sample label space through the fully connected layer:
[0028] Input the bag features into the fully connected layer, and the fully connected layer performs a linear transformation on the bag features through the linear combination of the internal weight matrix and the bias term to obtain a new feature representation.
[0029] Furthermore, in step 4, when outputting the feature prediction probability of the bag through the classifier:
[0030] Input the new feature representation into the sigmoid classifier to generate the corresponding prediction probability from 0 to 1 for each new feature representation.
[0031] Compared with the prior art, the beneficial effects of the present invention are: The present invention classifies by focusing on key examples, and improves the robustness of the model in real scenarios under the same image example classification accuracy rate to adapt to different types of image classification tasks in reality, so as to improve the drawbacks of the current image classification algorithm and provide a solution for adapting to different classification scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic diagram of the method flow of the present invention;
[0033] Figure 2 It is a schematic diagram of the preprocessing process of the present invention;
[0034] Figure 3 It is a structural diagram of the image classification network model of the present invention;
[0035] Figure 4Structural diagram of the spatial state convolutional encoder of the present invention;
[0036] Figure 5 Schematic diagram of the state space branch structure of the present invention;
[0037] Figure 6 Structural diagram of the dynamic long-term example fusion module of the present invention. Detailed implementation manners
[0038] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the described embodiments.
[0039] As Figure 1 and 3 shown, the image classification method based on multi-instance learning disclosed by the present invention includes the following steps:
[0040] Step 1, perform a preprocessing operation on the input image to mask the irrelevant region images in the input image and obtain a preprocessed image;
[0041] Step 2, use an object detection algorithm to extract each region of interest in the preprocessed image, and merge the regions of interest in the preprocessed image into a bag. One preprocessed image corresponds to one bag, as Figure 2 shown;
[0042] Step 3, input the optimized multi-instance image classification network in units of bags. The multi-instance image classification network includes a feature extraction network and a feature fusion module. The feature semantic information related to the example image classification task in the bag is extracted by the feature extraction network as example features, and then the example features are fused by the feature fusion module to obtain bag features;
[0043] Step 4, map the bag features to the sample label space through a fully connected layer, and then output the feature prediction probability of the bag through a classifier. The classification result with the largest prediction probability is used as the classification result of the input image.
[0044] Further, in step 1, when performing the preprocessing operation on the input image:
[0045] Use the two-stage object detection algorithm R-CNN to crop the input image and remove the region images irrelevant to classification. The region images irrelevant to classification include the edge region of the input image and the region where the image noise is greater than a preset noise threshold.
[0046] Further, in step 2, the region of interest is the region image related to the example image classification task.
[0047] The Region of Interest (ROI) refers to the area in an image that has specific attributes or features. The ROI is the area that needs to be further analyzed, processed, or recognized, rather than the entire image. In the present invention, it specifically refers to the area containing information related to the classification task.
[0048] Furthermore, in step 3, after the multi-instance image classification network is constructed, it needs to be optimized. The optimization steps are as follows:
[0049] First, extract each region of interest in the preprocessed instance images and merge the regions of interest into an instance bag;
[0050] Then, input the instance bag into the multi-instance image classification network to obtain the instance bag features;
[0051] Then, map the instance bag features to the sample label space through a fully connected layer, and then output the feature prediction probability of the instance bag through a classifier;
[0052] Finally, use the cross-entropy function to calculate the network prediction error based on the label of the instance image and the feature prediction probability of the instance bag, and then optimize the hyperparameters in the multi-instance image classification model according to the network prediction error.
[0053] Furthermore, in step 3, when optimizing the hyperparameters in the multi-instance image classification model according to the network prediction error:
[0054] Use the Adam optimizer to update the hyperparameters in the multi-instance image classification model through the gradient descent algorithm according to the gradient information of the cross-entropy function, so as to achieve the optimization of the hyperparameters.
[0055] Furthermore, in step 3, when the feature extraction network extracts the feature semantic information related to the instance image classification task in the bag as the instance feature:
[0056] First, input each bag into the multi-instance image classification network in turn through forward propagation;
[0057] Then, the feature extraction network of the multi-instance image classification network extracts features from each bag to obtain the feature semantic information related to the instance image classification task. The feature semantic information is a high-dimensional tensor;
[0058] Take the obtained feature semantic information as the instance feature of the corresponding bag, as Figure 3 shown.
[0059] Furthermore, in step 3, the feature extraction network is composed of multiple spatial state convolutional encoders stacked together, which is used to learn the feature semantic information of different scales of the instance images in the bag.
[0060] Further, in step 3, the spatial state convolution encoder includes a splitting module, a first extraction branch, a second extraction branch, a splicing module, a third convolution module, and a residual connection module, as Figure 4 shown;
[0061] The splitting module is used to split the input packet and distribute it to the first extraction branch and the second extraction branch respectively;
[0062] The first extraction branch and the second extraction branch are independent of each other. The first extraction branch is a common convolution branch, which is used to extract local features of the input packet; the second extraction branch is a state space model branch, which is used to capture long-distance dependencies of the input packet and obtain the correlation between examples in the packet;
[0063] The splicing module is used to splice the local features extracted by the first extraction branch and the correlation captured by the second extraction branch;
[0064] The third convolution module is used to perform 1×1 convolution processing on the features spliced by the splicing module;
[0065] The residual connection module is used to perform residual connection on the features after 1×1 convolution processing and the input packet, and output feature semantic information.
[0066] Adopting two independent extraction branches, compared with the single-branch structure, effectively improves the feature extraction ability of the encoder; solves the problem that the ordinary convolution operation is limited in long-distance modeling ability, reduces the number of parameters and the amount of calculation, and at the same time retains important feature information, and can better apply to the feature extraction of some special scenarios.
[0067] Further, as Figure 4 shown, the first extraction branch is a common convolution branch, including a 3×3 convolution module and two 1×1 convolution modules; the local feature extraction ability is realized by the common convolution operation of the sequentially connected 3×3 convolution module and two 1×1 convolution modules.
[0068] Further, as Figure 5 shown, the second extraction branch is a state space model branch, including a first layer normalization module (LN), a first linear layer (Linear), a DW convolution module (DWConv), an SS2D module (SS2D), a second layer normalization module (LN), a second linear layer (Linear), and a third linear layer (Linear);
[0069] The first layer normalization module is used to limit the data range and eliminate the adverse effects caused by singular sample data;
[0070] The first linear layer is used to realize the linear transformation of features;
[0071] DW convolution is used to retain spatial information and improve local perception ability;
[0072] The SS2D module is used to process visual data. It traverses the spatial domain through a four-way scanning mechanism, solves the problems faced by applying sequence models to two-dimensional visual data, and aims to perform efficient visual representation learning with linear time complexity;
[0073] The second normalization module is used to limit the data range and eliminate the adverse effects caused by singular sample data;
[0074] The second linear layer is used to implement the linear transformation of features;
[0075] The third linear layer is used to implement the linear transformation of features.
[0076] Furthermore, in step 3, the feature fusion module includes a long short-term memory module, a multi-head self-attention module, a dynamic multi-instance normalization module, and a one-dimensional convolution module;
[0077] The long short-term memory module is used to receive example features, capture the associated feature information between each example feature, and send the captured associated feature information to both the multi-head self-attention module and the adder;
[0078] The multi-head self-attention module is used to receive the associated feature information, learn the relationship between different associated feature information through the multi-head self-attention mechanism to obtain attention features, and then assign different weights to each attention feature;
[0079] The dynamic multi-instance normalization module is used to receive the attention features and use the statistics to achieve the consistency of the feature representation at the bag level and the example level for the attention features, obtain the normalized features, and send the normalized features to the adder;
[0080] The adder is used to add the associated feature information and the features to obtain a feature vector;
[0081] The one-dimensional convolution module is used to perform one-dimensional convolutional layer fusion on the feature vector obtained by the adder to obtain the bag-level bag features.
[0082] The feature fusion module is used to fuse the example features in different numbers of bags into bag features, and obtain the feature representation of the input image corresponding to the bag.
[0083] Furthermore, in step 4, when mapping the bag features to the sample label space through the fully connected layer:
[0084] The bag features are input into the fully connected layer, and the fully connected layer performs a linear transformation on the bag features through the linear combination of the internal weight matrix and bias term to obtain a new feature representation.
[0085] Further, in step 4, when the feature prediction probability of the output packet is obtained by the classifier:
[0086] The new feature representation is input into the sigmoid classifier to generate corresponding prediction probabilities ranging from 0 to 1 for each new feature representation, which is used to classify the input image.
[0087] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as a limitation on the present invention itself. Various changes in form and detail may be made without departing from the spirit and scope of the present invention defined by the appended claims.
Claims
1. An image classification method based on multi-instance learning, characterized in that: The steps include: Step 1, preprocessing the input image, shielding the irrelevant area image in the input image, and obtaining a preprocessed image; Step 2, using the target detection algorithm to extract the various regions of interest in the preprocessed image, and merging the various regions of interest in the preprocessed image into one package; Step 3, inputting the optimized multi-instance image classification network in units of packages, the multi-instance image classification network includes a feature extraction network and a feature fusion module, the feature extraction network extracts feature semantic information related to the example image classification task in the package as the example feature, and then the feature fusion module fuses the example features to obtain the package feature; Step 4: Map the packet features to the sample label space through the fully connected layer, and then output the feature prediction probability of the packet through the classifier, and take the classification result with the largest prediction probability as the classification result of the input image.
2. The image classification method based on multi-instance learning according to claim 1, characterized in that: In step 1, when preprocessing the input image: The two-stage target detection algorithm R-CNN is used to crop the input image and remove the area image irrelevant to the classification. The area image irrelevant to the classification includes the edge area of the input image and the area where the image noise is greater than the preset noise threshold.
3. The image classification method based on multi-instance learning according to claim 1, characterized in that: In step 2, the region of interest is a region image related to the example image classification task.
4. The image classification method based on multi-instance learning according to claim 1, characterized in that: In step 3, the multi-instance image classification network needs to be optimized after construction. The optimization steps are: Firstly, each region of interest in the preprocessed sample image is extracted, and each region of interest is merged into a sample package; Then the sample bag is input into the multi-instance image classification network to obtain the sample bag features; Then, the sample package features are mapped to the sample label space through the fully connected layer, and then the feature prediction probability of the sample package is output through the classifier; Finally, the network prediction error is calculated using the cross entropy function based on the label of the example image and the feature prediction probability of the example package, and the hyperparameters in the multi-example image classification model are optimized based on the network prediction error.
5. The image classification method based on multi-instance learning according to claim 4, characterized in that: In step 3, when optimizing the hyperparameters in the multi-instance image classification model based on the network prediction error: The Adam optimizer is used to update the hyperparameters in the multi-instance image classification model through the gradient descent algorithm according to the gradient information of the cross entropy function, thereby optimizing the hyperparameters.
6. The image classification method based on multi-instance learning according to claim 1, characterized in that: In step 3, the feature extraction network consists of multiple stacked spatial state convolutional encoders to learn the feature semantic information of different scales of example images in the bag.
7. The image classification method based on multi-instance learning according to claim 6, characterized in that: In step 3, the spatial state convolution encoder includes a first extraction branch and a second extraction branch which are independent of each other; the first extraction branch is a common convolution branch, which is used for local feature extraction; the second extraction branch is a state space model branch, which is used for capturing long-distance dependencies of features.
8. The image classification method based on multi-instance learning according to claim 1, characterized in that: In step 3, the feature fusion module includes a long short-term memory module, a multi-head self-attention module, a dynamic multi-instance normalization module, and a one-dimensional convolution module; The long short-term memory module is used to receive example features and capture the associated feature information between each example feature, and simultaneously send the captured associated feature information to the multi-head self-attention module and the adder; The multi-head self-attention module is used to receive related feature information, and learn the relationship between different related feature information through the multi-head self-attention mechanism to obtain attention features, and then assign different weights to each attention feature; The dynamic multi-instance normalization module is used to receive the attention features, and use the statistics to achieve the consistency of the feature representation at the package level and the instance level for the attention features, obtain the normalized features, and send the normalized features to the adder; The adder is used to add the associated feature information and the features to obtain a feature vector; The one-dimensional convolution module is used to perform one-dimensional convolution layer fusion on the feature vector obtained by the adder to obtain packet-level packet features.
9. The image classification method based on multi-instance learning according to claim 1, characterized in that: In step 4, when mapping the packet features to the sample label space through the fully connected layer: The packet features are input into the fully connected layer, which performs a linear transformation on the packet features through a linear combination of the internal weight matrix and bias terms to obtain a new feature representation.
10. The image classification method based on multi-instance learning according to claim 7, characterized in that: In step 4, when the feature prediction probability of the classifier output package is: The new feature representation is input into the sigmoid classifier to generate a corresponding prediction probability of 0 to 1 for each new feature representation.
Citation Information
Patent Citations
Medical image classification method based on deep multi-instance learning and self-attention
CN112598024A
Cancer pathological image survival prognosis model construction method based on deep learning
CN113947607A
Facial expression recognition method and device, equipment and medium
CN117218695A
Thyroid cell pathology full-slide image classification method
CN117994783A
High-resolution remote sensing image target detection method based on multi-scale network
CN118485927A
Cited By
Vehicle re-identification model construction method based on multi-image learning
CN122347784A