Multi-instance image classification method based on pre-check feature fusion
Through the multi-instance image classification method based on pre-check feature fusion, the problems of poor detection effect and insufficient computing resources in the prior art image classification are solved, and efficient multi-instance image classification is realized, and the accuracy and speed are improved.
Patent Information
- Application Number
- CN202510300356.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The existing multi-instance image classification method has poor detection effects in scenarios where there are many instances, small target sizes, and insufficient computing resources, and it is difficult to achieve lightweight calculations and real-time image analysis.
Using a multi-instance image classification method based on pre-check feature fusion, multi-instance imaging is obtained through a multi-instance image acquisition tool, region pre-checking is used to perform area pre-checking, single-instance feature extraction is used to perform multi-instance feature weighting and fusion is used to perform multi-instance feature compression, and finally joint prediction of multi-instance classification is performed through instance fusion view and spatial fusion view.
The image size and computing volume in multi-instance scenarios are reduced, the effectiveness of multi-instance image information is improved, the problems of sparse information and high computing requirements are solved, and the accuracy and speed of multi-instance image classification are improved.
Smart Images

Figure CN120164031A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning multi-instance image classification, and in particular to a multi-instance image classification method based on pre-check feature fusion. Background Art
[0002] With the large-scale application of deep learning technology in the field of computer vision, deep learning-based multi-instance learning algorithms for images have become mainstream in multi-instance learning tasks. Due to the large number of images in multi-instance tasks, existing multi-instance image classification methods often suffer from the problem of information sparsity. At the same time, multi-instance tasks also increase the computational complexity. These problems pose great challenges to existing multi-instance learning methods in improving information effectiveness and reducing computational complexity, resulting in a relatively low accuracy of multi-instance image classification. Especially in scenarios such as remote sensing imaging and medical imaging where the recognition target area is small or blurred, large-sized images and small-sized targets exacerbate the information sparsity problem of multi-instance images, causing a lot of useless background images in multi-instance images to participate in the calculation.
[0003] Therefore, traditional multi-instance image classification methods have poor detection effects in scenarios with a large number of instances, small target sizes, and insufficient computing resources, and it is difficult to achieve lightweight computing and real-time image analysis. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-instance image classification method based on pre-check feature fusion to solve the problems that traditional multi-instance image classification methods have poor detection effects in scenarios with a large number of instances, small target sizes, and insufficient computing resources, and it is difficult to achieve lightweight computing and real-time image analysis.
[0005] To achieve the above purpose, the present invention provides a multi-instance image classification method based on pre-check feature fusion, and the multi-instance image classification method based on pre-check feature fusion includes the following steps:
[0006] Obtain multi-instance imaging through a multi-instance image acquisition tool;
[0007] Input the multi-instance imaging into a pre-check module for regional pre-detection to obtain the pre-check area of each instance;
[0008] Use a pre-check feature encoder to extract single-instance features from the pre-check area to obtain pre-check features;
[0009] Use an instance fusion module to perform multi-instance feature weighting and fusion on the pre-check features to obtain an instance fusion view;
[0010] Use a spatial fusion module to perform spatial feature compression on the pre-check features to obtain a spatial fusion view;
[0011] Perform joint prediction for multi-instance classification on the instance fusion view and the spatial fusion view to obtain the target category.
[0012] Among them, the specific content of the step "obtain multi-instance imaging through a multi-instance image acquisition tool" is:
[0013] For the recognition target, perform multi-instance image acquisition using multiple imaging and multi-angle imaging methods in the time and space dimensions;
[0014] Perform representative image screening on the multi-instance images, select images with different imaging times and different imaging angles, and delete similar images with adjacent imaging times and the same imaging angles to improve data representativeness.
[0015] Among them, the specific content of the step "input the multi-instance imaging into the pre-inspection module for regional pre-detection" is:
[0016] Set the aspect ratio, width and height dimensions, and center offset distance of the prior box for the multi-instance imaging. The calculation method of the center offset distance is that the horizontal and vertical coordinate differences between the geometric center point of the image and the geometric center point of the prior box form a matrix containing only two elements. Divide the horizontal coordinate difference in the matrix by the width of the image, divide the vertical coordinate difference in the matrix by the height of the image, and then calculate the Frobenius norm of the matrix to obtain the center offset distance;
[0017] Use the pre-inspection module to perform regional detection using the multi-stage detection model HTC, and select the prediction region with the largest difference in confidence and center offset distance as the pre-inspection region.
[0018] Among them, in the step "use the pre-inspection feature encoder to perform single-instance feature extraction on the pre-inspection region to obtain pre-inspection features", the pre-inspection feature encoder includes a single-instance feature extractor and a single-instance feature mixer;
[0019] The single-instance feature extractor includes a feature dimensionality reduction module and a feature aggregation module. The feature dimensionality reduction module is composed of a convolutional layer with a stride of 2, a random pooling layer, and a mixed pooling layer in parallel to achieve feature dimensionality reduction. The feature aggregation module is composed of a linear mapping layer, a state space model, and a dot product attention layer in series to achieve feature aggregation;
[0020] The single-instance feature mixer is used to splice the dimensionality-reduced features and the aggregated features in the channel dimension, and use grouped convolution with a group number of 2 and pointwise convolution with a group number of 1 to achieve channel compression and channel information mixing for the dimensionality-reduced features and the aggregated features.
[0021] Among them, the specific content of the step "use the instance fusion module to perform multi-instance feature weighting and fusion on the pre-inspection features to obtain an instance fusion view" is:
[0022] Arrange the multi-instance features in the instance order, convert the sequence number into a binary number to obtain a sequence code, accumulate the sequence code with the corresponding instance features and map them into a one-dimensional sequence to obtain multi-instance sequence features;
[0023] Input the multi-instance sequence features into a 2D state space model and a multi-layer perceptron to predict the weights w1 to w corresponding to each instance n , multiply each instance by the corresponding weight w i and obtain an instance fusion view through a one-dimensional convolutional layer.
[0024] Among them, the 2D state space model uses a bidirectional scanning mechanism in the channel dimension to achieve cross-instance feature mixing, and uses a bidirectional scanning mechanism in the spatial dimension to achieve cross-spatial position feature mixing.
[0025] Among them, the specific content of the step "compress the pre-inspection features using a spatial fusion module to obtain a spatial fusion view" is:
[0026] Map each instance of the pre-inspection features onto the same spatial plane and perform horizontal splicing to obtain a spatial splicing feature map with a height of h and a width of n×w, realizing instance rearrangement in the spatial dimension;
[0027] Use a horizontal and vertical asymmetric convolutional block for the spatial splicing feature map to obtain a spatial fusion view.
[0028] Among them, the specific content of the horizontal and vertical asymmetric convolutional block is: first perform ordinary convolution on the features, and then perform asymmetric horizontal convolution with a length of n×d and asymmetric vertical convolution with a length of d, where n is the number of instances and d is the convolutional kernel length parameter.
[0029] Among them, the specific content of the step "perform joint prediction of multi-instance classification on the instance fusion view and the spatial fusion view" is:
[0030] Perform global average pooling and multi-layer perceptron prediction on the instance fusion view and the spatial fusion view in a single-view manner and a dual-view manner to obtain an instance fusion view prediction value, a spatial fusion view prediction value, and a dual-view prediction value;
[0031] Concatenate the instance fusion view prediction value, the spatial fusion view prediction value, and the dual-view prediction value into a one-dimensional vector, and use a multi-layer perceptron to achieve feature dimensionality reduction of the one-dimensional vector to realize multi-view joint prediction.
[0032] A multi-instance image classification method based on pre-inspection feature fusion of the present invention includes obtaining multi-instance images of an object; using a pre-inspection module to perform regional pre-detection on each instance image; using a pre-inspection feature encoder to extract single-instance features from the pre-inspected regions to obtain pre-inspection features; using an instance fusion module to perform feature weighting and fusion on each instance to obtain an instance fusion view; using a spatial fusion module to map each instance onto the same spatial plane and perform spatial feature compression to obtain a spatial fusion view; and using the instance fusion view and the spatial fusion view to perform joint prediction of multi-instance classification. Through pre-inspection feature fusion, multi-instance classification learning of images can be achieved, reducing the image size and computational amount in multi-instance scenarios, improving the effectiveness of multi-instance image information, so as to solve the problems of information sparsity and high computational requirements in multi-instance image learning tasks, and improving the accuracy and speed of multi-instance image classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0034] Figure 1 is a flowchart of the steps of the multi-instance image classification method based on pre-inspection feature fusion provided by the present invention.
[0035] Figure 2 is a structural diagram of the multi-instance pre-inspection feature fusion network provided by the present invention.
[0036] Figure 3 is a structural diagram of the pre-inspection feature encoder provided by the present invention.
[0037] Figure 4 is a schematic diagram of the scanning direction of the 2D state space model provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The following will describe in detail the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0039] Please refer to Figures 1 to 4 , the present invention provides a multi-instance image classification method based on pre-inspection feature fusion. The multi-instance image classification method based on pre-inspection feature fusion includes the following steps:
[0040] Step 101: Obtain multi-instance images of the target.
[0041] The computer device performs multi-instance image acquisition on the recognition target in the time and space dimensions using multiple imaging and multi-angle imaging methods; performs representative image screening on the multi-instance images, selects images with different imaging times and different imaging angles, and deletes similar images with adjacent imaging times and the same imaging angles to improve data representativeness.
[0042] Step 102: Use the pre-inspection module to perform regional pre-detection on each instance image.
[0043] The computer device sets the aspect ratio, width and height dimensions, and center offset distance of the prior box for multi-instance imaging. The calculation method of the center offset distance is that the horizontal and vertical coordinate differences between the geometric center point of the image and the geometric center point of the prior box form a matrix containing only two elements. Divide the horizontal coordinate difference in the matrix by the width of the image, divide the vertical coordinate difference in the matrix by the height of the image, and then calculate the Frobenius norm of the matrix to obtain the center offset distance; the pre-inspection module uses the multi-stage detection model HTC for regional detection and selects the prediction region with the largest difference between the confidence level and the center offset distance as the pre-inspection region.
[0044] Step 103: Use the pre-inspection feature encoder to perform single-instance feature extraction on the pre-inspection region to obtain pre-inspection features.
[0045] The computer device uses a single-instance feature extractor and a single-instance feature mixer to achieve pre-inspection feature extraction and mixing; the single-instance feature extractor includes a feature dimensionality reduction module and a feature aggregation module; the feature dimensionality reduction module is composed of a convolutional layer with a stride of 2, a random pooling layer, and a mixed pooling layer in parallel to achieve feature dimensionality reduction; the feature aggregation module is composed of a linear mapping layer, a state space model, and a dot product attention layer in series to achieve feature aggregation; the single-instance feature mixer splices the dimensionality-reduced features and the aggregated features in the channel dimension, and uses grouped convolution with a group number of 2 and pointwise convolution with a group number of 1 to achieve channel compression and channel information mixing for the dimensionality-reduced features and the aggregated features.
[0046] Step 104: Use the instance fusion module to perform feature weighting and fusion on each instance to obtain an instance fusion view.
[0047] The computer device arranges the multi-instance features in the instance order, converts the sequence number into a binary number to obtain a sequence encoding, accumulates the sequence encoding with the corresponding instance features and maps them into a one-dimensional sequence to obtain multi-instance sequence features; inputs the multi-instance sequence features into a 2D state space model and a multi-layer perceptron to predict the weight corresponding to each instance, multiplies each instance by the corresponding weight, and obtains an instance fusion view through a one-dimensional convolutional layer.
[0048] Step 105: Use the spatial fusion module to map each instance onto the same spatial plane and perform spatial feature compression to obtain a spatially fused view.
[0049] The computer device can map each instance of the pre-inspection features onto the same spatial plane and perform horizontal splicing to obtain a spatially spliced feature map with a height of h and a width of n×w, achieving instance rearrangement in the spatial dimension; use a horizontally and vertically asymmetric convolution block on the spatially spliced feature map to obtain a spatially fused view.
[0050] Step 106: Use the instance fusion view and the spatial fusion view for joint prediction of multi-instance classification.
[0051] The computer device performs global average pooling and multi-layer perceptron prediction on the instance fusion view and the spatial fusion view in a single-view manner and a dual-view manner to obtain the instance fusion view prediction value, the spatial fusion view prediction value, and the dual-view prediction value, splice the three prediction values into a one-dimensional vector, and use a multi-layer perceptron to achieve feature dimensionality reduction of the one-dimensional vector, realizing multi-view joint prediction.
[0052] In one embodiment, as Figure 2 shown, a specific network structure of a multi-instance image classification method based on pre-inspection feature fusion is provided. The multi-instance pre-inspection feature fusion network can obtain multi-instance images of the target; use the pre-inspection module to perform regional pre-detection on each instance image; use the pre-inspection feature encoder to extract single-instance features from the pre-inspected regions to obtain pre-inspection features; use the instance fusion module to perform feature weighting and fusion on each instance to obtain an instance fusion view; use the spatial fusion module to map each instance onto the same spatial plane and perform spatial feature compression to obtain a spatially fused view; use the instance fusion view and the spatial fusion view for joint prediction of multi-instance classification.
[0053] In one embodiment, a multi-instance image classification method based on pre-inspection feature fusion may further include obtaining multi-instance imaging through a multi-instance image acquisition tool, including: performing multi-instance image acquisition on the recognition target in the time and space dimensions using multiple imaging and multi-angle imaging methods; performing representative image screening on the multi-instance images, selecting images with different imaging times and different imaging angles, and deleting similar images with adjacent imaging times and the same imaging angles to improve representativeness.
[0054] In one embodiment, a multi-instance image classification method based on pre-check feature fusion may further include setting the aspect ratio, width and height dimensions, and center offset distance of the prior box for multi-instance imaging. The calculation method of the center offset distance is that the horizontal and vertical coordinate differences between the geometric center point of the image and the geometric center point of the prior box form a matrix containing only two elements. Divide the horizontal coordinate difference in the matrix by the width of the image, divide the vertical coordinate difference in the matrix by the height of the image, and then calculate the Frobenius norm of the matrix to obtain the center offset distance. Use the pre-check module to perform region detection using the multi-stage detection model HTC, and select the prediction region with the largest difference between the confidence level and the center offset distance as the pre-check region.
[0055] In one embodiment, as Figure 3 shown, a multi-instance image classification method based on pre-check feature fusion may further include a pre-check feature encoder. The pre-check feature encoder includes a single-instance feature extractor and a single-instance feature mixer. The single-instance feature extractor includes a feature dimensionality reduction module and a feature aggregation module. The feature dimensionality reduction module is composed of a convolutional layer with a stride of 2, a random pooling layer, and a mixed pooling layer in parallel to achieve feature dimensionality reduction. The feature aggregation module is composed of a linear mapping layer, a state space model, and a dot product attention layer in series to achieve feature aggregation. The single-instance feature mixer concatenates the dimensionality-reduced feature and the aggregated feature in the channel dimension, and uses grouped convolution with a group number of 2 and pointwise convolution with a group number of 1 to perform channel compression and channel information mixing on the dimensionality-reduced feature and the aggregated feature.
[0056] In one embodiment, a multi-instance image classification method based on pre-check feature fusion may further include arranging the multi-instance features in the instance order, converting the sequence number into a binary number to obtain a sequence encoding, adding the sequence encoding to the corresponding instance feature and mapping it into a one-dimensional sequence to obtain a multi-instance sequence feature. Input the multi-instance sequence feature into a 2D state space model and a multi-layer perceptron to predict the weights w1 to w n for each instance, and multiply each instance by the corresponding weight w i and obtain an instance fusion view through a one-dimensional convolutional layer.
[0057] In one embodiment, a multi-instance image classification method based on pre-check feature fusion may further include mapping each instance of the pre-check feature to the same spatial plane and performing horizontal splicing to obtain a spatial splicing feature map with a height of h and a width of n×w, realizing instance rearrangement in the spatial dimension. Use a horizontally and vertically asymmetric convolutional block on the spatial splicing feature map to obtain a spatial fusion view.
[0058] In one embodiment, a multi-instance image classification method based on pre-check feature fusion may further include performing global average pooling and multi-layer perceptron prediction on the instance fusion view and the spatial fusion view in a single-view manner and a dual-view manner to obtain an instance fusion view prediction value, a spatial fusion view prediction value, and a dual-view prediction value; concatenating the instance fusion view prediction value, the spatial fusion view prediction value, and the dual-view prediction value into a one-dimensional vector, and using a multi-layer perceptron to implement feature dimensionality reduction of the one-dimensional vector to achieve multi-view joint prediction.
[0059] In one embodiment, as Figure 4 shown, a multi-instance image classification method based on pre-check feature fusion may further include a 2D state space model. The 2D state space model uses a bidirectional scanning mechanism in the channel dimension to achieve cross-instance feature mixing and a bidirectional scanning mechanism in the spatial dimension to achieve cross-spatial location feature mixing.
[0060] In one embodiment, a multi-instance image classification method based on pre-check feature fusion may further include a horizontal and vertical asymmetric convolution block. The horizontal and vertical asymmetric convolution block first performs ordinary convolution on the features, and then performs asymmetric horizontal convolution with a length of n×d and asymmetric vertical convolution with a length of d, where n is the number of instances and d is the convolution kernel length parameter.
[0061] In summary, using the multi-instance image classification method based on pre-check feature fusion provided by this technical solution can reduce the image size and computation amount in a multi-instance scenario, improve the effectiveness of multi-instance image information, so as to solve the problems of information sparsity and high computational requirements in the multi-instance image learning task, and improve the accuracy and speed of multi-instance image classification.
[0062] The above-disclosed is only a preferred embodiment of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand the entire or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A multi-instance image classification method based on pre-inspection feature fusion, characterized in that: The steps include: Multi-instance imaging is obtained through a multi-instance image acquisition tool; Inputting the multiple instance images into a pre-inspection module for region pre-inspection to obtain a pre-inspection region for each instance; Extracting single instance features from the pre-check area using a pre-check feature encoder to obtain pre-check features; The pre-check features are weighted and fused by an instance fusion module to obtain an instance fusion view; The pre-check feature is compressed by a spatial fusion module to obtain a spatial fusion view; The instance fusion view and the spatial fusion view are jointly predicted for multi-instance classification to obtain a target category.
2. The multi-instance image classification method based on pre-inspection feature fusion as claimed in claim 1, characterized in that: The specific contents of the step "obtaining multi-instance imaging by a multi-instance image acquisition tool" are: For the identification target, multiple instances of images are collected in time and space dimensions using multiple imaging and multi-angle imaging methods; Representative images are screened for the multiple instance images, images with different imaging times and different imaging angles are selected, and similar images with adjacent imaging times and the same imaging angles are deleted to improve data representativeness.
3. The multi-instance image classification method based on pre-check feature fusion as claimed in claim 1, characterized in that: The specific content of the step "inputting the multi-instance imaging into the pre-detection module for regional pre-detection" is: The aspect ratio, width and height of the prior frame and the center offset distance are set for the multi-instance imaging. The center offset distance is calculated by forming a matrix containing only two elements with the horizontal and vertical coordinate differences of the geometric center point of the image and the geometric center point of the prior frame, dividing the horizontal coordinate difference in the matrix by the image width, dividing the vertical coordinate difference in the matrix by the image height, and then calculating the Frobenius norm of the matrix to obtain the center offset distance; The pre-inspection module is used to perform region detection using a multi-stage detection model HTC, and a predicted region with the largest difference in confidence and center offset distance is selected as a pre-inspection region.
4. The multi-instance image classification method based on pre-check feature fusion as claimed in claim 1, characterized in that: In the step of "using a pre-check feature encoder to extract a single instance feature from the pre-check area to obtain a pre-check feature", the pre-check feature encoder includes a single instance feature extractor and a single instance feature mixer; The single instance feature extractor includes a feature dimension reduction module and a feature aggregation module. The feature dimension reduction module is composed of a convolution layer with a stride of 2, a random pooling layer and a mixed pooling layer in parallel to achieve feature dimension reduction. The feature aggregation module is composed of a linear mapping layer, a state space model and a dot product attention layer in a serial manner to achieve feature aggregation. The single instance feature mixer is used to splice the reduced dimension features and the aggregated features in the channel dimension, and use grouped convolution with a group number of 2 and point-by-point convolution with a group number of 1 to achieve channel compression and channel information mixing for the reduced dimension features and the aggregated features.
5. The multi-instance image classification method based on pre-check feature fusion as claimed in claim 1, characterized in that: The specific content of the step "using the instance fusion module to perform multi-instance feature weighting and fusion on the pre-check features to obtain an instance fusion view" is: Arranging the multiple instance features in the instance order, converting the sequence number into a binary number to obtain a sequence code, accumulating the sequence code and the corresponding instance feature and mapping them into a one-dimensional sequence to obtain a multiple instance sequence feature; The multi-instance sequence features are input into the 2D state space model and the multi-layer perceptron to predict the weights w1 to w corresponding to each instance. n , associate each instance with the corresponding weight w i Multiply and pass through a one-dimensional convolutional layer to get the instance fusion view.
6. The multi-instance image classification method based on pre-check feature fusion as claimed in claim 5, characterized in that: The 2D state space model uses a bidirectional scanning mechanism in the channel dimension to achieve feature mixing across instances, and uses a bidirectional scanning mechanism in the spatial dimension to achieve feature mixing across spatial positions.
7. The multi-instance image classification method based on pre-check feature fusion as claimed in claim 1, characterized in that: The specific content of the step "using the spatial fusion module to compress the pre-check features to obtain a spatial fusion view" is: Mapping each instance of the pre-check feature onto the same spatial plane and performing horizontal splicing to obtain a spatial splicing feature map with a height of h and a width of n×w, thereby achieving instance rearrangement in the spatial dimension; The spatial splicing feature map is subjected to horizontal and vertical asymmetric convolution blocks to obtain a spatial fusion view.
8. The multi-instance image classification method based on pre-check feature fusion as claimed in claim 7, characterized in that: The specific content of the horizontal and vertical asymmetric convolution block is: first perform ordinary convolution on the features, then perform asymmetric horizontal convolution of length n×d and asymmetric vertical convolution of length d, where n is the number of instances and d is the convolution kernel length parameter.
9. The multi-instance image classification method based on pre-check feature fusion as claimed in claim 1, characterized in that: The specific content of the step "jointly predicting the instance fusion view and the spatial fusion view for multi-instance classification" is: Performing global average pooling and multi-layer perceptron prediction on the instance fusion view and the spatial fusion view in a single-view mode and a dual-view mode to obtain an instance fusion view prediction value, a spatial fusion view prediction value and a dual-view prediction value; The instance fusion view prediction value, the spatial fusion view prediction value and the dual-view prediction value are concatenated into a one-dimensional vector, and a multi-layer perceptron is used to reduce the feature dimension of the one-dimensional vector to achieve multi-view joint prediction.
Citation Information
Patent Citations
Sonar image target identification method based on instance segmentation
CN110084234A
Target detection system and acquisition method
CN114220126A
Image instance segmentation method based on deep learning
CN115131556A
Detection method using fusion network based on attention mechanism, and terminal device
US11222217B1