Weak supervision target detection method and device and medium

The method improves weakly-supervised object detection by using a modified VGG16 network with region pooling and a prototype relation network to enhance feature extraction and pseudo-label generation, addressing local focus and missed detection issues, thereby increasing detection accuracy and robustness.

CN120318491APending Publication Date: 2025-07-15HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510387268.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing weakly supervised object detection methods have low accuracy in object detection, especially in the problems of local focus and missed detection, and the quality of the pseudo-label is limited, which affects the accuracy and robustness of the model.

Method used

The improved VGG16 network, region pooling layer and full connection layer are used to construct a feature vector extraction model, combined with multi-branch learning module, prototype relational network module and classification task execution module, region segmentation and merging are performed through selective search algorithms, and pseudo-labels are optimized using non-maximum suppression method to improve the accuracy of the detection model.

Benefits of technology

Through collaborative optimization, high-quality pseudo-labels are generated to alleviate missed detection problems and accurately adjust the candidate box position, which significantly improves the accuracy and comprehensiveness of target detection under weak supervision conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318491A_ABST
    Figure CN120318491A_ABST
Patent Text Reader

Abstract

The invention discloses a weak supervision target detection method and device and a medium, and relates to the technical field of target detection, and the method comprises the steps: obtaining a to-be-detected picture and a to-be-detected target category set; preprocessing the to-be-tested picture to obtain a preprocessed to-be-tested picture; performing region segmentation and merging on the preprocessed to-be-detected picture by adopting a selective search algorithm to obtain a to-be-classified region set; inputting the preprocessed to-be-detected picture and the to-be-classified region set into a feature vector extraction model to obtain a feature vector of each to-be-classified region; and inputting the feature vector of each to-be-classified region and the to-be-detected target category set into a classification positioning module to obtain one or more to-be-positioned regions of each to-be-detected category of the to-be-detected picture and final absolute coordinates of each to-be-positioned region on each to-be-detected category. According to the invention, the accuracy of target detection during weak target supervision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of object detection, and in particular, to a weakly-supervised object detection method, device, and medium. Background Art

[0002] As one of the core tasks in the field of computer vision, object detection aims to achieve accurate localization and classification of target objects in images, and is a key technology supporting advanced applications such as robot vision, face recognition, image retrieval, and autonomous driving. In recent years, with the breakthrough progress of convolutional neural network technology and the successive release of large-scale datasets, under the condition of fully supervised learning with sufficient training data, the performance of object detection models has approached the theoretical limit.

[0003] However, with the continuous expansion of the application scenarios of object detection technology, the requirements for the form of image data and annotation specifications in different scenarios are becoming increasingly different, resulting in a high technical barrier between datasets in various fields. The construction of traditional object detection datasets requires annotators to accurately label the bounding boxes of all target instances in the images. With the increase in the number of categories and instances in the dataset, the time cost and economic investment of this dense annotation method increase exponentially, which has become the main bottleneck restricting the development of object detection technology.

[0004] To address this challenge, the weakly-supervised object detection (WSOD) technology has emerged. This technology trains the model by using image-level classification labels instead of instance-level bounding box annotations, significantly reducing the data annotation cost. The current mainstream weakly-supervised object detection methods generally adopt the multiple instance learning (MIL) framework. In this framework, each image is regarded as a "bag", and the candidate instance regions in the image are regarded as "instances". As long as there is at least one positive sample instance in the "bag", the "bag" is labeled as a positive sample; only when all instances are negative samples, the "bag" is labeled as a negative sample. This learning mechanism enables the model to infer instance-level information through image-level labels.

[0005] However, due to the lack of accurate position annotations under weakly-supervised conditions and the relatively loose constraints of multiple instance learning, the existing methods have two significant defects: First, the local focus problem, that is, the model tends to focus on the prominent regions of objects (such as faces, animal heads, etc.) and cannot accurately locate the complete range of objects; second, the missed detection problem, when there are multiple targets of the same category in the image, the model is difficult to achieve comprehensive detection.

[0006] In response to the above problems, general solutions mainly screen positive samples through multi-instance learning, select high-quality samples as pseudo-labels in the candidate instance regions, and gradually improve the detection performance by combining iterative optimization algorithms. However, in the process of sample screening, these methods mostly adopt simple linear discrimination strategies, failing to fully consider the subtle differences and their mutual relationships among objects of the same class, resulting in limited quality of pseudo-labels, which in turn affects the accuracy and robustness of the detection model. Summary of the Invention

[0007] The objective of the present application is to provide a weakly supervised object detection method, device, and medium to solve the problem of low accuracy in object detection in weakly supervised scenarios.

[0008] To achieve the above objective, the present application provides the following solutions:

[0009] In a first aspect, the present application provides a weakly supervised object detection method, including:

[0010] Obtain a to-be-detected image and a set of to-be-detected target categories; the set of to-be-detected target categories includes multiple to-be-detected categories, and the to-be-detected category is the category of the to-be-detected target included in the to-be-detected image;

[0011] Preprocess the to-be-detected image to obtain a preprocessed to-be-detected image;

[0012] Use the selective search algorithm to perform region segmentation and merging on the preprocessed to-be-detected image to obtain a set of regions to be classified; the set of regions to be classified includes multiple regions to be classified and the original absolute coordinates of each region to be classified, and the region to be classified is a candidate instance region of the to-be-detected image;

[0013] Input the preprocessed to-be-detected image and the set of regions to be classified into a feature vector extraction model to obtain the feature vectors of each region to be classified; the feature vector extraction model is constructed based on an improved VGG16 network, a region pooling layer, and a fully connected layer;

[0014] Input the feature vectors of each region to be classified and the set of to-be-detected target categories into a classification and localization module to obtain one or more regions to be located for each to-be-detected category of the to-be-detected image and the final absolute coordinates of each region to be located in each to-be-detected category; the classification and localization module includes a multi-branch learning module, a classification task execution module, and a localization module in an object detection model, the object detection model is obtained by weakly supervised training of an object detection network, and the object detection network includes: a multi-instance learning detection module, a multi-branch learning module, a prototype relationship network module, a classification task execution module, and a localization module; the object detection network is constructed based on a fully connected layer.

[0015] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the weakly supervised object detection method described in any one of the above.

[0016] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the weakly supervised object detection method described in any one of the above is implemented.

[0017] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:

[0018] The present application discloses a weakly supervised object detection method, device, and medium. First, an image to be tested and a set of target categories to be tested are obtained; the set of target categories to be tested includes multiple categories to be tested, and the category to be tested is the category of the target to be tested included in the image to be tested; then, the image to be tested is preprocessed to obtain the preprocessed image to be tested; secondly, the selective search algorithm is used to perform region segmentation and merging on the preprocessed image to be tested to obtain a set of regions to be classified; the set of regions to be classified includes multiple regions to be classified and the original absolute coordinates of each region to be classified, and the region to be classified is a candidate instance region of the image to be tested; subsequently, the preprocessed image to be tested and the set of regions to be classified are input into a feature vector extraction model to obtain the feature vectors of each region to be classified; the feature vector extraction model is constructed based on an improved VGG16 network, a regional pooling layer, and a fully connected layer; finally, the feature vectors of each region to be classified and the set of target categories to be tested are input into a classification and localization module to obtain one or more regions to be localized for each category to be tested of the image to be tested and the final absolute coordinates of each region to be localized on each category to be tested; the classification and localization module includes a multi-branch learning module, a classification task execution module, and a localization module in an object detection model, and the object detection model is obtained by weakly supervised training of an object detection network, and the object detection network includes: a multi-instance learning detection module, a multi-branch learning module, a prototype relationship network module, a classification task execution module, and a localization module; the object detection network is constructed based on a fully connected layer. The present application uses an object detection model obtained by weakly supervised training of an object detection network to classify and localize the regions to be localized for each category to be tested of the image to be tested, improving the accuracy of object detection in weakly supervised object detection. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 Schematic diagram of the weak supervision object detection method provided in an embodiment of the present application;

[0021] Figure 2 Schematic diagram of the feature vector extraction model structure;

[0022] Figure 3 Schematic diagram of the object detection model training architecture;

[0023] Figure 4 Schematic diagram of the prototype relationship network structure;

[0024] Figure 5 Schematic diagram of the structure of a computer device provided in an embodiment of the present application. Detailed implementation manners

[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0026] The purpose of the present application is to provide a weak supervision object detection method, device and medium, aiming to improve the accuracy of object detection in weak supervision.

[0027] To make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0028] In an exemplary embodiment, as Figure 1 shown, a weak supervision object detection method is provided, including:

[0029] Step 1: Obtain the picture to be tested and the set of target categories to be tested.

[0030] Among them, the set of target categories to be tested includes multiple categories to be tested, and the category to be tested is the category of the target to be tested included in the picture to be tested.

[0031] Step 2: Preprocess the picture to be tested to obtain the preprocessed picture to be tested.

[0032] As an optional implementation manner, step 2 includes:

[0033] Step 21: Adjust the size of the picture to be tested according to a preset scale to obtain the picture to be tested with the size adjusted.

[0034] Specifically, randomly adjust the shortest side of the image to be tested to one of five preset scales, namely {480, 576, 688, 864, 1200} pixels, while restricting the longest side not to exceed 2000 pixels to ensure the rationality of the image size and avoid waste of computing resources caused by overly large images.

[0035] Step 22: Perform a random horizontal flipping operation on the image to be tested with adjusted size to obtain the flipped image to be tested.

[0036] Step 23: Perform standardization processing on the flipped image to be tested to obtain the preprocessed image to be tested.

[0037] Specifically, calculate the mean value for each of the RGB three channels of the image, and after converting the image into an array, subtract the corresponding mean value from each channel to achieve data standardization processing.

[0038] Step 3: Use the selective search algorithm to perform region segmentation and merging on the preprocessed image to be tested to obtain a set of regions to be classified.

[0039] Among them, the set of regions to be classified includes multiple regions to be classified and the original absolute coordinates of each region to be classified, and the region to be classified is a candidate instance region of the image to be tested.

[0040] Step 4: Input the preprocessed image to be tested and the set of regions to be classified into the feature vector extraction model to obtain the feature vectors of each region to be classified.

[0041] Among them, the feature vector extraction model is constructed based on the improved VGG16 network, the regional pooling layer, and the fully connected layer.

[0042] As an optional implementation manner, Step 4 includes:

[0043] Step 41: Input the preprocessed image to be tested into the backbone feature extraction module in the feature vector extraction model to obtain the image-level feature map of the image to be tested; the backbone feature extraction module is obtained by training the improved VGG16 network.

[0044] As an optional implementation manner, as Figure 2 shown, the improved VGG16 network includes: a first convolutional module (conv1 layer), a second convolutional module (conv2 layer), a third convolutional module (conv3 layer), a fourth convolutional module (conv4 layer), and a fifth convolutional module (conv5 layer) connected in sequence.

[0045] Both the first convolutional module and the second convolutional module include a 3×3 convolutional layer, a ReLU activation function, a 3×3 convolutional layer, a ReLU activation function, and a 2×2 max pooling layer connected in sequence.

[0046] Specifically, the first convolution module is responsible for extracting low-level features from the input image. The input is a 3-channel image. First, it undergoes a convolution operation through a 3×3 convolution layer, outputting a feature map with 64 channels. Then, a ReLU activation function is used for non-linear transformation. After that, another convolution operation through a 3×3 convolution layer is performed, and the output is still a feature map with 64 channels. Finally, through a 2×2 max pooling layer, the spatial size of the feature map is reduced to half of the original.

[0047] The second convolution module further extracts features based on the first convolution module. The input is a feature map with 64 channels. First, it undergoes a convolution operation through a 3×3 convolution layer, outputting a feature map with 128 channels. Then, a ReLU activation function is used for non-linear transformation. After that, another convolution operation through a 3×3 convolution layer is performed, and the output is still a feature map with 128 channels. Finally, through a 2×2 max pooling layer, the spatial size of the feature map is reduced to half of the original again.

[0048] The third convolution module includes a 3×3 convolution layer, a ReLU activation function, a 3×3 convolution layer, a ReLU activation function, a 3×3 convolution layer, a ReLU activation function, and a 2×2 max pooling layer connected in sequence.

[0049] Specifically, the third convolution module is responsible for extracting higher-level features. The input is a feature map with 128 channels. First, it undergoes a convolution operation through a 3×3 convolution layer, outputting a feature map with 256 channels. Then, a ReLU activation function is used for non-linear transformation. After that, two convolution operations through 3×3 convolution layers are performed, and the output of each is a feature map with 256 channels. Finally, through a 2×2 max pooling layer, the spatial size of the feature map is reduced to half of the original.

[0050] The fourth convolution module includes a 3×3 convolution layer, a ReLU activation function, a 3×3 convolution layer, a ReLU activation function, a 3×3 convolution layer, and a ReLU activation function connected in sequence.

[0051] Specifically, the fourth convolution module further extracts features based on the third convolution module. The input is a feature map with 256 channels. First, it undergoes a convolution operation through a 3×3 convolution layer, outputting a feature map with 512 channels. Then, a ReLU activation function is used for non-linear transformation. After that, two convolution operations through 3×3 convolution layers are performed, and the output of each is a feature map with 512 channels. Different from the previous layers, the fourth convolution module does not use a pooling operation, so the spatial size of the feature map remains unchanged.

[0052] The fifth convolutional module includes a 3×3 dilated convolutional layer, a ReLU activation function, a 3×3 dilated convolutional layer, a ReLU activation function, a 3×3 dilated convolutional layer, and a ReLU activation function that are connected in sequence.

[0053] Specifically, the fifth convolutional module is the last part of the improved VGG16 network. The input is a feature map with 512 channels. First, a 3×3 dilated convolutional layer performs a convolution operation with a dilation rate of 2 and a padding of 2, and the output is still a feature map with 512 channels. Then, a ReLU activation function is used for non-linear transformation, and then two 3×3 dilated convolutional layers perform convolution operations. Each convolution uses a dilated convolution with a dilation rate of 2, and the output is still a feature map with 512 channels. The use of dilated convolution enables the network to expand the receptive field and capture more extensive context information while keeping the spatial size of the feature map unchanged. The number of channels of the output feature map (i.e., the image-level feature map) of the entire network is 512, and the spatial size is 1 / 8 of the input image.

[0054] Step 42: Input the image-level feature maps of each region to be classified and the test image into the region pooling module (i.e., Figure 2 the RoI Pooling layer in

[0055] the feature vector extraction model) to obtain the feature maps of multiple sub-regions of each region to be classified; the region pooling module is obtained by training the region pooling layer. Specifically, using the region pooling module, the features of each candidate instance region (region to be classified or first sample region) are extracted from the image-level feature map output by the fifth convolutional module according to the positional correspondence. The channel dimension of the feature map extracted from each candidate instance region is 512, but the width and height are different. Then, the width and height of the feature map of each candidate instance region are divided into 7×7 sub-regions, and a max pooling operation is performed within each sub-region, and the output is a feature map with a fixed size of 7×7×512.

[0056] Step 43: Input the feature maps of all sub-regions of each region to be classified into the feature vector determination module in the feature vector extraction model to obtain the feature vectors of each region to be classified; the feature vector determination module is obtained by training two fully connected layers (i.e., Figure 2 fully connected layer 1 and fully connected layer 2 in

[0057] Specifically, each candidate instance region is further processed by two fully connected layers, and finally a 4096-dimensional feature vector of each candidate instance region is obtained.

[0058] Step 5: Input the feature vectors of each region to be classified and the set of target categories to be measured into the classification and localization module, and obtain one or more regions to be localized for each target category to be measured in the image to be measured and the final absolute coordinates of each region to be localized for each target category to be measured.

[0059] Among them, the classification and localization module includes a multi-branch learning module, a classification task execution module, and a localization module in the object detection model. The object detection model is obtained by weakly supervised training of an object detection network. The object detection network includes: a multi-instance learning detection module, a multi-branch learning module, a prototype relationship network module, a classification task execution module, and a localization module; the object detection network is constructed based on a fully connected layer.

[0060] As an optional implementation manner, Step 5 includes:

[0061] Step 511: Input the feature vectors of each region to be classified and the set of target categories to be measured into the multi-branch learning module, and obtain the k-th learning sub-branch score matrix of the image to be measured, where k is 1, 2, or 3; any element in any current row and any current column of the k-th learning sub-branch score matrix of the image to be measured represents the score of the k-th learning sub-branch indicating that the target in the region to be classified corresponding to the current column belongs to the background or the target category corresponding to the current row. The k-th learning sub-branch score is the score output by the k-th learning sub-branch in the multi-branch learning module; each learning sub-branch in the multi-branch learning module is constructed based on a fully connected layer.

[0062] Step 512: Input the feature vectors of each region to be classified and the set of target categories to be measured into the classification task execution module, and obtain the classification task score matrix of the image to be measured; any element in any current row and any current column of the classification task score matrix of the image to be measured represents the score of the classification task indicating that the target in the region to be classified corresponding to the current column belongs to the background or the target category corresponding to the current row. The classification task score is the score output by the classification task execution module; the classification task execution module is constructed based on a fully connected layer.

[0063] Step 513: Perform pixel-by-pixel averaging on each k-th learning sub-branch score matrix and the classification task score matrix of the image to be measured, and obtain the final score matrix of the image to be measured; any element in any current row and any current column of the final score matrix of the image to be measured represents the final score of the classification task indicating that the target in the region to be classified corresponding to the current column belongs to the background or the target category corresponding to the current row.

[0064] Step 514: Use the non-maximum suppression method to process all regions to be classified in each row of the final score matrix of the image to be measured, and obtain one or more regions to be localized for each target category corresponding to each row.

[0065] Step 515: Using the positioning module, respectively determine the predicted offset coordinates of each to-be-positioned region of the to-be-tested image in each to-be-tested category based on the feature vectors of each to-be-positioned region of each to-be-tested category of the to-be-tested image, and determine the final absolute coordinates of each to-be-positioned region of the to-be-tested image in each to-be-tested category based on the predicted offset coordinates and the original final absolute coordinates of each to-be-positioned region of the to-be-tested image in each to-be-tested category; the positioning module is constructed based on a fully connected layer.

[0066] As an alternative implementation, as Figure 3 shown, in Step 5, the determination process of the classification and positioning module includes:

[0067] Step 521: Obtain multiple sample images and the set of image-level labels for each sample image; the set of image-level labels includes the image-level labels of multiple sample categories, and the image-level label is the inclusion situation of the target of the sample image for the sample category; the inclusion situation is inclusion or non-inclusion.

[0068] Specifically, the set of image-level labels of any sample image is expressed as Y = [y1, y2, …, y C , where C represents the total number of sample categories containing the target in all sample images, y c = 1 indicates that the sample image contains the c-th sample category, and y c = 0 indicates that the sample image does not contain the c-th sample category, and c = 1, 2, …, C.

[0069] Step 522: Respectively determine the feature vectors of multiple first sample regions of each sample image.

[0070] As an alternative implementation, Step 522 includes:

[0071] Step 5221: Respectively preprocess each sample image to obtain multiple preprocessed sample images.

[0072] Specifically, perform a random horizontal flipping operation on the sample image to increase data diversity, thereby improving the adaptability of the target detection model to different perspectives and directions. Calculate the mean value for each of the RGB three channels of the sample image, and after converting the sample image into an array, subtract the corresponding mean value from each channel to achieve data normalization processing, thereby enhancing the training stability and convergence speed of the target detection model.

[0073] Step 5222: Adopt the selective search algorithm to respectively perform region segmentation and merging on each preprocessed sample image to obtain the set of first sample regions of each preprocessed sample image; the set of first sample regions includes multiple first sample regions and the original absolute coordinates of each first sample region, and the first sample region is the candidate instance region of the sample image.

[0074] Specifically, the first sample region set of any sample picture is denoted as R.

[0075] Step 5223: Input each preprocessed sample picture and the corresponding first sample region set into the feature vector extraction model respectively to obtain the feature vectors of each first sample region.

[0076] Step 523: Construct an object detection network based on multiple fully connected layers.

[0077] Step 524: Use the feature vectors of multiple first sample regions of each sample picture and the picture-level labels to perform multiple weak supervision trainings on the object detection network to obtain an object detection model.

[0078] As an alternative implementation, during the process of step 524, the training process at any current training times includes:

[0079] Step 52401: Determine any sample picture as the current sample picture.

[0080] Step 52402: Input the feature vectors of each first sample region of the current sample picture into the classification branch and the detection branch of the multi-instance learning detection module in the object detection network at the current training times respectively to obtain the classification branch score matrix and the detection branch score matrix Both the classification branch and the detection branch are fully connected layers with an output of 20 dimensions; the element at any current row and any current column in the classification branch score matrix represents the classification branch score that the object in the first sample region corresponding to the current column belongs to the sample category corresponding to the current row, and the classification branch score is the score output by the classification branch; the element at any current row and any current column in the detection branch score matrix represents the detection branch score of the first sample region corresponding to the current column for the object of the sample category corresponding to the current row, and the detection branch score is the score output by the detection branch.

[0081] Specifically, It represents a matrix of C rows and |R| columns; |R| represents the number of first sample regions in the first sample region set.

[0082] Step 52403: Input the classification branch score matrix X cls and the detection branch score matrix X det into the softmax function of the multi-instance learning detection module in the object detection network at the corresponding current training times respectively to obtain the normalized classification branch score matrix θ cls (X cls ) and the normalized detection branch score matrix θdet (X det )。

[0083] Specifically, the classification branch score matrix and the detection branch score matrix respectively correspond to a softmax function.

[0084] Step 52404: Perform an element-wise product calculation on the normalized classification branch score matrix θ cls (X cls ) and the normalized detection branch score matrix θ det (X det ) of the current sample image at the current training iteration to obtain the classification and detection comprehensive score matrix X of the current sample image at the current training iteration.

[0085] Specifically, the calculation formula for the classification and detection comprehensive score matrix X is:

[0086] X = θ cls (X cls ) ⊙ θ det (X det )。

[0087] Wherein, ⊙ represents an element-wise product.

[0088] Step 52405: Accumulate all the elements in each row of the classification and detection comprehensive score matrix X of the current sample image at the current training iteration to obtain the image category scores of each sample category of the current sample image at the current training iteration.

[0089] Specifically, the calculation formula for the image category score is:

[0090]

[0091] Wherein, φ c represents the image category score of the c-th sample category; X c,r represents the element in the c-th row and r-th column of the classification and detection comprehensive score matrix X, that is, the classification and detection comprehensive score of the first sample region corresponding to the r-th column for the sample category corresponding to the c-th row.

[0092] Step 52406: Determine the multi-label classification loss at the current training iteration based on the image category scores of each sample category of all sample images at the current training iteration and the image-level label set.

[0093] Specifically, the calculation formula for the multi-label classification loss is:

[0094]

[0095] Wherein, represents the multi-label classification loss; N represents the total number of sample images; denote L of the nth sample image midn ; L midn represents the multi-label classification sub-loss.

[0096] Step 52407: Input the feature vectors of each first sample region of the current sample image and each sample category into the multi-branch learning module in the object detection network at the current training iteration, to obtain the kth learning sub-branch score matrix of the current sample image at the current training iteration The element at any current row and any current column in the kth learning sub-branch score matrix of the current sample image represents the score of the object in the first sample region corresponding to the current column belonging to the background or the kth learning sub-branch corresponding to the sample category of the current row; each kth learning sub-branch in the multi-branch learning module is a fully connected layer with an output of 21 dimensions.

[0097] Specifically, represents a matrix of C + 1 rows and |R| columns.

[0098] Step 52408: Determine the first sample region with the highest k - 1 learning sub-branch score in any current row of the (k - 1)th learning sub-branch score matrix X of the current sample image at the current training iteration as the reference region of the sample category corresponding to the current row in the kth learning sub-branch of the current sample image at the current training iteration k-1 where, when k = 1, the kth learning sub-branch score matrix is the classification detection comprehensive score matrix.

[0099] Specifically, represents the reference region of the cth sample category in the kth learning sub-branch.

[0100] Step 52409: Calculate the intersection over union (IoU) between each first sample region of the current sample image and the reference region of the sample category corresponding to the current row in the kth learning sub-branch of the current sample image at the current training iteration respectively, and based on each IoU and the first preset IoU threshold determine the initial instance-level pseudo-labels of each first sample region for each sample category in the kth learning sub-branch of the current sample image at the current training iteration, and filter each first sample region based on the initial instance-level pseudo-labels to obtain the second sample region set of each sample category in the kth learning sub-branch of the current sample image at the current training iteration; the second sample region set includes multiple second sample regions.

[0101] Specifically, if the bounding box of the r1th first sample region in the first sample region set R intersects with and the IoU is greater than Then, label the r1-th first sample region as the c-th sample category, that is, the initial instance-level pseudo-label of the r1-th first sample region for the c-th sample category Otherwise, label the r1-th first sample region as the background category, that is, the initial instance-level pseudo-label of this first sample region for the c-th sample category

[0102] The calculation formula for the second sample region set is as follows:

[0103]

[0104] Wherein, represents the second sample region set of the c-th sample category in the k-th learning sub-branch; r r2 represents the r2-th second sample region in; represents the initial instance-level pseudo-label of the r2-th second sample region for the c-th sample category.

[0105] Step 52410: Average the feature vectors of all second sample regions corresponding to each sample category in the k-th learning sub-branch of the current sample picture at the current training times to obtain the prototype features of each sample category in the k-th learning sub-branch of the current sample picture at the current training times.

[0106] Specifically, the calculation formula for the prototype features is as follows:

[0107]

[0108] Wherein, represents the prototype feature of the c-th sample category in the k-th learning sub-branch; represents the second sample region set the number of second sample regions in; η(r r2 ) is the feature vector of the r2-th second sample region r in the second sample region set r2 .

[0109] Step 52411: Sort all the k-th learning sub-branch scores in any current row in the k-th learning sub-branch score matrix of the current sample picture at the current training times in descending order, and determine the third sample region set of the sample category corresponding to the current row in the k-th learning sub-branch of the current sample picture at the current training times based on the first sample regions corresponding to the k-th learning sub-branch scores ranked in the top preset ratio α The third sample region set includes multiple third sample regions.

[0110] Specifically, Represents the set of third sample regions corresponding to the samples in the c-th row of the k-th learning sub-branch.

[0111] Step 52412: Concatenate the prototype features of each sample category in the k-th learning sub-branch of the current sample image at the current training iteration with the feature vectors of each third sample region to obtain the concatenated features of each sample category and each third sample region in the k-th learning sub-branch of the current sample image at the current training iteration.

[0112] Specifically, the calculation formula for the concatenated features is:

[0113]

[0114] Where, Represents the concatenated feature of the c-th sample category and the r3-th third sample region in the set of third sample regions in the k-th learning sub-branch; concat(·,·) represents the concatenation operation; η(r r3 ) represents the r3-th third sample region r in the set of third sample regions r3 's feature vector; Represents the number of third sample regions in the set of third sample regions .

[0115] Step 52413: Input the concatenated features of each sample category and each third sample region in the k-th learning sub-branch of the current sample image at the current training iteration into the k-th relationship sub-branch of the prototype relationship network module in the object detection network at the current training iteration to obtain the similarity scores between the feature vectors of each third sample region and the prototype features of each sample category in the k-th relationship sub-branch of the current sample image at the current training iteration; each k-th relationship sub-branch of the prototype relationship network module includes a fully connected layer with an output of 256 dimensions and a fully connected layer with an output of 1 dimension. The structure of the prototype relationship network module is as Figure 4 shown.

[0116] Specifically, the calculation formula for the similarity scores is:

[0117]

[0118] Where, Represents the similarity score between the feature vector of the r3-th third sample region and the prototype feature of the c-th sample category in the k-th relationship sub-branch; σ(·) represents the Sigmoid activation function; Represents the k-th relationship sub-branch.

[0119] Step 52414: Based on the similarity scores being greater than the preset similarity score threshold τscore For each third sample region, determine the set of fourth sample regions of each sample category in the k-th relational sub-branch of the current sample image at the current training iteration; the set of fourth sample regions includes multiple fourth sample regions.

[0120] Specifically, the calculation formula for the set of fourth sample regions is:

[0121]

[0122] Where, represents the set of fourth sample regions of the c-th sample category in the k-th relational sub-branch; r r4 represents the r4-th fourth sample region in represents the similarity score between the feature vector of the r4-th fourth sample region of the c-th sample category in the k-th relational sub-branch and the prototype feature.

[0123] Step 52415: Use the non-maximum suppression method to process the set of fourth sample regions of each sample category in the k-th relational sub-branch of the current sample image at the current training iteration, to obtain the set of fifth sample regions of each sample category in the k-th relational sub-branch of the current sample image at the current training iteration The set of fifth sample regions includes multiple fifth sample regions.

[0124] Specifically, represents the set of fifth sample regions of the c-th sample category in the k-th relational sub-branch.

[0125] Step 52416: Calculate the intersection over union (IoU) between each first sample region of the current sample image at the current training iteration and the fifth sample regions in the set of fifth sample regions of each sample category in the k-th relational sub-branch of the current sample image at the current training iteration, and based on each IoU and the second preset IoU threshold determine the optimized instance-level pseudo-labels of each first sample region of the current sample image at the current training iteration for each sample category or the background in the k-th relational sub-branch.

[0126] Specifically, if the IoU between the bounding box of the r1-th first sample region in the set of first sample regions R and the fifth sample region is greater than then label the r1-th first sample region as the c-th sample category, that is, the optimized instance-level pseudo-label of the r1-th first sample region for the c-th sample category Otherwise, label the r1-th first sample region as the background category, that is, the optimized instance-level pseudo-label of this first sample region for the c-th sample category

[0127] Step 52417: Determine the loss of the k-th learning sub-branch at the current training iteration based on the optimized instance-level pseudo-labels of each first sample region in the k-th relationship sub-branch for each sample category or background and the k-th learning sub-branch score matrix of all sample images at the current training iteration.

[0128] Specifically, the calculation formula for the loss of the k-th learning sub-branch is:

[0129]

[0130] Where, represents the loss of the k-th learning sub-branch; represents the of the n-th sample image, representing the sub-loss of the k-th learning sub-branch; represents the loss function weight of the k-th learning sub-branch, defined as the highest score among all the scores of the k - 1-th learning sub-branch corresponding to the c-th row in the k - 1-th learning sub-branch score matrix; represents the element in the c-th row and r1-th column of the k-th learning sub-branch score matrix, that is, the score of the k-th learning sub-branch of the first sample region corresponding to the r1-th column for the sample category corresponding to the c-th row.

[0131] Step 52418: Determine the loss of the k-th relationship sub-branch at the current training iteration based on the similarity scores between the feature vectors of each third sample region of each sample category in the k-th relationship sub-branch of all sample images at the current training iteration and the prototype features, and the initial instance-level pseudo-labels of each third sample region for each sample category in the k-th learning sub-branch of all sample images at the current training iteration Determine the loss of the k-th relationship sub-branch at the current training iteration.

[0132] Specifically, the calculation formula for the loss of the k-th relationship sub-branch is:

[0133]

[0134]

[0135] Where, represents the loss of the k-th relationship sub-branch; represents the of the n-th sample image, representing the sub-loss of the k-th relationship sub-branch.

[0136] Step 52419: Input the feature vectors of each first sample region of the current sample image and each sample category into the classification task execution module in the object detection network at the current training iteration to obtain the classification task score matrix X of the current sample image at the current training iteration 4; In the classification task score matrix of the current sample image, the element in any current row and any current column represents the classification task score indicating whether the target in the first sample region corresponding to the current column of the current sample image belongs to the background or the sample category corresponding to the current row; the classification task execution module is a fully connected layer with an output of 21 dimensions.

[0137] Step 52420: Determine the reference region of the sample category corresponding to the current row in the classification task execution module of the current sample image at the current training iteration as the first sample region with the highest score in the third learning sub-branch score matrix of the current sample image at the current training iteration in any current row.

[0138] Step 52421: Calculate the intersection over union (IoU) between each first sample region of the current sample image and the reference region of the sample category corresponding to the current row in the classification task execution module of the current sample image at the current training iteration, and based on each IoU and the third preset IoU threshold Determine the initial instance-level pseudo-labels of each first sample region in the classification task execution module of the current sample image at the current training iteration for each sample category or the background.

[0139] Step 52422: Based on the initial instance-level pseudo-labels of each first sample region in the classification task execution module of all sample images at the current training iteration for each sample category or the background and the classification task score matrix of all sample images at the current training iteration, determine the classification task loss at the current training iteration.

[0140] Specifically, the calculation formula for the classification task loss is:

[0141]

[0142] where represents the classification task loss; represents L of the nth sample image cls ; L cls represents the classification task sub-loss; represents the loss function weight of the classification task execution module, defined as the highest score among all the scores of the third learning sub-branch corresponding to the cth row in the third learning sub-branch score matrix; represents the initial instance-level pseudo-label of the r1th first sample region in the classification task execution module for the cth sample category or the background; represents the element in the cth row and r1th column of the classification task score matrix, that is, the classification task score of the first sample region corresponding to the r1th column for the sample category corresponding to the cth row.

[0143] Step 52423: Perform per-pixel averaging on the score matrices of each k-th learning sub-branch and the classification task score matrix of the current sample image to obtain the final score matrix of the current sample image; an element in any current row and any current column of the final score matrix of the current sample image represents the final score that the target in the first sample region corresponding to the current column belongs to the background or the sample category corresponding to the current row.

[0144] Step 52424: Use the non-maximum suppression method to process all the first sample regions in each row of the final score matrix of the current sample image to obtain one or more sixth sample regions corresponding to the sample categories of each row.

[0145] Step 52425: Calculate the intersection over union (IoU) between each first sample region and each sixth sample region of the current sample image respectively, and determine one or more seventh sample regions of each sample category of the current sample image at the current training iteration based on each IoU and the fourth preset IoU threshold.

[0146] Specifically, when there are multiple sixth sample regions corresponding to the sample category of any current row, calculate the IoU between any current first sample region and each sixth sample region respectively to obtain multiple IoUs corresponding to the current first sample region, and determine the IoU value with the maximum value as the IoU for screening of the current first sample region. If the IoU for screening of the current first sample region is greater than the fourth preset IoU threshold, then determine the current first sample region as the seventh sample region corresponding to the sample category of the current row.

[0147] Step 52426: Determine the sixth sample regions with the largest IoU of each sample category as the localization reference regions of each sample category of the current sample image at the current training iteration.

[0148] Step 52427: Based on the original absolute coordinates of each seventh sample region of each sample category of the current sample image at the current training iteration and the original absolute coordinates of the localization reference region, determine the original deviation coordinates of each seventh sample region of each sample category of the current sample image at the current training iteration.

[0149] Step 52428: Use the localization module in the object detection network at the current training iteration to determine the predicted offset coordinates of each seventh sample region of the current sample image on each sample category respectively based on the feature vectors of each seventh sample region of each sample category of the current sample image; the localization module is a fully connected layer with an output of 84 dimensions.

[0150] Step 52429: Determine the localization loss at the current training iteration based on the predicted offset coordinates and the original offset coordinates of each seventh sample region of all sample images on each sample category.

[0151] Specifically, the calculation formula for the localization loss is as follows:

[0152]

[0153] Wherein, represents the localization loss; represents the L of the nth sample image loc ; L loc localization sub-loss; |R7,c| represents the number of the seventh sample regions corresponding to the cth sample category; represents a custom function with respect to the independent variable g; represents the predicted offset coordinate of the parameter i of the r7th seventh sample region corresponding to the cth sample category on the cth sample category, x0 represents the x-axis component of the center coordinate of the seventh sample region, y0 represents the y-axis component of the center coordinate of the seventh sample region, w represents the width of the seventh sample region, and h represents the height of the seventh sample region; represents the original offset coordinate of the parameter i of the r7th seventh sample region corresponding to the cth sample category; otherwise represents others; |g| represents the absolute value of g.

[0154] Step 52430: Based on the multi-label classification loss, the kth learning sub-branch loss, the kth relationship sub-branch loss, the classification task loss, and the localization loss at the current training iteration, determine the total loss at the current training iteration.

[0155] Specifically, the calculation formula for the total loss is as follows:

[0156]

[0157] Wherein, L 总 represents the total loss.

[0158] Step 52431: Determine whether the training stop condition is satisfied; the training stop condition is that the total loss at the current training iteration is less than the preset loss or the current training iteration reaches the maximum training iteration.

[0159] Step 52432: If so, determine the object detection network at the current training iteration as the object detection model, thereby determining the classification and localization module.

[0160] Step 52433: If not, update the object detection network at the current training iteration to the object detection network at the next training iteration, and return to Step 52402.

[0161] Step 525: Based on the multi-branch learning module, the classification task execution module, and the localization module in the object detection model, determine the classification and localization module.

[0162] This application obtains a non-linear similarity metric by introducing a prototype relationship network, thereby screening out the bounding boxes of more potential high-quality candidate instance regions and their feature vectors. Since more positive sample instances are obtained, the classification and localization model can learn the diversity of each type of instance from a richer set of samples, and thus capture a wider range of feature vectors. This mechanism effectively alleviates the common problem of missed detection in traditional weakly supervised object detection, and generates higher-quality instance-level pseudo-labels through iterative optimization, further improving the detection performance of the classification and localization model. In addition, the improved multi-branch classification results can generate more reliable pseudo-labels for training the localization module. Under the supervision of high-quality pseudo-labels, the localization module can more precisely adjust the position of the candidate box, thereby alleviating the local focus problem to a certain extent. Through this collaborative optimization method, this application achieves more accurate object localization and more comprehensive object detection under weakly supervised conditions, significantly improving the overall detection performance.

[0163] In an exemplary embodiment, a computer device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the weakly supervised object detection method.

[0164] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the weakly supervised object detection method is implemented.

[0165] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the weakly supervised object detection method is implemented.

[0166] In an exemplary embodiment, a computer device is provided, and the computer device can be a server or a terminal, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a weakly supervised object detection method is implemented.

[0167] Those skilled in the art can understand that Figure 5 the structure shown in Figure 5 is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0168] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memories can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0169] The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0170] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0171] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0172] Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, based on the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A weakly supervised object detection method, characterized in that, The weakly supervised object detection method includes: Obtain a to-be-detected picture and a set of to-be-detected target categories; the set of to-be-detected target categories includes multiple to-be-detected categories, and the to-be-detected category is the category of the to-be-detected target included in the to-be-detected picture; Preprocess the to-be-detected picture to obtain a preprocessed to-be-detected picture; Adopt a selective search algorithm to perform region segmentation and merging on the preprocessed to-be-detected picture to obtain a set of regions to be classified; the set of regions to be classified includes multiple regions to be classified and the original absolute coordinates of each region to be classified, and the region to be classified is a candidate instance region of the to-be-detected picture; Input the preprocessed to-be-detected picture and the set of regions to be classified into a feature vector extraction model to obtain the feature vectors of each region to be classified; the feature vector extraction model is constructed based on an improved VGG16 network, a regional pooling layer, and a fully connected layer; Input the feature vectors of each region to be classified and the set of to-be-detected target categories into a classification and localization module to obtain one or more regions to be localized for each to-be-detected category of the to-be-detected picture and the final absolute coordinates of each region to be localized on each to-be-detected category; the classification and localization module includes a multi-branch learning module, a classification task execution module, and a localization module in an object detection model, the object detection model is obtained by weakly supervised training of an object detection network, and the object detection network includes: a multi-instance learning detection module, a multi-branch learning module, a prototype relationship network module, a classification task execution module, and a localization module; the object detection network is constructed based on a fully connected layer.

2. The weakly supervised object detection method according to claim 1, wherein Preprocess the to-be-detected picture to obtain a preprocessed to-be-detected picture, including: Adjust the size of the to-be-detected picture according to a preset scale to obtain a to-be-detected picture with adjusted size; Perform a random horizontal flipping operation on the to-be-detected picture with adjusted size to obtain a flipped to-be-detected picture; Perform a normalization process on the flipped to-be-detected picture to obtain a preprocessed to-be-detected picture.

3. The weakly supervised object detection method according to claim 1, wherein Input the preprocessed to-be-detected picture and the set of regions to be classified into a feature vector extraction model to obtain the feature vectors of each region to be classified, including: Input the preprocessed to-be-detected picture into the backbone feature extraction module in the feature vector extraction model to obtain a picture-level feature map of the to-be-detected picture; the backbone feature extraction module is obtained by training an improved VGG16 network; Input each region to be classified and the picture-level feature map of the to-be-detected picture into the regional pooling module in the feature vector extraction model to obtain the feature maps of multiple sub-regions of each region to be classified; the regional pooling module is obtained by training a regional pooling layer; Input the feature maps of all sub-regions of each region to be classified into the feature vector determination module in the feature vector extraction model to obtain the feature vectors of each region to be classified; the feature vector determination module is obtained by training two fully connected layers.

4. The weakly supervised object detection method according to claim 3, wherein The improved VGG16 network includes: a first convolutional module, a second convolutional module, a third convolutional module, a fourth convolutional module, and a fifth convolutional module connected in sequence; Both the first convolution module and the second convolution module include a 3×3 convolution layer, a ReLU activation function, a 3×3 convolution layer, a ReLU activation function, and a 2×2 max pooling layer connected in sequence; The third convolution module includes a 3×3 convolution layer, a ReLU activation function, a 3×3 convolution layer, a ReLU activation function, a 3×3 convolution layer, a ReLU activation function, and a 2×2 max pooling layer connected in sequence; The fourth convolution module includes a 3×3 convolution layer, a ReLU activation function, a 3×3 convolution layer, a ReLU activation function, a 3×3 convolution layer, and a ReLU activation function; The fifth convolution module includes a 3×3 dilated convolution layer, a ReLU activation function, a 3×3 dilated convolution layer, a ReLU activation function, a 3×3 dilated convolution layer, and a ReLU activation function.

5. The weakly supervised object detection method according to claim 1, wherein Inputting the feature vectors of each region to be classified and the set of target categories to be measured into the classification and localization module, obtaining one or more regions to be localized for each category to be measured of the image to be measured and the final absolute coordinates of each region to be localized on each category to be measured, including: Inputting the feature vectors of each region to be classified and the set of target categories to be measured into the multi-branch learning module, obtaining the score matrix of the k-th learning sub-branch of the image to be measured, where k is 1, 2, or 3; any element in any current row and any current column of the score matrix of the k-th learning sub-branch of the image to be measured represents the score of the k-th learning sub-branch indicating that the target in the region to be classified corresponding to the current column belongs to the background or the category to be measured corresponding to the current row; the score of the k-th learning sub-branch is the score output by the k-th learning sub-branch in the multi-branch learning module; each learning sub-branch in the multi-branch learning module is constructed based on a fully connected layer; Inputting the feature vectors of each region to be classified and the set of target categories to be measured into the classification task execution module, obtaining the classification task score matrix of the image to be measured; any element in any current row and any current column of the classification task score matrix of the image to be measured represents the classification task score indicating that the target in the region to be classified corresponding to the current column belongs to the background or the category to be measured corresponding to the current row; the classification task score is the score output by the classification task execution module; the classification task execution module is constructed based on a fully connected layer; Performing per-pixel averaging on the score matrices of each k-th learning sub-branch and the classification task score matrix of the image to be measured, obtaining the final score matrix of the image to be measured; any element in any current row and any current column of the final score matrix of the image to be measured represents the final score indicating that the target in the region to be classified corresponding to the current column belongs to the background or the category to be measured corresponding to the current row; Using the non-maximum suppression method to process all regions to be classified in each row of the final score matrix of the image to be measured, obtaining one or more regions to be localized for each category to be measured corresponding to each row; Using a positioning module, based on the feature vectors of each region to be located of each category to be measured in the image to be measured, determine the predicted offset coordinates of each region to be located in the image to be measured for each category to be measured, and based on the predicted offset coordinates and the original final absolute coordinates of each region to be located in the image to be measured for each category to be measured, determine the final absolute coordinates of each region to be located in the image to be measured for each category to be measured; the positioning module is constructed based on a fully connected layer.

6. The weakly supervised object detection method according to claim 5, wherein The determination process of the classification and positioning module includes: Obtain multiple sample images and the set of image-level labels for each sample image; the set of image-level labels includes image-level labels of multiple sample categories, and the image-level label is the inclusion situation of the target of the sample image for the sample category; the inclusion situation is inclusion or non-inclusion. Respectively determine the feature vectors of multiple first sample regions of each sample image. Construct the object detection network based on multiple fully connected layers. Use the feature vectors of multiple first sample regions of each sample image and the image-level labels to perform multiple weak supervision trainings on the object detection network to obtain the object detection model. Based on the multi-branch learning module, classification task execution module and positioning module in the object detection model, determine the classification and positioning module.

7. The weakly supervised object detection method according to claim 6, wherein Respectively determining the feature vectors of multiple first sample regions of each sample image includes: Preprocess each sample image respectively to obtain multiple preprocessed sample images. Adopt the selective search algorithm to perform region segmentation and merging on each preprocessed sample image respectively to obtain the set of first sample regions of each preprocessed sample image; the set of first sample regions includes multiple first sample regions and the original absolute coordinates of each first sample region, and the first sample region is a candidate instance region of the sample image. Respectively input each preprocessed sample image and the corresponding set of first sample regions into the feature vector extraction model to obtain the feature vectors of each first sample region.

8. The weakly supervised object detection method according to claim 7, wherein In the process of using the feature vectors of multiple first sample regions of each sample image and the image-level labels to perform multiple weak supervision trainings on the object detection network to obtain the object detection model, the training process at any current training times includes: Determine any sample image as the current sample image. Respectively input the feature vectors of each first sample region of the current sample image into the classification branch and the detection branch of the multi-instance learning detection module in the object detection network at the current training times to obtain the classification branch score matrix and the detection branch score matrix of the current sample image at the current training times; both the classification branch and the detection branch are fully connected layers with an output of 20 dimensions; the element in any current row and any current column of the classification branch score matrix represents the classification branch score that the target in the first sample region corresponding to the current column belongs to the sample category corresponding to the current row, and the classification branch score is the score output by the classification branch; the element in any current row and any current column of the detection branch score matrix represents the detection branch score of the first sample region corresponding to the current column for the target of the sample category corresponding to the current row, and the detection branch score is the score output by the detection branch. Input the classification branch score matrix and the detection branch score matrix of the current sample image at the current training iteration into the softmax function of the multi-instance learning detection module in the object detection network at the current training iteration respectively, to obtain the normalized classification branch score matrix and the normalized detection branch score matrix of the current sample image at the current training iteration; Perform an element-wise product calculation on the normalized classification branch score matrix and the normalized detection branch score matrix of the current sample image at the current training iteration, to obtain the classification-detection comprehensive score matrix of the current sample image at the current training iteration; Accumulate all the elements in each row of the classification-detection comprehensive score matrix of the current sample image at the current training iteration respectively, to obtain the image category scores of each sample category of the current sample image at the current training iteration; Based on the image category scores of each sample category of all sample images at the current training iteration and the image-level label set, determine the multi-label classification loss at the current training iteration; Input the feature vectors of each first sample region of the current sample image and each sample category into the multi-branch learning module in the object detection network at the current training iteration, to obtain the k-th learning sub-branch score matrix of the current sample image at the current training iteration; any element in any current row and any current column of the k-th learning sub-branch score matrix of the current sample image represents the score of the k-th learning sub-branch that the object in the first sample region corresponding to the current column of the current sample image belongs to the background or the sample category corresponding to the current row; each k-th learning sub-branch in the multi-branch learning module is a fully connected layer with an output of 21 dimensions; Determine the reference region of the sample category corresponding to the current row in the k-th learning sub-branch of the current sample image at the current training iteration as the first sample region with the highest k-1 learning sub-branch score in any current row of the (k-1)-th learning sub-branch score matrix of the current sample image at the current training iteration; where, when k = 1, the k-th learning sub-branch score matrix is the classification-detection comprehensive score matrix; Calculate the intersection over union (IoU) between each first sample region of the current sample image and the reference region of the sample category corresponding to the current row in the k-th learning sub-branch of the current sample image at the current training iteration respectively, and determine the initial instance-level pseudo-labels of each first sample region for each sample category based on each IoU and the first preset IoU threshold, and filter each first sample region based on the initial instance-level pseudo-labels, to obtain the set of second sample regions of each sample category in the k-th learning sub-branch of the current sample image at the current training iteration; the set of second sample regions includes multiple second sample regions; Average the feature vectors of all the second sample regions corresponding to each sample category in the k-th learning sub-branch of the current sample image at the current training iteration, to obtain the prototype features of each sample category in the k-th learning sub-branch of the current sample image at the current training iteration; Sort all the k-th learning sub-branch scores in any current row of the k-th learning sub-branch score matrix of the current sample image at the current training iteration in descending order, and based on the first sample regions corresponding to the k-th learning sub-branch scores ranked in the top preset proportion, determine the set of third sample regions of the sample category corresponding to the current row in the k-th learning sub-branch of the current sample image at the current training iteration; the set of third sample regions includes multiple third sample regions. Concatenate the prototype features of each sample category in the k-th learning sub-branch of the current sample image at the current training iteration with the feature vectors of each third sample region to obtain the concatenated features of each sample category and each third sample region in the k-th learning sub-branch of the current sample image at the current training iteration. Input the concatenated features of each sample category and each third sample region in the k-th learning sub-branch of the current sample image at the current training iteration into the k-th relationship sub-branch of the prototype relationship network module in the object detection network at the current training iteration respectively, to obtain the similarity scores between the feature vectors of each third sample region and the prototype features of each sample category in the k-th relationship sub-branch of the current sample image at the current training iteration; each k-th relationship sub-branch of the prototype relationship network module includes a fully connected layer with an output of 256 dimensions and a fully connected layer with an output of 1 dimension. Based on the third sample regions with similarity scores greater than the preset similarity score threshold, determine the set of fourth sample regions of each sample category in the k-th relationship sub-branch of the current sample image at the current training iteration; the set of fourth sample regions includes multiple fourth sample regions. Use the non-maximum suppression method to process the set of fourth sample regions of each sample category in the k-th relationship sub-branch of the current sample image at the current training iteration, to obtain the set of fifth sample regions of each sample category in the k-th relationship sub-branch of the current sample image at the current training iteration; the set of fifth sample regions includes multiple fifth sample regions. Calculate the intersection over union between each first sample region of the current sample image at the current training iteration and the fifth sample regions in the set of fifth sample regions of each sample category in the k-th relationship sub-branch of the current sample image at the current training iteration respectively, and based on each intersection over union and the second preset intersection over union threshold, determine the optimized instance-level pseudo-labels of each first sample region in the k-th relationship sub-branch of the current sample image at the current training iteration for each sample category or the background. Based on the optimized instance-level pseudo-labels of each first sample region for each sample category or the background and the k-th learning sub-branch score matrix in the k-th relationship sub-branch of all sample images at the current training iteration, determine the k-th learning sub-branch loss at the current training iteration. Based on the similarity scores between the feature vectors of each third sample region of each sample category in the k-th relationship sub-branch of all sample images at the current training iteration and the prototype features, and the initial instance-level pseudo-labels of each third sample region for each sample category in the k-th learning sub-branch of all sample images at the current training iteration, determine the loss of the k-th relationship sub-branch at the current training iteration; Input the feature vectors of each first sample region of the current sample image and each sample category into the classification task execution module in the object detection network at the current training iteration to obtain the classification task score matrix of the current sample image at the current training iteration; the element in any current row and any current column in the classification task score matrix of the current sample image represents the classification task score that the target in the first sample region corresponding to the current column of the current sample image belongs to the background or the sample category corresponding to the current row; the classification task execution module is a fully connected layer with an output of 21 dimensions; Determine the reference region of the sample category corresponding to the current row in the classification task execution module of the current sample image at the current training iteration as the first sample region with the highest score in the third learning sub-branch score matrix of the current sample image at the current training iteration in any current row; Calculate the intersection over union (IoU) between each first sample region of the current sample image and the reference region of the sample category corresponding to the current row in the classification task execution module of the current sample image at the current training iteration, and determine the initial instance-level pseudo-labels of each first sample region of the current sample image for each sample category or the background in the classification task execution module at the current training iteration based on each IoU and the third preset IoU threshold; Based on the initial instance-level pseudo-labels of each first sample region of all sample images for each sample category or the background in the classification task execution module at the current training iteration and the classification task score matrix of all sample images at the current training iteration, determine the classification task loss at the current training iteration; Perform per-pixel averaging on the score matrices of each k-th learning sub-branch and the classification task score matrix of the current sample image to obtain the final score matrix of the current sample image; the element in any current row and any current column in the final score matrix of the current sample image represents the final score that the target in the first sample region corresponding to the current column belongs to the background or the sample category corresponding to the current row; Use the non-maximum suppression method to process all first sample regions in each row of the final score matrix of the current sample image to obtain one or more sixth sample regions of the sample categories corresponding to each row; Calculate the IoU between each first sample region of the current sample image and each sixth sample region respectively, and determine one or more seventh sample regions of each sample category of the current sample image at the current training iteration based on each IoU and the fourth preset IoU threshold; Determine the sixth sample region with the largest IoU of each sample category as the localization reference region of each sample category of the current sample image at the current training iteration; Based on the original absolute coordinates of each seventh sample region of each sample category in the current sample image at the current training iteration and the original absolute coordinates of the positioning reference region, determine the original deviation coordinates of each seventh sample region of each sample category in the current sample image at the current training iteration; Using the positioning module in the object detection network at the current training iteration, respectively based on the feature vectors of each seventh sample region of each sample category in the current sample image, determine the predicted offset coordinates of each seventh sample region in the current sample image for each sample category; the positioning module is a fully connected layer with an output of 84 dimensions; Based on the predicted offset coordinates and the original offset coordinates of each seventh sample region of all sample images for each sample category, determine the positioning loss at the current training iteration; Based on the multi-label classification loss, the k-th learning sub-branch loss, the k-th relationship sub-branch loss, the classification task loss, and the positioning loss at the current training iteration, determine the total loss at the current training iteration; Judge whether the training stop condition is satisfied; the training stop condition is that the total loss at the current training iteration is less than the preset loss or the current training iteration reaches the maximum training iteration; If so, determine the object detection network at the current training iteration as the object detection model, thereby determining the classification and positioning module; If not, update the object detection network at the current training iteration to the object detection network at the next training iteration, and return "Input the feature vectors of each first sample region of the current sample image into the classification branch and the detection branch of the multi-instance learning detection module in the object detection network at the current training iteration, respectively, to obtain the classification branch score matrix and the detection branch score matrix of the current sample image at the current training iteration".

9. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the weakly supervised object detection method according to any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the weakly supervised object detection method according to any one of claims 1-8.