Weak supervision target detection method based on misclassification correction and related equipment
By designing a misclassification correction-driven label allocation module in the weakly supervised visual object detection method, identifying and correcting the misclassification situation, and reassigning tags to improve the model training effect, the misclassification problem in weakly supervised object detection is solved, and the classification performance is significantly improved.
Patent Information
- Application Number
- CN202510092114.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The weakly supervised visual object detection method has shortcomings in the misclassification problem, resulting in limited improvement in classification performance.
By designing a label allocation module driven by misclassification correction, identifying the misclassification situation and correcting it, and reassigning the tags for object detection model training. This module uses a convolutional neural network to extract regional features and initially allocate category tags through a classifier, and combines the confidence difference between categories to perform misclassification corrections and label reallocation.
The classification performance of weakly supervised object detection method is significantly improved, and the model's learning ability of various categories is enhanced by correcting the error category labels in the training stage.
Smart Images

Figure CN120014234A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a weakly supervised target detection method based on misclassification correction and related equipment. Background Art
[0002] Driven by national policies, artificial intelligence technology is accelerating its penetration into all walks of life, especially in industrial manufacturing, smart transportation, medical health, culture and education. As an important part of the field of artificial intelligence, computer vision, especially advanced learning and cognitive technologies represented by visual object detection, has become one of the key driving forces for promoting the intelligent upgrading of industries.
[0003] Although deep learning-based visual object detection models have made great progress in accuracy, their high reliance on large-scale manually labeled data has become a significant bottleneck. This labeling process is time-consuming, labor-intensive, and costly, which directly limits the efficient use of large-scale data and seriously restricts the development potential of related technologies.
[0004] To reduce the reliance on costly manual labeling, researchers have proposed weakly supervised visual object detection techniques. However, such methods often overlook a key problem in their design: misclassification. Even if the detector can accurately locate the target area, it cannot accurately determine its category, thus affecting the improvement of classification performance. Summary of the invention
[0005] In view of the above problems, the present invention provides a weakly supervised target detection method, system, electronic device and storage medium based on misclassification correction, aiming to significantly improve the classification performance of the weakly supervised target detection method by correcting the erroneous category labels generated in the training phase.
[0006] According to a first aspect of an embodiment of the present disclosure, a weakly supervised target detection method based on misclassification correction is provided, the method comprising the following steps:
[0007] Use convolutional neural network to extract regional features of target candidate locations;
[0008] Assign category labels to the extracted regional features through a classifier;
[0009] According to the confidence differences between categories, a misclassification correction-driven label assignment module is designed to identify and correct misclassifications, and reallocate labels for target detection model training.
[0010] In some embodiments, the convolutional neural network is based on the VGGNet neural network, and a spatial pyramid pooling layer is added at the output end of the VGGNet to generate convolutional features corresponding to the multi-scale combined target candidate area; and two fully connected layers are added at the output end of the spatial pyramid pooling layer of the VGGNet to generate the feature vector required by the classifier.
[0011] In some embodiments, the extracted regional features are assigned category labels by a classifier, and the training process includes the following steps:
[0012] The regional feature vectors are respectively input into two fully connected layers in the classifier to obtain a first matrix and a second matrix;
[0013] Input the first matrix and the second matrix into the category dimension Softmax layer and the region dimension Softmax layer respectively, and obtain a third matrix representing the confidence score of each candidate region belonging to each category and a fourth matrix representing the normalized contribution of each candidate region to the inclusion of each category in the image;
[0014] Performing an element-by-element multiplication operation on the third matrix and the fourth matrix to generate a candidate region score matrix;
[0015] Adding the candidate region score matrix along the candidate region dimension to obtain the confidence that the image contains each category;
[0016] The confidence that the image contains each category and the true category are input into the multi-class cross entropy loss function.
[0017] In some embodiments, during the training of the misclassification correction driven label assignment module:
[0018] Iterate the following steps until the maximum number of iterations is reached:
[0019] Input the region feature vector into the Softmax layer branch k containing the fully connected layer and the category dimension to obtain the candidate region score matrix
[0020] The candidate region score from the previous branch In the above example, select the candidate region with the highest confidence in each positive class c. The corresponding confidence level is Select the candidate region with the highest confidence among all negative classes c′ The corresponding confidence level is
[0021] like It is regarded as a potential misclassification of the category pair cc′;
[0022] If the training process is less than 3 / 10, when a potential misclassification occurs for a category pair cc′, it is recorded in the misclassification memory B and b is updated. cc′ =b cc′ +1, where b cc′ Represents the number of times the category pair cc′ appears; constructs a positive class memory library P to record the number of times each category appears as a positive class in previous iterations;
[0023] If the training process reaches three-tenths, calculate the frequency of potential misclassification of each category pair: F = B / P; sort the elements in F in descending order to obtain the sorted category pair frequency vector o, where o i ≥o i+1 , and the corresponding ranking matrix S, where s cc′ Represents the frequency ranking of category pair cc′; calculates the difference between adjacent elements to obtain the frequency difference vector d; finds the largest element in d and records its ranking N;
[0024] If the training process exceeds three tenths, when a potential misclassified pair cc′ appears, if the condition s is satisfied cc′ ≤N, the candidate region with the highest negative score will be obtained is marked as a positive example; otherwise, the positive example is the region with the highest positive score
[0025] Calculate the intersection-over-union ratio of other candidate regions and positive examples. Candidate regions with an intersection-over-union ratio greater than 0.5 are also selected as positive examples, regions with an intersection-over-union ratio less than 0.1 are ignored, and the remaining regions are set as negative examples as supervision information θ k ;
[0026] Will and θ k Input into the multi-class cross entropy loss function;
[0027] After the iteration is completed, calculate The average value of is taken as the final prediction result of the confidence of each candidate region.
[0028] According to a second aspect of an embodiment of the present disclosure, a weakly supervised target detection system based on misclassification correction is provided, the system comprising:
[0029] A regional feature acquisition unit, used for extracting regional features of target candidate locations using a convolutional neural network;
[0030] A category label acquisition unit, used for assigning category labels to the extracted regional features through a classifier;
[0031] The label assignment module design unit is used to design a misclassification correction driven label assignment module according to the confidence difference between categories, so as to identify the misclassification situation and correct it, and reallocate labels for target detection model training.
[0032] In some embodiments, the convolutional neural network in the regional feature acquisition unit is based on the VGGNet neural network, and a spatial pyramid pooling layer is added at the output end of the VGGNet to generate convolutional features corresponding to the multi-scale combined target candidate region; and two fully connected layers are added at the output end of the spatial pyramid pooling layer of the VGGNet to generate the feature vector required by the classifier.
[0033] In some embodiments, the category label acquisition unit assigns the category label to the extracted regional features through a classifier, and the training process includes the following steps:
[0034] The regional feature vectors are respectively input into two fully connected layers in the classifier to obtain a first matrix and a second matrix;
[0035] Input the first matrix and the second matrix into the category dimension Softmax layer and the region dimension Softmax layer respectively, and obtain a third matrix representing the confidence score of each candidate region belonging to each category and a fourth matrix representing the normalized contribution of each candidate region to the inclusion of each category in the image;
[0036] Performing an element-by-element multiplication operation on the third matrix and the fourth matrix to generate a candidate region score matrix;
[0037] Adding the candidate region score matrix along the candidate region dimension to obtain the confidence that the image contains each category;
[0038] The confidence that the image contains each category and the true category are input into the multi-class cross entropy loss function.
[0039] In some embodiments, in the label assignment module design unit, during the process of training the misclassification correction driven label assignment module:
[0040] Iterate the following steps until the maximum number of iterations is reached:
[0041] Input the region feature vector into the Softmax layer branch k containing the fully connected layer and the category dimension to obtain the candidate region score matrix
[0042] The candidate region score from the previous branch In the above example, select the candidate region with the highest confidence in each positive class c. The corresponding confidence level is Select the candidate region with the highest confidence among all negative classes c′ The corresponding confidence level is
[0043] like It is regarded as a potential misclassification of the category pair cc′;
[0044] If the training process is less than 3 / 10, when a potential misclassification occurs for a category pair cc′, it is recorded in the misclassification memory B and b is updated. cc′ =b cc′ +1, where b cc′ Represents the number of times the category pair cc′ appears; constructs a positive class memory library P to record the number of times each category appears as a positive class in previous iterations;
[0045] If the training process reaches three-tenths, calculate the frequency of potential misclassification of each category pair: F = B / P; sort the elements in F in descending order to obtain the sorted category pair frequency vector o, where o i ≥o i+1 , and the corresponding ranking matrix S, where s cc′ Represents the frequency ranking of category pair cc′; calculates the difference between adjacent elements to obtain the frequency difference vector d; finds the largest element in d and records its ranking N;
[0046] If the training process exceeds three tenths, when a potential misclassified pair cc′ appears, if the condition s is satisfied cc′ ≤N, the candidate region with the highest negative score will be obtained is marked as a positive example; otherwise, the positive example is the region with the highest positive score
[0047] Calculate the intersection-over-union ratio of other candidate regions and positive examples. Candidate regions with an intersection-over-union ratio greater than 0.5 are also selected as positive examples, regions with an intersection-over-union ratio less than 0.1 are ignored, and the remaining regions are set as negative examples as supervision information θ k ;
[0048] Will and θ k Input into the multi-class cross entropy loss function;
[0049] After the iteration is completed, calculate The average value of is taken as the final prediction result of the confidence of each candidate region.
[0050] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the above-mentioned weakly supervised target detection method based on misclassification correction are implemented.
[0051] According to a fourth aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the above-mentioned weakly supervised target detection method based on misclassification correction are implemented.
[0052] The embodiments of the present disclosure provide a weakly supervised target detection method, system, electronic device and storage medium based on misclassification correction, wherein a convolutional neural network is used to extract regional features corresponding to target candidate positions; a classifier classifies the extracted regional features and preliminarily assigns a category label to each candidate region; a label assignment module driven by misclassification correction is used to identify misclassification situations and correct them, and reallocate labels for model training, thereby more accurately learning the features of each category, providing a practical solution to the misclassification problem, and having great application value. The method of the present invention significantly improves the classification performance of the weakly supervised target detection method by correcting the erroneous category labels generated in the training phase.
[0053] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the description, serve to explain the principles of the present invention.
[0055] Figure 1 is a flow chart of a weakly supervised target detection method based on misclassification correction in an embodiment of the present invention;
[0056] Figure 2 It is an overall block diagram of the weakly supervised target detection method based on misclassification correction in an embodiment of the present invention;
[0057] Figure 3 is a schematic diagram of the structure of a weakly supervised target detection system based on misclassification correction in an embodiment of the present invention;
[0058] Figure 4 It is a schematic diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION
[0059] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0060] It should be mentioned before discussing the exemplary embodiments in more detail that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0061] The embodiments of the present invention provide the following embodiments for a weakly supervised target detection method, system, electronic device and storage medium based on misclassification correction:
[0062] like Figure 1 As shown in FIG. 1 , the weakly supervised target detection method based on misclassification correction includes the following steps:
[0063] S1, using convolutional neural network to extract regional features of target candidate locations;
[0064] S2, assigning a category label to the extracted regional features through a classifier; specifically, classifying the extracted regional features through a classifier, and preliminarily assigning a category label to each candidate region;
[0065] S3. According to the confidence differences between categories, a misclassification correction-driven label assignment module is designed to identify misclassification situations and correct them, and to reallocate labels for target detection model training. Specifically, according to the confidence differences between categories, a misclassification correction-driven label assignment module is designed to identify misclassification situations and correct them, and to reallocate labels for model training, so as to learn the characteristics of each category more accurately.
[0066] Overall, such as Figure 2 As shown in the figure, the convolutional neural network is based on the VGGNet neural network, and a spatial pyramid pooling layer is added at the output of the VGGNet to generate convolutional features corresponding to the multi-scale combined target candidate areas; and two fully connected layers are added at the output of the spatial pyramid pooling layer of the VGGNet to generate the feature vector required by the classifier.
[0067] In a specific embodiment, the convolutional neural network in S1 is modified from the original version of VGGNet (Simonyan, Karen, and Andrew Zisserman. "Very deep convolutional networks for large-scale image recognition." ArXiv. 2014.).
[0068] Specifically, the specific method of establishing a convolutional neural network is as follows:
[0069] 1) Add a spatial pyramid pooling layer at the output of VGGNet to generate convolutional features corresponding to the target candidate regions provided by the multi-scale combination grouping method;
[0070] 2) Add two fully connected layers at the output of the spatial pyramid pooling layer.
[0071] The establishment of the convolutional neural network realizes the generation of feature vectors of the target candidate region and provides input for the classifier in S2.
[0072] In S2, the extracted regional features are assigned category labels by a classifier. The training process includes the following steps:
[0073] The regional feature vectors are respectively input into two fully connected layers in the classifier to obtain a first matrix and a second matrix;
[0074] Input the first matrix and the second matrix into the category dimension Softmax layer and the region dimension Softmax layer respectively, and obtain a third matrix representing the confidence score of each candidate region belonging to each category and a fourth matrix representing the normalized contribution of each candidate region to the inclusion of each category in the image;
[0075] Performing an element-by-element multiplication operation on the third matrix and the fourth matrix to generate a candidate region score matrix;
[0076] Adding the candidate region score matrix along the candidate region dimension to obtain the confidence that the image contains each category;
[0077] The confidence that the image contains each category and the true category are input into the multi-class cross entropy loss function.
[0078] In a specific embodiment, the classifier is modified from the original version of WSDDN (Bilen, Hakan, and Andrea Vedaldi. "Weakly supervised deep detection networks." Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016.).
[0079] The specific method of establishing the classifier is as follows:
[0080] S21, input the feature vectors generated in S1 into two fully connected layers respectively to generate two matrices X cls and X det ;
[0081] S22, input the two matrices generated in S21 into the Softmax layer on the category dimension and the Softmax layer on the region dimension respectively, and obtain two matrices σ(X cls ) and σ(X det ), representing the confidence score of each candidate region belonging to each category, and the normalized contribution of each candidate region to the image containing each category;
[0082] S23, perform element-by-element multiplication on the two matrices generated by S22 to generate a candidate region score matrix
[0083] S24, adding the candidate region score matrix in S23 along the candidate region dimension to obtain the confidence that the image contains each category;
[0084] S25, inputting the image category confidence and category truth generated by S24 into a multi-class cross entropy loss function;
[0085] If it is in the training phase, S21 to S25 are performed; if it is in the application phase, S21 to S25 are skipped.
[0086] At this point, the establishment of the classifier is completed, and the classifier for classifying images is trained to provide supervision information for the next branch.
[0087] The misclassification correction driven label assignment module described in S3 includes three branches with the same structure:
[0088] The construction method of the misclassification correction driven label assignment module is as follows:
[0089] S31, input the feature vector generated in step 1) into branch k having a fully connected layer and a Softmax layer on the category dimension to generate a candidate region score matrix
[0090] S32, candidate region score from the previous branch In the above example, select the candidate region with the highest confidence in each positive class c. The corresponding confidence level is Select the candidate region with the highest confidence among all negative classes The corresponding confidence level is If the difference between the two satisfies It is regarded as a potential misclassification of the category pair cc′;
[0091] Step S33: If the training process is less than 3 / 10, when a potential misclassification of the class pair cc′ occurs, it is recorded in the misclassification memory bank B: cc′ =bcc′ +1, where b cc′ Represents the number of times the class pair cc′ appears; construct a positive class memory P to record the number of times each class appears as a positive class in previous iterations. If the training process reaches three-tenths, calculate the frequency of potential misclassification of each class pair: F = B / P; then sort the elements in F in descending order to obtain the sorted class pair frequency vector o, where o i ≥o i+1 , and the corresponding ranking matrix S, where s cc′ Represents the frequency ranking of category pair cc′; calculates the difference between adjacent elements to obtain the frequency difference vector d; finds the largest element in d and records its ranking N. If the training process exceeds three tenths, when a potential misclassified pair cc′ appears, if the condition s is met cc′ ≤N, the candidate region with the highest negative score will be obtained is marked as a positive example; otherwise, the positive example is the region with the highest positive score
[0092] S34, calculate the intersection-over-union ratio of other candidate regions and positive examples. Candidate regions with an intersection ratio greater than 0.5 are also selected as positive examples, regions with an intersection ratio less than 0.1 are ignored, and the remaining regions are set as negative examples as supervision information θ k ;
[0093] S35, and θ k Input into the multi-class cross entropy loss function;
[0094] S36, if in the training phase, repeat S31 to S35 3 times; if in the application phase, only repeat step S31 3 times;
[0095] S37, calculation The average value of is taken as the final prediction result of the confidence of each candidate region.
[0096] This completes the establishment of the label allocation module driven by misclassification correction.
[0097] Another embodiment is used to illustrate a weakly supervised target detection system based on misclassification correction, such as Figure 3 As shown, the system 300 includes:
[0098] A regional feature acquisition unit 310 is used to extract regional features of target candidate locations using a convolutional neural network;
[0099] A category label acquisition unit 320 is used to assign a category label to the extracted regional features through a classifier;
[0100] The label assignment module design unit 330 is used to design a misclassification correction driven label assignment module according to the confidence difference between categories, so as to identify misclassification situations and correct them, and reallocate labels for target detection model training.
[0101] The convolutional neural network in the regional feature acquisition unit 310 is based on the VGGNet neural network, and a spatial pyramid pooling layer is added at the output end of the VGGNet to generate convolutional features corresponding to the multi-scale combined target candidate region; and two fully connected layers are added at the output end of the spatial pyramid pooling layer of the VGGNet to generate the feature vector required by the classifier.
[0102] The category label acquisition unit 320 assigns the category label to the extracted regional features through a classifier. The training process includes the following steps:
[0103] The regional feature vectors are respectively input into two fully connected layers in the classifier to obtain a first matrix and a second matrix;
[0104] Input the first matrix and the second matrix into the category dimension Softmax layer and the region dimension Softmax layer respectively, and obtain a third matrix representing the confidence score of each candidate region belonging to each category and a fourth matrix representing the normalized contribution of each candidate region to the inclusion of each category in the image;
[0105] Performing an element-by-element multiplication operation on the third matrix and the fourth matrix to generate a candidate region score matrix;
[0106] Adding the candidate region score matrix along the candidate region dimension to obtain the confidence that the image contains each category;
[0107] The confidence that the image contains each category and the true category are input into the multi-class cross entropy loss function.
[0108] In the label assignment module design unit 330, during the process of training the label assignment module driven by the misclassification correction:
[0109] Iterate the following steps until the maximum number of iterations is reached:
[0110] Input the regional feature vector into the Softmax layer branch k containing the fully connected layer and the category dimension to obtain the candidate region score matrix φ k ;
[0111] The candidate region score φ from the previous branch k-1 In the above example, select the candidate region with the highest confidence in each positive class c. The corresponding confidence level is Select the candidate region with the highest confidence among all negative classes c′ The corresponding confidence level is
[0112] like It is regarded as a potential misclassification of the category pair cc′;
[0113] If the training process is less than 3 / 10, when a potential misclassification occurs for a category pair cc′, it is recorded in the misclassification memory B and b is updated. cc′ =b cc′ +1, where b cc′ Represents the number of times the category pair cc′ appears; constructs a positive class memory library P to record the number of times each category appears as a positive class in previous iterations;
[0114] If the training process reaches three-tenths, calculate the frequency of potential misclassification of each category pair: F = B / P; sort the elements in F in descending order to obtain the sorted category pair frequency vector o, where o i ≥o i+1 , and the corresponding ranking matrix S, where s cc′ Represents the frequency ranking of category pair cc′; calculates the difference between adjacent elements to obtain the frequency difference vector d; finds the largest element in d and records its ranking N;
[0115] If the training process exceeds three tenths, when a potential misclassified pair cc′ appears, if the condition s is satisfied cc′ ≤N, the candidate region with the highest negative score will be obtained is marked as a positive example; otherwise, the positive example is the region with the highest positive score
[0116] Calculate the intersection-over-union ratio of other candidate regions and positive examples. Candidate regions with an intersection-over-union ratio greater than 0.5 are also selected as positive examples, regions with an intersection-over-union ratio less than 0.1 are ignored, and the remaining regions are set as negative examples as supervision information θ k ;
[0117] After the iteration is completed, φ k and θ k Input into the multi-class cross entropy loss function;
[0118] After the iteration is completed, calculate φ k The average value of is taken as the final prediction result of the confidence of each candidate region.
[0119] In addition to the upper module, the unit 300 may also include other components, however, since these components are irrelevant to the content of the embodiment of the present disclosure, their illustration and description are omitted here.
[0120] The other specific working processes of the weakly supervised target detection system 300 based on misclassification correction refer to the description of the above-mentioned embodiment of the weakly supervised target detection method based on misclassification correction, which will not be repeated here.
[0121] Another embodiment is used to illustrate that the system of the present invention can also be used with the help of Figure 4 The architecture of the computing device shown is implemented. Figure 4 The architecture of the computing device is shown. Figure 4 As shown, a computer system 410, a system bus 430, one or more CPUs 440, an input / output 420, a memory 450, etc. The memory 450 can store various data or files used for computer processing and / or communication and program instructions executed by the CPU including the weakly supervised target detection method based on misclassification correction of the embodiment. Figure 4 The architecture shown is only exemplary and can be adjusted according to actual needs when implementing different devices. Figure 4 One or more components in. The memory 450, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as the program instructions / modules corresponding to the weakly supervised target detection method based on misclassification correction in the embodiment of the present invention (for example, the regional feature acquisition unit 310, the category label acquisition unit 320 and the label assignment module design unit 330 in the weakly supervised target detection system 300 based on misclassification correction). One or more CPUs 440 execute various functional applications and data processing of the system of the present invention by running the software programs, instructions and modules stored in the memory 450, that is, to implement the above-mentioned weakly supervised target detection method based on misclassification correction, which includes the following steps:
[0122] Use convolutional neural network to extract regional features of target candidate locations;
[0123] Assign category labels to the extracted regional features through a classifier;
[0124] According to the confidence differences between categories, a misclassification correction-driven label assignment module is designed to identify and correct misclassifications, and reallocate labels for target detection model training.
[0125] Of course, the processor of the server provided in the embodiment of the present invention is not limited to executing the method operations described above, but can also execute relevant operations in the weakly supervised target detection method based on misclassification correction provided in any embodiment of the present invention.
[0126] The memory 450 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 450 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 450 may further include a memory remotely arranged relative to one or more CPUs 440, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0127] The input / output 420 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The input / output 420 may also include a display device such as a display screen.
[0128] The embodiment of the present invention also provides a non-temporary computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by the processor, the weakly supervised target detection method based on misclassification correction described in the above embodiment is implemented. The computer-readable storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media (non-exhaustive list) include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.
[0129] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0130] The program code contained on the storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the foregoing.
[0131] In addition, other specific working processes of a non-temporary computer-readable storage medium refer to the description of the above-mentioned weakly supervised target detection method embodiment based on misclassification correction, and will not be repeated here.
[0132] In summary, the technical solutions provided by the above embodiments provide a method, system, electronic device and storage medium for weakly supervised target detection based on misclassification correction, wherein a convolutional neural network is used to extract regional features corresponding to target candidate positions; a classifier classifies the extracted regional features and preliminarily assigns a category label to each candidate region; a label assignment module driven by misclassification correction is used to identify misclassification situations and correct them, and reallocate labels for model training, thereby more accurately learning the features of each category, providing a practical solution to the misclassification problem, and having great application value. The method of the present invention significantly improves the classification performance of the weakly supervised target detection method by correcting the erroneous category labels generated in the training phase.
[0133] In this document, the terms "comprises," "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, such that a step or method that includes a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such step or method.
[0134] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the scope of protection of the present invention.
Claims
1. A weakly supervised target detection method based on misclassification correction, characterized in that: The method comprises the following steps: Use convolutional neural network to extract regional features of target candidate locations; Assign category labels to the extracted regional features through a classifier; According to the confidence differences between categories, a misclassification correction-driven label assignment module is designed to identify and correct misclassifications, and reallocate labels for target detection model training.
2. The weakly supervised target detection method based on misclassification correction according to claim 1, characterized in that: The convolutional neural network is based on the VGGNet neural network, and a spatial pyramid pooling layer is added to the output end of the VGGNet to generate convolutional features corresponding to the multi-scale combined target candidate regions; And add two fully connected layers at the output of the spatial pyramid pooling layer of VGGNet to generate the feature vector required by the classifier.
3. The weakly supervised target detection method based on misclassification correction according to claim 1, characterized in that: The extracted regional features are assigned category labels by the classifier. The training process includes the following steps: The regional feature vectors are respectively input into two fully connected layers in the classifier to obtain a first matrix and a second matrix; Input the first matrix and the second matrix into the category dimension Softmax layer and the region dimension Softmax layer respectively, and obtain a third matrix representing the confidence score of each candidate region belonging to each category and a fourth matrix representing the normalized contribution of each candidate region to the inclusion of each category in the image; Performing an element-by-element multiplication operation on the third matrix and the fourth matrix to generate a candidate region score matrix; Adding the candidate region score matrix along the candidate region dimension to obtain the confidence that the image contains each category; The confidence that the image contains each category and the true category are input into the multi-class cross entropy loss function.
4. The weakly supervised target detection method based on misclassification correction according to claim 1, characterized in that: During the training of the misclassification correction driven label assignment module: Iterate the following steps until the maximum number of iterations is reached: Input the regional feature vector into the Softmax layer branch k containing the fully connected layer and the category dimension to obtain the candidate region score matrix φ k ; The candidate region score φ from the previous branch k-1 In the above example, select the candidate region with the highest confidence in each positive class c. The corresponding confidence level is Select the candidate region with the highest confidence among all negative classes c′ The corresponding confidence level is like It is regarded as a potential misclassification of the category pair cc′; If the training process is less than 3 / 10, when a potential misclassification occurs for the category pair cc′, it is recorded in the misclassification memory B and b is updated. cc′ =b cc′ +1, where b cc′ Represents the number of times the category pair cc′ appears; constructs a positive class memory library P to record the number of times each category appears as a positive class in previous iterations; If the training process reaches three-tenths, calculate the frequency of potential misclassification of each category pair: F = B / P; sort the elements in F in descending order to obtain the sorted category pair frequency vector o, where o i ≥o i+1 , and the corresponding ranking matrix S, where s cc′ represents the frequency ranking of category pair cc′; Calculate the difference between adjacent elements to obtain the frequency difference vector d; find the largest element in d and record its rank N; If the training process exceeds three tenths, when a potential misclassified pair cc′ appears, if the condition s is satisfied cc′ ≤N, the candidate region with the highest negative score will be obtained is marked as a positive example; otherwise, the positive example is the region with the highest positive score Calculate the intersection-over-union ratio of other candidate regions and positive examples. Candidate regions with an intersection-over-union ratio greater than 0.5 are also selected as positive examples, regions with an intersection-over-union ratio less than 0.1 are ignored, and the remaining regions are set as negative examples as supervision information θ k ; φ k and θ k Input into the multi-class cross entropy loss function; After the iteration is completed, calculate φ k The average value of is taken as the final prediction result of the confidence of each candidate region.
5. A weakly supervised target detection system based on misclassification correction, characterized in that: The system comprises: A regional feature acquisition unit, used for extracting regional features of target candidate locations using a convolutional neural network; A category label acquisition unit, used for assigning category labels to the extracted regional features through a classifier; The label assignment module design unit is used to design a misclassification correction driven label assignment module according to the confidence difference between categories, so as to identify the misclassification situation and correct it, and reallocate labels for target detection model training.
6. The weakly supervised target detection system based on misclassification correction according to claim 5, characterized in that: The convolutional neural network in the region feature acquisition unit is based on the VGGNet neural network, and a spatial pyramid pooling layer is added to the output end of the VGGNet to generate convolutional features corresponding to the multi-scale combined target candidate region; And add two fully connected layers at the output of the spatial pyramid pooling layer of VGGNet to generate the feature vector required by the classifier.
7. The weakly supervised target detection system based on misclassification correction according to claim 5, characterized in that: In the category label acquisition unit, the extracted regional features are assigned category labels by a classifier. The training process includes the following steps: The regional feature vectors are respectively input into two fully connected layers in the classifier to obtain a first matrix and a second matrix; Input the first matrix and the second matrix into the category dimension Softmax layer and the region dimension Softmax layer respectively, and obtain a third matrix representing the confidence score of each candidate region belonging to each category and a fourth matrix representing the normalized contribution of each candidate region to the inclusion of each category in the image; Performing an element-by-element multiplication operation on the third matrix and the fourth matrix to generate a candidate region score matrix; Adding the candidate region score matrix along the candidate region dimension to obtain the confidence that the image contains each category; The confidence that the image contains each category and the true category are input into the multi-class cross entropy loss function.
8. The weakly supervised target detection system based on misclassification correction according to claim 5, characterized in that: In the label assignment module design unit, during the process of training the label assignment module driven by misclassification correction: Iterate the following steps until the maximum number of iterations is reached: Input the region feature vector into the Softmax layer branch k containing the fully connected layer and the category dimension to obtain the candidate region score matrix The candidate region score from the previous branch In the above example, select the candidate region with the highest confidence in each positive class c. The corresponding confidence level is Select the candidate region with the highest confidence among all negative classes c′ The corresponding confidence level is like It is regarded as a potential misclassification of the category pair cc′; If the training process is less than 3 / 10, when a potential misclassification occurs for the category pair cc′, it is recorded in the misclassification memory B and b is updated. cc′ =b cc′ +1, where b cc′ Represents the number of times the category pair cc′ appears; constructs a positive class memory library P to record the number of times each category appears as a positive class in previous iterations; If the training process reaches three-tenths, calculate the frequency of potential misclassification of each category pair: F = B / P; sort the elements in F in descending order to obtain the sorted category pair frequency vector o, where o i ≥o i+1 , and the corresponding ranking matrix S, where s cc′ represents the frequency ranking of category pair cc′; Calculate the difference between adjacent elements to obtain the frequency difference vector d; find the largest element in d and record its rank N; If the training process exceeds three tenths, when a potential misclassified pair cc′ appears, if the condition s is satisfied cc′ ≤N, the candidate region with the highest negative score will be obtained is marked as a positive example; otherwise, the positive example is the region with the highest positive score Calculate the intersection-over-union ratio of other candidate regions and positive examples. Candidate regions with an intersection-over-union ratio greater than 0.5 are also selected as positive examples, regions with an intersection-over-union ratio less than 0.1 are ignored, and the remaining regions are set as negative examples as supervision information θ k ; Will and θ k Input into the multi-class cross entropy loss function; After the iteration is completed, calculate The average value of is taken as the final prediction result of the confidence of each candidate region.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the weakly supervised target detection method based on misclassification correction as described in any one of claims 1 to 4 are implemented.
10. A non-transitory computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instructions are executed by the processor, the steps of the weakly supervised target detection method based on misclassification correction as described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Weak supervision target detection method based on spatial attention guiding feature erasing
CN117974976A
Contact net bolt weak supervision detection model training method, detection method and system
CN118552807A
Sound sorting system and method capable of increasing and correcting sound class
CN1889172A
Method and apparatus for training, classification model, mobile terminal, and readable storage medium
US20190377972A1