Object Detection Method, Apparatus, Terminal, and Computer-Readable Storage Medium
By using multiple initial classification network branches in the detection head of the target detection network to classify and predict multiple transformation proposal characteristics, the problem of poor generalization in the small sample target detection task is solved, and the accuracy of target detection is improved.
Patent Information
- Application Number
- CN202410372819.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-03-28
AI Technical Summary
The existing target detection network has poor generalization in small sample object detection tasks, resulting in a reduced accuracy of target detection.
By introducing multiple initial classification network branches into the detection head network, the various transformation proposal characteristics are classified and predicted separately, and network training is carried out based on the prediction results to enhance the feature extraction and classification prediction capabilities of the classification network.
By fully exploring the commonalities among the characteristics of multiple transformation proposals, the accuracy of object detection is improved, especially in small sample object detection tasks.
Smart Images

Figure CN118334308B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an object detection method, device, terminal, and computer-readable storage medium. Background Art
[0002] Currently, for an object detection network used to perform a small-sample object detection task, it is usually to learn general knowledge from a basic category for detecting a new category corresponding to a small sample. However, in the process of network training or machine learning of the object detection network, only the label information of a single sample is used for training, resulting in poor generalization of the classification network, thereby reducing the accuracy of object detection. Summary of the Invention
[0003] This application expects to provide an object detection method, device, terminal, and computer-readable storage medium, which can improve the accuracy of object detection.
[0004] The technical solution of this application is implemented as follows:
[0005] In a first aspect, this application provides an object detection method, and the method includes:
[0006] Performing object region detection and feature processing on an input image to obtain region proposal features;
[0007] Based on the region proposal features, performing region position correction through a regression network in a detection head network to obtain corrected proposal features;
[0008] Performing classification prediction based on the corrected proposal features through a classification network in the detection head network to obtain an object detection result; wherein,
[0009] The classification network is obtained by training at least two initial classification network branches; the at least two initial classification network branches are used to respectively perform classification prediction on at least two transformed proposal features for network training based on the obtained classification prediction results; the at least two transformed proposal features are obtained by performing at least two feature transformation processes on the same sample corrected proposal features.
[0010] In a second aspect, this application provides an object detection device, and the device includes:
[0011] A feature processing module, configured to perform object region detection and feature processing on an input image to obtain region proposal features;
[0012] A feature correction module, configured to perform region position correction based on the region proposal features through a regression network in a detection head network to obtain corrected proposal features;
[0013] A prediction module, configured to perform classification prediction based on the corrected proposal features through a classification network in the detection head network to obtain a target detection result; wherein, the classification network is obtained by training at least two initial classification network branches; the at least two initial classification network branches are configured to perform classification prediction on at least two transformed proposal features respectively, so as to perform network training based on the obtained classification prediction results; the at least two transformed proposal features are obtained by performing at least two feature transformation processes on the same sample corrected proposal feature.
[0014] In a third aspect, the present application provides a terminal, including a memory and a processor; wherein,
[0015] The memory is configured to store executable instructions;
[0016] The processor is configured to implement the target detection method provided in the embodiment of the present application when executing the executable instructions stored in the memory.
[0017] In a fourth aspect, the present application provides a computer-readable storage medium, storing executable instructions, which are used to cause a processor to implement the target detection method provided in the embodiment of the present application when executed.
[0018] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the target detection method provided in the embodiment of the present application is implemented.
[0019] The present application provides a target detection method, device, terminal and computer-readable storage medium, which perform target area detection and feature processing on an input image to obtain area proposal features; perform area position correction on the area proposal features through a regression network in the detection head network to obtain corrected proposal features; perform classification prediction on the corrected proposal features through a classification network in the detection head network to obtain a target detection result; wherein, the classification network is obtained by training at least two initial classification network branches; the at least two initial classification network branches are configured to perform classification prediction on at least two transformed proposal features respectively, so as to perform network training based on the obtained classification prediction results; the at least two transformed proposal features are obtained by performing at least two feature transformation processes on the same sample corrected proposal feature. Since the at least two transformed proposal features come from the same sample proposal feature, by performing classification prediction on the at least two transformed proposal features through the at least two initial classification network branches respectively, and training the classification network in the detection head network according to the obtained classification prediction results, the commonalities between the at least two transformed proposal features can be fully exploited, the feature extraction ability and classification prediction ability of the classification network can be enhanced, and thus the accuracy of target detection can be improved. Description of the Drawings
[0020] Figure 1 Flow schematic of the object detection method provided by the embodiment of the present application Figure 1 ;
[0021] Figure 2 Flow schematic of the object detection method provided by the embodiment of the present application Figure 2 ;
[0022] Figure 3 Schematic diagram of the process of feature transformation processing provided by the embodiment of the present application;
[0023] Figure 4 Schematic diagram of an optional network structure of the dual-path classification module provided by the embodiment of the present application;
[0024] Figure 5 Flow schematic of the object detection method provided by the embodiment of the present application Figure 3 ;
[0025] Figure 6 Schematic diagram of an optional structure of the regression network provided by the embodiment of the present application;
[0026] Figure 7 Schematic diagram of an optional network structure of the two-stage object detection network provided by the embodiment of the present application;
[0027] Figure 8 Schematic diagram of an optional network structure of the object detection network constructed in the training stage provided by the embodiment of the present application;
[0028] Figure 9 Schematic diagram of an optional structure of the object detection device provided by the embodiment of the present application;
[0029] Figure 10 Schematic diagram of an optional structure of the terminal provided by the embodiment of the present application. Detailed implementation manners
[0030] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0031] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0032] In the following description, the terms "first / second / third" only distinguish similar objects and do not represent a specific order for the objects. Understandably, "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0034] Currently, few-shot object detection aims to detect new classes by learning general knowledge from base classes. For the object detection network used in the few-shot object detection task, it usually uses the information between different instances, such as reducing the feature distance of different samples of the same class and increasing the feature distance of different samples of different classes to train the network model, so as to improve the discriminative ability of the model for samples. However, for a single sample, only simple cross-entropy and absolute error loss (L1 loss) are used for supervision, and deeper information of a single sample, such as transformation invariant information, etc. is not mined. In this way, it will cause insufficient generalization ability of the network and reduce the accuracy of object detection, especially few-shot object detection. Moreover, in the few-shot object detection network based on transfer learning, the model will first be trained on a large number of base classes; then it will be fine-tuned on a small number of new classes. In this process, the parameters of the backbone network and the region proposal network (RPN) in the object detection network are shared between the base classes and the new classes. When the detection head network in the object detection network is training on new classes, the network parameters are randomly initialized. Therefore, in the two-stage detection network of few-shot object detection, the parameters of the detection head network are not fully trained, which further reduces the accuracy of object detection.
[0035] The embodiments of the present application provide an object detection method, device, terminal and computer-readable storage medium, which can improve the accuracy of object detection, especially the accuracy of the few-shot object detection network. The object detection method of the embodiments of the present application can be as Figure 1 shown, including S101 - S103, as follows:
[0036] S101. Perform object region detection and feature processing on the input image to obtain region proposal features.
[0037] The object detection method provided by the embodiments of this application is applied to a terminal, on which an object detection network based on artificial intelligence is configured to detect information such as the position and classification of an object of interest (i.e., the object to be detected by object detection) from an input image through the object detection network. In some embodiments, the object detection network may include a two-stage object detection network. In the first stage of object detection network processing, object region detection and feature processing are performed on the input image to obtain Region Proposal features.
[0038] In some embodiments, object region detection may be performed on the input image to obtain an object region; based on the features corresponding to the object region, Region Proposal features are determined. Here, the object region represents the region of the object of interest in the input image. Exemplarily, the object of interest may be a person, an animal, an item, etc. By performing object region detection on the input object, a detection box (or bounding box) corresponding to the object of interest is obtained as the object region. Based on the image features in the detection box, Region Proposal features are determined.
[0039] In some embodiments, the object detection network may include a backbone network, a region proposal network, a second pooling network, and a detection head network. Among them, the backbone network, the region proposal network, and the second pooling network are used to implement the object detection in the first stage to obtain Region Proposal features. The detection head network is used to implement the object detection in the second stage to obtain the final object detection result based on the Region Proposal features.
[0040] Among them, the backbone network is used for feature extraction, that is, converting the input image into a high-level feature representation. The backbone network may include a pre-trained convolutional neural network. By stacking multiple convolutional and pooling layers, the spatial dimension of the image is gradually reduced while the number of channels of the features is increased, thereby converting the input image into a feature map. Exemplarily, the backbone network may include a Visual Geometry Group (VGG) network, a Residual Network (ResNet), and a MobileNet, etc., which are specifically selected according to the actual situation and are not limited in the embodiments of this application.
[0041] Among them, the region proposal network is used to generate one or more candidate regions corresponding to the object of interest, and then filter the one or more candidate regions to determine a target region corresponding to the object of interest. Exemplarily, the object region may be a detection box corresponding to the object of interest.
[0042] Among them, the second pooling network is used to perform pooling processing on the feature part corresponding to the target region detected by the RPN in the feature map corresponding to the entire input image for the target region, so as to align the feature sizes and convert it into a preset size, such as a 7×7 feature, as the region proposal feature. That is to say, the region proposal feature is the feature part corresponding to the target region in the feature map corresponding to the input image. Exemplarily, the second pooling network may include a RoI-align (Region of Interest Align) network, which is specifically selected according to the actual situation, and is not limited in the embodiments of the present application.
[0043] Among them, the detection head network is used to perform class prediction based on the region proposal feature to obtain the class information corresponding to the target region, and correct the position of the target region to obtain the position information of the target region, and combine the class information and the position information to obtain the target detection result.
[0044] In some embodiments, based on the above target detection network structure, for S101, the input image can be subjected to feature extraction through the backbone network to obtain the initial image features; through the region proposal network, target region detection can be performed based on the initial image features to obtain the target region; through the second pooling network, feature pooling can be performed based on the target region to obtain the region proposal feature.
[0045] S102. Through the regression network in the detection head network, perform region position correction based on the region proposal feature to obtain the corrected proposal feature.
[0046] In the embodiments of the present application, the detection head network includes a regression network and a classification network. Among them, the regression network is used to adjust the target region corresponding to the region proposal feature through a regression method to make it closer to the real position of the target object in the input image; Exemplarily, adjust the detection frame obtained by the target detection in the first stage to obtain the corrected target region; and based on the features corresponding to the corrected target region, obtain the corrected proposal feature. Thus, target positioning is achieved, and the position information of the target object in the input image is obtained; and the corrected proposal feature is input into the classification network, so that the classification network is used to perform classification prediction based on the corrected proposal feature to obtain the class information.
[0047] S103. Through the classification network in the detection head network, perform classification prediction based on the corrected proposal feature to obtain the target detection result; where the classification network is obtained by training at least two initial classification network branches; the at least two initial classification network branches are used to perform classification prediction on at least two transformed proposal features respectively, so as to perform network training based on the obtained classification prediction results; the at least two transformed proposal features are obtained by performing at least two feature transformation processes on the same sample corrected proposal feature.
[0048] In the embodiments of the present application, the classification network performs classification prediction based on the corrected proposal features output by the regression network, and outputs the category information corresponding to the corrected proposal features. Exemplarily, it outputs at least one probability corresponding to at least one preset category for the corrected proposal features. Combining the category information output by the classification network and the position information corresponding to the corrected target region, the final target prediction result is obtained. Exemplarily, the target prediction result may include the center point position corresponding to the detection box, the length and width dimensions of the detection box, the probability of the detection box corresponding to the preset category, etc., which are specifically selected according to the actual situation and are not limited in the embodiments of the present application.
[0049] In the embodiments of the present application, the classification network is obtained by training at least two initial classification network branches. For the training process of the classification network, an initial classification network is constructed by at least two initial classification network branches with shared parameters and the same structure, and the target detection network part in the first stage is used to detect and process the features of the target region for the sample image to obtain the sample region proposal features; through the regression network, based on the sample region proposal features, the region position is corrected to obtain the sample corrected proposal features. Here, it should be noted that the target detection network part in the first stage and the regression network may be untrained networks, that is, the target detection network part in the first stage and the regression network are trained together with the initial classification network; or, at least one of the target detection network part in the first stage and the regression network may also be a trained network, that is, the trained network is used to process the features of the sample image before the initial classification network to obtain the sample corrected proposal features input to the initial classification network. It is specifically selected according to the actual situation and is not limited in the embodiments of the present application.
[0050] In the embodiments of the present application, for the same sample corrected proposal feature, at least two feature transformation processes are performed to obtain the transformed proposal features corresponding to each feature transformation process, so as to obtain at least two transformed proposal features corresponding to the same sample corrected proposal. In some embodiments, the at least two feature transformation processes may include performing at least two masking processes on the sample corrected proposal feature, or performing at least two cropping and filling processes, or performing at least two scaling and filling or at least two scaling and cropping processes, etc., which are specifically selected according to the actual situation and are not limited in the embodiments of the present application. Through each initial classification network branch in the at least two initial classification network branches, classification prediction is performed on one of the at least two transformed proposal features to obtain the classification prediction result corresponding to the transformed proposal feature, so as to realize performing classification prediction on the at least two transformed proposal features respectively through the at least two initial classification network branches to obtain at least two classification prediction results.
[0051] It can be understood that since at least two transformed proposal features come from the same sample proposal feature, by separately classifying and predicting at least two transformed proposal features through at least two initial classification network branches, and training the classification network in the detection head network according to the obtained classification prediction results, the commonalities between at least two transformed proposal features can be fully exploited, enhancing the feature extraction ability and classification prediction ability of the classification network, thereby improving the accuracy of object detection.
[0052] In some embodiments, before classifying and predicting based on the corrected proposal features through the classification network in the detection head network to obtain the object detection result, it can be as Figure 2 shown, through the training process of S201 - S204, train the classification network in the detection head network as follows:
[0053] S201. Determine the initial classification network; the initial classification network includes at least two initial classification network branches with shared parameters.
[0054] In the embodiments of the present application, first, an initial classification network is constructed. The initial classification network includes at least two initial classification network branches with the same network structure, and the parameters of at least two initial classification network branches are shared. That is to say, during the training process, when adjusting the network parameters according to the training results, the network parameters of at least two initial classification network branches are synchronously adjusted in the same way.
[0055] S202. Use at least two masks to perform feature transformation processing on the sample corrected proposal features corresponding to each sample image in the sample image set to obtain at least two transformed proposal features corresponding to the sample corrected proposal features.
[0056] In the embodiments of the present application, the initial classification network is trained using a sample image set, and the sample image set includes at least one sample image. For each sample image in the sample image set, object region detection and feature processing are performed on each sample image to obtain the region proposal features corresponding to each sample image, and region position correction is performed based on the region proposal features corresponding to each sample image to obtain the corrected proposal features corresponding to each sample image. Here, object region detection and feature processing on each sample image can be performed through the network part in the first stage of the object detection network, such as through the backbone network and RPN; region position correction based on the region proposal features corresponding to each sample image can be performed through the regression network in the object detection network; where at least one of the backbone network, RPN, and regression network can be a trained network model or an untrained network model, and they are co - trained with the initial classification network. Specifically, it is selected according to the actual situation, and the embodiments of the present application do not make limitations.
[0057] In the embodiments of the present application, for the sample corrected proposal features corresponding to each sample image, at least two masks are used to perform feature transformation processing on them, and the same sample corrected proposal feature is transformed into at least two different transformed proposal features. Exemplarily, taking the example of using two masks to perform feature transformation processing on a sample corrected proposal feature, the two obtained transformed proposal features (transformed proposal feature 1 and transformed proposal feature 2) can be as Figure 3 shown. It can be seen that transformed proposal feature 1 and transformed proposal feature 2 respectively perform masking processing on different parts of the same feature to obtain different transformed proposal features for different initial classification network branches to process, thereby realizing explicit feature expansion and in-depth mining of single-sample features, being able to improve the value and role of a single sample during the training process, and improving the feature extraction ability and classification prediction ability of the trained network model.
[0058] S203. Through at least two initial classification network branches, respectively perform classification prediction based on at least two transformed proposal features corresponding to the sample corrected proposal features, and obtain at least two classification prediction results.
[0059] In the embodiments of the present application, through each initial network branch in at least two initial classification network branches, perform classification prediction on one of the at least two transformed proposal features, that is, each of the at least two transformed proposal features is input into each initial classification network branch one by one, and each initial classification network branch performs classification prediction based on one transformed proposal feature to obtain one classification prediction result, thereby obtaining at least two classification prediction results.
[0060] Exemplarily, the classification prediction result may include at least one probability corresponding to the transformed proposal feature and at least one preset category.
[0061] S204. Determine the training loss based on at least two classification prediction results, and according to the training loss, perform iterative training on at least two initial classification network branches until the training is completed, and determine any one of the at least two obtained classification network branches as the classification network.
[0062] In the embodiments of the present application, since the transformed proposal features processed by at least two initial classification network branches are different, the possible classification prediction results may also be different. That is to say, at least two classification prediction results represent the different classification prediction results of at least two initial classification network branches for the same sample correction proposal features from the same sample image. Based on at least two classification prediction results and in combination with the true class information (ground truth) of each sample image, the training loss can be determined. In some embodiments, the training loss corresponding to the current sample image used in the current round of training can be used to make the same synchronous adjustment to at least two initial classification network branches. Based on the at least two initial classification network branches with adjusted parameters, the next sample image in the sample images is continued to be used for the next training until the training is completed, and at least two classification network branches are obtained. Here, the network structures of at least two classification network branches are the same and the parameters are also the same. Therefore, any one of the at least two classification network branches can be determined as the classification network to be applied in the detection head network for classifying and predicting the correction proposal features corresponding to the input image.
[0063] In some embodiments, the process of determining the training loss based on at least two classification prediction results may include:
[0064] Based on at least two classification prediction results and the true classification results corresponding to each sample image, at least two classification prediction errors are determined; the relative entropy loss between at least two classification prediction results is determined; and the training loss is determined according to at least two classification prediction errors and the relative entropy loss.
[0065] Among them, for each sample image, according to each classification prediction result in at least two classification prediction results and the true classification result corresponding to the sample image, the classification prediction error corresponding to each classification prediction result is determined, so as to obtain at least two classification prediction errors. The relative entropy loss (Kullback-Leibler divergence, KL-divergence) between at least two classification prediction results is determined. Here, the relative entropy loss represents the difference between at least two classification prediction results. In combination with at least two classification prediction errors and the relative entropy loss, the training loss is determined. In this way, through network adjustment using the training loss, the error between the classification prediction result of the initial classification network branch and the true classification result can be gradually reduced, and the at least two classification prediction results of at least two initial classification network branches can gradually tend to be consistent.
[0066] In some embodiments, the sum of at least two classification prediction errors and the relative entropy loss can be used as the training loss, or, according to the actual project requirements, at least two classification prediction errors and the relative entropy loss are weighted and combined to obtain the training loss, and the specific selection is made according to the actual situation, which is not limited in the embodiments of the present application.
[0067] It can be understood that in the current detection head network, only the cross-entropy loss is used to supervise the training process of the classification network, which is far from sufficient for mining image information. In the embodiments of the present application, by combining at least two classification prediction errors and the relative entropy loss to obtain the training loss for supervising the training of the classification network, image information can be fully mined for network training, the training effect can be improved, and further the accuracy of object detection using the classification network can be improved.
[0068] In some embodiments, the initial classification network includes: a dual-branch classification module (Dual Branch Classification Module, DBCM); at least two initial classification network branches include the first initial classification network branch and the second initial classification network branch in the dual-branch classification module. At least two transformed proposal features include the first transformed proposal feature and the second transformed proposal feature. Among them, the first initial classification network branch is used to perform classification prediction on the first transformed proposal feature to obtain the first classification prediction result; the second initial classification network branch is used to perform classification prediction on the second transformed proposal feature to obtain the second classification prediction result; the determination of the relative entropy loss between at least two classification prediction results includes:
[0069] The relative entropy loss is determined by calculating the KL-divergence loss between the first classification prediction result and the second classification prediction result.
[0070] Exemplarily, the network structure of the dual-branch classification module can be as Figure 4 shown, and it can also include a first convolution module, a second convolution module, and a loss calculation module. Among them, the first convolution module is used to perform convolution processing on the first transformed proposal feature, convert it into a feature format that can be processed by the first initial classification network branch, and output it to the first initial network branch. The first initial classification network branch performs classification prediction based on the features output by the first convolution module to obtain the first classification prediction result. The second convolution module is used to perform convolution processing on the second transformed proposal feature, convert it into a feature format that can be processed by the second initial classification network branch, and output it to the second initial network branch. The second initial classification network branch performs classification prediction based on the features output by the second convolution module to obtain the second classification prediction result; the loss calculation module is used to obtain the first prediction classification error according to the first classification prediction result and the true classification result; obtain the second prediction classification error according to the second classification prediction result and the true classification result; that is to say, at least two classification prediction errors can include the first prediction classification error and the second prediction classification error. The loss calculation module is also used to calculate the KL-divergence loss between the first classification prediction result and the second classification prediction result, and combine the first prediction classification error, the second prediction classification error, and the KL-divergence loss to obtain the training loss.
[0071] It can be understood that by using at least two masks, feature transformation processing is performed on the sample correction proposal features corresponding to each sample image in the sample image set to obtain at least two transformed proposal features corresponding to the sample correction proposal features, realizing explicit feature expansion and in-depth mining of individual sample features, which can improve the value and role of a single sample during the training process; and by using at least two initial classification network branches, classification predictions are respectively performed based on at least two transformed proposal features corresponding to the sample correction proposal features, and the final training loss is obtained by calculating at least two classification prediction errors and relative entropy losses of the at least two obtained classification prediction results. Adjusting the network parameters according to the training loss can strengthen and unify the recognition ability of the network model for different transformation forms of the same sample, thereby improving the feature extraction ability and classification prediction ability of the trained classification network.
[0072] In some embodiments, the regression network in the above detection head network includes: a regression network branch independent of categories, and the process in S102 can be as Figure 5 shown, and is implemented by performing the process of S1021 - S1022 as follows:
[0073] S1021: Based on the region proposal features, a set of position offsets are predicted through the regression network branch independent of categories.
[0074] In the embodiments of the present application, the regression network branch is used to perform network inference according to the difference in the intersection over union between the current target region and the region corresponding to the real target object, and predict the position offsets to narrow the gap between the target region and the region corresponding to the real target object through the position offsets. Here, the regression network branch in the regression network of the detection head network is a regression network branch independent of categories. The regression network branch independent of categories only predicts a set of position offsets for any target, rather than predicting a set of position offsets for all preset categories.
[0075] In this way, for the training process of small samples of new categories, the regression network branch independent of categories can inherit the network parameters obtained by training on the base categories instead of randomly initializing the network parameters, thereby alleviating the data insufficiency in the regression training task, improving the network training effect, and further improving the accuracy of the trained detection head network for target detection.
[0076] S1022: Based on the position offsets, position offset correction is performed on the target region position corresponding to the region proposal features to obtain the corrected proposal features.
[0077] In the embodiments of the present application, the target region position includes the position information of the target region corresponding to the region proposal feature. Exemplarily, the target region position may include the two-dimensional coordinates of the center point of the target region. The position offset may include the offsets corresponding to the target region position in each coordinate dimension. In this way, based on the position offset, the position offset correction is performed on the target region position corresponding to the region proposal feature, and the corrected proposal feature is obtained according to the feature part corresponding to the new target region position obtained by the position offset correction in the feature map of the entire input image.
[0078] It can be understood that for the regression network branch independent of categories, the network parameters can be transferred between the base categories and the new categories, thereby improving the training effect of the regression network, enhancing the accuracy of the regression network for target localization, and further improving the accuracy of target detection based on the target region located by the regression network.
[0079] In some embodiments, in order to improve the region position regression (such as bounding box regression) effect of the regression network branch independent of categories, the regression network may further include: a confidence network branch; the process of performing the position offset correction on the target region position corresponding to the region proposal feature based on the position offset to obtain the corrected proposal feature may include:
[0080] Predicting the confidence corresponding to the target region position through the confidence network branch; the confidence represents the matching degree between the target region position and the real region position; in the case where the confidence is less than the preset confidence threshold, performing the position offset correction on the target region position based on the position offset to determine the corrected region position; and obtaining the corrected proposal feature based on the corrected region position.
[0081] In the embodiments of the present application, the confidence network branch is used to predict the matching degree between the target region position and the real region position corresponding to the target object in the input image as the confidence corresponding to the target region position. Among them, the matching degree is proportional to the confidence, and the higher the matching degree, the greater the confidence.
[0082] In some embodiments, the matching degree between the target region position and the real region position may be represented by the intersection over union. That is to say, the intersection over union between the target region position and the real region position can be predicted through the confidence network branch as the confidence corresponding to the target region position. The higher the intersection over union, the higher the matching degree between the target region position and the real region position.
[0083] In some embodiments, when the confidence level is greater than or equal to a preset confidence threshold, it indicates that the position of the target region is accurate enough and no position offset correction is required. Therefore, the position of the target region is not corrected using the position offset amount, and the region proposal feature is directly determined as the corrected proposal feature. When the confidence level is less than the preset confidence threshold, it indicates that the position of the target region is quite different from the position of the true region. It is necessary to perform position offset correction on the position of the target region based on the position offset amount, determine the corrected region position, and obtain the corrected proposal feature based on the feature part corresponding to the corrected region position in the entire feature map.
[0084] In some embodiments, the confidence network branch in the regression network of the detection head network can be trained through the following process:
[0085] Determine the initial confidence network branch and the sample target region position corresponding to the sample region proposal feature of the sample image; determine the first intersection over union (IoU) between the sample target region position and the sample true region position corresponding to the sample image; predict the second IoU between the sample target region position and the sample true region position through the initial confidence network branch; and iteratively train the initial confidence network branch according to the first IoU and the second IoU until the training is completed to obtain the confidence network branch.
[0086] Among them, the sample target region position is the position of the target region obtained by detecting the target region in the sample image; the sample true region position is the true region position corresponding to the target object in the sample image, that is, the position of the true target object annotated in the sample image. In this way, using the sample true region position as the supervision signal to predict the IoU between the sample target region position and the sample true region position to train the initial confidence network branch can enable the trained confidence network branch to predict the accuracy of the target region position and improve the accuracy of target detection.
[0087] In some embodiments, the regression network further includes: a first pooling network. The two input ends of the first pooling network are respectively connected to the output end of the confidence network branch and the output end of the category-independent regression network branch in the regression network; the first pooling network is used to obtain the confidence output by the confidence network branch and the position offset output by the category-independent regression network branch. In this way, the process of correcting the position of the target region based on the position offset to determine the position of the corrected region can be executed by the first pooling network: through the first pooling network, when the confidence is less than the preset confidence threshold, the position of the corrected region is determined by combining the position offset with the target region position corresponding to the region proposal feature. The obtaining of the corrected proposal feature based on the position of the corrected region can be executed by the first pooling network: through the first pooling network, pooling processing is performed on the feature corresponding to the position of the corrected region, that is, the feature part corresponding to the position of the corrected region in the feature map corresponding to the input image, to obtain the corrected proposal feature.
[0088] In some embodiments, the corrected proposal feature can be obtained by formula (1) as follows:
[0089]
[0090] In formula (1), f is the region proposal feature; Reg(f) is a set of position offsets predicted by the regression network branch based on the region proposal feature; C is the target region position; that is, the region position corresponding to the region proposal feature; RoI represents pooling processing, is the corrected proposal feature.
[0091] It should be noted that formula (1) is applied to the case where the confidence is less than the preset confidence threshold. When the confidence is greater than or equal to the preset confidence threshold, the first pooling network directly uses the region proposal feature f as the corrected proposal feature That is, when the confidence is greater than or equal to the preset confidence threshold, f and are the same.
[0092] Exemplarily, in the target detection network, the structure of the regression network in the detection head network can be as Figure 6 shown, where the first pooling network can include a RoI-align network, which is specifically selected according to the actual situation, and is not limited in the embodiments of the present application. Based on Figure 6 this, the region proposal feature can also be expanded through a third convolution module, and the expanded region proposal features are respectively input into the category-independent regression network branch and the confidence network branch for network inference, which is specifically selected according to the actual situation, and is not limited in the embodiments of the present application.
[0093] Exemplarily, the network structure of a two-stage target detection network provided by the embodiments of the present application can be asFigure 7 As shown in the figure. Among them, the regression network in the detection head network can be implemented as a Feature Refinement Module (FRM). For the object detection in the first stage, the backbone network is used to extract features from the input image to obtain the feature map corresponding to the input image; the region proposal network is used to detect the target region of the object from the feature map corresponding to the input image to obtain the target region corresponding to the target object; the second pooling network is used to perform feature processing on the feature part corresponding to the target region in the feature map to obtain the region proposal feature. For the object detection in the second stage, the third convolutional module in the feature correction module is used to flatten the region proposal feature and output it to the regression network branch and the confidence network branch that are independent of the category in the feature correction module respectively; a set of position offsets are predicted through the regression network branch that is independent of the category; the confidence is predicted through the confidence network branch; through the first pooling network, when the confidence is less than the preset confidence threshold, the position of the target region corresponding to the region proposal feature is corrected by using the position offset to obtain the corrected proposal feature. Through the classification network, classification prediction is performed based on the corrected proposal feature to obtain the final object detection result, such as the detection box and the position of the detection box in the input image where the classification prediction probability of the preset category "sheep" is greater than the preset threshold.
[0094] Exemplarily, based on Figure 7 the object detection network shown in the figure, the network structure constructed in its training stage can be as Figure 8As shown in the figure. Among them, in the training stage for a new category, the category-agnostic initial regression network branch is initialized with the network parameters obtained by inheriting the training of the base category during the training stage. For one round of training, the sample image passes through the backbone network, RPN, and the second pooling network to obtain the sample region proposal features. Through the category-agnostic initial regression network branch, a set of sample position offsets are predicted based on the sample region proposal features; through the initial confidence network, the second intersection over union (IoU) of the sample target region position corresponding to the sample region proposal features is predicted; the first IoU between the sample target region position and the sample ground-truth region position corresponding to the sample image is determined, and based on the first IoU and the second IoU, the training loss corresponding to the initial confidence network in this training is obtained, and the network parameters of the initial confidence network are adjusted. Through the first pooling network, when the second IoU is greater than the preset confidence threshold, the sample corrected region position is obtained according to the sample position offset and the sample target region position, and based on the feature part corresponding to the sample corrected region position in the sample feature map of the sample image, the sample corrected proposal features are obtained. Among them, the training loss corresponding to the category-agnostic initial regression network branch can be obtained according to the difference between the sample corrected region position and the sample ground-truth region position, and the network parameters of the category-agnostic initial regression network branch are adjusted. Using two different masks, the sample corrected proposal features are subjected to feature transformation processing to obtain the first transformed proposal features and the second transformed proposal features. Through the first convolutional module and the first initial classification network branch in the DBCM module, the first transformed proposal features are classified and predicted to obtain the first classification prediction result; through the second convolutional module and the second initial classification network branch in the DBCM module, the second transformed proposal features are classified and predicted to obtain the second classification prediction result. Through the loss calculation module in the DBCM module, the classification prediction loss and the relative entropy loss are calculated based on the first classification prediction result and the second classification prediction result to obtain the training loss corresponding to the DBCM module, and the network parameters of the DBCM module are adjusted. In this way, through the above process of iterative training, a category-agnostic regression network branch, a confidence network branch, and at least two classification network branches can be obtained. Among them, the category-agnostic regression network branch and the confidence network branch are regression networks obtained through training, and any one of the at least two classification network branches is used as the classification network, so as to combine the regression network and the classification network to obtain the detection head network.
[0095] It should be noted that Figure 8 Only an example of a case where the regression network (feature correction module) and the classification network (dual-path classification module) in the detection head network are trained simultaneously is shown. In actual applications, the regression network and the classification network can also be independently trained, and specific selection is made according to the actual situation, which is not limited in the embodiments of the present application.
[0096] The embodiment of the present application further provides an object detection device. Figure 9 It is a schematic structural diagram of the object detection device provided by the embodiment of the present application. As Figure 9 shown, the object detection device 1 includes: a feature processing module 11, a feature correction module 12, and a prediction module 13. Among them:
[0097] The feature processing module 11 is configured to perform object region detection and feature processing on the input image to obtain region proposal features;
[0098] The feature correction module 12 is configured to perform region position correction based on the region proposal features through a regression network in the detection head network to obtain corrected proposal features;
[0099] The prediction module 13 is configured to perform classification prediction based on the corrected proposal features through a classification network in the detection head network to obtain an object detection result; wherein, the classification network is obtained by training at least two initial classification network branches; the at least two initial classification network branches are respectively configured to perform classification prediction on at least two transformed proposal features to perform network training based on the obtained classification prediction results; the at least two transformed proposal features are obtained by performing at least two feature transformation processes on the same sample corrected proposal feature.
[0100] In some embodiments, the object detection device 1 further includes: a training module, and the training module is configured to determine an initial classification network; the initial classification network includes the at least two initial classification network branches with shared parameters; use at least two masks to perform feature transformation processes on the sample corrected proposal features corresponding to each sample image in the sample image set to obtain at least two transformed proposal features corresponding to the sample corrected proposal features; respectively perform classification prediction on the at least two transformed proposal features corresponding to the sample corrected proposal features through the at least two initial classification network branches to obtain at least two classification prediction results; determine a training loss based on the at least two classification prediction results, and perform iterative training on the at least two initial classification network branches according to the training loss until the training is completed, and determine any one of the at least two classification network branches obtained as the classification network.
[0101] In some embodiments, the training module is further configured to determine at least two classification prediction errors based on the at least two classification prediction results and the true category information corresponding to each sample image; determine the relative entropy loss between the at least two classification prediction results; and determine the training loss according to the at least two classification prediction errors and the relative entropy loss.
[0102] In some embodiments, the initial classification network includes: a dual-path classification module; the at least two initial classification network branches include a first initial classification network branch and a second initial classification network branch in the dual-path classification module; the first initial classification network branch is configured to perform classification prediction on a first transformed proposal feature corresponding to the sample correction proposal feature to obtain a first classification prediction result; the second initial classification network branch is configured to perform classification prediction on a second transformed proposal feature corresponding to the sample correction proposal feature to obtain a second classification prediction result; the training module is further configured to determine the relative entropy loss by calculating the KL divergence loss between the first classification prediction result and the second classification prediction result.
[0103] In some embodiments, the regression network includes: a class-agnostic regression network branch, and the feature correction module 12 is further configured to, through the class-agnostic regression network branch, predict a set of position offsets based on the region proposal feature; and perform position offset correction on the target region position corresponding to the region proposal feature based on the position offsets to obtain the correction proposal feature.
[0104] In some embodiments, the regression network further includes: a confidence network branch; the feature correction module 12 is further configured to, through the confidence network branch, predict the confidence corresponding to the target region position; the confidence characterizes the matching degree between the target region position and the true region position; in the case where the confidence is less than a preset confidence threshold, perform position offset correction on the target region position based on the position offsets to determine a corrected region position; and obtain the correction proposal feature based on the corrected region position.
[0105] In some embodiments, the feature correction module 12 is further configured to, in the case where the confidence is greater than or equal to the preset confidence threshold, determine the region proposal feature as the correction proposal feature.
[0106] In some embodiments, the training module is further configured to determine an initial confidence network branch and a sample target region position corresponding to the sample region proposal feature of the sample image; determine a first intersection over union between the sample target region position and the sample true region position corresponding to the sample image; predict a second intersection over union between the sample target region position and the sample true region position through the initial confidence network branch; and perform iterative training on the initial confidence network branch according to the first intersection over union and the second intersection over union until the training is completed to obtain the confidence network branch.
[0107] In some embodiments, the regression network further includes: a first pooling network; the first pooling network is configured to obtain the confidence level output by the confidence network branch and the position offset output by the class-agnostic regression network branch; the feature correction module 12 is further configured to, through the first pooling network, when the confidence level is less than a preset confidence threshold, combine the position offset with the target region position to determine the corrected region position; and through the first pooling network, perform pooling processing on the features corresponding to the corrected region position to obtain the corrected proposal features.
[0108] In some embodiments, the target detection device 1 is applied to a target detection network, and the target detection network includes: a backbone network, a region proposal network, a second pooling network, and the detection head network; the feature processing module 11 is further configured to, through the backbone network, perform feature extraction on the input image to obtain initial image features; through the region proposal network, perform target region detection based on the initial image features to obtain target regions; and through the second pooling network, perform feature pooling based on the target regions to obtain the region proposal features.
[0109] It should be noted that the description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0110] An embodiment of the present application further provides a terminal Figure 10 which is an optional structural schematic diagram of the terminal provided by the embodiment of the present application. As Figure 10 shown, the terminal 3 includes: a memory 32 and a processor 33. Among them, the memory 32 and the processor 33 are connected through a communication bus 34; the memory 32 is used to store executable instructions; the processor 33 is configured to, when executing the executable instructions stored in the memory 32, implement the target detection method provided by the embodiment of the present application.
[0111] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, where the executable instructions, when executed by the above-mentioned processor, will cause the above-mentioned processor to execute the target detection method provided by the embodiment of the present application.
[0112] In some embodiments, the computer-readable storage medium may be a ferroelectric random access memory (FRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disc, or a memory such as a CD-ROM; it may also be various devices including one or any combination of the above memories.
[0113] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0114] As an example, the executable instructions may or may not correspond to a file in the file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a hypertext markup language (HTML) document, stored in a single file dedicated to the program being discussed, or, stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or portions of code). As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one site, or, on multiple computing devices distributed at multiple sites and interconnected by a communication network.
[0115] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, system, or computer program product. Therefore, the present application may take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) that contain computer-usable program code.
[0116] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in one or more flows and / or blocks Figure 1 one or more flows and / or blocks Figure 1 or in a plurality of blocks.
[0117] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one or more flows and / or blocks Figure 1 one or more flows and / or blocks Figure 1 or in a plurality of blocks.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or blocks Figure 1 one or more flows and / or blocks Figure 1 or in a plurality of blocks.
[0119] As mentioned above, it is only a preferred embodiment of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A target detection method, characterized in that: include: Perform target region detection and feature processing on the input image to obtain region proposal features; By using a regression network in the detection head network, the region position is corrected based on the region proposal feature to obtain a corrected proposal feature; Through the classification network in the detection head network, classification prediction is performed based on the correction proposal features to obtain the target detection result; wherein, The classification network is obtained by training at least two initial classification network branches; the at least two initial classification network branches are used to respectively perform classification prediction on at least two conversion proposal features, so as to perform network training based on the obtained classification prediction results; the at least two conversion proposal features are obtained by performing at least two feature conversion processes on the same sample correction proposal feature; Wherein, the method further comprises: Determine an initial classification network; the initial classification network includes at least two initial classification network branches sharing parameters; Using at least two masks, perform feature conversion processing on the sample correction proposal feature corresponding to each sample image in the sample image set to obtain at least two conversion proposal features corresponding to the sample correction proposal feature; By using each of the at least two initial classification network branches, a classification prediction is performed on one of the at least two conversion proposal features corresponding to the sample correction proposal feature to obtain at least two classification prediction results; Determine a training loss based on the at least two classification prediction results, perform iterative training on the at least two initial classification network branches according to the training loss until the training is completed, and determine any one of the at least two classification network branches obtained as the classification network; Wherein, determining the training loss based on the at least two classification prediction results includes: Determining at least two classification prediction errors based on the at least two classification prediction results and the true category information corresponding to each sample image; Determining a relative entropy loss between the at least two classification prediction results; the relative entropy loss characterizes a difference between the at least two classification prediction results; The training loss is determined according to the at least two classification prediction errors and the relative entropy loss.
2. The method according to claim 1, characterized in that The regression network includes: a category-independent regression network branch, and the regression network in the detection head network corrects the region position based on the region proposal feature to obtain a corrected proposal feature, including: Predicting a set of position offsets based on the region proposal features through the category-independent regression network branch; Based on the position offset, the position of the target area corresponding to the area proposal feature is corrected to obtain the corrected proposal feature.
3. The method according to claim 2, characterized in that The regression network further includes: a confidence network branch; the position offset correction of the target area position corresponding to the area proposal feature based on the position offset to obtain the corrected proposal feature includes: Predicting the confidence corresponding to the target area position through the confidence network branch; the confidence represents the degree of matching between the target area position and the real area position; When the confidence is less than a preset confidence threshold, the position of the target area is corrected based on the position offset, a correction area position is determined, and the correction proposal feature is obtained based on the correction area position; Or, when the confidence is greater than or equal to the preset confidence threshold, the region proposal feature is determined as the revised proposal feature.
4. The method according to claim 3, characterized in that The method further comprises: Determine the initial confidence network branch and the sample target region position corresponding to the sample region proposal feature of the sample image; Determine a first intersection-over-union ratio between the sample target area position and the sample real area position corresponding to the sample image; Predicting a second intersection-over-union ratio between the sample target region position and the sample true region position through the initial confidence network branch; The initial confidence network branch is iteratively trained according to the first intersection-over-union ratio and the second intersection-over-union ratio until the training is completed to obtain the confidence network branch.
5. The method according to claim 3 or 4, characterized in that: The regression network further includes: a first pooling network; the first pooling network is used to obtain the confidence output by the confidence network branch and the position offset output by the category-independent regression network branch; when the confidence is less than a preset confidence threshold, the position offset of the target area is corrected based on the position offset to determine the position of the correction area, including: By using the first pooling network, when the confidence is less than a preset confidence threshold, the position offset and the target area position are combined to determine the correction area position; The obtaining of the correction proposal feature based on the correction area position includes: The features corresponding to the position of the correction area are pooled through the first pooling network to obtain the correction proposal features.
6. The method according to claim 1, or claim 3, or claim 4, characterized in that: The target detection method is applied to a target detection network, which includes: a backbone network, a region proposal network, a second pooling network, and the detection head network; the target region detection and feature processing of the input image to obtain the region proposal feature includes: Extracting features of the input image through a backbone network to obtain initial image features; By using a region proposal network, target region detection is performed based on the initial image features to obtain the target region; Through the second pooling network, feature pooling is performed based on the target area to obtain the area proposal feature.
7. A target detection device, characterized in that: include: The feature processing module is used to detect and process the target area of the input image to obtain the region proposal features; A feature correction module, used for correcting the region position based on the region proposal feature through a regression network in the detection head network to obtain a corrected proposal feature; A prediction module, configured to perform classification prediction based on the revised proposal features through a classification network in the detection head network to obtain a target detection result; wherein the classification network is obtained by training at least two initial classification network branches; the at least two initial classification network branches are used to perform classification prediction on at least two conversion proposal features respectively, so as to perform network training based on the obtained classification prediction results; the at least two conversion proposal features are obtained by performing at least two feature conversion processes on the revised proposal features of the same sample; Wherein, the target detection device also includes a training module, and the training module is used to determine an initial classification network; the initial classification network includes the at least two initial classification network branches that share parameters; using at least two masks, feature conversion processing is performed on the sample correction proposal feature corresponding to each sample image in the sample image set to obtain at least two conversion proposal features corresponding to the sample correction proposal feature; through each initial network branch in the at least two initial classification network branches, one of the at least two conversion proposal features corresponding to the sample correction proposal feature is classified and predicted to obtain at least two classification prediction results; based on the at least two classification prediction results, the training loss is determined, and according to the training loss, the at least two initial classification network branches are iteratively trained until the training is completed, and any one of the at least two classification network branches obtained is determined as the classification network; Among them, the training module is also used to determine at least two classification prediction errors based on the at least two classification prediction results and the true category information corresponding to each sample image; determine the relative entropy loss between the at least two classification prediction results; the relative entropy loss characterizes the difference between the at least two classification prediction results; and determine the training loss based on the at least two classification prediction errors and the relative entropy loss.
8. A terminal, characterized in that: include: A memory and a processor; wherein, The memory is used to store executable instructions; The processor is used to implement the method according to any one of claims 1 to 6 when executing the executable instructions stored in the memory.
9. A computer-readable storage medium, characterized in that: Executable instructions are stored, and when the executable instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Image processing method and device and related equipment
CN115984537A