Network Training and Object Detection Methods, Devices, Electronic Devices, and Storage Media
By calculating the matching loss between the anchor box area and the pseudo-notation area in the target detection network, it is used to train the target detection network, and the problem of matching the pseudo-label box to the suboptimal anchor box is solved, and the training efficiency of the target detection network is improved.
Patent Information
- Application Number
- CN202210431662.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-04-22
AI Technical Summary
In the scenario where the object detection network is trained based on the semi-supervised training method, since some sample images use inaccurate pseudo-label boxes, this method of matching based on the interleaving ratio will make the pseudo-label boxes match the suboptimal anchor boxes, affecting the training effect and training efficiency of the object detection network.
A network training method is proposed, by obtaining sample images without labels and the object detection network to be trained, processing the sample images, determining the pseudo-label information, and calculating matching losses based on the pseudo-label information and anchor box information, which is used to train the object detection network.
By using the network loss obtained by using the matching loss between the anchor box area and the pseudo-notation area, training the target detection network can reduce the matching inconsistency and inaccurate matching when the anchor box area matches the pseudo-notation area, and improving the training efficiency of the target detection network.
Smart Images

Figure CN114722958B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a network training and object detection method, apparatus, electronic device, and storage medium. Background Art
[0002] With the wide application of deep learning in object detection technology, the object detection technology has formed a paradigm in which features are extracted by a deep network, and the object category and location are obtained by classifying and regressing the features. Currently, the training method of the object detection network mainly matches the anchor box with the label box by means of the Intersection over Union (IoU), and then, taking the anchor box as a reference, learns the relative position between the object label box and the anchor box.
[0003] However, in the scenario of training an object detection network based on a semi-supervised training method, since some sample images use inaccurate pseudo-label boxes, this method of matching according to the IoU will cause the pseudo-label boxes to be matched to sub-optimal anchor boxes. For example, the matched anchor box may not be on the object, or the matched anchor box may not completely cover the entire object, which affects the training effect of the object detection network and reduces the training efficiency. Summary of the Invention
[0004] The present disclosure provides a technical solution for network training and object detection.
[0005] According to an aspect of the present disclosure, there is provided a network training method, including: obtaining a sample image without label annotation and an object detection network to be trained; processing the sample image through the object detection network to obtain a detection result of the sample image, where the detection result includes a predicted region where an object is located in the sample image; determining, according to the detection result, pseudo-label information corresponding to the sample image, where the pseudo-label information includes a pseudo-annotation region where the object is located in the sample image; determining a matching loss between the pseudo-annotation region and an anchor box region represented by the anchor box information according to the pseudo-label information and the anchor box information preset in the object detection network; determining a network loss of the object detection network according to the matching loss, and training the object detection network based on the network loss.
[0006] In a possible implementation, the pseudo-label information includes a plurality of pseudo-annotation regions. Among them, determining the network loss of the target detection network according to the matching loss includes: for each pseudo-annotation region in the plurality of pseudo-annotation regions, determining the first weight corresponding to each of the plurality of anchor regions according to the matching loss between the pseudo-annotation region and the plurality of anchor regions, and the first weight of the same anchor region is negatively correlated with the matching loss; determining the region loss corresponding to the pseudo-annotation region according to the first weight corresponding to each of the plurality of anchor regions and the matching loss corresponding to each of the plurality of anchor regions; determining the network loss of the target detection network according to the obtained region losses corresponding to the plurality of pseudo-annotation regions.
[0007] In a possible implementation, the pseudo-label information includes the confidence corresponding to the pseudo-annotation region, and the confidence characterizes the reliability of the pseudo-annotation region. Among them, determining the network loss of the target detection network according to the obtained region losses corresponding to the plurality of pseudo-annotation regions includes: determining the second weight corresponding to each of the plurality of pseudo-annotation regions according to the confidence corresponding to each of the plurality of pseudo-annotation regions, and the second weight is positively correlated with the confidence; determining the network loss of the target detection network according to the second weight corresponding to each of the plurality of pseudo-annotation regions and the region loss corresponding to each of the plurality of pseudo-annotation regions.
[0008] In a possible implementation, the matching loss includes at least one of the following: the classification loss between the pseudo-annotation region and the anchor region; the first position loss between the pseudo-annotation region and the predicted region corresponding to the anchor region; the second position loss between the pseudo-annotation region and the anchor region.
[0009] In a possible implementation, the anchor information includes the classification score predicted by the target detection network for the anchor region, and the pseudo-label information includes the pseudo-label value corresponding to the pseudo-annotation region, and the pseudo-label value is used to indicate whether the target is included in the pseudo-annotation region, and the classification score characterizes the probability that the target is included in the anchor region; among them, determining the matching loss between the pseudo-annotation region and the anchor region characterized by the anchor information according to the pseudo-label information and the anchor information preset in the target detection network includes:
[0010] Determine the classification loss between the pseudo-labeled region and the anchor box region according to the classification score corresponding to the anchor box region and the pseudo-label value corresponding to the pseudo-labeled region; and / or, determine the first position loss between the pseudo-labeled region and the prediction region corresponding to the anchor box region according to the position information of the prediction region corresponding to the anchor box region and the position information of the pseudo-labeled region; and / or, determine the second position loss between the pseudo-labeled region and the anchor box region according to the position information of the anchor box region and the position information of the pseudo-labeled region; determine the matching loss between the pseudo-labeled region and the anchor box region according to at least one of the classification loss, the first position loss, and the second position loss.
[0011] In a possible implementation, the detection result includes prediction scores corresponding to multiple prediction regions respectively, and the prediction score represents the probability that the target is included in the prediction region; wherein, determining the pseudo-label information corresponding to the sample image according to the detection result includes: clustering the multiple prediction regions according to the prediction scores corresponding to the multiple prediction regions respectively to obtain a clustering result of the multiple prediction regions; determining a screening threshold according to the clustering result of the multiple prediction regions, and screening out pseudo-labeled regions from the multiple prediction regions based on the screening threshold; determining the pseudo-label information corresponding to the pseudo-labeled region according to the detection result corresponding to the pseudo-labeled region.
[0012] In a possible implementation, clustering the multiple prediction regions according to the prediction scores corresponding to the multiple prediction regions respectively to obtain a clustering result of the multiple prediction regions includes: clustering the multiple prediction regions according to the prediction scores corresponding to the multiple prediction regions respectively by using a Gaussian mixture model to obtain a clustering result of the multiple prediction regions; wherein, the clustering result includes positive sample prediction regions and negative sample prediction regions among the multiple prediction regions, and the accuracy of the positive sample prediction regions is higher than that of the negative sample prediction regions.
[0013] In a possible implementation, the clustering result includes positive sample prediction regions among the multiple prediction regions, wherein, determining a screening threshold according to the clustering result of the multiple prediction regions, and screening out pseudo-labeled regions from the multiple prediction regions based on the screening threshold includes: determining the prediction score corresponding to the peak value of the Gaussian distribution corresponding to the positive sample prediction region as the screening threshold; determining the positive sample prediction regions with prediction scores greater than or equal to the screening threshold as the pseudo-labeled regions.
[0014] In a possible implementation, processing the sample image through the target detection network to obtain the detection result of the sample image includes: extracting features from the sample image to obtain classification features and regression features; performing decoding processing on the regression features and the classification features to obtain an initial prediction region and a prediction score corresponding to the initial prediction region; determining an offset corresponding to the initial prediction region according to the regression features and the classification features; and adjusting the initial prediction region according to the offset to obtain a prediction region.
[0015] In a possible implementation, determining the offset corresponding to the initial prediction region according to the regression features and the classification features includes: splicing and convolving the classification features and the regression features to obtain multi-scale bias features; performing scale alignment on the multi-scale bias features to obtain a target bias feature with scale alignment; and performing decoding processing on the target bias feature to obtain the offset corresponding to the initial prediction region.
[0016] According to one aspect of the present disclosure, a target detection method is provided, including: processing a to-be-detected image through a target detection network to obtain a target detection result of the to-be-detected image, where the target detection result includes a region where a target is located in the to-be-detected image; and the target detection network is trained according to the above network training method.
[0017] In a possible implementation, the target includes a vehicle, and processing the to-be-detected image through the target detection network to obtain the target detection result of the to-be-detected image includes: processing the to-be-detected image through the target detection network to obtain a vehicle region where the vehicle is located in the to-be-detected image; and after obtaining the vehicle region, the method further includes: performing vehicle recognition on the vehicle region to obtain a vehicle recognition result of the vehicle in the vehicle region.
[0018] According to one aspect of the present disclosure, there is provided a network training device, including: an acquisition module configured to acquire sample images without labeled annotations and a target detection network to be trained; a detection module configured to process the sample images through the target detection network to obtain detection results of the sample images, where the detection results include predicted regions where targets are located in the sample images; a pseudo-label determination module configured to determine pseudo-label information corresponding to the sample images according to the detection results, where the pseudo-label information includes pseudo-annotation regions where targets are located in the sample images; a matching loss determination module configured to determine a matching loss between the pseudo-annotation regions and anchor box regions represented by the anchor box information in the target detection network according to the pseudo-label information and the anchor box information preset in the target detection network; and a training module configured to determine a network loss of the target detection network according to the matching loss and train the target detection network based on the network loss.
[0019] In a possible implementation manner, the pseudo-label information includes multiple pseudo-annotation regions. Wherein, the training module includes: a weight determination sub-module configured to, for each pseudo-annotation region among the multiple pseudo-annotation regions, determine respective first weights corresponding to the multiple anchor box regions according to the matching loss between the pseudo-annotation region and the multiple anchor box regions, and the first weights of the same anchor box region are negatively correlated with the matching loss; a region loss determination sub-module configured to determine a region loss corresponding to the pseudo-annotation region according to the respective first weights corresponding to the multiple anchor box regions and the respective matching losses corresponding to the multiple anchor box regions; and a network loss determination sub-module configured to determine the network loss of the target detection network according to the obtained region losses corresponding to the multiple pseudo-annotation regions.
[0020] In a possible implementation manner, the pseudo-label information includes a confidence level corresponding to the pseudo-annotation region, and the confidence level represents the reliability degree of the pseudo-annotation region. Wherein, determining the network loss of the target detection network according to the obtained region losses corresponding to the multiple pseudo-annotation regions includes: determining respective second weights corresponding to the multiple pseudo-annotation regions according to the confidence levels corresponding to the multiple pseudo-annotation regions, and the second weights are positively correlated with the confidence levels; and determining the network loss of the target detection network according to the respective second weights corresponding to the multiple pseudo-annotation regions and the respective region losses corresponding to the multiple pseudo-annotation regions.
[0021] In a possible implementation manner, the matching loss includes at least one of the following: a classification loss between the pseudo-annotation region and the anchor box region; a first position loss between the pseudo-annotation region and a predicted region corresponding to the anchor box region; and a second position loss between the pseudo-annotation region and the anchor box region.
[0022] In a possible implementation, the anchor box information includes the classification scores predicted by the target detection network for the anchor box regions, and the pseudo-label information includes the pseudo-label values corresponding to the pseudo-labeled regions. The pseudo-label values are used to indicate whether the target is included in the pseudo-labeled regions, and the classification scores represent the probabilities that the target is included in the anchor box regions. Among them, the matching loss determination module includes: a classification loss determination sub-module, configured to determine the classification loss between the pseudo-labeled region and the anchor box region according to the classification scores corresponding to the anchor box region and the pseudo-label values corresponding to the pseudo-labeled region; and / or, a first position loss determination sub-module, configured to determine the first position loss between the pseudo-labeled region and the predicted region corresponding to the anchor box region according to the position information of the predicted region corresponding to the anchor box region and the position information of the pseudo-labeled region; and / or, a second position loss determination sub-module, configured to determine the second position loss between the pseudo-labeled region and the anchor box region according to the position information of the anchor box region and the position information of the pseudo-labeled region; and determine the matching loss between the pseudo-labeled region and the anchor box region according to at least one of the classification loss, the first position loss, and the second position loss.
[0023] In a possible implementation, the detection results include the prediction scores corresponding to multiple prediction regions respectively, and the prediction scores represent the probabilities that the target is included in the prediction regions. Among them, the pseudo-label determination module includes: a clustering sub-module, configured to cluster the multiple prediction regions according to the prediction scores corresponding to the multiple prediction regions respectively to obtain the clustering results of the multiple prediction regions; a screening sub-module, configured to determine a screening threshold according to the clustering results of the multiple prediction regions, and based on the screening threshold, screen out the pseudo-labeled regions from the multiple prediction regions; and a root pseudo-label determination sub-module, configured to determine the pseudo-label information corresponding to the pseudo-labeled regions according to the detection results corresponding to the pseudo-labeled regions.
[0024] In a possible implementation, the clustering of the multiple prediction regions according to the prediction scores corresponding to the multiple prediction regions respectively to obtain the clustering results of the multiple prediction regions includes: clustering the multiple prediction regions according to the prediction scores corresponding to the multiple prediction regions respectively by using a Gaussian mixture model to obtain the clustering results of the multiple prediction regions; where the clustering results include positive sample prediction regions and negative sample prediction regions among the multiple prediction regions, and the accuracy of the positive sample prediction regions is higher than that of the negative sample prediction regions.
[0025] In a possible implementation, the clustering result includes positive sample prediction regions among the multiple prediction regions. Wherein, determining a screening threshold according to the clustering result of the multiple prediction regions and screening out pseudo-labeled regions from the multiple prediction regions based on the screening threshold includes: determining the prediction score corresponding to the peak of the Gaussian distribution corresponding to the positive sample prediction region as the screening threshold; determining the positive sample prediction regions with prediction scores greater than or equal to the screening threshold as the pseudo-labeled regions.
[0026] In a possible implementation, the detection module includes: a feature extraction sub-module for extracting features from the sample image to obtain classification features and regression features; a decoding sub-module for performing decoding processing on the regression features and the classification features to obtain an initial prediction region and a prediction score corresponding to the initial prediction region; an offset determination sub-module for determining an offset corresponding to the initial prediction region according to the regression features and the classification features; an adjustment sub-module for adjusting the initial prediction region according to the offset to obtain a prediction region.
[0027] In a possible implementation, determining an offset corresponding to the initial prediction region according to the regression features and the classification features includes: splicing and convolving the classification features and the regression features to obtain multi-scale bias features; performing scale alignment on the multi-scale bias features to obtain a target bias feature with scale alignment; performing decoding processing on the target bias feature to obtain an offset corresponding to the initial prediction region.
[0028] According to one aspect of the present disclosure, there is provided an object detection device, including: an image detection module for processing a to-be-detected image through an object detection network to obtain an object detection result of the to-be-detected image, where the object detection result includes a region where an object is located in the to-be-detected image; wherein, the object detection network is trained according to the above network training method.
[0029] In a possible implementation, the object includes a vehicle, and the image detection module is specifically configured to process the to-be-detected image through the object detection network to obtain a vehicle region where the vehicle is located in the to-be-detected image; wherein, after obtaining the vehicle region, the device further includes: a vehicle recognition module for performing vehicle recognition on the vehicle region to obtain a vehicle recognition result of the vehicle in the vehicle region.
[0030] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above method.
[0031] According to one aspect of the present disclosure, there is provided a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the above method is implemented.
[0032] In the embodiments of the present disclosure, by using the network loss obtained from the matching loss between the anchor box region and the pseudo-label region to train the object detection network, it is possible to reduce the influence of inconsistent or inaccurate matching when the anchor box region matches the pseudo-label region on the training effect of the object detection network, and improve the training efficiency of the object detection network.
[0033] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. Description of the Drawings
[0034] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0035] Figure 1 A schematic diagram showing an anchor box matching according to the related art.
[0036] Figure 2 A flowchart showing a network training method according to an embodiment of the present disclosure.
[0037] Figure 3 A schematic diagram showing a pseudo-label box according to the related art.
[0038] Figure 4 A schematic diagram showing a processing of an object detection network according to an embodiment of the present disclosure.
[0039] Figure 5 A block diagram showing a network training apparatus according to an embodiment of the present disclosure.
[0040] Figure 6 A block diagram showing an electronic device 1900 according to an embodiment of the present disclosure. Detailed Embodiments
[0041] The following will detail various exemplary embodiments, features, and aspects of the present disclosure with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.
[0042] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein need not be construed as superior or better than other embodiments.
[0043] As used herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" as used herein means any one or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent any one or more elements selected from the set composed of A, B, and C.
[0044] In addition, to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail in order to highlight the gist of the present disclosure.
[0045] It is known that in the semi-supervised training method, usually a network model is pre-trained using labeled sample data first, then the pre-trained network model is used to process unlabeled sample data to obtain the detection results of the unlabeled sample data, and then relatively accurate detection results are selected from these detection results based on a preset screening threshold as the pseudo-labels of these unlabeled sample data, and these pseudo-labels are used to continue training the network model until the network model converges.
[0046] As described above, the current training method of the object detection network mainly matches the anchor boxes with the label boxes by means of the intersection over union (IoU), and then takes the anchor boxes as a reference to learn the relative positions of the object label boxes and the anchor boxes. In the scenario of training an object detection network based on the semi-supervised training method, there is noise in the pseudo-label boxes selected from the detection results output by the object detection network, or in other words, the pseudo-label boxes obtained in different training stages are inconsistent. Then, the method of matching anchor boxes based on IoU will not only cause inconsistent pseudo-label boxes to be matched with inconsistent anchor boxes, but also cause the pseudo-label boxes to be matched with sub-optimal anchor boxes. For example, the matched anchor boxes may not be on the labeled object, or the matched anchor boxes may not completely cover the entire labeled object. Figure 1 A schematic diagram of an anchor box matching according to the related art is shown, as Figure 1 shown. When the sample image adopts the pseudo-label box as shown in Figure 1 (i.e., the solid line box in Figure 1 ), the anchor boxes matched based on the IoU method (i.e., Figure 1The two dashed boxes in it cannot completely cover the object, which will affect the training effect of the target detection network and reduce the training efficiency of the target detection network.
[0047] Figure 2 FIG. shows a flowchart of a network training method according to an embodiment of the present disclosure. The network training method can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be implemented by a processor calling computer-readable instructions stored in a memory, or the method can be executed by a server. As Figure 2 shown, the network training method includes:
[0048] In step S11, obtain a sample image without label annotation and a target detection network to be trained.
[0049] Among them, the sample image can be a large number of images collected by an image acquisition device (such as a camera), or an open-source image set obtained. The sample image contains detectable targets, and the detectable targets include, for example, human bodies, faces, vehicles, objects, etc. It should be understood that the present disclosure embodiment does not limit the acquisition method of the sample image and the types of targets in the sample image.
[0050] In a possible implementation manner, a target detection network can be pre-set. The target detection network can be a network pre-trained by using some sample images with label annotation. The target detection network can be used to detect the region where at least one target is located in the image (that is, the position where the target is located).
[0051] Among them, the target detection network can adopt, for example, RCNN (Region Convolutional Neural Networks), Fast RCNN (Fast Region Convolutional Neural Networks), Faster RCNN (Faster Region Convolutional Neural Networks), etc. Of course, the target detection network can also be a neural network independently developed by those skilled in the art. It should be understood that the present disclosure embodiment does not limit the network structure of the target detection network.
[0052] In step S12, the sample image is processed by a target detection network to obtain the detection result of the sample image.
[0053] Among them, processing the sample image by the target detection network to obtain the detection result of the sample image can be understood as inputting the sample image into the target detection network and outputting the detection result of the sample image. It should be understood that the targets in the sample image can include at least one type, and each type of target in the sample image can include at least one. Therefore, each sample image can correspond to at least one prediction region.
[0054] In a possible implementation manner, the detection result can include at least one of the prediction region where the target is located in the sample image, the prediction score corresponding to the prediction region, and the confidence corresponding to the prediction region. Among them, the prediction score corresponding to the prediction region can represent the probability that the prediction region contains a target, the confidence corresponding to the prediction region can represent the reliability of the prediction region, and the prediction region can be understood as the region where the target predicted by the target detection network is located in the sample image, and the prediction region can be represented by a prediction box (or detection box), that is, the prediction region is equivalent to the prediction box.
[0055] It can be known that there are preset anchor boxes in the target detection network. The anchor box is a set of reference boxes with different sizes preset at different pixel positions on the image, mainly used to train the target detection network to learn the relative positions between the anchor boxes and the label boxes (or target boxes, annotation boxes, etc.), so that the trained target detection network can detect the targets in the image, or detect the regions where the targets in the image are located. The region indicated by the anchor box is also the anchor box region.
[0056] In addition, during the process of the target detection network processing the sample image, it will judge whether each anchor box contains a target or a background, or in other words, judge whether each anchor box covers a target, and perform bounding box regression (or coordinate correction) on the anchor boxes covering the target to obtain the prediction box. That is, the prediction box can be obtained by performing bounding box regression on the anchor box; among them, judging whether it is a target or a background is equivalent to performing binary classification on each anchor box. Therefore, each anchor box will obtain 2 classification scores, respectively representing the probability of being a target and the probability of being a background. It can be understood that the prediction score corresponding to the prediction region can be the 2 classification scores of the anchor box region corresponding to the prediction region, or the score representing the probability of being a target among the 2 classification scores. The embodiments of the present disclosure do not limit this.
[0057] Among them, the bounding box regression will obtain information such as the coordinates (such as the center coordinates or vertex coordinates) and sizes (height h and width w) of the predicted boxes corresponding to each anchor box. Based on this, the position information of the predicted region may include the center coordinates or vertex coordinates of the predicted region, and the embodiments of the present disclosure do not limit this. The object detection network can also output a confidence for each predicted box, which is used to characterize the reliability determined by the object detection network for each predicted box.
[0058] In step S13, according to the detection result, the pseudo-label information corresponding to the sample image is determined, and the pseudo-label information includes the pseudo-labeled region where the object is located in the sample image.
[0059] As described above, the detection result may include at least one of the predicted region where the object is located in the sample image, the predicted score of the predicted region, and the confidence of the predicted region. In a possible implementation manner, a screening threshold may be preset in advance, and the predicted regions in the detection results of each sample image whose predicted scores are greater than or equal to the screening threshold are determined as the pseudo-labeled regions of the sample image; then, according to the detection results corresponding to the pseudo-labeled regions, the pseudo-label information is determined.
[0060] Among them, the pseudo-labeled region can be understood as a relatively accurate predicted region in the detection result, and the pseudo-labeled region is equivalent to a pseudo-labeled box (or called a pseudo-label box) marked for the sample image. After screening out the pseudo-labeled regions from the predicted regions, it can be considered that there is a high probability that the object is contained in the pseudo-labeled regions. Therefore, a pseudo-label value can be set for the pseudo-labeled regions, and the pseudo-label value is used to indicate whether the object is contained in the pseudo-labeled regions. For example, the pseudo-label value "1" can be set to indicate that the object is contained in the pseudo-labeled region, and the pseudo-label value "0" can be set to indicate that the object is not contained in the pseudo-labeled region. The above pseudo-label information may include at least one of the position information of the pseudo-labeled region, the pseudo-label value corresponding to the pseudo-labeled region, and the confidence corresponding to the pseudo-labeled region.
[0061] It should be understood that the position information of the pseudo-labeled region, that is, the position information of the predicted region whose predicted score is greater than or equal to the screening threshold. Therefore, the position information of the pseudo-labeled region may include the center coordinates or vertex coordinates of the pseudo-labeled region; the confidence corresponding to the pseudo-labeled region, that is, the confidence corresponding to the predicted region whose predicted score is greater than or equal to the screening threshold. After screening out the pseudo-labeled regions from the predicted regions, the pseudo-label information corresponding to the pseudo-labeled region can be determined according to the detection result corresponding to the pseudo-labeled region.
[0062] In step S14, according to the pseudo-label information and the anchor box information preset in the object detection network, the matching loss between the pseudo-labeled region and the anchor box region represented by the anchor box information is determined.
[0063] As described above, anchor boxes are preset in the object detection network. An anchor box is a set of reference boxes with different sizes preset at different pixel positions on an image, mainly used to train the object detection network to learn the relative positions between the anchor boxes and the labeled boxes (or target boxes), so that the trained object detection network can detect the objects in the image, or detect the regions where the objects are located in the image. The region indicated by the anchor box is also the anchor box region.
[0064] In addition, during the process of the object detection network processing the sample image, it will determine whether there is an object or background in each anchor box region. Determining whether there is an object or background is equivalent to performing binary classification on each anchor box region. Therefore, each anchor box will obtain two classification scores, respectively representing the probability of being an object and the probability of being a background. Based on this, the anchor box information in the embodiments of the present disclosure may include the classification scores predicted by the object detection network for the anchor box region. The classification scores may represent the probability that the anchor box region includes an object, and may also include the probability that the anchor box region includes a background; the anchor box information may also include the position information of the anchor box region, and the position information of the anchor box region may include the center coordinates or vertex coordinates of the anchor box region.
[0065] In a possible implementation manner, the matching loss may include at least one of the following: the classification loss between the pseudo-labeled region and the anchor box region, the first position loss between the pseudo-labeled region and the predicted region corresponding to the anchor box region, and the second position loss between the pseudo-labeled region and the anchor box region. Among them, the classification loss may be determined according to the pseudo-label value of the pseudo-labeled region and the classification score corresponding to the anchor box region. The first position loss may be determined according to the position information of the pseudo-labeled region and the position information of the predicted region corresponding to the anchor box region. The second position loss may be determined according to the position information of the pseudo-labeled region and the position information of the anchor box region. The predicted region corresponding to the anchor box region may be understood as the predicted region obtained by performing regression on the anchor box region.
[0066] In step S15, according to the matching loss, determine the network loss of the object detection network, and based on the network loss, train the object detection network.
[0067] It can be understood that in each training stage, the object detection network can process multiple sample images to obtain the detection results of the multiple sample images. There are multiple anchor box regions set in each sample image. Each sample image may correspond to at least one predicted region, and each sample image may correspond to at least one pseudo-labeled region. Therefore, the pseudo-label information may include multiple pseudo-labeled regions.
[0068] For each pseudo-labeled region in each sample image, the matching loss between each pseudo-labeled region and each anchor box region in the sample image can be obtained through step S14; in a possible implementation, the cumulative value or average value of the matching losses between multiple pseudo-labeled regions and multiple anchor box regions can be determined as the network loss of the object detection network, and the embodiments of the present disclosure do not limit this.
[0069] Among them, training the object detection network based on the network loss can be understood as optimizing the network parameters of the object detection network based on the network loss through backpropagation and other methods; it should be understood that the above steps S11 to S14 can be iteratively executed multiple rounds until the object detection network reaches a preset condition, for example, the network loss converges or the network loss reaches a specified value (such as 0), etc., and the embodiments of the present disclosure do not limit this.
[0070] In the embodiments of the present disclosure, by using the network loss obtained from the matching loss between the anchor box region and the pseudo-labeled region to train the object detection network, it is possible to reduce the influence of situations such as inconsistent matching and inaccurate matching when the anchor box region matches the pseudo-labeled region on the training effect of the object detection network, and improve the training efficiency of the object detection network.
[0071] As described above, the matching loss may include at least one of the following: the classification loss between the pseudo-labeled region and the anchor box region, the first position loss between the pseudo-labeled region and the predicted region corresponding to the anchor box region, and the second position loss between the pseudo-labeled region and the anchor box region; in a possible implementation, in step S14, according to the pseudo-label information and the anchor box information preset in the object detection network, determining the matching loss between the pseudo-labeled region and the anchor box region represented by the anchor box information includes:
[0072] Step S141: Determine the classification loss between the pseudo-labeled region and the anchor box region according to the classification score corresponding to the anchor box region and the pseudo-label value corresponding to the pseudo-labeled region.
[0073] In a possible implementation, the cross-entropy loss between the pseudo-labeled region and the anchor box region can be calculated according to the classification score and the pseudo-label value as the classification loss. Of course, other classification losses can also be used, such as KL divergence, etc., and the embodiments of the present disclosure do not limit this.
[0074] Step S142: Determine the first position loss between the pseudo-labeled region and the predicted region corresponding to the anchor box region according to the position information of the predicted region corresponding to the anchor box region and the position information corresponding to the pseudo-labeled region.
[0075] As described above, the position information of the prediction region may include the center coordinates or vertex coordinates of the prediction region, and the position information of the pseudo-labeled region may include the center coordinates or vertex coordinates of the pseudo-labeled region. In a possible implementation manner, based on loss calculation methods such as L1 loss, L2 loss, smooth L1 loss, etc., the first position loss between the prediction region and the pseudo-labeled region can be obtained according to the position information of the prediction region and the position information of the pseudo-labeled region. The embodiments of the present disclosure do not limit this.
[0076] Step S143: Determine the second position loss between the pseudo-labeled region and the anchor box region according to the position information corresponding to the anchor box region and the position information corresponding to the pseudo-labeled region.
[0077] In a possible implementation manner, the second position loss may be the nth power of a specified constant, where n may be the norm of the distance between the pseudo-labeled region and the anchor box region. The second position loss can play a role in stabilizing the training effect in the initial stage of network training, or in other words, it can reduce the learning weight of the anchor box for the pseudo-labeled box with a relatively large distance in the initial stage of training, thereby stabilizing the training effect of the network. Among them, the specified constant can be an empirical value obtained through experimental tests. The embodiments of the present disclosure do not limit the specific value of the specified constant. Formula (1) shows a calculation method of a second position loss C d of the present disclosure.
[0078]
[0079] Among them, i represents the i-th pseudo-label region, j represents the j-th anchor box region, p j represents the position information of the j-th anchor box region, g i represents the position information of the i-th pseudo-label region, d(p j , g i ) represents the distance between the i-th pseudo-label region and the j-th anchor box region, ‖‖ 2 represents the two-norm, and the specified constant is 10.
[0080] As described above, the position information corresponding to the anchor box region may include the center coordinates or vertex coordinates of the anchor box region, and the position information of the pseudo-labeled region may include the center coordinates or vertex coordinates of the pseudo-labeled region. After knowing the coordinates of the two regions, the distance between the pseudo-labeled region and the anchor box region can be calculated based on well-known distance formulas in the art, such as Euclidean distance, cosine distance, etc. The embodiments of the present disclosure do not limit this.
[0081] In a possible implementation, based on the above loss calculation methods such as L1 loss, L2 loss, smooth L1 loss, etc., the second position loss between the pseudo-labeled region and the anchor box region can also be determined according to the position information corresponding to the anchor box region and the position information corresponding to the pseudo-labeled region. The embodiments of the present disclosure do not limit this.
[0082] Step S144: Determine the matching loss between the pseudo-labeled region and the anchor box region according to at least one of the classification loss, the first position loss, and the second position loss.
[0083] As described above, the matching loss may include at least one of the classification loss, the first position loss, and the second position loss. In a possible implementation, the matching loss may be a weighted sum of the classification loss, the first position loss, and the second position loss. For example, a matching loss shown in formula (2) is as follows.
[0084] C ij =λ c C c +λ r C r +λ d C d (2)
[0085] Wherein, C ij represents the matching loss between the i-th pseudo-label region and the j-th anchor box region, C c represents the classification loss, C r represents the first position loss, C d represents the second position loss, and λ c 、λ r and λ d respectively represent the weight parameters corresponding to the three losses, and the weight parameters may be parameter values optimized synchronously during the network training process.
[0086] In the embodiments of the present disclosure, the matching loss between the pseudo-labeled region and the anchor box region can be effectively determined according to the pseudo-label information and the anchor box information.
[0087] Considering that directly taking the accumulated value or average value of the matching losses between multiple pseudo-labeled regions and multiple anchor box regions as the network loss of the object detection network will cause the anchor box region to learn each pseudo-labeled region to the same extent. Since there is noise in some pseudo-labeled regions, this may reduce the learning effect of the object detection network. In a possible implementation, in step S15, determining the network loss of the object detection network according to the matching loss includes:
[0088] Step S151: For each of the multiple pseudo-labeled regions, determine the first weight corresponding to each of the multiple anchor regions according to the matching loss between the pseudo-labeled region and the multiple anchor regions. The first weight of the same anchor region is negatively correlated with the matching loss.
[0089] Among them, for each pseudo-labeled region, the matching loss between the pseudo-labeled region and each anchor region can be obtained according to the above step S14. It should be understood that the smaller the matching loss, the more matching the anchor region is to the pseudo-labeled region. On the contrary, the larger the matching loss, the less matching the anchor region is to the pseudo-labeled region. Then, for each pseudo-labeled region, based on the principle that the first weight is negatively correlated with the matching loss, that is, the smaller the matching loss, the larger the first weight; the larger the matching loss, the smaller the first weight, the first weight corresponding to each of the multiple anchor regions can be determined according to the matching loss between the pseudo-labeled region and the multiple anchor regions. Among them, the specific values and specific allocation methods of the first weight in the embodiments of the present disclosure are not limited.
[0090] Step S152: Determine the region loss corresponding to the pseudo-labeled region according to the first weight corresponding to each of the multiple anchor regions and the matching loss corresponding to each of the multiple anchor regions.
[0091] In a possible implementation manner, determining the region loss corresponding to the pseudo-labeled region according to the first weight corresponding to each of the multiple anchor regions and the matching loss corresponding to each of the multiple anchor regions may include: performing a weighted sum of the matching losses corresponding to the multiple anchor regions according to the first weight corresponding to each of the multiple anchor regions to obtain the region loss corresponding to the pseudo-labeled region; or may also include: determining the weighted average value of the matching losses corresponding to the multiple anchor regions according to the first weight corresponding to each of the multiple anchor regions to obtain the region loss corresponding to the pseudo-labeled region. The embodiments of the present disclosure do not limit this.
[0092] Step S153: Determine the network loss of the target detection network according to the region losses corresponding to the obtained multiple pseudo-labeled regions.
[0093] In a possible implementation manner, determining the network loss of the target detection network according to the region losses corresponding to the obtained multiple pseudo-labeled regions may include: determining the cumulative value or average value between the region losses corresponding to the multiple pseudo-labeled regions as the network loss of the target detection network. The embodiments of the present disclosure do not limit this.
[0094] In the embodiments of the present disclosure, for each pseudo-labeled region, the weight distribution of multiple anchor box regions can be performed according to the matching loss between the pseudo-labeled region and the multiple anchor box regions, and the region loss corresponding to each pseudo-labeled region can be determined according to the assigned first weight and the matching loss. Then, the network loss can be determined according to the region losses of the multiple pseudo-labeled regions. In this way, the anchor box regions can learn more with emphasis on the pseudo-labeled regions with small matching losses according to the different weights, realizing the soft matching between the anchor boxes and the pseudo-labeled boxes, and further improving the overall learning effect of the object detection network.
[0095] As described above, since in the semi-supervised training method, the pseudo-labeled boxes predicted by the object detection network for the sample images without labeled tags have noise. To reduce the influence of the pseudo-labeled boxes with noise on the network training, the weighted evaluation of the region losses corresponding to the multiple pseudo-labeled regions can be performed based on the confidence levels corresponding to the multiple pseudo-labeled regions respectively to determine the network loss. For example, for the pseudo-labeled regions with high confidence levels, the weights of their region losses are increased, while for the pseudo-labeled regions with low confidence levels, the weights of their region losses are decreased, so as to reduce the noise influence brought by the pseudo-labeled regions with lower confidence levels on the network training and improve the robustness of the object detection network.
[0096] As described above, the pseudo-label information includes the confidence level corresponding to the pseudo-labeled region, and the confidence level represents the reliability degree of the pseudo-labeled region. In a possible implementation manner, in step S153, according to the obtained region losses corresponding to the multiple pseudo-labeled regions, determining the network loss of the object detection network includes:
[0097] Determining the second weight corresponding to each of the multiple pseudo-labeled regions according to the confidence level corresponding to each of the multiple pseudo-labeled regions, where the second weight is positively correlated with the confidence level; determining the network loss of the object detection network according to the second weight corresponding to each of the multiple pseudo-labeled regions and the region loss corresponding to each of the multiple pseudo-labeled regions.
[0098] It should be understood that the higher the confidence level of the pseudo-labeled region, the higher the reliability degree of the pseudo-labeled region. On the contrary, the lower the confidence level of the pseudo-labeled region, the lower the reliability degree of the pseudo-labeled region. Then, based on the principle that the second weight is positively correlated with the confidence level, that is, the larger the confidence level, the larger the second weight; the smaller the confidence level, the smaller the second weight, the second weight corresponding to each of the multiple pseudo-labeled regions can be determined according to the confidence level corresponding to each of the multiple pseudo-labeled regions. Among them, the embodiments of the present disclosure do not limit the specific values and specific distribution methods of the second weight.
[0099] In a possible implementation, determining the network loss of the object detection network according to the second weights corresponding to multiple pseudo-labeled regions and the region losses corresponding to the multiple pseudo-labeled regions may include: performing a weighted sum of the region losses corresponding to the multiple pseudo-labeled regions according to the second weights corresponding to the multiple pseudo-labeled regions to obtain the network loss of the object detection network; or may further include: determining the weighted average of the region losses corresponding to the multiple pseudo-labeled regions according to the second weights corresponding to the multiple pseudo-labeled regions to obtain the network loss of the object detection network. The embodiments of the present disclosure do not limit this.
[0100] In the embodiments of the present disclosure, by performing weight assignment on the region loss of the pseudo-labeled region based on the confidence corresponding to the pseudo-labeled region to determine the network loss, the noise impact brought by the pseudo-labeled region with a relatively low confidence to network training can be reduced, and the robustness of the object detection network can be improved.
[0101] According to the embodiments of the present disclosure, it can enable the object detection network to dynamically match the best anchor box for each pseudo-labeled box to learn during training. Moreover, when implementing the dynamic soft matching between the anchor box and the pseudo-labeled box, the confidence of the pseudo-labeled box is taken into account, so that the object detection network can dynamically select a more accurate and reliable pseudo-labeled box to learn, greatly alleviating the noise impact brought by inaccurate pseudo-labeled boxes to network training.
[0102] As described above, in the semi-supervised training method, relatively accurate prediction boxes are usually selected from the detection results of the object detection network based on a preset screening threshold as pseudo-label boxes. However, this method will result in inconsistent pseudo-label boxes obtained in different training stages. For example, in the initial stage of training, due to the low accuracy of the object detection network, the classification scores of some prediction boxes may be below the screening threshold and be filtered out, resulting in false negatives. As training progresses, the reliability and accuracy of the object detection network continue to increase, and these previously filtered-out prediction boxes will be used as pseudo-label boxes to participate in network training, thus resulting in the situation where the pseudo-label boxes of the same sample image before and after training are inconsistent. During the network training process, the inconsistency of the pseudo-label boxes of the same image may affect the learning effect and convergence effect of the object detection network.
[0103] Based on the above problems, in a possible implementation, in step S13, determining the pseudo-label information corresponding to the sample image according to the detection result includes:
[0104] Step S131: Clustering the multiple prediction regions according to the prediction scores corresponding to the multiple prediction regions to obtain the clustering result of the multiple prediction regions.
[0105] As described above, the detection results include prediction scores corresponding to multiple prediction regions respectively, and the prediction scores represent the probability that the target is included in the prediction region. According to the prediction scores corresponding to multiple prediction regions respectively, clustering the multiple prediction regions can cluster the multiple prediction regions into positive sample prediction regions and negative sample prediction regions. The classification accuracy of the positive sample prediction regions is higher than that of the negative sample prediction regions. Or rather, the probability that the target is included in the positive sample prediction regions is higher than that of the negative sample prediction regions. This facilitates dynamically screening out more accurate and reliable pseudo-labeled regions from the positive sample prediction regions with higher classification accuracy later.
[0106] In a possible implementation manner, clustering the multiple prediction regions according to the prediction scores corresponding to the multiple prediction regions respectively to obtain the clustering results of the multiple prediction regions may include: clustering the multiple prediction regions according to the prediction scores corresponding to the multiple prediction regions respectively through a Gaussian mixture model (GMM) to obtain the clustering results of the multiple prediction regions; wherein, the clustering results include positive sample prediction regions and negative sample prediction regions among the multiple prediction regions, and the accuracy of the positive sample prediction regions is higher than that of the negative sample prediction regions. Through this method, the prediction regions can be clustered efficiently and accurately, thus facilitating dynamically determining the screening threshold based on the clustering results later.
[0107] Among them, the Gaussian mixture model (GMM) is an extension of the Gaussian probability density function, which accurately quantifies the variable distribution with multiple Gaussian probability density functions (normal distribution curves), decomposes the variable distribution into several statistical models based on Gaussian probability density function distributions, and is a commonly used clustering algorithm. Formula (3) shows a two-component Gaussian mixture model P(b), where N p represents the Gaussian distribution corresponding to the positive sample prediction region, N n represents the Gaussian distribution corresponding to the negative sample prediction region, b is the prediction score corresponding to the prediction region, μ n , μ p are the means of the two Gaussian distributions respectively, ρ n , ρ p are the variances of the two Gaussian distributions respectively, w n , w p represent the mixing parameters of the two Gaussian distributions respectively, and the sum of w n and w p is 1.
[0108] P(b) = w n N n (b|μ n , ρ n ) + w p N p (b|μ p , ρ p ) (3)
[0109] Among them, the expectation maximization algorithm (EM algorithm) can be used to solve the means, variances, and mixing coefficients of the two Gaussian distributions in the Gaussian mixture model shown in formula (1). It should be understood that the embodiments of the present disclosure do not limit the solution process of this Gaussian mixture model.
[0110] Among them, after solving the mixing coefficients w n , w p , the sample attributes of the prediction region can be determined according to the magnitude relationship between w n and w p , that is, it is determined whether the prediction region is a positive sample prediction region or a negative sample prediction region, thereby realizing the clustering of the prediction region. For example, if w n is greater than w p , it can be determined that the prediction region corresponding to the classification score b belongs to the negative sample prediction region; if w n is less than w p , it can be determined that the prediction region corresponding to the classification score b belongs to the positive sample prediction region.
[0111] Among them, the accuracy of the positive sample prediction region is higher than that of the negative sample prediction region. It can be understood that the classification accuracy of the positive sample prediction region is higher than that of the negative sample prediction region; it can be understood that with multiple rounds of iteration of network training, the classification score output by the object detection network and the accuracy of the prediction region increase synchronously. The classification accuracy of the positive sample prediction region is higher than that of the negative sample prediction region, that is, the classification accuracy and localization accuracy of the positive sample prediction region are both higher than those of the negative sample prediction region.
[0112] It should be understood that the above clustering of the prediction region using the Gaussian mixture model is an implementation provided by the embodiments of the present disclosure. In fact, those skilled in the art can use any clustering algorithm in the art, such as the K-means clustering algorithm, to cluster the prediction region, and the embodiments of the present disclosure do not limit this.
[0113] Step S132: Determine a screening threshold according to the clustering results of multiple prediction regions, and based on the screening threshold, screen out pseudo-labeled regions from the multiple prediction regions.
[0114] As described above, the clustering results include positive sample prediction regions among multiple prediction regions. In one possible implementation, a screening threshold is determined according to the clustering results of the multiple prediction regions. For example, it may include: determining the screening threshold according to the prediction scores of the positive sample prediction regions. Among them, determining the screening threshold according to the prediction scores of the positive sample prediction regions may include, for example: taking the average value of the prediction scores corresponding to the positive sample prediction regions as the screening threshold; or may also include: taking the median value of the prediction scores corresponding to the positive sample prediction regions as the screening threshold, etc. The embodiments of the present disclosure do not limit this. In this way, it is possible to dynamically determine the screening threshold based on the prediction scores of the positive sample prediction regions, so that the screening threshold can adapt to different training stages. For example, the screening threshold in the initial stage of training may be relatively small, and the screening threshold in the later stage of training may be relatively large. This is beneficial to reducing the influence of inconsistent pseudo-labeled regions in different training stages on the network training effect and convergence effect.
[0115] In one possible implementation, based on the screening threshold, pseudo-labeled regions are screened out from the multiple prediction regions. For example, it may include: determining the positive sample prediction regions with prediction scores greater than or equal to the screening threshold as the pseudo-labeled regions. In this way, more accurate and reliable pseudo-labeled regions can be screened out from the positive sample prediction regions based on the screening threshold, which is beneficial to improving the learning effect and training efficiency of the target detection network.
[0116] As described above, the prediction regions can be clustered by the Gaussian mixture model, and the positive sample prediction regions correspond to Gaussian distributions. In one possible implementation, according to the clustering results of the multiple prediction regions, a screening threshold is determined, and based on the screening threshold, pseudo-labeled regions are screened out from the multiple prediction regions, which may include: taking the prediction score corresponding to the peak value of the Gaussian distribution corresponding to the positive sample prediction region as the screening threshold; determining the positive sample prediction regions with prediction scores greater than or equal to the screening threshold as the pseudo-labeled regions. In this way, it is possible to dynamically and conveniently determine the screening threshold based on the Gaussian distribution corresponding to the positive sample prediction regions, so that the screening threshold can adapt to different training stages, which is beneficial to reducing the influence of inconsistent pseudo-labeled regions in different training stages on the network training effect and convergence effect.
[0117] As described above, the EM algorithm can be used to solve the means and variances of the two Gaussian distributions in the Gaussian mixture model shown in the above formula (1), and then the Gaussian distribution P corresponding to the positive sample prediction region is obtained. p (b)=N p (b|μ p ,ρ p ) and the prediction score corresponding to the peak value of the Gaussian distribution corresponding to the positive sample prediction region can be determined. Based on this, the screening threshold τ can be expressed as formula (4).
[0118] τ = argmax b P p (b) (4)
[0119] Step S133: Determine the pseudo-label information corresponding to the pseudo-labeled region according to the detection result corresponding to the pseudo-labeled region.
[0120] As described above, the detection result may include at least one of the predicted region where the target is located in the sample image, the prediction score of the predicted region, and the confidence of the predicted region. The pseudo-label information may include at least one of the position information of the pseudo-labeled region, the pseudo-labeled value corresponding to the pseudo-labeled region, and the confidence corresponding to the pseudo-labeled region.
[0121] It should be understood that the position information of the pseudo-labeled region, that is, the position information of the predicted region whose prediction score is greater than or equal to the screening threshold. Therefore, the position information of the pseudo-labeled region may include the center coordinates or vertex coordinates of the pseudo-labeled region; the confidence corresponding to the pseudo-labeled region, that is, the confidence corresponding to the predicted region whose prediction score is greater than or equal to the screening threshold. After screening out the pseudo-labeled regions from the predicted regions, the pseudo-label information corresponding to the pseudo-labeled region can be determined according to the detection result corresponding to the pseudo-labeled region.
[0122] In the embodiments of the present disclosure, all predicted regions can be divided into two categories by clustering. One category is the accurate positive sample predicted regions, which can be used as pseudo-label regions to train the object detection network, and the other category is the relatively inaccurate negative sample predicted regions, which can be discarded; among them, determining the screening threshold according to the clustering result can dynamically adjust the screening threshold during the training process, so that the screening threshold adapts to different training stages, thereby reducing the impact of inconsistent pseudo-labeled regions in different training stages on the network training effect and convergence effect.
[0123] According to the embodiments of the present disclosure, a dynamic threshold selection algorithm based on the Gaussian Mixture Model (GMM) can be implemented. This dynamic threshold selection algorithm can dynamically adjust the screening threshold during the training process. A smaller screening threshold is used at the beginning of the training to increase the recall rate and reduce false negatives (i.e., screening out the predicted boxes containing the target), and the screening threshold is increased at the later stage of the training to increase the accuracy of the pseudo-labeled boxes and reduce false positives (i.e., screening out the predicted boxes without containing the target), thereby greatly alleviating the impact of inconsistent pseudo-labeled boxes before and after training on network training.
[0124] It is known that in the process of a target detection network processing a sample image, the target detection network not only needs to predict the category of each object in the sample image (that is, determine whether it is a target or a background), but also needs to regress the coordinates of each object (that is, regress the predicted bounding box). The detection result includes the classification result (that is, the predicted score) and the regression result (that is, the predicted region) output by the target detection network.
[0125] In the scenario of training a target detection network based on a semi-supervised training method, the pseudo-labeled bounding boxes obtained through the screening threshold will simultaneously supervise the training of both the classification and regression tasks. Therefore, there may be a situation where the classification result and the regression result do not match (or are inconsistent). Among them, the non-matching of the classification result and the regression can be simply understood as that the predicted bounding box indicated by the classification score, which is likely to contain the target, may not cover the largest region where the target is located, or in other words, does not completely cover the actual region where the target is located.
[0126] Specifically, a pseudo-labeled bounding box with a good classification result is not necessarily a pseudo-labeled bounding box with a good regression result, and a pseudo-labeled bounding box with a good regression result is not necessarily a pseudo-labeled bounding box with a good classification result. Figure 3 A schematic diagram of a pseudo-labeled bounding box according to the related art is shown. As Figure 3 In the figure, the dotted box and the solid box are two pseudo-labeled bounding boxes with the target being a person (Person). Although the predicted score Score of the solid box is as high as 0.92, the IoU between the solid box and the target box (the target box represents the actual region where the target is located) is only 0.4. Such a pseudo-labeled bounding box is beneficial to the training of classification but not to the training of regression. Although the predicted score of the dotted box is only 0.34, the IoU between the dotted box and the target box is as high as 0.96. Using the dotted box is not beneficial to the training of classification but is beneficial to the training of regression. The non-matching of the classification result and the regression result on the pseudo-labeled bounding box will affect the training effect of the target detection network.
[0127] Based on the above problems, in a possible implementation manner, in step S12, the target detection network processes the sample image to obtain the detection result of the sample image, including:
[0128] Feature extraction is performed on the sample image to obtain classification features and regression features; decoding processing is performed on the regression features and the classification features to obtain an initial predicted region and a predicted score corresponding to the initial predicted region; an offset amount corresponding to the initial predicted region is determined according to the regression features and the classification features; and the initial predicted region is adjusted according to the offset amount to obtain a predicted region.
[0129] Among them, obtaining the initial prediction region, that is, obtaining information such as the initial coordinates (x, y) and the initial size (width w and height h) of the initial prediction region, then the offset corresponding to the initial prediction region may include at least one of the offset corresponding to the coordinates and the offset corresponding to the size. Among them, the offset corresponding to the size may include at least one of the offset corresponding to the height and the offset corresponding to the width, and the offset corresponding to the coordinates may include at least one of the offset corresponding to the abscissa and the offset corresponding to the ordinate.
[0130] In a possible implementation manner, adjusting the initial prediction region according to the offset to obtain the prediction region may include: adjusting at least one of the initial coordinates and the initial size of the initial prediction region according to at least one of the offset corresponding to the coordinates and the offset corresponding to the size to obtain the target coordinates and the target size of the prediction region. For example, the target coordinates of the prediction region can be obtained by adding the offset corresponding to the coordinates to the initial coordinates of the initial prediction region; and / or, the target size of the prediction region can be obtained by adding the offset corresponding to the size to the initial size of the initial prediction region.
[0131] As described above, the network structure of the object detection network in the embodiments of the present disclosure is not limited. In a possible implementation manner, the object detection network may at least include an encoding network layer, a decoding network layer, and an offset extraction layer. The encoding network layer (or feature extraction layer) is used to extract classification features and regression features from the sample image. The decoding network layer is used to perform decoding processing on the regression features and classification features to obtain the initial prediction region and the prediction score corresponding to the initial prediction region. The offset extraction layer is used to determine the offset corresponding to the initial prediction region according to the regression features and classification features.
[0132] In the embodiments of the present disclosure, the initial prediction region can be adjusted based on the offset to obtain a prediction region in which the classification result and the regression result are consistently matched. In this way, the classification result and the regression result of the pseudo-labeled region screened from the prediction region based on the screening threshold are also consistently matched, which is beneficial to improving the training effect and training efficiency of the network. Among them, the consistent matching of the classification result and the regression result can be simply understood as that the prediction box with a high probability of containing the target indicated by the classification score will cover the largest region where the target is located. For example, the prediction score of the same prediction region is 0.92 and the IoU is 0.96.
[0133] In a possible implementation manner, determining the offset corresponding to the initial prediction region according to the regression features and classification features includes: performing splicing processing and convolution processing on the classification features and the regression features to obtain multi-scale bias features; performing scale alignment on the multi-scale bias features to obtain the target bias features with scale alignment; performing decoding processing on the target bias features to obtain the offset corresponding to the initial prediction region.
[0134] Among them, splicing and convolutional processing are performed on the classification features and regression features to obtain multi-scale bias features, which may include: splicing the classification features and regression features in channels to obtain spliced features; then performing convolutional processing on the spliced features to obtain bias features, where the number of channels of the bias features is less than that of the spliced features, and the bias features may be multi-scale features. The multi-scale features can be understood as features including multiple different channels and each channel having a different scale.
[0135] Among them, scale alignment is performed on the multi-scale bias features to obtain target bias features with scale alignment, which may include: upsampling the small-scale bias features and / or downsampling the large-scale bias features to obtain target bias features with scale alignment. It can be understood that the scale of the target bias features can be the same as the intermediate scale in the bias features, or the same as the largest scale in the bias features, or the same as the smallest scale in the bias features. The embodiments of the present disclosure do not limit this.
[0136] It should be understood that the decoding network layer for decoding the target bias features and the decoding network layer for decoding the classification features and regression features described above may be different network layers. The embodiments of the present disclosure do not limit the network structures of the two decoding network layers.
[0137] Figure 4 Shows a processing schematic diagram of an object detection network according to an embodiment of the present disclosure. As Figure 4 shown, the process of processing the sample image through the object detection network may include:
[0138] Feature extraction is performed on the sample image to obtain 4 classification features with 256 channels and 4 regression features with 256 channels; decoding processing is performed on the regression features and classification features to obtain an initial prediction region and a prediction score corresponding to the initial prediction region; the classification features and regression features are spliced in channels to obtain a spliced feature with 2048 channels; convolutional processing is performed on the spliced feature to obtain a bias feature with 64 channels; scale alignment is performed on the multi-scale bias features to obtain target bias features; decoding processing is performed on the target bias features to obtain an offset corresponding to the initial prediction region; according to the offset, the initial prediction region is adjusted to obtain a prediction region.
[0139] In the embodiments of the present disclosure, the offset corresponding to the initial prediction region can be effectively determined by using the object detection network, so as to adjust the initial prediction region based on the offset to obtain a prediction region that can cover the target as much as possible, which is beneficial to obtaining a prediction region where the classification result and the regression result are consistently matched.
[0140] According to an embodiment of the present disclosure, a target detection network is further provided. The target detection method includes: processing an image to be detected through the target detection network to obtain a target detection result of the image to be detected, where the target detection result includes the region where the target is located in the image to be detected; wherein, the target detection network is trained according to the above network training method. In this way, a target detection network with better accuracy and robustness can be used to accurately detect the region where the target is located in the image to be detected, improving the target detection accuracy.
[0141] That is to say, the target detection network trained by the above network training method can be deployed to implement the target detection of the image to be detected. The image to be detected can be, for example, an image collected by an image acquisition device (such as a camera), and the image may include the target to be detected, such as a human body, a face, a vehicle, an object, etc., and the embodiments of the present disclosure do not limit this.
[0142] Among them, processing the image to be detected through the target detection network to obtain the target detection result of the image to be detected can be understood as inputting the image to be processed into the target detection network for processing to obtain the target detection result output by the target detection network. The target detection result may include the region where the target is located in the image to be processed (i.e., the position where the target is located), and the region where the target is located can be indicated by a prediction box.
[0143] As described above, the target may include a vehicle. In a possible implementation manner, processing the image to be detected through the target detection network to obtain the target detection result of the image to be detected includes: processing the image to be detected through the target detection network to obtain the vehicle region where the vehicle is located in the image to be detected; wherein, after obtaining the vehicle region, the method further includes: performing vehicle recognition on the vehicle region to obtain the vehicle recognition result of the vehicle in the vehicle region. In this way, a target detection network with better accuracy and robustness can be used to accurately detect the vehicle region where the vehicle is located in the image to be detected, and at the same time improve the accuracy of vehicle recognition in the vehicle region.
[0144] It should be understood that those skilled in the art can use the known image recognition technology in the art to implement vehicle recognition on the vehicle region to obtain the vehicle recognition result of the vehicle in the vehicle region. The vehicle recognition result may include, for example, the license plate number, color, vehicle type, etc., and the embodiments of the present disclosure do not limit this.
[0145] The network training method according to an embodiment of the present disclosure can be applied to the field of industrial defect detection. It is known that when using an object detection network based on deep learning for defect detection, a large amount of manually labeled sample data is required to train the object detection network. Due to the limited cost and efficiency of manual labeling, this also increases the training cost and time of the object detection network. According to the network training method of the embodiment of the present disclosure, it is possible to efficiently implement semi-supervised training of the object detection network by using an image set with only 10% or even 1% labeled images.
[0146] The network training method according to an embodiment of the present disclosure can be applied to the field of remote sensing change detection. Due to the low resolution, large and blurred images, and different annotation standards varying with different scenarios in remote sensing images, professional annotators are usually required to manually annotate remote sensing images. And due to the low image resolution, both the annotation efficiency and cost are greatly increased, and mislabeling and missing labeling often occur. Therefore, manual annotation often has more noise. According to the network training method of the embodiment of the present disclosure, it can well adapt to this scenario. In an image set with only partial annotation, inaccurate annotation, or noisy annotation, etc., it is possible to dynamically determine a screening threshold and screen out relatively accurate and reliable pseudo annotation boxes, realize the dynamic matching of anchor boxes and pseudo annotation boxes, and efficiently implement semi-supervised training for the object detection network, improving the robustness of the object detection network after training.
[0147] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0148] In addition, the present disclosure also provides a network training device, an object detection device, an electronic device, a computer-readable storage medium, and a program. The above can all be used to implement any network training method and object detection method provided by the present disclosure. For the corresponding technical solutions and descriptions, please refer to the corresponding records in the method part and will not be elaborated further.
[0149] Figure 5 The block diagram showing the network training device according to an embodiment of the present disclosure is as Figure 5 shown, and the device includes:
[0150] An acquisition module 101, configured to acquire a sample image without label annotation and an object detection network to be trained;
[0151] A detection module 102, configured to process the sample image through the object detection network to obtain a detection result of the sample image, where the detection result includes a predicted region where an object is located in the sample image;
[0152] The pseudo-label determination module 103 is configured to determine the pseudo-label information corresponding to the sample image according to the detection result, where the pseudo-label information includes the pseudo-labeled region where the target is located in the sample image;
[0153] The matching loss determination module 104 is configured to determine the matching loss between the pseudo-labeled region and the anchor region represented by the anchor box information according to the pseudo-label information and the preset anchor box information in the target detection network;
[0154] The training module 105 is configured to determine the network loss of the target detection network according to the matching loss, and train the target detection network based on the network loss.
[0155] In a possible implementation manner, the pseudo-label information includes multiple pseudo-labeled regions. Wherein, the training module 105 includes: a weight determination sub-module, configured to, for each pseudo-labeled region among the multiple pseudo-labeled regions, determine the first weight corresponding to each of the multiple anchor regions according to the matching loss between the pseudo-labeled region and the multiple anchor regions, and the first weight of the same anchor region is negatively correlated with the matching loss; a region loss determination sub-module, configured to determine the region loss corresponding to the pseudo-labeled region according to the first weight corresponding to each of the multiple anchor regions and the matching loss corresponding to each of the multiple anchor regions; a network loss determination sub-module, configured to determine the network loss of the target detection network according to the region losses corresponding to the obtained multiple pseudo-labeled regions.
[0156] In a possible implementation manner, the pseudo-label information includes the confidence corresponding to the pseudo-labeled region, and the confidence represents the reliability degree of the pseudo-labeled region. Wherein, the determining the network loss of the target detection network according to the region losses corresponding to the obtained multiple pseudo-labeled regions includes: determining the second weight corresponding to each of the multiple pseudo-labeled regions according to the confidence corresponding to each of the multiple pseudo-labeled regions, and the second weight is positively correlated with the confidence; determining the network loss of the target detection network according to the second weight corresponding to each of the multiple pseudo-labeled regions and the region loss corresponding to each of the multiple pseudo-labeled regions.
[0157] In a possible implementation manner, the matching loss includes at least one of the following: the classification loss between the pseudo-labeled region and the anchor region; the first position loss between the pseudo-labeled region and the predicted region corresponding to the anchor region; the second position loss between the pseudo-labeled region and the anchor region.
[0158] In a possible implementation, the anchor box information includes the classification scores predicted by the target detection network for the anchor box regions, and the pseudo-label information includes the pseudo-label values corresponding to the pseudo-labeled regions. The pseudo-label values are used to indicate whether the target is included in the pseudo-labeled regions, and the classification scores represent the probabilities that the target is included in the anchor box regions. Among them, the matching loss determination module 104 includes: a classification loss determination sub-module, configured to determine the classification loss between the pseudo-labeled region and the anchor box region according to the classification score corresponding to the anchor box region and the pseudo-label value corresponding to the pseudo-labeled region; and / or, a first position loss determination sub-module, configured to determine the first position loss between the pseudo-labeled region and the predicted region corresponding to the anchor box region according to the position information of the predicted region corresponding to the anchor box region and the position information of the pseudo-labeled region; and / or, a second position loss determination sub-module, configured to determine the second position loss between the pseudo-labeled region and the anchor box region according to the position information of the anchor box region and the position information of the pseudo-labeled region; and determine the matching loss between the pseudo-labeled region and the anchor box region according to at least one of the classification loss, the first position loss, and the second position loss.
[0159] In a possible implementation, the detection results include the prediction scores corresponding to the respective prediction regions. The prediction scores represent the probabilities that the target is included in the prediction regions. Among them, the pseudo-label determination module 103 includes: a clustering sub-module, configured to cluster the multiple prediction regions according to the prediction scores corresponding to the respective prediction regions to obtain the clustering results of the multiple prediction regions; a screening sub-module, configured to determine a screening threshold according to the clustering results of the multiple prediction regions, and based on the screening threshold, screen out the pseudo-labeled regions from the multiple prediction regions; and a root pseudo-label determination sub-module, configured to determine the pseudo-label information corresponding to the pseudo-labeled regions according to the detection results corresponding to the pseudo-labeled regions.
[0160] In a possible implementation, the step of clustering the multiple prediction regions according to the prediction scores corresponding to the respective prediction regions to obtain the clustering results of the multiple prediction regions includes: clustering the multiple prediction regions according to the prediction scores corresponding to the respective prediction regions by using a Gaussian mixture model to obtain the clustering results of the multiple prediction regions; where the clustering results include positive sample prediction regions and negative sample prediction regions among the multiple prediction regions, and the accuracy of the positive sample prediction regions is higher than that of the negative sample prediction regions.
[0161] In a possible implementation, the clustering result includes positive sample prediction regions among the multiple prediction regions. Wherein, determining a screening threshold according to the clustering result of the multiple prediction regions and screening out pseudo-labeled regions from the multiple prediction regions based on the screening threshold includes: determining the prediction score corresponding to the peak of the Gaussian distribution corresponding to the positive sample prediction region as the screening threshold; determining positive sample prediction regions with a prediction score greater than or equal to the screening threshold as the pseudo-labeled regions.
[0162] In a possible implementation, the detection module 102 includes: a feature extraction sub-module for extracting features from the sample image to obtain classification features and regression features; a decoding sub-module for performing decoding processing on the regression features and the classification features to obtain initial prediction regions and prediction scores corresponding to the initial prediction regions; an offset determination sub-module for determining an offset corresponding to the initial prediction region according to the regression features and the classification features; an adjustment sub-module for adjusting the initial prediction region according to the offset to obtain a prediction region.
[0163] In a possible implementation, determining an offset corresponding to the initial prediction region according to the regression features and the classification features includes: concatenating and convolving the classification features and the regression features to obtain multi-scale bias features; performing scale alignment on the multi-scale bias features to obtain target bias features with scale alignment; performing decoding processing on the target bias features to obtain an offset corresponding to the initial prediction region.
[0164] In the embodiments of the present disclosure, by using the network loss obtained from the matching loss between the anchor box region and the pseudo-labeled region to train the target detection network, it is possible to reduce the influence of situations such as inconsistent matching and inaccurate matching when the anchor box region matches the pseudo-labeled region on the training effect of the target detection network, and improve the training efficiency of the target detection network.
[0165] The embodiments of the present disclosure further provide a target detection device, including: an image detection module for processing a to-be-detected image through a target detection network to obtain a target detection result of the to-be-detected image, where the target detection result includes the region where the target is located in the to-be-detected image; wherein, the target detection network is trained according to the above network training method.
[0166] In a possible implementation, the target includes a vehicle. The image detection module is specifically configured to process the image to be detected through the target detection network to obtain the vehicle area where the vehicle is located in the image to be detected. After obtaining the vehicle area, the device further includes: a vehicle recognition module, configured to perform vehicle recognition on the vehicle area to obtain a vehicle recognition result of the vehicle within the vehicle area.
[0167] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0168] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0169] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the above methods.
[0170] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of the electronic device, the processor in the electronic device executes the above methods.
[0171] The electronic device can be provided as a terminal, a server or other forms of devices.
[0172] Figure 6 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 6 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to execute the above methods.
[0173] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (Mac OS X TM ), the multi-user and multi-process computer operating system (Unix TM ), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ) or the like.
[0174] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.
[0175] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0176] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, (but is not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0177] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0178] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0179] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0180] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when the instructions are executed by the processor of the computer or other programmable data processing apparatus, an apparatus is created that implements the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, so that, the computer-readable medium storing the instructions comprises a manufacture, which comprises instructions for implementing various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0181] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, so that a series of operation steps are executed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, whereby the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0182] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur in a different order than noted in the figures. For example, two consecutive boxes may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or can be implemented by a combination of dedicated hardware and computer instructions.
[0183] The computer program product can be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is embodied as a computer storage medium. In another alternative embodiment, the computer program product is embodied as a software product, such as a Software Development Kit (SDK), etc.
[0184] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or resemblances can be referred to each other. For the sake of brevity, they will not be elaborated herein.
[0185] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.
[0186] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0187] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A network training method, characterized in that, it includes: Obtain sample images without labeled annotations and a target detection network to be trained; Process the sample images through the target detection network to obtain the detection results of the sample images, where the detection results include the predicted regions where the targets are located in the sample images; According to the detection results, determine the pseudo-label information corresponding to the sample images, where the pseudo-label information includes the pseudo-annotation regions where the targets are located in the sample images, and the pseudo-annotation regions are selected from the predicted regions; According to the pseudo-label information and the preset anchor box information in the target detection network, determine the matching loss between the pseudo-annotation regions and the anchor box regions represented by the anchor box information; According to the matching loss, determine the network loss of the target detection network, and based on the network loss, train the target detection network; wherein, the anchor box information includes the classification scores predicted by the target detection network for the anchor box regions, the pseudo-label information includes the pseudo-annotation values corresponding to the pseudo-annotation regions, the pseudo-annotation values are used to indicate whether the target is included in the pseudo-annotation regions, and the classification scores represent the probabilities that the target is included in the anchor box regions; wherein, the determining the matching loss between the pseudo-annotation regions and the anchor box regions represented by the anchor box information according to the pseudo-label information and the preset anchor box information in the target detection network includes: Determine the classification loss between the pseudo-annotation regions and the anchor box regions according to the classification scores corresponding to the anchor box regions and the pseudo-annotation values corresponding to the pseudo-annotation regions; and / or, Determine the first position loss between the pseudo-annotation regions and the predicted regions corresponding to the anchor box regions according to the position information of the predicted regions corresponding to the anchor box regions and the position information of the pseudo-annotation regions; and / or, Determine the second position loss between the pseudo-annotation regions and the anchor box regions according to the position information of the anchor box regions and the position information of the pseudo-annotation regions; Determine the matching loss between the pseudo-annotation regions and the anchor box regions according to at least one of the classification loss, the first position loss, and the second position loss.
2. The method according to claim 1, characterized in that, the pseudo-label information includes multiple pseudo-annotation regions, wherein, the determining the network loss of the target detection network according to the matching loss includes: For each pseudo-annotation region among the multiple pseudo-annotation regions, determine the first weight corresponding to each of the multiple anchor box regions according to the matching loss between the pseudo-annotation region and the multiple anchor box regions, and the first weight of the same anchor box region is negatively correlated with the matching loss; Determine the region loss corresponding to the pseudo-annotation region according to the first weights corresponding to the multiple anchor box regions and the matching losses corresponding to the multiple anchor box regions; Determine the network loss of the target detection network according to the obtained region losses corresponding to the multiple pseudo-annotation regions.
3. The method according to claim 2, characterized in that, The pseudo-label information includes the confidence corresponding to the pseudo-labeled region, and the confidence characterizes the reliability of the pseudo-labeled region. Among them, determining the network loss of the target detection network according to the region losses corresponding to the obtained multiple pseudo-labeled regions includes: Determining the second weight corresponding to each of the multiple pseudo-labeled regions according to the confidence corresponding to each of the multiple pseudo-labeled regions, where the second weight is positively correlated with the confidence; Determining the network loss of the target detection network according to the second weight corresponding to each of the multiple pseudo-labeled regions and the region loss corresponding to each of the multiple pseudo-labeled regions.
4. The method according to any one of claims 1 to 3, characterized in that, The matching loss includes at least one of the following: The classification loss between the pseudo-labeled region and the anchor box region; The first position loss between the pseudo-labeled region and the predicted region corresponding to the anchor box region; The second position loss between the pseudo-labeled region and the anchor box region.
5. The method according to claim 1, characterized in that, The detection result includes the prediction scores corresponding to multiple predicted regions, and the prediction scores characterize the probability that the target is included in the predicted region; Among them, determining the pseudo-label information corresponding to the sample image according to the detection result includes: Clustering the multiple predicted regions according to the prediction scores corresponding to the multiple predicted regions to obtain the clustering result of the multiple predicted regions; Determining a screening threshold according to the clustering result of the multiple predicted regions, and based on the screening threshold, screening out pseudo-labeled regions from the multiple predicted regions; Determining the pseudo-label information corresponding to the pseudo-labeled region according to the detection result corresponding to the pseudo-labeled region.
6. The method according to claim 5, characterized in that, Clustering the multiple predicted regions according to the prediction scores corresponding to the multiple predicted regions to obtain the clustering result of the multiple predicted regions, includes: Using a Gaussian mixture model to cluster the multiple predicted regions according to the prediction scores corresponding to the multiple predicted regions to obtain the clustering result of the multiple predicted regions; Among them, the clustering result includes positive sample predicted regions and negative sample predicted regions among the multiple predicted regions, and the accuracy of the positive sample predicted regions is higher than that of the negative sample predicted regions.
7. The method according to claim 5 or 6, characterized in that, The clustering result includes positive sample predicted regions among the multiple predicted regions. Among them, determining a screening threshold according to the clustering result of the multiple predicted regions, and based on the screening threshold, screening out pseudo-labeled regions from the multiple predicted regions, includes: Determining the prediction score corresponding to the peak value of the Gaussian distribution corresponding to the positive sample predicted region as the screening threshold; Determining the positive sample predicted regions with prediction scores greater than or equal to the screening threshold as the pseudo-labeled regions.
8. The method according to any one of claims 1 to 3, characterized in that, Processing the sample image through the target detection network to obtain a detection result of the sample image, including: Performing feature extraction on the sample image to obtain classification features and regression features; Performing decoding processing on the regression features and the classification features to obtain an initial prediction region and a prediction score corresponding to the initial prediction region; Determining an offset corresponding to the initial prediction region according to the regression features and the classification features; Adjusting the initial prediction region according to the offset to obtain a prediction region.
9. The method according to claim 8, wherein, the determining an offset corresponding to the initial prediction region according to the regression features and the classification features includes: Performing splicing processing and convolution processing on the classification features and the regression features to obtain multi-scale bias features; Performing scale alignment on the multi-scale bias features to obtain a target bias feature with scale alignment; Performing decoding processing on the target bias feature to obtain an offset corresponding to the initial prediction region.
10. A target detection method, wherein, the method includes: Processing a to-be-detected image through a target detection network to obtain a target detection result of the to-be-detected image, where the target detection result includes a region where a target is located in the to-be-detected image; wherein, the target detection network is trained according to the network training method of any one of claims 1-9.
11. The method according to claim 10, wherein, the target includes a vehicle, and the processing the to-be-detected image through the target detection network to obtain the target detection result of the to-be-detected image includes: Processing the to-be-detected image through the target detection network to obtain a vehicle region where a vehicle is located in the to-be-detected image; wherein, after obtaining the vehicle region, the method further includes: Performing vehicle recognition on the vehicle region to obtain a vehicle recognition result of the vehicle in the vehicle region.
12. A network training device, wherein, including: An acquisition module, configured to acquire a sample image without label annotation and a to-be-trained target detection network; A detection module, configured to process the sample image through the target detection network to obtain a detection result of the sample image, where the detection result includes a prediction region where a target is located in the sample image; A pseudo-label determination module, configured to determine pseudo-label information corresponding to the sample image according to the detection result, where the pseudo-label information includes a pseudo-annotation region where a target is located in the sample image, and the pseudo-annotation region is selected from the prediction region; A matching loss determination module, configured to determine a matching loss between the pseudo-annotation region and an anchor region represented by the anchor box information according to the pseudo-label information and the anchor box information preset in the target detection network; A training module, configured to determine a network loss of the target detection network according to the matching loss, and train the target detection network based on the network loss; Among them, the anchor box information includes the classification scores predicted by the target detection network for the anchor box regions, the pseudo-label information includes the pseudo-label values corresponding to the pseudo-labeled regions, the pseudo-label values are used to indicate whether the target is included in the pseudo-labeled regions, and the classification scores represent the probabilities that the target is included in the anchor box regions; Among them, determining the matching loss between the pseudo-labeled region and the anchor box region represented by the anchor box information according to the pseudo-label information and the preset anchor box information in the target detection network includes: Determining the classification loss between the pseudo-labeled region and the anchor box region according to the classification score corresponding to the anchor box region and the pseudo-label value corresponding to the pseudo-labeled region; and / or, Determining a first position loss between the pseudo-labeled region and the predicted region corresponding to the anchor box region according to the position information of the predicted region corresponding to the anchor box region and the position information of the pseudo-labeled region; and / or, Determining a second position loss between the pseudo-labeled region and the anchor box region according to the position information of the anchor box region and the position information of the pseudo-labeled region; Determining the matching loss between the pseudo-labeled region and the anchor box region according to at least one of the classification loss, the first position loss, and the second position loss.
13. A target detection device, Characterized in that, The device includes: An image detection module, configured to process a to-be-detected image through a target detection network to obtain a target detection result of the to-be-detected image, where the target detection result includes regions where targets are located in the to-be-detected image; Among them, the target detection network is trained according to the network training method of any one of claims 1-9.
14. An electronic device, Characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Among them, the processor is configured to call the instructions stored in the memory to execute the method of any one of claims 1 to 11.
15. A computer-readable storage medium, on which computer program instructions are stored, Characterized in that, When the computer program instructions are executed by a processor, the method of any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
3D target detection method and device, electronic equipment and storage medium
CN112163541A
Twin network target tracking method based on dynamic label distribution and mobile equipment
CN113255611A