Method for processing image target classification and localization tasks using a neural network

The novel binary target detection neural network architecture addresses performance inconsistencies by integrating multi-dimensional joint matching and dynamic learning weights, enhancing classification and localization accuracy for resource-constrained devices.

CN114841307BActive Publication Date: 2025-07-15BEIJING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210197861.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-01
Publication Date
2025-07-15
Estimated Expiration
2042-03-01

AI Technical Summary

Technical Problem

The existing binary object detection neural network has inconsistent performance in target positioning and classification tasks, resulting in a decrease in detection accuracy.

Method used

A binary object detection neural network including backbone network, shared feature pool network, classification decoupling network and positioning decoupling network is built. Through multi-dimensional joint matching object detection task consistency training and synchronous optimization, an improved anchor box sampling strategy and correlation constraint loss function are adopted, combined with a dynamic learning weighted target loss function to improve network information capacity and detection accuracy.

Benefits of technology

The binarized object detection neural network has been improved in consistency in classification and positioning tasks, improved detection accuracy and robustness, and is suitable for edge computing devices with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114841307B_ABST
    Figure CN114841307B_ABST
Patent Text Reader

Abstract

The present invention provides a method for processing image target classification and localization tasks using a neural network. The method includes: constructing a binary object detection neural network, which includes a backbone network, a shared feature pool network, a classification decoupling network, and a localization decoupling network; performing consistency training of the binary object detection neural network for object detection tasks based on multi-dimensional joint matching; and synchronously optimizing the classification and localization tasks of the binary object detection neural network. The present invention solves the task inconsistency problem of Anchor sampling in the binary object detection neural network through an improved Anchor sampling strategy and a new loss function algorithm based on correlation constraints, and synchronously optimizes the classification and localization tasks of the binary object detection neural network through an object loss function with dynamically learnable weights, which can improve the quality of detection boxes, improve the detection accuracy of the binary object detection neural network, and the robustness of the algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural networks, and in particular, to a method for processing classification and positioning tasks using a binary object detection neural network. Background Art

[0002] The binary quantization of an object detection neural network refers to compressing a neural network in 32-bit floating-point format into a 1-bit fixed-point format to reduce storage and computational costs. Binarizing the weights and activations of an object detection neural network can reduce the storage by 32 times and the computational cost by 64 times. These features make the binary object detection neural network particularly suitable for deployment on low-cost edge computing devices with limited resources.

[0003] Currently, the binary object detection neural networks in the prior art include Bi-Det, Auto-BiDet, and LWS-Det. Among them, Bi-Det mainly removes redundant information in the binary object detection neural network through the information bottleneck theory, that is, restricts the amount of information in the high-level feature map and maximizes the mutual information between the feature map and the object detection head. On the basis of Bi-Det, Auto-BiDet adds the function of controlling the compression level of the information bottleneck according to the characteristics of the input data, that is, adopts a lower compression level for pictures with low complexity and a higher compression level for pictures with high complexity, realizing the dynamic compression of the amount of information in the high-level feature map. LWS-Det (Layer-wise Searching for 1-bit Detectors) introduces angular and amplitude loss functions to increase the capacity of the binary object detection neural network. In the 1-bit quantization layer, this method uses differentiable binarization search to minimize the angular error in the student-teacher guidance network framework, and learns the scale factor by minimizing the amplitude loss in the same student-teacher guidance network framework, so as to increase the network capacity of the binary object detection neural network and improve the performance of the binary object detection neural network.

[0004] Although the binary object detection neural network can effectively reduce the storage and computational costs, due to the limited information capacity of the binary neural network itself, there is a serious problem of unbalanced extraction of object localization and object classification feature information in the existing binary object detection neural networks in the current art (i.e., the neural network shows inconsistent performance in localization and classification tasks). Compared with the full-precision object detection neural network, the binary object detection neural network will face a significant decrease in detection accuracy when deployed and applied in actual scenarios. Taking the benchmark object detection neural network SSD300-VGG16 as an example, the accuracy of the typical binary object detection neural network BiDet based on it is only 66.0% (mAP) on the PASCAL VOC dataset. Compared with the accuracy of the corresponding benchmark full-precision object detection neural network of 74.3% on the PASCAL VOC dataset, the accuracy has decreased by 8.3%. In addition, another further optimized binary object detection neural network AutoBiDet achieved 14.3% (mAP@[.5,.95]) on the COCO dataset. Compared with the accuracy of the corresponding full-precision neural network of 23.2% (mAP@[.5,.95]) on the COCO dataset, the accuracy has decreased by 8.9%. Summary of the Invention

[0005] An embodiment of the present invention provides a method for processing image object classification and localization tasks using a neural network to achieve better performance in the consistency of classification and localization tasks of a binary object detection neural network.

[0006] To achieve the above object, the present invention adopts the following technical solutions.

[0007] A method for processing image object classification and localization tasks using a neural network includes:

[0008] Construct a binary object detection neural network, where the binary object detection neural network includes a backbone network, a shared feature pool network, a classification decoupling network, and a localization decoupling network;

[0009] Perform consistency training of object detection tasks based on multi-dimensional joint matching on the binary object detection neural network;

[0010] Perform synchronous optimization of classification and localization tasks on the binary object detection neural network.

[0011] Preferably, the construction of the binary object detection neural network, where the binary object detection neural network includes a backbone network, a shared feature pool network, a classification decoupling network, and a localization decoupling network, includes:

[0012] Construct a binary object detection neural network including a backbone network and a shared feature pool network, perform network feature decoupling branch processing on the shared feature pool network to obtain two sets of feature decoupling branch networks including a series of feature decoupling blocks;

[0013] Use one of the feature decoupling branch networks to perform classification task feature learning to obtain a classification decoupling network, and use the other feature decoupling branch network to perform localization task feature learning to obtain a localization decoupling network. The feature decoupling blocks in the feature decoupling branch network are connected to the shared feature pool network, and each feature decoupling block learns specific features by applying a decoupling code to a specific layer of the shared feature pool network.

[0014] Preferably, the consistency training of the binary object detection neural network for object detection tasks based on multi-dimensional joint matching includes:

[0015] Design an improved anchor box sampling strategy. This anchor box sampling strategy comprehensively considers multi-modal information such as the position information and semantic information of the anchor box, and modifies the intersection over union (IOU) between the ground truth label and the anchor box through the confidence score Conf_score of the detection box Anchor , to obtain the modified intersection over union (IOU) Amendment , as shown in formula (5):

[0016]

[0017] σ and Th r take constant values, where σ is a hyperparameter used to adjust the intensity of IOU modification, and Th r is the confidence score screening threshold;

[0018] The anchor box sampling strategy adopts a new correlation constraint loss function L relevance , L relevance By increasing the linear correlation between the confidence score Conf_score of the detection box and the modified intersection over union (IOU) between it and the corresponding ground truth label Amendment , to reduce the gap between Conf_score and the ground truth label, and increase the consistency of the performance evaluation indicators for classification and localization tasks, as shown in formula (6):

[0019] L relevance = |Conf_score - IOU Amendment | (6)

[0020] Preferably, the synchronous optimization of the classification and localization tasks for the binary object detection neural network includes:

[0021] Design a target loss function with dynamically learnable weights, and calculate the relative change values a of the loss objective functions for the classification and localization tasks respectively cls (t - 1) and a loc (t - 1), as shown in Formulas (7) and (8), where t represents the training time, and the distillation temperature T is added to the softmax layer, as shown in Formulas (9) and (10), to obtain the dynamic weight values λ cls (t) and λ loc (t), and finally obtain the target loss function L of the dynamically learnable weights for object detection loss (t), as shown in Formula (11). This target loss function synchronously optimizes the object detection classification and localization tasks by dynamically learning weights:

[0022]

[0023] L loss (t) = λ cls (t)L cls (t) + λ loc (t)L loc (t) + L relevance (11)

[0024] where t represents the number of neural network training iterations, L cls (t - 1) and L loc (t - 1) represent the classification and localization Loss values at the (t - 1)-th iteration respectively; a cls (t - 1) and a loc (t - 1) represent the relative change values of the Loss values for the classification and localization tasks respectively; T represents the distillation temperature to control the softness of different task weights; K represents the number of tasks, and in the object detection network, K = 2, that is, it includes the classification task and the localization task; λ cls (t) and λ loc (t) represent the dynamic weight values of the classification and localization loss functions respectively

[0025] It can be seen from the technical solutions provided by the embodiments of the present invention described above that the embodiments of the present invention improve the network information capacity of the binary object detection neural network by adding a classification decoupling network and a localization decoupling network, and avoid the problem of unbalanced extraction of classification and localization feature information; by designing and improving the Anchor sampling and a new loss function algorithm based on correlation constraints, the task inconsistency problem of Anchor sampling in the binary object detection neural network is solved, and the binary object detection neural network is synchronously optimized for the classification and localization tasks through the target loss function with dynamically learnable weights, which can improve the quality of the detection box, improve the detection accuracy of the binary object detection network and the robustness of the algorithm

[0026] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 It is a schematic diagram of a novel structure of a binarized object detection neural network provided for an embodiment of the present invention.

[0029] Figure 2 It is a training flow chart of the structure and model of the binarized object detection neural network provided for an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0031] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any and all combinations of one or more of the associated listed items.

[0032] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which this invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless defined as here.

[0033] For the convenience of understanding the embodiments of the present invention, the following will further explain with several specific embodiments in conjunction with the accompanying drawings, and each embodiment does not constitute a limitation to the embodiments of the present invention.

[0034] The embodiments of the present invention provide a new binarized object detection neural network that can be deployed for actual scenarios, perform consistency training on the object detection task of multi-dimensional joint matching for this network, and synchronously optimize the classification and localization tasks of this network, so that the classification performance and localization performance of the finally obtained detection boxes are both relatively excellent, greatly reducing the time and computational expenses of the binarized object detection neural network, and can be better deployed on edge devices with limited hardware resources such as embedded and mobile devices.

[0035] Specifically, in the constructed binarized object detection neural network, the neural network with a multi-level structure automatically learns task-sharing features and task-specific features in an end-to-end manner, thereby eliminating the inconsistency of features in classification and localization, and effectively improving the representation information capacity of the binarized object detection neural network. At the same time, in order to further solve the task inconsistency problem of anchor sampling, perform consistency training on the binarized object detection neural network for multi-dimensional joint matching, and optimize and retain high-quality (both correctly classified and correctly localized) detection boxes by introducing an improved anchor sampling strategy and a new loss function based on correlation constraints. Finally, synchronously optimize the binarized object detection neural network through an object loss function with dynamically learnable weights, and finally obtain detection boxes with relatively excellent classification performance and localization performance.

[0036] The processing flow of a method for processing image object classification and localization tasks using a neural network provided by the embodiments of the present invention is as Figure 2 shown, including the following processing steps:

[0037] Step S10: Construct a binarized object detection neural network.

[0038] The new network structure of a binarized object detection neural provided by the embodiments of the present invention is as Figure 1As shown in the figure. It includes a backbone network and a shared feature pool network, and performs network feature decoupling branch processing on the shared feature pool network to obtain a number of feature decoupling branch networks including a series of feature decoupling blocks and corresponding task detection heads (heads). Figure 1 In the embodiment shown, two sets of feature decoupling branch networks are used. One of the feature decoupling branch networks is used to learn classification task features to obtain a classification decoupling network, and the other feature decoupling branch network is used to learn localization task features to obtain a localization decoupling network. The feature decoupling blocks in the feature decoupling branch network are connected to the shared feature pool network, and each feature decoupling block learns specific features by applying a decoupling code to a specific layer of the shared feature pool network.

[0039] First, binarize the baseline network

[0040] The object detection neural network extracts multi-scale object detection shared features. Specifically, the VGG16 network structure (other network structures can also be selected, such as ResNet, MobileNet, etc.) is selected as the backbone network. The shared feature pool network is used as a global feature pool after the backbone network. If the features of this global feature pool are directly used for classification or localization, it will cause information mismatch or conflict in the performance of different tasks. Therefore, the binarized object detection neural network proposed in the present invention uses several independent feature decoupling branch networks to learn classification task features and localization task features respectively. Specifically, Figure 1 In the embodiment, two feature decoupling branch networks, as well as a classification decoupling network and a localization decoupling network are used. A series of feature decoupling blocks in the feature decoupling branch network are connected to the shared feature pool network branch. Each feature decoupling block learns specific features by applying a decoupling code to a specific layer of the shared feature pool network, where the decoupling code is a feature selector automatically learned in an end-to-end manner during the training process of the entire object detection network; during the training process of the object detection network, the learning of the decoupling code does not directly affect the learning processes of the backbone network and the shared feature pool network (that is, there is no direct loss function constraint relationship). Therefore, the features of the shared feature pool network layer and the feature decoupling network can learn together, aiming to maximize the generalization ability of the shared features in classification and localization tasks, and the feature decoupling code can also maximize the overall classification or localization performance of the object detection network.

[0041] The specific design and training method of the binarized object detection neural network proposed by the present invention are as follows: First, the present invention defines the shared features output by the j-th convolutional layer of the shared feature pool network as For classification or localization tasks, it is through the features in the shared feature pool network Apply decoupling codes to filter features. Each feature channel has a decoupling code. The feature decoupling code of the j-th building block of the classification decoupling network is called The feature decoupling code of the j-th building block of the localization decoupling network is called Then, the first decoupling block of the network feature decoupling branch takes only the features of the shared network layer as input. However, for subsequent decoupling blocks, the input is the shared features of the current layer and the specific task features of the previous layer Or connection, where Or will pass through a 3*3 convolution f (j) and be passed to the current layer; then respectively pass through two 1*1 convolutions and Or and and then through a sigmoid function, the feature decoupling code will be obtained Or The decoupling code is a mask signal between [0,1] learned through a backpropagation self-supervised method, as shown in formulas (1)(2). Finally, the decoupling code and the corresponding shared layer features of the shared feature pool network of the corresponding layer are multiplied pixel by pixel to obtain the task-aware features of the corresponding layer and When the value of the decoupling code is 1, the features of the feature decoupling branch and the features of the feature sharing layer will be equal, as shown in formulas (3)(4), where · represents pixel-by-pixel multiplication operation.

[0042]

[0043] represents the output shared features of the j-th building block of the shared feature pool network; and respectively represent the classification task features and localization task features of the (j-1)-th building block; and are two convolutional layers of the j-th building block of the classification decoupling network; and are two convolutions of the j-th building block of the localization decoupling network; f (j) is the shared 3*3 convolutional layer of the j-th building block of the two decoupling networks; and are the classification feature decoupling code and the localization feature decoupling code respectively; and are the classification task-aware features and the localization task-aware features respectively. The above convolutional layers can also be 1*1 convolutional layers or convolutional layers of other sizes, which does not affect the implementation effect of the method of the present invention.

[0044] Step S20: Train the binarized object detection neural network through multi-dimensional joint matching.

[0045] The present invention proposes a method for training the consistency of object detection tasks based on multi-dimensional joint matching. The main task of this method is to perform multi-dimensional joint matching learning in multiple processing stages of the classification and localization tasks of the object detection neural network, so that the classification performance and localization performance of the finally obtained detection boxes are both excellent.

[0046] First, a predefined anchor sampling strategy is designed; this strategy aims to optimize anchor sampling and comprehensively considers multi-modal information such as the position information and semantic information of the anchor, that is, not only considering the intersection over union (IOU) between the anchor and the GT (Ground Truth, true label or true detection box) Anchor , but also fully considering the richness of the semantic information contained in the anchor itself. As shown in formula (5), where σ is a hyperparameter used to adjust the intensity of the corrected IOU, Th r is the confidence score screening threshold. The specific values of σ and Th r can be set according to different training data sets. For example, on the PASCAL VOC data set, σ takes a constant value of 2, and Th r takes a constant value of 0.1 to achieve better detection results; the IOU between the GT and the anchor is corrected through the confidence score Conf_score of the detection box Anchor to obtain the corrected IOU Amendment . The purpose is to correct some detection boxes that were originally defined as negative samples but have rich semantics into positive samples; at the same time, correct the detection boxes that were originally defined as positive samples but have less semantic information into negative samples, effectively reducing the misguidance of interference samples on the training process and improving the accuracy of the training results.

[0047]

[0048] Among them, Conf_score is the confidence score of the detection box, σ is a hyperparameter used to control the degree of correction of the Conf_score of the detection box, IOU Anchor is the intersection over union between the anchor and the GT, and Th r is the threshold used to determine whether the confidence of the detection box is too low, that is, the confidence screening threshold.

[0049] Then, to address the phenomenon that detection boxes are incorrectly suppressed (targets are missed) during the NMS (Non Maximum Suppression) post - processing, that is, the NMS algorithm first sorts the detection boxes according to the confidence scores, and the detection boxes with high confidence scores are more likely to be retained. However, some detection boxes with high IOU scores and sub - high confidence scores are extremely likely to be incorrectly suppressed. Therefore, the algorithm adopts a new correlation - constraint loss function L relevance ,L relevance By increasing the linear correlation between the confidence score Conf_score of the detection box and the corrected intersection - over - union IOU between it and the corresponding GT Amendment to minimize the gap between the two and increase the consistency of the performance evaluation indicators for classification and localization tasks. Specifically, since the value ranges of both the confidence score Conf_score and the IOU Amendment are in [0, 1], the absolute - value difference between the two is directly used to measure the distance between them. As shown in formula (6), it simply and efficiently realizes the improvement of the linear correlation between the Conf_score of the retained detection boxes and the corrected IOU Amendment , thereby greatly improving the consistency between the confidence score, the measurement index of the detection - box classification effect, and the IOU, the measurement index of the localization effect.

[0050] L relevance = |Conf_score - IOU Amendment | (6)

[0051] where Conf_score represents the confidence score of the detection box, and IOU detected-box represents the intersection - over - union between the detection box and the corresponding GT.

[0052] Step S30: Synchronously optimize the classification and localization tasks of the binary object - detection neural network.

[0053] To effectively avoid the task - inconsistency phenomenon (many false alarms and missed detections) in the results of the binary object - detection neural network and achieve the goal of synchronously optimizing the classification and localization tasks of object detection. Therefore, the weighting method of the loss objective function during network training no longer uses a fixed value but is dynamically adjusted according to the learning effects and difficulties of the classification and localization tasks. To achieve this goal, we propose a dynamic - weight learning strategy. The purpose of this strategy is to make the classification and localization tasks learn at a similar speed. The specific process is as follows: First, calculate the relative change values a cls (t - 1) and a loc (t - 1) of the loss - objective - function (Loss function) values of the classification and localization tasks respectively, as shown in formulas (7)(8), where t represents the training time. To make the output acls (t - 1) and a loc (t - 1) has better learning effect. The distillation temperature T is added to the softmax layer to improve the performance of distillation. As shown in formulas (9) and (10), the dynamic weight values λ of the classification and localization loss functions are obtained. cls (t) and λ loc (t). Finally, the loss objective function L of object detection with dynamically learnable weights is obtained. loss (t), as shown in formula (11). This function realizes the synchronous optimization of object detection classification and localization tasks by dynamically learning weights, achieves the task consistency goal of detection results, and enhances the stability of network training.

[0054]

[0055] L loss (t) = λ cls (t)L cls (t) + λ loc (t)L loc (t) + L relevance (11)

[0056] Where t represents the number of neural network training iterations, L cls (t - 1) and L loc (t - 1) represent the classification and localization Loss values at the (t - 1)-th iteration respectively; a cls (t - 1) and a loc (t - 1) represent the relative change values of the Loss values of the classification and localization tasks respectively; T represents the distillation temperature to control the softness of different task weights; K represents the number of tasks. In the object detection network, K = 2, that is, it includes the classification task and the localization task; λ cls (t) and λ loc (t) represent the dynamic weight values of the classification and localization loss functions respectively.

[0057] The binary object detection neural network of the present invention, with its advantages of extremely high model compression rate and extremely low computational complexity, can be applied to devices with limited computing resources, such as embedded devices and mobile devices based on mobile phones, etc., for classification and localization task processing. The classification and localization task consistency of the binary object detection neural network implemented by the present invention greatly improves the detection accuracy, can ensure that the neural network can have detection frames with high classification confidence and high localization accuracy at the same time, and can replace the full-precision object detection neural network algorithm to meet the requirements of high-precision and low-cost object detection algorithms in actual application scenarios.

[0058] In summary, the embodiments of the present invention construct a binary object detection neural network through the effective combination of a benchmark binary object detection neural network and a network feature decoupling branch to solve the problem of unbalanced feature information extraction in classification and localization tasks due to insufficient representation ability of the binary object detection neural network; and perform task consistency training of multi-dimensional joint matching on the constructed neural network to solve the problem of task inconsistency in Anchor sampling of the binary object detection neural network; finally, through an object loss function with dynamically learnable weights, synchronously optimize the classification and localization tasks of the binary object detection neural network; ultimately achieve a detection result with better performance of task consistency in the classification and localization tasks of the binary object detection neural network.

[0059] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.

[0060] From the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0061] Each embodiment in this specification is described in a progressive manner. For the same or similar parts between the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0062] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for processing image target classification and localization tasks using a neural network, characterized in that, Including: Construct a binary object detection neural network, which includes a backbone network, a shared feature pool network, a classification decoupling network, and a localization decoupling network; Use image data to perform consistency training on the binary object detection neural network for object detection tasks based on multi-dimensional joint matching; Synchronously optimize the classification and localization tasks of the binary object detection neural network to obtain a trained binary object detection neural network; Deploy the trained binary object detection neural network on an edge device with limited hardware resources to process object classification and localization tasks in the input image; The construction of the binary object detection neural network, which includes a backbone network, a shared feature pool network, a classification decoupling network, and a localization decoupling network, includes: Construct a binary object detection neural network including a backbone network and a shared feature pool network, perform network feature decoupling branch processing on the shared feature pool network to obtain two groups of feature decoupling branch networks including a series of feature decoupling blocks; Use one of the feature decoupling branch networks to perform classification task feature learning to obtain a classification decoupling network, and use the other feature decoupling branch network to perform localization task feature learning to obtain a localization decoupling network. The feature decoupling blocks in the feature decoupling branch network are connected to the shared feature pool network, and each feature decoupling block learns specific features by applying a decoupling code to a specific layer of the shared feature pool network; The consistency training of the binary object detection neural network for object detection tasks based on multi-dimensional joint matching includes: Design an improved anchor box sampling strategy that comprehensively considers the multi-modal information of the position information and semantic information of the anchor box, and corrects the intersection over union (IOU) between the ground truth label and the anchor box through the confidence score Conf_score of the detection box Anchor , and obtain the corrected intersection over union (IOU) Amendment , and the specific formula (5) is shown as follows: σ and Th r take constant values, where σ is a hyperparameter used to adjust the intensity of the intersection over union correction, and Th r is the confidence score screening threshold value; The anchor box sampling strategy adopts a new correlation constraint loss function L relevance , L relevance By increasing the linear correlation between the confidence score Conf_score of the detection box and the corrected intersection over union IOU between it and the corresponding ground truth label Amendment to reduce the gap between Conf_score and the ground truth label and increase the consistency of the performance evaluation indicators for the classification and localization tasks. The specific formula (6) is shown as follows: L relevance = |Conf_score - IOU Amendment | (6).

2. The method according to claim 1, wherein The synchronous optimization of the classification and localization tasks of the binary object detection neural network includes: Design a target loss function with dynamically learnable weights, which realizes the synchronous optimization of object detection classification and localization tasks by dynamically learning weights: calculate the relative change values a cls (t - 1) and a loc (t - 1) respectively. The specific formulas (7) and (8) are as follows: Where t represents the training time, and the distillation temperature T is added to the softmax layer. The specific formulas (9) and (10) are as follows: Obtain the dynamic weight value λ of the classification and localization loss functions cls (t) and λ loc (t), and finally obtain the objective loss function L of the dynamically learnable weights for object detection loss (t), and the specific formula (11) is as follows: L loss (t) = λ cls (t)L cls (t) + λ loc (t)L loc (t) + L relevance (11) where t represents the number of neural network training iterations, L cls (t - 1) and L loc (t - 1) represent the classification and localization Loss values at the (t - 1)-th iteration respectively; a cls (t - 1) and a loc (t - 1) represent the relative change values of the Loss values for the classification and localization tasks respectively; T represents the distillation temperature to control the softening degree of different task weights, K represents the number of tasks, in the object detection network K = 2, that is, it includes the classification task and the localization task, λ cls (t) and λ loc (t) represent the dynamic weight values of the classification and localization loss functions respectively.

Citation Information

Patent Citations

  • Target detection method for optimizing classification and positioning tasks

    CN112149664A

  • Unmanned aerial vehicle target tracking method based on anchor frame matching and Siamese network

    CN113807188A