A target detection method and system based on a DQA mechanism and PointSR-R
Patent Information
- Application Number
- CN202610892913.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-21
- Publication Date
- 2026-09-25
AI Technical Summary
第一,时间集成编码器中的固定集成系数无法根据各伪框的实际质量进行自适应调整,这导致原型库更新受低质量伪框污染,且
的人工搜索过程繁琐低效;第二,PointSR方案仅能适用于水平框检测,未提供面向旋转目标检测的技术方案,且现有的旋转目标检测方案缺乏对伪框空间重叠、同类目标几何一致性以及包内候选框结构关联性的约束机制
[0020]第一,实现了锚点原型库更新的自适应优化,提升了伪框生成的质量与鲁棒性。本申请使用DQA机制,摒弃了现有技术中依赖人工经验设定的全局固定集成系数,转而采用逐样本自适应的动态集成系数,通过对每个伪框进行定位一致性和分类可靠性的双维度质量评估,能够自动识别并赋予高质量伪框更高的更新权重,同时抑制低质量伪框对原型库的污染。这种设计不仅消除了繁琐的人工超参数搜索过程,降低了模型部署成本,还使得原型库能够根据训练过程中的实际样本质量进行精准调整,显著提高了伪框生成的空间准确性,并增强了模型对手动标注扰动的鲁棒性。
Smart Images

Figure CN122821402A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this application relate to the field of target detection technology, and in particular to a target detection method and system based on DQA mechanism and PointSR-R. Background Technology
[0002] With the rapid development of drone technology and the continuous expansion of low-altitude application scenarios, target detection from the drone's perspective has become a core technological requirement in fields such as smart cities, traffic inspection, logistics transportation, and emergency rescue. In actual drone aerial photography scenarios, targets are usually characterized by dense distribution, drastic scale changes, and arbitrary orientation, and are often accompanied by severe occlusion, which places extremely high demands on the accuracy and robustness of target detection.
[0003] Chinese invention patent CN120279444A proposes a UAV-based target detection method and system based on self-regularized point supervision (hereinafter referred to as the PointSR scheme). The self-regularized point-supervised target detection framework in the PointSR scheme comprises three core components: the first is a temporal ensemble encoder (TE Encoder), which aggregates the aspect ratio information of historical pseudo-boundaries using an exponential moving average (EMA) method, with a fixed ensemble coefficient. The PointSR scheme employs a weighted average of the current pseudo-boundary aspect ratio and historical prototype values to construct a category-aware anchor point prototype library, enabling dynamic adjustment of anchor point shapes. Secondly, an Information Sample Collector (IS Collector) filters high-uncertainty negative samples by calculating the loss value of a Multi-Instance Learning (MIL) classifier and removes redundancy using non-maximum suppression, collecting information-rich negative samples. Thirdly, a pseudo-boundary optimization module uses informative negative samples to suppress and fine-tune coarse-grained pseudo-boundaries, outputting the final horizontal pseudo-boundary. Using only point annotations, the PointSR scheme effectively improves the accuracy and robustness of target detection from a UAV perspective, representing the current advanced level of point-supervised target detection technology.
[0004] Although the temporal ensemble encoder in the PointSR scheme effectively utilizes prior information from historical pseudoboxes to dynamically adjust anchor point shapes, its anchor point prototype library update mechanism has a critical flaw: it uses globally uniform fixed coefficients. The prototype library is updated indiscriminately for all pseudo-boundaries. This coefficient can only be set manually and cannot be adaptively adjusted according to the actual quality of each pseudo-boundary during training. When a pseudo-boundary has a large positional deviation from the target center indicated by the point label or has a low classification confidence, its contribution to the prototype library update is the same as that of high-quality pseudo-boundaries. This causes low-quality pseudo-boundaries to introduce incorrect prior information into the prototype library, causing the temporal ensemble encoder to converge to a suboptimal solution, ultimately limiting further improvement in the accuracy of point-supervised target detection.
[0005] Furthermore, the PointSR solution only addresses the generation and optimization of horizontal bounding boxes. However, target detection from a UAV perspective typically exhibits characteristics such as dense, elongated, and arbitrarily oriented distribution. Therefore, compared to horizontal boxes, rotated boxes can more accurately describe the target's attitude and contour, making them more suitable for precise localization in aerial photography scenarios. Moreover, existing point-supervised rotated target detection (P-ROD), when only point annotations are available, lacks direct supervision of the rotated box's angle and shape, leading to difficulties in estimating pseudo-box orientation and unstable pseudo-label quality.
[0006] It's also important to note that existing solutions still have shortcomings in dense scenes. First, different target pseudo-boundaries are prone to overlap. While this overlap is acceptable for horizontal bounding boxes, it's generally undesirable for rotated bounding boxes, as the introduction of angles allows for a more accurate alignment with the target's actual direction. Second, the geometric consistency of objects within the same category is not fully utilized. For target detection from an aerial perspective, objects of the same category should maintain a consistent shape and size in the image, while objects of different categories should have different shapes and sizes. For example, in the same image, cars should have the same shape and size, while cars and bicycles should have different shapes and sizes. Third, the internal relational constraints of the proposal package are insufficient, affecting the quality of the final rotated pseudo-boundaries. Each candidate box is generated individually through dithering, lacking explicit constraints between them. If the dithering range is too small, the generated candidate boxes will be close to the initial coarse pseudo-boundary, affecting subsequent optimization. If the dithering range is too large, it will introduce background interference, resulting in a large amount of noise.
[0007] In summary, the PointSR scheme suffers from the following core defects that this application aims to address. First, the fixed integration coefficient in the time-integrated encoder... The inability to adaptively adjust based on the actual quality of each pseudo-box results in the prototype library updates being polluted by low-quality pseudo-boxes, and First, the manual search process is cumbersome and inefficient. Second, the PointSR scheme is only applicable to horizontal bounding box detection and does not provide a technical solution for rotating target detection. Furthermore, existing rotating target detection schemes lack constraints on the spatial overlap of pseudo-bounding boxes, the geometric consistency of similar targets, and the structural correlation of candidate boxes within the bag. Summary of the Invention
[0008] To address the aforementioned technical issues, embodiments of this application propose a target detection method and system based on the DQA mechanism and PointSR-R. This method integrates localization consistency assessment, classification reliability assessment, and a three-layer spatial structure regularization mechanism. By performing a two-dimensional quality assessment on each pseudo-box and adaptively generating integration coefficients, it achieves a systematic improvement in high-quality anchor point prototype library updates and rotation target detection accuracy.
[0009] To achieve the above objectives, embodiments of this application propose a target detection method based on DQA mechanism and PointSR-R, comprising: acquiring target images to be detected via a UAV and generating feature vectors of candidate regions through a region proposal mechanism; constructing a PointSR-R target detection network model based on a weakly supervised learning target detection network, wherein the PointSR-R target detection network model includes a temporal ensemble encoder with embedded DQA mechanism, a coarse rotation pseudo-box generation module, a spatial optimization module, and a pseudo-box refinement module, wherein the spatial optimization module constrains the rotation pseudo-box generation process from three levels of relationships: target level, category level, and bag level, through spatial structure regularization; and refining the feature vectors of candidate regions based on the PointSR-R target detection network model. Point-supervised target detection is performed to obtain point-supervised target detection results from the perspective of a UAV. Among them, the DQA mechanism is embedded in the prototype library update stage of the temporal ensemble encoder. For the candidate bounding boxes corresponding to each point label, the quality of pseudo-boxes is evaluated from two dimensions: localization consistency and classification reliability. Localization consistency is measured by Gaussian mixture model to measure the deviation of the center distribution of candidate bounding boxes from the point label position. Classification reliability is measured by Softmax normalization of the classification scores of candidate bounding boxes and calculation of information entropy to measure prediction uncertainty. Based on the quality evaluation indicators of the two dimensions, a sample-by-sample adaptive ensemble coefficient is calculated to replace the empirically set fixed temporal ensemble coefficient, thereby updating the prototype library of anchor points of each category and reducing the impact of low-quality pseudo-boxes on the prototype library update.
[0010] To achieve the above objectives, embodiments of this application also propose a target detection system based on DQA mechanism and PointSR-R, comprising: a target image acquisition and feature map extraction module, used to acquire the target image to be detected through a UAV and generate feature vectors of candidate regions through a region proposal mechanism; a PointSR-R target detection network model construction module, used to construct a PointSR-R target detection network model based on a weakly supervised learning target detection network, wherein the PointSR-R target detection network model is provided with a time ensemble encoder with embedded DQA mechanism, a coarse rotation pseudo-box generation module, a spatial optimization module, and a pseudo-box refinement module, the spatial optimization module constraining the rotation pseudo-box generation process from three levels of relationships: target level, category level, and bag level through spatial structure regularization; and a target detection execution module, used to perform target detection based on P... The ointSR-R object detection network model performs point-supervised object detection on the feature vectors of candidate regions to obtain point-supervised object detection results from the perspective of a UAV. The DQA mechanism is embedded in the prototype library update stage of the temporal ensemble encoder. For each candidate bounding box corresponding to a point label, the quality of pseudo-boxes is evaluated from two dimensions: localization consistency and classification reliability. Localization consistency measures the deviation of the center distribution of the candidate bounding box from the point label position using a Gaussian mixture model. Classification reliability measures prediction uncertainty by performing Softmax normalization on the classification scores of the candidate bounding boxes and calculating information entropy. Based on the two quality evaluation metrics, a sample-by-sample adaptive ensemble coefficient is calculated to replace the empirically set fixed temporal ensemble coefficient, thereby updating the prototype library of category anchors and reducing the impact of low-quality pseudo-boxes on the prototype library update.
[0011] To achieve the above objectives, embodiments of this application also propose a computer-readable storage medium storing a computer program that, when executed by a processor, enables a target detection method based on the DQA mechanism and PointSR-R as described above.
[0012] Optionally, the DQA mechanism is embedded in the prototype library update phase of the time-integrated encoder, and its execution flow includes: Candidate box extraction, in the first In the coarse-grained pseudo-frame generation stage, for the first... Individual tags First, the depth features of each candidate box are extracted through the RoI pooling layer and two shared fully connected layers. Then, the depth features of each candidate box are input into the MIL dual-stream classifier, and the softmax operations of the classification stream and detection stream are performed respectively to obtain the class prediction score and detection confidence score of each candidate box. The top candidates are selected in descending order of comprehensive score. Each candidate suggestion box constitutes a... The proposal package , serving as the basic input unit for subsequent two-dimensional quality assessment; Positioning consistency assessment, quantifying from the positioning level. and The degree of spatial agreement between them is measured by a Gaussian mixture model. The central distribution relative to The deviation in position is used to obtain positioning quality assessment indicators; Classification reliability assessment introduces information entropy at the classification level to supplement the assessment dimensions, through... The classification scores are normalized using Softmax and the information entropy is calculated to measure the prediction uncertainty, thus obtaining the classification quality evaluation index. The adaptive fusion coefficients are generated sample-by-sample by substituting the localization quality assessment index and the classification quality assessment index into the linear fusion formula for calculation. Corresponding sample-by-sample adaptive integration coefficients ; Prototype library differential updates, using Replace the fixed-time integration coefficient set by experience Perform differential weighted updates on the aspect ratios of the target objects stored in the prototype library for each type of anchor point.
[0013] Optionally, quantification can be performed at the positioning level. and The degree of spatial fit between them, through the The classification scores are Softmax normalized and information entropy is calculated to measure prediction uncertainty, resulting in classification quality evaluation metrics, including: by center coordinates As the mean parameter of the corresponding Gaussian component, the detection confidence score of the detection stream output is used. As a mixing ratio ,by Compared to center Constructing a shared covariance matrix from coordinate deviations ; ; based on and Establish a Gaussian mixture model, and Substituting the values into a Gaussian mixture model to calculate its probability density value, we obtain the classification quality assessment index. ; ; in, This represents the established Gaussian mixture model. The higher the value, the more it indicates The more the space is close to The higher the accuracy of the indicated target location, the better the positioning quality. Introducing information entropy at the classification level to supplement the evaluation dimensions, through... The classification scores are Softmax normalized and information entropy is calculated to measure prediction uncertainty, resulting in classification quality evaluation metrics, including: right Corresponding category predicted score Perform Softmax normalization to obtain the probability distributions for each category, the th Probability distribution of each category Represented as: ; in, Indicates the total number of categories; exist Each candidate suggestion box and The average information entropy across each category is calculated to obtain the classification quality assessment index. ; ; in, For a preset small constant, The higher the value, the better. The more chaotic the category attribution judgment, the lower the semantic confidence and the worse the quality of the pseudobox.
[0014] Optionally, the positioning quality assessment index and the classification quality assessment index are substituted into the linear fusion formula to calculate... Corresponding sample-by-sample adaptive integration coefficients This can be achieved through the following formula: ; in, Incorporating positive contributions into the integration Incorporating negative contributions into the integration; use Replace the fixed-time integration coefficient set by experience The aspect ratio of the target points stored in the prototype library for each type of anchor point is updated using a differential weighted method, which is achieved by the following formula: ; in, The aspect ratio of the current pseudo-frame. For the first Historical prototype values for each category For the first The updated prototype values for each category, The larger, The higher the contribution weight to the prototype library update, the more the shape prior carried by the high-quality pseudo-box will dominate the anchor point adjustment. The smaller the size, the more historical prototype values are preserved, and the risk of low-quality pseudo-frames introducing incorrect aspect ratios is effectively suppressed.
[0015] Optionally, the updated prototype library is invoked immediately at the start of the next training iteration. Based on the corrected aspect ratio prior, the shape of the candidate anchor points is adaptively adjusted to provide higher quality candidate proposal boxes for the coarse pseudo-box generation stage. Candidate box extraction, localization consistency evaluation, classification reliability evaluation, sample-by-sample adaptive ensemble coefficient generation, and prototype library differential update are executed sequentially in each training iteration. This works in conjunction with the hard sample screening mechanism of the information negative sample collector and the confidence suppression strategy of the pseudo-box refinement module. As the training rounds progress, a positive feedback loop in which the quality of pseudo-boxes and the accuracy of the prototype library mutually promote each other is gradually formed, ultimately achieving a systematic improvement in the accuracy and robustness of point-supervised target detection.
[0016] Optionally, based on the PointSR-R target detection network model, point-supervised target detection is performed on the feature vectors of the candidate regions to obtain point-supervised target detection results from the UAV's perspective, including: The feature vector of the candidate region is input into the coarse rotation pseudobox generation module of the PointSR-R object detection network model to obtain the coarse rotation pseudobox; the coarse rotation pseudobox generation module integrates an angle acquisition module and a MIL prediction head; Based on the spatial optimization module of the PointSR-R object detection network model, target-level overlap constraints, category-level geometric consistency constraints, and packet-level autoregressive constraints are applied to the coarse rotated pseudoboxes to obtain target-level overlap constraint loss, category-level geometric consistency constraint loss, and packet-level autoregressive constraint loss. We obtain the spatial domain constraint loss by weighting the target-level overlap constraint loss and the category-level geometric consistency constraint loss, and then combine the spatial domain constraint loss and the bag-level autoregressive constraint loss to obtain the joint optimization loss. The coarse rotated pseudobox is optimized based on the joint optimization loss to obtain the candidate box after spatial optimization; The pseudo-box refinement module based on the PointSR-R object detection network model refines the candidate boxes after spatial optimization and outputs the final pseudo-labels.
[0017] Optionally, based on the spatial optimization module of the PointSR-R object detection network model, target-level overlap constraints, category-level geometric consistency constraints, and packet-level autoregressive constraints are applied to the coarsely rotated pseudo-boundaries to obtain target-level overlap constraint loss, category-level geometric consistency constraint loss, and packet-level autoregressive constraint loss, including: For all coarse rotated pseudoboxes in the same image Perform 2D Gaussian modeling to obtain the Gaussian distribution corresponding to each coarse rotated pseudo-box. The specific modeling formula is as follows: ; ; in, For points within the coarsely rotated pseudo-frame, , , These represent the height, width, and rotation angle of the coarse rotated pseudo-frame, respectively. The width rotation factor, It is the height rotation factor; The Batachalia coefficient is used to measure the degree of overlap between different Gaussian distributions. The specific formula is as follows: ; in, Indicates the first Gaussian distribution With the Gaussian distribution The degree of overlap between them, i.e. and The Batachalia coefficient between them It is a square matrix function; For nearest coarse-rotated pseudo-boundaries with a center distance less than a preset threshold, the Batachalia coefficient is calculated to obtain the BC matrix. The diagonal elements of the BC matrix are ignored. Finally, the effective values in the BC matrix are averaged to obtain the target-level efficient overlap loss. , , This represents the number of valid values in the BC matrix. For each coarse rotated pseudo-frame, the corresponding Gaussian distribution For targets belonging to the same category, the KL divergence is used to measure the shape differences between different Gaussian distributions, ignoring terms related to the mean, i.e., ignoring the position information of the coarse rotated pseudo-box, and only retaining the parts reflecting the differences in width, height, aspect ratio, and orientation distribution. The specific formula is as follows: ; in, express and KL divergence between them; The KL divergence matrix for each category is calculated, and the average value of all KL divergence matrices is taken to obtain the category-level geometric consistency loss. , , The total number of KL divergence values in the matrix; The package consisting of the coarse rotated pseudobox and the candidate boxes generated by its jitter. Features are extracted via RoI Align and two shared fully connected layers, and then the extracted features are input into a regression layer to obtain an adaptive package. The coarse rotated pseudo-box generated in the previous stage As a regression reference, a smoothed L1 loss is used to obtain the packet-level autoregressive constraint loss. ; ; in, Indicates L1 loss, Indicates the target quantity. This indicates the number of proposals in the proposal package corresponding to each target. Indicates the first The proposal package corresponding to the target is the first One proposal, Indicates the first The preliminary candidate boxes for each target were generated in the previous stage.
[0018] Optionally, the target-level overlap constraint loss and the category-level geometric consistency constraint loss are weighted and combined to obtain the spatial domain constraint loss. The spatial domain constraint loss and the bag-level autoregressive constraint loss are then combined to obtain the joint optimization loss, which is achieved through the following formula: ; ; in, For spatial domain constraint loss, These are the preset weighting coefficients. To jointly optimize losses.
[0019] This application proposes a target detection method and system based on the DQA mechanism and PointSR-R, which achieves the following improvements compared with the PointSR scheme and other existing technologies.
[0020] First, this application achieves adaptive optimization of the anchor prototype library update, improving the quality and robustness of pseudo-boundary generation. Using a DQA mechanism, it abandons the globally fixed ensemble coefficients set by manual experience in existing technologies, instead employing sample-by-sample adaptive dynamic ensemble coefficients. By performing a dual-dimensional quality assessment of each pseudo-boundary—consistent localization and reliable classification—it automatically identifies and assigns higher update weights to high-quality pseudo-boundaries while suppressing the contamination of the prototype library by low-quality pseudo-boundaries. This design not only eliminates the tedious manual hyperparameter search process and reduces model deployment costs, but also allows the prototype library to be precisely adjusted based on the actual sample quality during training, significantly improving the spatial accuracy of pseudo-boundary generation and enhancing the model's robustness to manual annotation perturbations.
[0021] Secondly, this application addresses the spatial overlap issue of rotated pseudo-boundaries in dense scenes, enhancing the independence of target localization. It introduces a spatial optimization module into the PointSR-R framework, particularly the target-level overlap constraint. By performing Gaussian modeling on the coarse rotated pseudo-boundaries and utilizing the Patacheria coefficient to measure the degree of overlap between distributions, the spatial distribution of adjacent target pseudo-boundaries is explicitly constrained. This design effectively reduces unreasonable overlap between different target rotated pseudo-boundaries in dense scenes (such as parking lots and congested roads), ensuring the spatial independence of each target and thus achieving more accurate target localization in complex contexts.
[0022] Third, it enhances the geometric consistency of similar targets and suppresses the drift of pseudo-boundary shapes. Considering that similar targets typically have similar geometric features from the UAV's perspective, this application proposes a category-level geometric consistency constraint. By calculating the KL divergence between the Gaussian distributions of targets of the same category, it ensures that targets of the same category maintain consistency in scale and shape during training. This design effectively utilizes prior geometric knowledge of the category, preventing unreasonable geometric deformation or drift of pseudo-boundaries during optimization, and improving the geometric stability of pseudo-labels.
[0023] Fourth, the structural correlation within the proposal package is optimized, improving the estimation accuracy of rotation angle and scale. This application establishes explicit regression relationships between candidate boxes within the proposal package through package-level autoregressive constraints. Unlike the existing method of generating candidate boxes by independent jittering, this application ensures that candidate boxes can converge to high-quality coarse rotated pseudo-boxes. This design effectively solves the problems of insufficient optimization space due to excessively small jitter range or noise introduced by excessively large jitter range, improving the overall structural quality of candidate boxes within the proposal package. This allows the model to more accurately estimate the rotation angle and scale of the target, ultimately generating high-quality rotated pseudo-labels.
[0024] Fifth, it achieves an effective extension from horizontal bounding box detection to rotated bounding box detection, adapting to complex aerial photography scenarios. While retaining the advantages of the original self-regularization framework, this application successfully extends it to the field of rotating target detection. By combining coarse rotating pseudo-bounding box generation, spatial structure regularization, and pseudo-bounding box refinement, it overcomes the shortcomings of existing point-supervised methods that are only applicable to horizontal bounding boxes or have inaccurate rotating bounding box estimations. It can better adapt to the complex characteristics of dense, slender, and arbitrarily oriented targets under UAV perspectives, significantly improving the performance of rotating target detection under weak supervision conditions. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies of this application will be briefly introduced below. The following drawings are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. The drawings described herein are only used to explain this application and are not intended to limit this application.
[0026] Figure 1 This is a flowchart of a target detection method based on DQA mechanism and PointSR-R provided in one embodiment of this application; Figure 2 This is a flowchart of a prototype library update based on the DQA mechanism provided in one embodiment of this application; Figure 3 This is a structural diagram of the PointSR-R target detection network model provided in one embodiment of this application; Figure 4 This is a flowchart of a PointSR-R-based target detection network model provided in one embodiment of this application, which performs point-supervised target detection on the feature vectors of candidate regions to obtain the point-supervised target detection results from the perspective of an unmanned aerial vehicle (UAV). Figure 5 This is a comparison of the detection results of different detection models provided in one embodiment of this application. Figure 1 ; Figure 6 This is a comparison of the detection results of different detection models provided in one embodiment of this application. Figure 2 ; Figure 7 This is a structural diagram of a target detection system based on DQA mechanism and PointSR-R provided in another embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. Those skilled in the art will understand that many technical details have been presented in the embodiments of this application to facilitate better understanding. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the following embodiments is for ease of description and should not constitute any limitation on the specific implementation of this application. The following embodiments can be combined with and referenced by each other without contradiction.
[0028] Based on the PointSR scheme, this application proposes two progressive levels of technical improvements.
[0029] The first level involves improving the temporal ensemble encoder based on the DQA mechanism. This application proposes and employs a dynamic quality alignment method, which, while retaining the basic framework of PointSR's temporal ensemble encoder—updating the anchor prototype library through historical pseudo-box aggregation—reduces the original globally fixed ensemble coefficients. Replace with sample-by-sample adaptive dynamic ensemble coefficients By performing a dual-dimensional quality assessment of each pseudo-box in terms of positioning consistency and classification reliability, personalized integrated weights are automatically generated, fundamentally solving the problems of quality contamination from fixed coefficients and manual parameter tuning.
[0030] The second level involves expanding the bounding box. This application proposes the PointSR-R framework and the SOM module. Based on the improvements of DQA, the PointSR framework is extended from horizontal bounding box detection to rotated bounding box detection. The SOM module is added, which fills the technical gap in the PointSR scheme in the field of rotated target detection through three-layer spatial structure regularization of target-level overlap constraints, category-level geometric consistency constraints and bag-level autoregressive constraints.
[0031] One embodiment of this application proposes a target detection method based on DQA mechanism and PointSR-R. The implementation details of the target detection method based on DQA mechanism and PointSR-R proposed in this embodiment are described in detail below. The following implementation details are provided for ease of understanding only and are not necessary for implementing this solution.
[0032] The specific process of the target detection method based on DQA mechanism and PointSR-R proposed in this embodiment can be described as follows: Figure 1 As shown, it includes: Step 11: Acquire the target image to be detected using a drone and generate feature vectors of candidate regions using a region proposal mechanism.
[0033] Specifically, the prerequisite for achieving point-supervised target detection from a UAV perspective is to acquire the target image to be detected through the UAV and generate feature vectors of candidate regions through a region proposal mechanism. Based on this, this embodiment will construct a PointSR-R target detection network model to progressively optimize the coarse rotation pseudo-box generation process, thereby obtaining more accurate rotation pseudo-labels and achieving high-precision and robust point-supervised target detection from a UAV perspective.
[0034] In one example, a convolutional neural network backbone is used to extract feature maps from the input image. Then, a region proposal mechanism is used to generate candidate regions centered on the point labels. The feature vector corresponding to each candidate region is obtained through RoIAlign, which provides input for the subsequent generation of coarse rotated pseudoboxes.
[0035] Step 12: Based on the weakly supervised learning object detection network, construct the PointSR-R object detection network model. The PointSR-R object detection network model is equipped with a time ensemble encoder with embedded DQA mechanism, a coarse rotating pseudo-box generation module, a spatial optimization module, and a pseudo-box refinement module. The spatial optimization module constrains the rotating pseudo-box generation process from three levels of relationships: target level, category level, and bag level, through spatial structure regularization.
[0036] Specifically, after obtaining the feature vectors of the candidate regions, a PointSR-R object detection network model can be constructed based on a weakly supervised learning object detection network. (The structure of the PointSR-R object detection network model is referenced here.) Figure 3 The system includes a time-integrated encoder with embedded DQA mechanism, a coarse-rotated pseudo-box generation module (CR-PBG), a spatial optimization module (SOM), and a pseudo-box refinement module (PBR). The CR-PBG module generates coarse-rotated pseudo-boxes based on point labels. The spatial optimization module constrains the pseudo-box generation process through spatial structure regularization at three levels: target level, category level, and bag level. The PBR further outputs the final pseudo-labels based on the spatial optimization results.
[0037] It is understandable that the basic framework of PointSR, such as backbone network feature extraction, region proposal mechanism, basic structure of MIL two-stream classifier, IS collector and basic pseudobox optimization module, is roughly the same as that of PointSR scheme, and will not be described in detail in this embodiment.
[0038] It is important to note that the DQA mechanism is embedded in the prototype library update stage of the temporal ensemble encoder. For each candidate bounding box corresponding to a point label, the quality of the pseudo-boxes is evaluated from two dimensions: localization consistency and classification reliability. Localization consistency measures the deviation of the center distribution of the candidate bounding boxes from the point label position using a Gaussian mixture model. Classification reliability measures the prediction uncertainty by performing Softmax normalization on the classification scores of the candidate bounding boxes and calculating information entropy. Based on the quality evaluation indicators of the two dimensions, a sample-by-sample adaptive ensemble coefficient is calculated to replace the empirically set fixed temporal ensemble coefficient, thereby updating the prototype library of anchor points of each category and reducing the impact of low-quality pseudo-boxes on the prototype library update.
[0039] This section first introduces the relevant content of adaptive updates of prototype libraries based on the DQA mechanism.
[0040] like Figure 2As shown, the DQA mechanism is embedded in the prototype library update stage of the time ensemble encoder, and its execution flow follows the training rounds. The basic iterative unit consists of four logically progressive sub-steps: candidate box extraction, two-dimensional quality assessment, adaptive coefficient generation, and prototype library differential update. The anchor point shape adjustment for the next iteration is driven by the closed-loop backflow of the prototype library, thereby achieving a synergistic improvement in pseudo-box quality and prototype library accuracy.
[0041] Candidate box extraction. In the first... In the coarse-grained pseudo-frame generation stage, for the first... Individual tags First, the depth features of each candidate box are extracted through the RoI pooling layer and two shared fully connected layers. Then, the depth features of each candidate box are input into the MIL dual-stream classifier, and the classification stream and detection stream are respectively subjected to softmax operations to obtain the class prediction score of each candidate box. and detection confidence score Select the top scorers in descending order of their overall scores. Each candidate suggestion box constitutes a... The proposal package This serves as the basic input unit for subsequent two-dimensional quality assessment. It should be noted that... The hyperparameters are preset and should be determined in the experiment based on the target density of the specific dataset and the constraints of computing resources, rather than being fixed to a specific value.
[0042] Positioning consistency assessment. Quantifying from a positioning perspective. and The degree of spatial agreement between them is measured by a Gaussian mixture model. The central distribution relative to The deviation in position is used to obtain the positioning quality assessment index.
[0043] Specifically, when conducting a positioning consistency assessment, with center coordinates As the mean parameter of the corresponding Gaussian component, the detection confidence score of the detection stream output is used. As a mixing ratio ( ),by Compared to center Constructing a shared covariance matrix from coordinate deviations .
[0044] .
[0045] based on and Establish a Gaussian mixture model, and Substituting the values into a Gaussian mixture model to calculate its probability density value, we obtain the classification quality assessment index. , ,in, This represents the established Gaussian mixture model. The higher the value, the more it indicates The more the space is close to The higher the accuracy of the indicated target location, the better the positioning quality. Classification reliability assessment. Localization bias alone is insufficient to fully characterize the semantic reliability of pseudo-boundaries; therefore, information entropy needs to be introduced at the classification level to supplement the assessment dimension. The classification scores are Softmax normalized and information entropy is calculated to measure prediction uncertainty, thus obtaining a classification quality evaluation index.
[0046] Specifically, when conducting a classification reliability assessment, for Corresponding category predicted score Perform Softmax normalization to obtain the probability distributions for each category, the th Probability distribution of each category Represented as: ; in, This indicates the total number of categories.
[0047] exist Each candidate suggestion box and The average information entropy across each category is calculated to obtain the classification quality assessment index. .
[0048] .
[0049] in, For a preset small constant, The higher the value, the better. The more chaotic the category attribution judgment, the lower the semantic confidence and the worse the quality of the pseudobox.
[0050] and They are naturally complementary in physical sense, together forming a two-dimensional three-dimensional characterization of the quality of the pseudoframe.
[0051] Sample-by-sample adaptive fusion coefficient generation. The localization quality assessment index and the classification quality assessment index are substituted into the linear fusion formula to calculate... Corresponding sample-by-sample adaptive integration coefficients , , Positive contributions are incorporated into the fusion; the better the positioning quality, the larger the coefficient. Negative contributions are included in the fusion; the higher the classification uncertainty, the smaller the coefficient. Because... For the probability density to take values, The values for information entropy are different in scale and range, and the formula itself does not automatically guarantee this. Strictly limited to Interval.
[0052] Prototype library differential updates, using Replace the fixed-time integration coefficient set by experience Perform differential weighted updates on the aspect ratios of the target objects stored in the prototype library for each type of anchor point.
[0053] use Replace the fixed-time integration coefficient set by experience The aspect ratio of the target points stored in the prototype library for each type of anchor point is updated using a differential weighted method, which is achieved by the following formula: ; in, The aspect ratio of the current pseudo-frame. For the first Historical prototype values for each category For the first The updated prototype values for each category, The larger, The higher the contribution weight to the prototype library update, the more the shape prior carried by the high-quality pseudo-box will dominate the anchor point adjustment. The smaller the value, the more historical prototype values are preserved, and the risk of low-quality pseudo-frames introducing incorrect aspect ratios is effectively suppressed, thereby statistically reducing the impact of low-quality pseudo-frames on prototype library updates.
[0054] It is important to note that the updated prototype library is invoked immediately at the start of the next training iteration. Based on the corrected aspect ratio prior, the shape of the candidate anchor points is adaptively adjusted, providing higher-quality candidate proposal boxes for the coarse pseudo-box generation stage. Candidate box extraction, localization consistency evaluation, classification reliability evaluation, sample-by-sample adaptive ensemble coefficient generation, and prototype library differential update are executed sequentially in each training iteration. This works in conjunction with the hard sample screening mechanism of the information negative sample collector and the confidence suppression strategy of the pseudo-box refinement module. As the training rounds progress, a positive feedback loop is gradually formed in which the quality of pseudo-boxes and the accuracy of the prototype library mutually promote each other, ultimately achieving a systematic improvement in the accuracy and robustness of point-supervised target detection.
[0055] To verify the effectiveness of the prototype library adaptive update method based on the DQA mechanism proposed in this embodiment, experiments were conducted on three UAV view target detection benchmark datasets: VisDrone, DroneVehicle, and UAVDT. The experiments employed a point-supervised setup, where only target point annotations were used to generate pseudo-boundaries during the training phase. These pseudo-boundaries were then used to train the subsequent detector, and the AP50 score of the detector was used as the benchmark score. To evaluate the final detection performance, the quality of the pseudo-boundary is evaluated by the mIoU between the generated pseudo-boundary and the real bounding box.
[0056] This embodiment is implemented based on the MMDetection toolbox, using a ResNet-50 pre-trained on the ImageNet dataset as the backbone network, and trained on an NVIDIA RTX 4090 GPU. The optimizer employs stochastic gradient descent, with a training strategy of 3×training schedule. The initial learning rate is set to 0.02, and the learning rate is decayed to 0.1 times its original value at the 24th and 33rd epochs, respectively. All comparison methods use the same backbone network and training settings to ensure the comparability of experimental results. Evaluation metrics include mean accuracy (AP50). ) and the average crossover ratio mIoU, where, This represents the average detection accuracy when the IoU threshold is 0.5. mIoU represents the average overlap between the generated pseudo-boundary and the real bounding box. The higher the mIoU, the better the quality of the pseudo-boundary.
[0057] Table 1: Comparative Experimental Results of DQA and Existing Methods
[0058] Table 1 shows that PointSR+DQA performs well on the VisDrone, DroneVehicle, and UAVDT drone-view target detection datasets. The accuracy rates reached 36.8%, 43.5%, and 40.2% respectively, representing improvements of 1.8, 1.0, and 1.5 percentage points compared to PointSR without DQA, and improvements of 5.1, 8.2, and 6.9 percentage points compared to the baseline method P2BNet. Meanwhile, PointSR+DQA achieved mIoU of 76.0%, 89.1%, and 96.9% on the three datasets, all higher than P2BNet and PointSR, indicating that DQA not only improved the final detection performance but also enhanced the spatial quality of the pseudo-boundaries generated from point annotations.
[0059] To further explain the reasons for the aforementioned performance improvement, this embodiment verifies it from two aspects. First, by using a fixed-time integration coefficient... Sensitivity experiments were conducted to verify that manually set fixed coefficients in PointSR have a significant impact on detection performance, thus demonstrating the necessity of introducing dynamic adaptive coefficients. Secondly, module ablation experiments were performed to verify whether introducing DQA under the same basic framework can further improve detection performance and pseudo-box quality.
[0060] Table 2: Fixed-Time Integration Coefficients Sensitivity test results
[0061] As shown in Table 2, the fixed-time integration coefficient The value of has a significant impact on detection performance. When When the AP50 values were 0.1, 0.3, 0.5, 0.7, and 0.9, respectively, they were 34.2%, 40.5%, 42.5%, 34.9%, and 33.9%, with the largest difference reaching 8.6 percentage points. This result indicates that the original PointSR fixed... Traditional methods rely on human experience for selection, and different values can lead to significant fluctuations in detection performance. In contrast, the DQA mechanism proposed in this embodiment dynamically generates per-sample ensemble coefficients based on the localization consistency and classification reliability of each pseudo-box, which can replace manually set fixed coefficients. This reduces the cost of hyperparameter search and improves the ease of model deployment.
[0062] Table 2 illustrates the necessity of DQA from the perspective of the sensitivity of fixed coefficient values, that is, the original fixed coefficients... It is difficult to stably adapt to changes in pseudoboxes of different training stages and quality. Furthermore, in order to verify whether the DQA mechanism itself can bring actual performance gains on the existing PointSR framework, this embodiment conducted module ablation experiments on the DroneVehicle dataset, comparing the detection performance of P2BNet, P2BNet with added temporal ensemble encoder, P2BNet with added information sample collector, PointSR, and PointSR+DQA. The results are shown in Table 3.
[0063] Table 3: Ablation Experiment Results of DQA Module
[0064] Table 3 shows that P2BNet's AP50 and mIoU are 35.26% and 86.0%, respectively. After introducing a time-integrated encoder on top of it... The accuracy rate increased to 36.92%, and the mIoU increased to 86.8%. After introducing the information sample collector, The accuracy was improved to 40.87%, and the mIoU was improved to 88.1%. Simultaneously, the introduction of a time-integrated encoder and an information sampler formed the PointSR. The mIoU and mIoU reached 42.53% and 88.9%, respectively. After further introducing the DQA mechanism of this embodiment based on PointSR, its... The mIoU was improved to 43.15% and 89.1% respectively, which shows that DQA can further optimize the prototype library update process on the basis of the existing self-regularized point supervision framework, improve the quality of pseudo-boxes and the final detection performance.
[0065] In summary, the fixed-coefficient sensitivity experiment shows that the manually set fixed-time integration coefficients in PointSR are effective. The impact on detection performance is significant, with issues including high costs for manual parameter tuning and large performance fluctuations with different values. Module ablation experiments further demonstrate that introducing DQA on top of PointSR improves both detection performance and pseudo-boundary quality. Therefore, the DQA mechanism proposed in this embodiment can effectively replace the original fixed-time integration coefficients through sample-by-sample pseudo-boundary quality evaluation and adaptive coefficient updates, making the update of the category anchor prototype library more stable and accurate, thereby effectively improving the performance of point-supervised target detection.
[0066] Step 13: Based on the PointSR-R target detection network model, perform point-supervised target detection on the feature vectors of the candidate regions to obtain the point-supervised target detection results from the perspective of the UAV.
[0067] Specifically, after constructing the PointSR-R target detection network model, point-supervised target detection can be performed on the feature vectors of candidate regions based on the PointSR-R target detection network model to obtain point-supervised target detection results from the perspective of UAVs.
[0068] In one example, based on the PointSR-R object detection network model, point-supervised object detection is performed on the feature vectors of candidate regions to obtain point-supervised object detection results from the UAV's perspective. This can be achieved through methods such as... Figure 4 The implementation of each sub-step shown includes the following sub-steps.
[0069] Sub-step 131: Input the feature vector of the candidate region into the coarse rotated pseudobox generation module of the PointSR-R object detection network model to obtain the coarse rotated pseudobox.
[0070] Specifically, after constructing the PointSR-R object detection network model, the feature vectors of candidate regions can be input into the coarse rotation pseudo-boundary generation module of the PointSR-R object detection network model to obtain coarse rotation pseudo-boundaries. The coarse rotation pseudo-boundary generation module integrates an angle acquisition module and a MIL prediction head.
[0071] The coarse rotation pseudo-boundary generation module first constructs the original view, rotated view, and flipped view, and inputs them into a shared backbone network to extract multi-view feature representations. Then, the angle acquisition module learns the target orientation information and outputs the angle prediction results corresponding to the candidate regions. Finally, the candidate proposal package with orientation information is input into the MIL prediction head to classify and score the candidate boxes, thereby selecting the candidate boxes with higher scores as coarse rotation pseudo-boundaries.
[0072] Each point label corresponds to a candidate proposal package, which consists of multiple candidate boxes with different scales, aspect ratios, and orientation information. The MIL prediction head scores the candidate boxes within the proposal package to obtain a coarsely rotated pseudo-box corresponding to the current target. This coarsely rotated pseudo-box already possesses information such as center position, width, height, and orientation angle; however, it may still suffer from issues such as target overlap, geometric drift of similar targets, and unstable quality of candidate boxes within the package. Therefore, further constraints are needed through a spatial optimization module.
[0073] Sub-step 132: Based on the spatial optimization module of the PointSR-R target detection network model, target-level overlap constraints, category-level geometric consistency constraints, and packet-level autoregressive constraints are applied to the coarsely rotated pseudo-boxes to obtain target-level overlap constraint loss, category-level geometric consistency constraint loss, and packet-level autoregressive constraint loss.
[0074] Specifically, the coarsely rotated pseudo-boundary is fed into the spatial optimization module, which performs target-level overlap constraints, category-level geometric consistency constraints, and packet-level autoregressive constraints on the coarsely rotated pseudo-boundary, respectively, to obtain target-level overlap constraint loss, category-level geometric consistency constraint loss, and packet-level autoregressive constraint loss.
[0075] The first is the target-level overlap constraint.
[0076] The spatial optimization module applies coarsely rotated pseudoboxes to all images in the same image. Perform 2D Gaussian modeling to obtain the Gaussian distribution corresponding to each coarse rotated pseudo-box. The specific modeling formula is as follows: ; ; in, For points within the coarsely rotated pseudo-frame, , , These represent the height, width, and rotation angle of the coarse rotated pseudo-frame, respectively. The width rotation factor, It is the height rotation factor.
[0077] The Batachalia coefficient is used to measure the degree of overlap between different Gaussian distributions. The specific formula is as follows: ; in, Indicates the first Gaussian distribution With the Gaussian distribution The degree of overlap between them, i.e. and The Batachalia coefficient between them It is a square matrix function.
[0078] Considering the large number of targets in UAV images, performing overlap calculations on all targets would significantly increase computational complexity. Therefore, this embodiment further sets a center distance threshold and calculates the Batachalia coefficient for nearest coarse-rotated pseudo-boxes with a center distance less than the preset threshold, thereby reducing complexity and obtaining the BC matrix. The diagonal elements of the BC matrix are ignored, i.e., the overlap of the distribution itself is ignored. Finally, the effective values in the BC matrix are averaged to obtain the target-level efficient overlap loss. , , This represents the number of valid values in the BC matrix.
[0079] This target-level overlap constraint can effectively reduce the overlap of rotational pseudo-boxes between different targets, and is especially suitable for dense target scenarios such as parking lots and road vehicles.
[0080] Secondly, there are category-level geometric consistency constraints.
[0081] The spatial optimization module applies coarsely rotated pseudoboxes to all images in the same image. Perform 2D Gaussian modeling to obtain the Gaussian distribution corresponding to each coarse rotated pseudo-box. For targets belonging to the same category, the KL divergence is used to measure the shape difference between different Gaussian distributions, and the specific formula is shown below: ; in, express and KL divergence between them This indicates taking the trace of the matrix.
[0082] Since category-level geometric consistency constraints focus on the stability of targets of the same category in terms of scale and shape, rather than their specific location, the KL divergence calculation ignores terms related to the Gaussian distribution mean, i.e., it ignores the positional information of the coarsely rotated pseudo-boxes, retaining only the parts reflecting the differences in width, height, aspect ratio, and orientation distribution. The KL divergence formula after ignoring the mean is as follows: .
[0083] Subsequently, the KL divergence matrix for each category is calculated, and the average value of all KL divergence matrices is taken to obtain the category-level geometric consistency loss. , , This represents the total number of KL divergence value matrices.
[0084] By using this category-level geometric consistency constraint, targets of the same category can maintain relative consistency in the scale and geometry of their rotated pseudo-boundaries, thereby suppressing significant and unreasonable geometric drift in pseudo-boundaries of the same type of targets.
[0085] Finally, there are package-level autoregressive constraints.
[0086] The package consisting of the coarse rotated pseudobox and the candidate boxes generated by its jitter. Features are extracted via RoI Align and two shared fully connected layers, and then the extracted features are input into a regression layer to obtain an adaptive package. The coarse rotated pseudo-box generated in the previous stage As a regression reference, a smoothed L1 loss is used to obtain the packet-level autoregressive constraint loss. , Represented as: ; in, Indicates L1 loss, Indicates the target quantity. This indicates the number of proposals in the proposal package corresponding to each target. Indicates the first The proposal package corresponding to the target is the first One proposal, Indicates the first The preliminary candidate boxes for each target were generated in the previous stage.
[0087] In this embodiment, the packet-level autoregressive constraint introduces an explicit regression relationship, causing the candidate boxes within the candidate proposal packet to gradually converge towards a better coarse-rotated pseudo-box, thereby solving the problem of independent candidate boxes within the packet and a lack of structural correlation in existing solutions. When the jitter range is too small, this constraint can prevent candidate boxes from being confined to a small area for a long time; when the jitter range is too large, it can suppress the introduction of background noise to a certain extent, thereby improving the overall quality of the candidate proposal packet.
[0088] Sub-step 133 involves weighting the target-level overlap constraint loss and the category-level geometric consistency constraint loss to obtain the spatial domain constraint loss, and then combining the spatial domain constraint loss and the package-level autoregressive constraint loss to obtain the joint optimization loss.
[0089] Sub-step 134: Optimize the coarse rotated pseudobox based on the joint optimization loss to obtain the candidate box after spatial optimization.
[0090] Specifically, after obtaining the target-level overlap constraint loss, the category-level geometric consistency constraint loss, and the packet-level autoregressive constraint loss, the target-level overlap constraint loss and the category-level geometric consistency constraint loss can be weighted and combined to obtain the spatial domain constraint loss. The spatial domain constraint loss and the packet-level autoregressive constraint loss can then be combined to obtain the joint optimization loss.
[0091] In one example, a weighted combination of the target-level overlap constraint loss and the category-level geometric consistency constraint loss is used to obtain the spatial domain constraint loss. Then, the spatial domain constraint loss and the bag-level autoregressive constraint loss are combined to obtain the joint optimization loss, which is achieved through the following formula: ; ; in, For spatial domain constraint loss, These are the preset weighting coefficients. To jointly optimize losses.
[0092] Joint optimization loss can simultaneously achieve non-overlapping targets, geometric consistency within categories, and adaptive convergence of candidate boxes within the bag, thereby improving the overall quality of coarsely rotated pseudoboxes.
[0093] Sub-step 135: Based on the PointSR-R object detection network model, the pseudo-box refinement module refines the spatially optimized candidate boxes and outputs the final pseudo-labels.
[0094] Specifically, after obtaining the spatially optimized candidate boxes, the pseudo-box refinement module of the PointSR-R object detection network model can be used to refine the spatially optimized candidate boxes and output the final pseudo-labels.
[0095] In one example, the candidate results constrained by the spatial optimization module are input into the pseudo-box refinement module. The MIL prediction head is then used to further filter and refine the candidate boxes, re-evaluating the category score and localization score of each candidate box, and thus outputting the final pseudo-label. The final pseudo-label includes the target center coordinates, width, height, and rotation angle information, which can be used as supervision information for the subsequent training of the rotated target detector.
[0096] In summary, this embodiment achieves high-quality rotated pseudo-labels through the collaborative work of a three-stage module: a coarse rotation pseudo-boundary generation module, a spatial optimization module, and a pseudo-boundary refinement module, even with only point supervision. Target-level overlap constraints reduce unreasonable overlap between different targets, category-level geometric consistency constraints improve the geometric stability of similar targets, and package-level autoregressive constraints improve the quality of candidate bounding boxes within candidate proposal packages, ultimately enhancing the reliability of rotated pseudo-labels and the performance of rotated target detection.
[0097] This embodiment proposes a target detection method and system based on DQA mechanism and PointSR-R, which achieves the following improvements compared with the PointSR scheme and other existing technologies.
[0098] First, adaptive optimization of the anchor prototype library update is achieved, improving the quality and robustness of pseudo-boundary generation. This embodiment uses a DQA mechanism, abandoning the globally fixed ensemble coefficients set by manual experience in existing technologies, and instead adopting a sample-by-sample adaptive dynamic ensemble coefficient. By performing a two-dimensional quality assessment of each pseudo-boundary based on localization consistency and classification reliability, it can automatically identify and assign higher update weights to high-quality pseudo-boundaries, while suppressing the contamination of the prototype library by low-quality pseudo-boundaries. This design not only eliminates the tedious manual hyperparameter search process and reduces model deployment costs, but also enables the prototype library to be accurately adjusted according to the actual sample quality during training, significantly improving the spatial accuracy of pseudo-boundary generation and enhancing the model's robustness to manual annotation perturbations.
[0099] Secondly, this embodiment addresses the spatial overlap issue of rotated pseudo-boundaries in dense scenes, enhancing the independence of target localization. This implementation introduces a spatial optimization module into the PointSR-R framework, particularly the target-level overlap constraint. By performing Gaussian modeling on the coarse rotated pseudo-boundaries and utilizing the Patacheria coefficient to measure the degree of overlap between distributions, it explicitly constrains the spatial distribution of adjacent target pseudo-boundaries. This design effectively reduces unreasonable overlap between different target rotated pseudo-boundaries in dense scenes (such as parking lots and congested roads), ensuring the spatial independence of each target and thus achieving more accurate target localization in complex contexts.
[0100] Third, it enhances the geometric consistency of similar targets and suppresses the drift of pseudo-boundary shapes. Considering that similar targets typically have similar geometric features from the UAV's perspective, this embodiment proposes a category-level geometric consistency constraint. By calculating the KL divergence between the Gaussian distributions of targets of the same category, it ensures that targets of the same category maintain consistency in scale and shape during training. This design effectively utilizes prior geometric knowledge of the category, preventing unreasonable geometric deformation or drift of pseudo-boundaries during optimization, and improving the geometric stability of pseudo-labels.
[0101] Fourth, the structural correlation within the proposal package is optimized, improving the estimation accuracy of rotation angle and scale. This embodiment establishes explicit regression relationships between candidate boxes within the proposal package through package-level autoregressive constraints. Unlike the existing method of generating candidate boxes by independent jittering, this embodiment ensures that candidate boxes can converge to high-quality coarse rotated pseudo-boxes. This design effectively solves the problems of insufficient optimization space due to excessively small jitter range or noise introduced by excessively large jitter range, improving the overall structural quality of candidate boxes within the proposal package. This allows the model to more accurately estimate the rotation angle and scale of the target, ultimately generating high-quality rotated pseudo-labels.
[0102] Fifth, it achieves an effective extension from horizontal bounding box detection to rotated bounding box detection, adapting to complex aerial photography scenarios. This embodiment, while retaining the advantages of the original self-regularization framework, successfully extends it to the field of rotating target detection. By combining coarse rotating pseudo-bounding box generation, spatial structure regularization, and pseudo-bounding box refinement, it overcomes the shortcomings of existing point-supervised methods that are only applicable to horizontal bounding boxes or have inaccurate rotating bounding box estimations. It can better adapt to the complex characteristics of dense, elongated, and arbitrarily oriented targets under UAV perspective, significantly improving the performance of rotating target detection under weak supervision conditions.
[0103] The steps described above are merely for clarity in describing the technical solution. In actual implementation, they can be combined into one step, or certain steps can be broken down into multiple steps, as long as they involve the same logical relationship, they are all within the scope of protection of this application. Any insignificant modifications or designs added to the algorithm or process, as long as they do not change the core of the algorithm or process, are also within the scope of protection of this application.
[0104] In one embodiment, to verify the effectiveness of the PointSR-R proposed in this application for point-supervised rotating target detection from a UAV perspective, we conducted experimental verification on two benchmark datasets, VisDrone and DroneVehicle. Next, this embodiment first introduces the experimental datasets and implementation details, and then presents detailed experimental results.
[0105] The VisDrone dataset is one of the most widely used drone vision datasets. It contains 10,209 images and 542,000 instances, capturing diverse and complex environments from urban to rural areas, day to night, and sunny to cloudy days, covering ten target categories: pedestrians, crowds, bicycles, cars, vans, trucks, tricycles, sunshade tricycles, buses, and motorcycles. The dataset is labeled using horizontal bounding boxes.
[0106] The DroneVehicle dataset is a dataset for multimodal vehicle recognition from the perspective of unmanned aerial vehicles (UAVs). It contains 56,878 images, capturing various scenes including urban roads, residential areas, parking lots, and highways under both daytime and nighttime lighting conditions, covering five vehicle categories: cars, trucks, buses, vans, and freight trucks. The dataset uses a rotated bounding box annotation format.
[0107] The datasets mentioned above all use box-level annotation. In order to obtain point annotations and simulate the inherent errors in human manual annotation, we first assume that the distribution of manually annotated points approximately follows a modified Gaussian distribution (RG) with a central ellipse constraint. Then, using the center of the box annotation as the mean, we sample according to the RG distribution to construct two point-supervised datasets, which are called "VisDrone (point supervision)" and "DroneVehicle (point supervision)" respectively.
[0108] Since this embodiment is geared towards the task of generating rotated pseudo-boundaries, further processing is required for datasets labeled with horizontal bounding boxes. This embodiment uses the visual basic model SAM to obtain the corresponding annotations. Specifically, the point label of each object is used as a point cue input to SAM to obtain the segmentation result of the corresponding object. Then, the minimum bounding rectangle of each segmentation result is constructed to obtain its rotated bounding box. Finally, these rough bounding boxes are manually adjusted to obtain the final accurate bounding box, which is then used as the ground truth annotation.
[0109] The experiments in this embodiment are based on the MMDetection toolkit and trained on an NVIDIA RTX 4090 GPU. The optimizer uses stochastic gradient descent and a 3× training schedule. The initial learning rate is set to 0.02, and the learning rate is decayed by a factor of 0.1 at the 24th and 33rd epochs. For all comparison models, we use a ResNet-50 pre-trained on ImageNet as the backbone network.
[0110] This embodiment uses the mean Intersection over Union (mIoU) to evaluate the quality of the predicted false boxes and the average precision (Average Precision) to evaluate the quality of the predicted false boxes. The accuracy of the model detection is evaluated using mIoU. mIoU refers to the degree of overlap between the predicted region and the ground truth labeled region; the closer the mIoU is to 1, the higher the accuracy. This refers to the average accuracy calculated when the Intersection over Union (IoU) threshold is 0.5. The closer the value is to 1, the more accurate the recognition and localization performance of the predicted bounding box.
[0111] The comparison results of this embodiment with five other state-of-the-art (SOTA) methods (including PLUG, Point2Rbox, PointOBB, PointOBBv2, and P2BNet-R) on two datasets are shown in Tables 4 and 5. Specifically, each method first generates pseudo-boxes on the VisDrone and DroneVehicle datasets, then trains the detector FOCS-R using these pseudo-boxes, and finally calculates the detection accuracy per class for different object categories. ), average detection accuracy ( (Mean) and mIoU. In addition, a Faster R-CNN-R and FOCS-R were trained using ground truth box-level annotations, and their detection performance was considered as the upper limit.
[0112] Table 4: Comparison of AP50 and mIoU of the detector trained with pseudo-boundaries on the VisDrone dataset for different object categories.
[0113] Table 5: Comparison of AP50 and mIoU of the detector trained with pseudo-boxes on the DroneVehicle dataset for different object categories.
[0114] Experimental results show that, under point-supervised conditions only, PointSR-R achieved 33.4% and 40.7% accuracy on VisDrone and DroneVehicle, respectively. It significantly outperforms existing state-of-the-art methods. Specifically, compared to P2BNet-R, PointSR-R... These improvements are 2.4% and 4.2% respectively. Furthermore, PointSR-R also outperforms existing methods in terms of mIoU on both datasets regarding the quality of rotated pseudoboxes. In addition, it is worth mentioning that... Figure 5 In some of the cases shown, the model trained based on point-level annotations in this embodiment has better detection results than the FOCS-R model trained based on box-level annotations.
[0115] Figure 6This demonstrates that the rotated pseudo-boundaries generated by PointSR-R proposed in this embodiment can accurately locate targets in UAV-view images and effectively distinguish densely packed objects. Although PointOBB can correctly predict the target angle, the generated pseudo-boundaries are relatively loose because it cannot accurately estimate the target scale. The pseudo-boundary generation mechanism of PointOBBv2 may not be suitable for UAV-view images, thus limiting CPM's ability to accurately estimate target orientation and size. In contrast, the model proposed in this embodiment can more accurately estimate the target's scale and orientation, thereby generating high-quality pseudo-boundaries based on point-level annotations.
[0116] Furthermore, ablation experiments were conducted on the DroneVehicle dataset in this embodiment. The ablation experiments verified the necessity of the three types of constraints in the Spatial Optimization Module (SOM): removing the target-level overlap constraint, the category-level geometric consistency constraint, or the bag-level autoregressive constraint all lead to a significant decrease in detection performance, reducing it by 7.9%, 9.1%, and 9.3%, respectively. The model achieves the best detection performance when all three constraints are retained simultaneously. Specific results are shown in the table below: Table 6 Ablation Experiment Data
[0117] Furthermore, Table 7 shows the different The performance of the FOCSR model trained with values, where express Medium-class geometric consistency loss The weight, The setting of the value has a significant impact on model performance. When At that time, the PointSR-R method proposed in this embodiment achieved the best detection performance. Experimental results verified the two losses. and The trade-off between these factors has a significant impact on the quality of counterfeit labels.
[0118] Table 7: Differences value pairs Impact
[0119] Furthermore, to verify the effectiveness of the Bhattacharyya coefficient proposed in this embodiment, Table 8 compares three metrics used to measure the overlap between two generated pseudo-boxes. Here, IoU represents the overlap quantified by the intersection-union ratio (IoU). Wasserstein and BC represent the overlap quantified by the Wasserstein distance between Gaussian distributions and the Bhattacharyya coefficient, respectively. Table 8 shows that the PointSR-R model using the proposed Bhattacharyya coefficient for overlap measurement in this embodiment achieves the best performance, with an accuracy of 40.7%.
[0120] Table 8: Effects of different overlap indices on Impact
[0121] Furthermore, to evaluate the efficiency of the proposed overlap quantization method, this embodiment sets different neighborhood thresholds. Experiments were conducted. Table 9 shows the results of training FOCS-R using pseudoboxes generated by PointSR-R under different conditions. Detection performance under various values. Results show that when... PointSR-R achieves optimal detection performance at that time. Furthermore, this embodiment uses frames per second (FPS) for quantitative evaluation. Impact on computational efficiency. As shown in Table 9, computational efficiency gradually decreases with increasing neighborhood radius, especially when the radius exceeds 120 pixels, where the decline becomes more pronounced. Therefore, a neighborhood radius of 120 pixels achieves the best balance between detection performance and computational efficiency and was adopted as the final setting.
[0122] Table 9: Different overlap thresholds And the impact of FPS
[0123] The experimental results above demonstrate that the spatial structure regularization mechanism proposed in this application can effectively improve the quality of rotating pseudo-tags and the final detection performance.
[0124] In summary, this application can generate high-quality rotated pseudo-boundaries under the condition of only point annotation, significantly improving the performance of rotating target detection, and can be widely applied in the field of UAV-view target detection technology.
[0125] Another embodiment of this application proposes a target detection system based on DQA mechanism and PointSR-R. The details of the target detection system based on DQA mechanism and PointSR-R proposed in this embodiment are described below. The following implementation details are provided for ease of understanding only and are not necessary for implementing this solution.
[0126] Figure 6 This is a schematic diagram of the structure of a target detection system based on DQA mechanism and PointSR-R proposed in this embodiment, including: target image acquisition and feature map extraction module 21, PointSR-R target detection network model construction module 22, and target detection execution module 23.
[0127] The target image acquisition and feature map extraction module 21 is used to acquire the target image to be detected by the UAV and generate the feature vector of the candidate region through the region proposal mechanism.
[0128] PointSR-R object detection network model construction module 22 is used to construct a PointSR-R object detection network model based on weakly supervised learning. The PointSR-R object detection network model is equipped with a time ensemble encoder with embedded DQA mechanism, a coarse rotation pseudo-box generation module, a spatial optimization module, and a pseudo-box refinement module. The spatial optimization module constrains the rotation pseudo-box generation process from three levels of relationship: target level, category level, and bag level through spatial structure regularization.
[0129] The target detection execution module 23 is used to perform point-supervised target detection on the feature vectors of candidate regions based on the PointSR-R target detection network model, and obtain the point-supervised target detection results from the perspective of the UAV.
[0130] The DQA mechanism is embedded in the prototype library update stage of the temporal ensemble encoder. For each candidate bounding box corresponding to a point label, the quality of the pseudo-boxes is evaluated from two dimensions: localization consistency and classification reliability. Localization consistency is measured by Gaussian mixture model to measure the deviation of the center distribution of the candidate bounding boxes from the point label position. Classification reliability is measured by Softmax normalization of the classification scores of the candidate bounding boxes and calculation of information entropy to measure prediction uncertainty. Based on the quality evaluation indicators of the two dimensions, a per-sample adaptive ensemble coefficient is calculated to replace the empirically set fixed temporal ensemble coefficient, thereby updating the prototype library of category anchors and reducing the impact of low-quality pseudo-boxes on the prototype library update.
[0131] It is worth noting that all modules involved in this embodiment are logical modules. In practical applications, a logical module can be a physical module, a part of a physical module, or an organic combination of multiple physical modules. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce modules that are not closely related to solving the technical problems proposed in this application. However, this does not mean that other modules are absent from this embodiment.
[0132] It is not difficult to see that this embodiment is a system embodiment corresponding to the above method embodiments, and this embodiment can be implemented in conjunction with the above method embodiments. The relevant technical details and technical effects mentioned in the above method embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiments.
[0133] Another embodiment of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, can implement a target detection method based on DQA mechanism and PointSR-R as described in the above method embodiments.
[0134] That is, those skilled in the art will understand that all or part of the steps in the above method embodiments can be implemented by a program instructing related hardware. The program is stored in a storage medium and includes several instructions to cause a device (such as a microcontroller, chip, etc.) or processor to execute all or part of the steps of the method described in the method embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0135] It will be understood by those skilled in the art that the above embodiments are specific implementations of this application, and in practical applications, various changes can be made in form, detail, and description without departing from the spirit and scope of this application. For those skilled in the art, several improvements and modifications can be made without departing from the principles of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.
Claims
1. A target detection method based on DQA mechanism and PointSR-R, characterized in that, include: The target image to be detected is acquired by a drone, and feature vectors of candidate regions are generated through a region proposal mechanism. Based on weakly supervised learning, a PointSR-R target detection network model is constructed. The PointSR-R target detection network model includes a temporal ensemble encoder with embedded DQA mechanism, a coarse rotating pseudo-box generation module, a spatial optimization module, and a pseudo-box refinement module. The spatial optimization module constrains the rotating pseudo-box generation process from three levels of relationships: target level, category level, and bag level, through spatial structure regularization. Based on the PointSR-R target detection network model, point-supervised target detection is performed on the feature vectors of candidate regions to obtain point-supervised target detection results from the perspective of UAVs. The DQA mechanism is embedded in the prototype library update stage of the temporal ensemble encoder. For each candidate bounding box corresponding to a point label, the quality of the pseudo-boxes is evaluated from two dimensions: localization consistency and classification reliability. Localization consistency is measured by Gaussian mixture model to measure the deviation of the center distribution of the candidate bounding boxes from the point label position. Classification reliability is measured by Softmax normalization of the classification scores of the candidate bounding boxes and calculation of information entropy to measure prediction uncertainty. Based on the quality evaluation indicators of the two dimensions, a per-sample adaptive ensemble coefficient is calculated to replace the empirically set fixed temporal ensemble coefficient, thereby updating the prototype library of anchor points of each category and reducing the impact of low-quality pseudo-boxes on the prototype library update.
2. The target detection method based on DQA mechanism and PointSR-R according to claim 1, characterized in that, The DQA mechanism is embedded in the prototype library update phase of the time-integrated encoder, and its execution flow includes: Candidate box extraction, in the first In the coarse-grained pseudo-frame generation stage, for the first... Individual tags First, the depth features of each candidate box are extracted through the RoI pooling layer and two shared fully connected layers. Then, the depth features of each candidate box are input into the MIL dual-stream classifier, and the softmax operations of the classification stream and detection stream are performed respectively to obtain the class prediction score and detection confidence score of each candidate box. The top candidates are selected in descending order of comprehensive score. Each candidate suggestion box constitutes a... The proposal package , serving as the basic input unit for subsequent two-dimensional quality assessment; Positioning consistency assessment, quantifying from the positioning level. and The degree of spatial agreement between them is measured by a Gaussian mixture model. The central distribution relative to The deviation in position is used to obtain positioning quality assessment indicators; Classification reliability assessment introduces information entropy at the classification level to supplement the assessment dimensions, through... The classification scores are normalized using Softmax and the information entropy is calculated to measure the prediction uncertainty, thus obtaining the classification quality evaluation index. The adaptive fusion coefficients are generated sample-by-sample by substituting the localization quality assessment index and the classification quality assessment index into the linear fusion formula for calculation. Corresponding sample-by-sample adaptive integration coefficients ; Prototype library differential updates, using Replace the fixed-time integration coefficient set by experience Perform differential weighted updates on the aspect ratios of the target objects stored in the prototype library for each type of anchor point.
3. The target detection method based on DQA mechanism and PointSR-R according to claim 2, characterized in that, Quantification from the perspective of positioning and The degree of spatial fit between them, through the The classification scores are Softmax normalized and information entropy is calculated to measure prediction uncertainty, resulting in classification quality evaluation metrics, including: by center coordinates As the mean parameter of the corresponding Gaussian component, the detection confidence score of the detection stream output is used. As a mixing ratio ,by Compared to center Constructing a shared covariance matrix from coordinate deviations ; ; based on and Establish a Gaussian mixture model, and Substituting the values into a Gaussian mixture model to calculate its probability density value, we obtain the classification quality assessment index. ; ; in, This indicates the established Gaussian mixture model. The higher the value, the more likely it is to indicate that... The more the overall space approaches The higher the accuracy of the indicated target location, the better the positioning quality. Introducing information entropy at the classification level to supplement the evaluation dimensions, through... The classification scores are Softmax normalized and information entropy is calculated to measure prediction uncertainty, resulting in classification quality evaluation metrics, including: right Corresponding category predicted score Perform Softmax normalization to obtain the probability distributions for each category, the th Probability distribution of each category Represented as: ; in, Indicates the total number of categories; exist Each candidate suggestion box and The average information entropy across each category is calculated to obtain the classification quality assessment index. ; ; in, For a preset small constant, The higher the value, the better. The more chaotic the category attribution judgment, the lower the semantic confidence and the worse the quality of the pseudobox.
4. The target detection method based on DQA mechanism and PointSR-R according to claim 3, characterized in that, Substituting the positioning quality assessment indicators and the classification quality assessment indicators into the linear fusion formula, calculate... Corresponding sample-by-sample adaptive integration coefficients This can be achieved through the following formula: ; in, Incorporating positive contributions into the integration Incorporating negative contributions into the integration; use Replace the fixed-time integration coefficient set by experience The aspect ratio of the target points stored in the prototype library for each type of anchor point is updated using a differential weighted method, which is achieved by the following formula: ; in, The aspect ratio of the current pseudo-frame. For the first Historical prototype values for each category For the first The updated prototype values for each category, The larger, The higher the contribution weight to the prototype library update, the more the shape prior carried by the high-quality pseudo-box will dominate the anchor point adjustment. The smaller the size, the more historical prototype values are preserved, and the risk of low-quality pseudo-frames introducing incorrect aspect ratios is effectively suppressed.
5. A target detection method based on DQA mechanism and PointSR-R according to any one of claims 1 to 4, characterized in that, The updated prototype library is invoked immediately at the start of the next training iteration. Based on the corrected aspect ratio prior, the shape of the candidate anchor points is adaptively adjusted to provide higher quality candidate proposal boxes for the coarse pseudo-box generation stage. Candidate box extraction, localization consistency evaluation, classification reliability evaluation, sample-by-sample adaptive ensemble coefficient generation, and prototype library differential update are executed sequentially in each training iteration. This works in conjunction with the hard sample screening mechanism of the information negative sample collector and the confidence suppression strategy of the pseudo-box refinement module. As the training rounds progress, a positive feedback loop in which the quality of pseudo-boxes and the accuracy of the prototype library mutually promote each other is gradually formed, ultimately achieving a systematic improvement in the accuracy and robustness of point-supervised target detection.
6. A target detection method based on DQA mechanism and PointSR-R according to any one of claims 1 to 4, characterized in that, Based on the PointSR-R object detection network model, point-supervised object detection is performed on the feature vectors of candidate regions to obtain point-supervised object detection results from the UAV perspective, including: The feature vector of the candidate region is input into the coarse rotation pseudobox generation module of the PointSR-R object detection network model to obtain the coarse rotation pseudobox; the coarse rotation pseudobox generation module integrates an angle acquisition module and a MIL prediction head; Based on the spatial optimization module of the PointSR-R object detection network model, target-level overlap constraints, category-level geometric consistency constraints, and packet-level autoregressive constraints are applied to the coarse rotated pseudoboxes to obtain target-level overlap constraint loss, category-level geometric consistency constraint loss, and packet-level autoregressive constraint loss. We obtain the spatial domain constraint loss by weighting the target-level overlap constraint loss and the category-level geometric consistency constraint loss, and then combine the spatial domain constraint loss and the bag-level autoregressive constraint loss to obtain the joint optimization loss. The coarse rotated pseudobox is optimized based on the joint optimization loss to obtain the candidate box after spatial optimization; The pseudo-box refinement module based on the PointSR-R object detection network model refines the candidate boxes after spatial optimization and outputs the final pseudo-labels.
7. The target detection method based on DQA mechanism and PointSR-R according to claim 6, characterized in that, Based on the spatial optimization module of the PointSR-R object detection network model, target-level overlap constraints, category-level geometric consistency constraints, and packet-level autoregressive constraints are applied to the coarsely rotated pseudo-boundaries, respectively, to obtain the target-level overlap constraint loss, category-level geometric consistency constraint loss, and packet-level autoregressive constraint loss, including: For all coarse rotated pseudoboxes in the same image Perform 2D Gaussian modeling to obtain the Gaussian distribution corresponding to each coarse rotated pseudo-box. The specific modeling formula is as follows: ; ; in, For points within the coarsely rotated pseudo-frame, , , These represent the height, width, and rotation angle of the coarse rotated pseudo-frame, respectively. The width rotation factor, It is the height rotation factor; The Batachalia coefficient is used to measure the degree of overlap between different Gaussian distributions. The specific formula is as follows: ; in, Indicates the first Gaussian distribution With the Gaussian distribution The degree of overlap between them, i.e. and The Batachalia coefficient between them It is a square matrix function; For nearest coarse-rotated pseudo-boundaries with a center distance less than a preset threshold, the Batachalia coefficient is calculated to obtain the BC matrix. The diagonal elements of the BC matrix are ignored. Finally, the effective values in the BC matrix are averaged to obtain the target-level efficient overlap loss. , , This represents the number of valid values in the BC matrix. For each coarse rotated pseudo-frame, the corresponding Gaussian distribution For targets belonging to the same category, the KL divergence is used to measure the shape differences between different Gaussian distributions, ignoring terms related to the mean, i.e., ignoring the position information of the coarse rotated pseudo-box, and only retaining the parts reflecting the differences in width, height, aspect ratio, and orientation distribution. The specific formula is as follows: ; in, express and KL divergence between them; The KL divergence matrix for each category is calculated, and the average value of all KL divergence matrices is taken to obtain the category-level geometric consistency loss. , , The total number of KL divergence values in the matrix; The package consisting of the coarse rotated pseudobox and the candidate boxes generated by its jitter. Features are extracted via RoI Align and two shared fully connected layers, and then the extracted features are input into a regression layer to obtain an adaptive package. The coarse rotated pseudo-box generated in the previous stage As a regression reference, a smoothed L1 loss is used to obtain the packet-level autoregressive constraint loss. ; ; in, Indicates L1 loss, Indicates the target quantity. This indicates the number of proposals in the proposal package corresponding to each target. Indicates the first The proposal package corresponding to the target is the first One proposal, Indicates the first The preliminary candidate boxes for each target were generated in the previous stage.
8. The target detection method based on DQA mechanism and PointSR-R according to claim 7, characterized in that, The spatial domain constraint loss is obtained by weighting the target-level overlap constraint loss and the category-level geometric consistency constraint loss. The spatial domain constraint loss is then combined with the bag-level autoregressive constraint loss to obtain the joint optimization loss, which is achieved through the following formula: ; ; in, For spatial domain constraint loss, These are the preset weighting coefficients. To jointly optimize losses.
9. A target detection system based on DQA mechanism and PointSR-R, characterized in that, include: The target image acquisition and feature map extraction module is used to acquire the target image to be detected through the UAV and generate feature vectors of candidate regions through a region proposal mechanism. The PointSR-R object detection network model building module is used to construct a PointSR-R object detection network model based on weakly supervised learning. The PointSR-R object detection network model includes a temporal ensemble encoder with embedded DQA mechanism, a coarse rotating pseudo-box generation module, a spatial optimization module, and a pseudo-box refinement module. The spatial optimization module constrains the rotating pseudo-box generation process from three levels of relationships: target level, category level, and bag level, through spatial structure regularization. The target detection execution module is used to perform point-supervised target detection on the feature vectors of candidate regions based on the PointSR-R target detection network model, and obtain the point-supervised target detection results from the perspective of the UAV. The DQA mechanism is embedded in the prototype library update stage of the temporal ensemble encoder. For each candidate bounding box corresponding to a point label, the quality of the pseudo-boxes is evaluated from two dimensions: localization consistency and classification reliability. Localization consistency is measured by Gaussian mixture model to measure the deviation of the center distribution of the candidate bounding boxes from the point label position. Classification reliability is measured by Softmax normalization of the classification scores of the candidate bounding boxes and calculation of information entropy to measure prediction uncertainty. Based on the quality evaluation indicators of the two dimensions, a per-sample adaptive ensemble coefficient is calculated to replace the empirically set fixed temporal ensemble coefficient, thereby updating the prototype library of category anchors and reducing the impact of low-quality pseudo-boxes on the prototype library update.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement a target detection method based on DQA mechanism and PointSR-R as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Unmanned aerial vehicle visual angle target detection method and system based on self-regularization point supervision
CN120279444A