An interactive target detection box labeling method based on a SAM large model

CN121305509BActive Publication Date: 2026-09-25福建汉特云智能科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511347401.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-09-25
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

然现有的分割模型在 box 检测场景中泛用性有限,对于自动检测的标注需求支持不够,无法有效输出高质量的目标框

Benefits of technology

一是通过静/动态属性、目标类别等丰富标签维度,增强模型对关键交通目标、异常动态和高风险区域的感知能力,降低模型漏检与误检风险。采用差异化随机点采样与点集标签生成策略,显著提升训练样本的空间分布合理性,使样本分布更加贴合自动驾驶感知需求,点采样与真值标注的结合,实现对点-框-目标之间的强几何监督,为后续点提示交互和自动框回归提供支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121305509B_ABST
    Figure CN121305509B_ABST
Patent Text Reader

Abstract

The application discloses an interactive target detection frame labeling method based on a SAM large model, and comprises the following steps: acquiring automatic driving scene image data, manually labeling a moving target, and constructing a frame-level detection training data set through random point sampling and label generation; inputting the frame-level detection training data set into a modified SAM model trunk to obtain intermediate image features and point features; fusing the intermediate image features and the point features according to the point features to obtain enhanced fusion features; inputting the enhanced fusion features into a regression branch of a detection head, performing regression classification through a multilayer perceptron of the detection head, and obtaining a prediction frame corresponding to the moving target; calculating an IOU loss between the prediction frame and a manual frame, constantly updating model parameters through a back propagation algorithm, and stopping until the model converges on the training data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of object detection bounding box annotation technology, and in particular to an interactive object detection bounding box annotation method based on the SAM large model. Background Technology

[0002] Object detection and automatic annotation, especially in the field of autonomous driving, has always been a labor-intensive task involving data labeling. Traditional annotation processes rely heavily on manual annotation of object bounding boxes or segmentation masks for each image, which is inefficient, costly, and difficult to guarantee inconsistency.

[0003] In recent years, with the rapid development of deep learning models, more and more semi-automated annotation tools have emerged. One typical example is the Segment Anything Model (SAM). The SAM model is based on large-scale visual Transformers, such as DINO, and integrates interactive annotations (such as points, boxes, and line segments) with image features to achieve semantic segmentation of images. This method represents a significant improvement over traditional manual annotation, greatly enhancing the data preparation efficiency for segmentation tasks. Taking the SAM model as an example, its main process is as follows: the user adds a small number of points or draws boxes in the image, and the model automatically outputs the segmentation results, making the annotation process more efficient. SAM uses point and image feature fusion, and through a transformer architecture, generates masked regions with prominent markings.

[0004] However, current mainstream semi-automatic annotation technologies such as SAM, while improving annotation efficiency through user interaction using points or boxes, still require frequent manual addition of points, boxes, or line segments to each image in practice. Even with improved efficiency compared to traditional methods, human involvement and intervention remain indispensable, making it difficult to achieve truly large-scale automatic annotation and limiting its application value in large-scale data production scenarios.

[0005] The core algorithm architecture and output format of SAM and similar models are essentially designed for "segmentation tasks." That is, the model's output is a pixel-level mask or region-based segmentation, rather than directly outputting rectangular boxes. For object detection common in autonomous driving—such as vehicles and people—the focus is usually more on the position and size of the box-level. However, existing segmentation models have limited generality in box detection scenarios, insufficient support for the annotation requirements of automatic detection, and cannot effectively output high-quality object boxes.

[0006] Currently used semi-automated segmentation tools typically require data to be in the form of a segmentable mask, and they have poor adaptability to special annotation methods such as background points and padding values. This makes them prone to annotation errors and even a decline in feature learning ability in actual annotation operations, especially in autonomous driving scenarios where the background and foreground are difficult to define precisely.

[0007] The purpose of this invention is to design an interactive object detection bounding box annotation method based on the SAM large model to address the problems existing in the prior art. Summary of the Invention

[0008] In view of this, the purpose of this invention is to propose an interactive object detection bounding box annotation method based on the SAM large model, which can solve the above-mentioned problems.

[0009] This invention provides an interactive object detection bounding box annotation method based on a large SAM model, comprising: Acquire image data of autonomous driving scenarios, manually annotate moving targets with bounding boxes, and construct a box-level detection training dataset through random point sampling and label generation; Input the box-level detection training dataset into the modified SAM model backbone to obtain intermediate image features and point features; The intermediate image features and point features are fused based on the point features to obtain enhanced fused features; The enhanced fusion features are input into the regression branch of the detector head, and the multilayer perceptron of the detector head is used for regression classification to obtain the predicted box corresponding to the moving target. Calculate the IOU loss between the predicted bounding box and the manual bounding box, and continuously update the model parameters through the backpropagation algorithm until the model converges on the training dataset.

[0010] The beneficial effects of this invention are: First, by enriching label dimensions such as static / dynamic attributes and target categories, the model's ability to perceive key traffic targets, abnormal dynamics, and high-risk areas is enhanced, reducing the risk of missed and false detections. A differentiated random point sampling and point set label generation strategy significantly improves the spatial distribution rationality of training samples, making the sample distribution more aligned with the perception needs of autonomous driving. The combination of point sampling and ground truth annotation achieves strong geometric supervision between points, boxes, and targets, providing support for subsequent point-based interactive prompts and automatic box regression.

[0011] Secondly, leveraging the advantages of the SAM backbone structure, the mask decoder is removed and the point type semantics are modified for autonomous driving scenarios, effectively improving the model's adaptability to point-level scene prompts. Through spatial location encoding and region masking mechanisms, invalid pixel / region interference is effectively isolated, allowing the model to focus more on key scene information during training and improving target segmentation and detection accuracy in complex backgrounds under autonomous driving. The point feature design considers the static / dynamic nature, category, and spatial attributes of the target, effectively enriching the point prompt information.

[0012] Third, bidirectional cross-attention enables dual-path cross-conditioning of global image context and point-level local information, significantly improving the model's response speed and accuracy to interactive point cues. Multi-branch fusion feature parallel regression significantly improves the detection accuracy and speed of multi-category targets (vehicles, pedestrians, obstacles, etc.) in autonomous driving scenarios, meeting real-time high-safety requirements. Joint screening using category confidence and bounding box confidence robustly and efficiently eliminates false targets, empty targets, and difficult-to-detect targets. Hungarian algorithm matching and redundancy removal ensure unique target allocation, achieving end-to-end global optimal matching without multi-stage processing, reducing duplicate bounding boxes, overlaps, and empty areas, resulting in high-precision and high-efficiency autonomous driving target detection outputs. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings required in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart of the method in this embodiment. Detailed Implementation

[0015] To facilitate understanding by those skilled in the art, the structure of the present invention will now be described in further detail with reference to the accompanying drawings. It should be understood that, unless otherwise specified, the order of the steps mentioned in this embodiment can be adjusted according to actual needs, and they can even be executed simultaneously or partially simultaneously.

[0016] like Figure 1 As shown, this embodiment of the invention provides an interactive object detection bounding box annotation method based on a large SAM model, including: S1 acquires image data of autonomous driving scenarios, manually annotates moving targets with bounding boxes, and constructs a box-level detection training dataset through random point sampling and label generation. S101 manually marks the moving targets in the autonomous driving scene image data to obtain manually marked training data; In this step, autonomous driving scenario data can be collected from multiple cameras, at multiple time periods, in various weather conditions, and on complex roads (such as intersections, tunnels, and at night). Manual bounding boxes are then used to label key categories such as moving vehicles, pedestrians, non-motorized vehicles, traffic signs, and obstacles. Each labeled target includes attributes such as category and static / dynamic state, which improves the sensitivity to key target detection and error scenarios.

[0017] S102 is based on manually labeled training data. Differentiated random point sampling and point supervision label generation are performed for different target regions to obtain box-level detection training data and then normalize it. In this step, to address the issue of poor detection performance for distant small targets and densely packed targets, a differentiated sampling strategy was adopted. The sampling density was appropriately increased for "high-risk areas" such as main lanes and intersections. For both static and dynamic targets, random points were sampled and labeled with type attributes, covering various typical traffic flow scenarios.

[0018] S1021 inputs image data into the UPN model to generate the corresponding uncertain heatmap; S1022 overlays the uncertainty heatmap with the prior risk area to generate an adaptive sampling probability map; S1023 randomly selects n pixels on each adaptive sampling probability map as the interaction point set. If a point falls into any manually labeled box, it is marked as a labeled point; otherwise, it is marked as an unlabeled point. For each marked point, S1024 generates a corresponding image point set supervision label based on the coordinate difference between the center point of its bounding box and the point, the bounding box scale parameter, the target category, and the static / dynamic state.

[0019] In this step, the UPN (Uncertainty Prediction Network) is a neural network model that analyzes the input image to predict which regions in the image are most difficult for the main model to judge and most prone to error (i.e., "high uncertainty"). For example, a region with high certainty: a complete car centered in the image, easily visible and reliable (low uncertainty). Regions with high uncertainty: a pedestrian half-obscured by a bus, a motorcycle only outlined in shadow, a distant traffic sign (high uncertainty). Manually labeled bounding boxes provide ground truth references (supervisory signals) for subsequent point-based supervised target generation. n points are randomly selected on each image, and labeled / unlabeled points are used to simulate real interaction / random prompts, establishing a geometric relationship between points and target existence. Unlabeled points replace background points to avoid ambiguity in background definition. This allows for learning robustness to point position perturbations. During inference, random points can be used to automatically generate bounding boxes. The controllable ratio of positive / unlabeled points improves sample distribution and attention alignment, avoiding training bias caused by mislabeled background points. Supervisory labels (center offset and box scale) are generated for each labeled point, formalizing the geometric relationship into a learnable regression target. The corresponding bounding box is predicted using the labeled point as a reference. Multiple points corresponding to the same bounding box achieve dense supervision, enhancing robustness to occluded / dense scenes.

[0020] S2 inputs the box-level detection training dataset into the modified SAM model backbone to obtain intermediate image features and point features; S201 removes the mask decoder of the SAM model and remaps the point type semantics to labeled / unlabeled, resulting in the modified SAM model backbone; In this step, the original SAM model includes an image encoder, a cue encoder, and a mask decoder. Its workflow is image embedding + cue embedding → mask decoder → mask, and its output is a mask, not a detection box. Since we need to directly regress the detection box from the point cue, instead of outputting a mask, we retain the part that encodes the image and points into fusionable features, remove the mask decoder, and remap the point type semantics to labeled / unlabeled. This effectively solves the problem of noise introduced by background points.

[0021] S202 inputs the image data from the box-level detection training dataset into the image encoder of the SAM model to obtain intermediate image features; S2021 normalizes and aligns the image data in the box-level detection training dataset, and records the effective region and the original size. S2022 inputs normalized and size-aligned image data into the image encoder to obtain downsampled grid features, and at the same time generates matching two-dimensional position codes and key target region image masks; S2023 flattens the downsampled grid features, two-dimensional position codes, and key target region image masks into a one-dimensional sequence, and aligns them according to the channel dimension to obtain intermediate image features, flattened position codes, and masks.

[0022] In this step, 2D positional encoding and image token masking enhance the spatial awareness of the image encoder, providing a stable positional signal for subsequent attention and avoiding interference from filler points. This step establishes a stable spatial encoding and pixel-level reversible mapping.

[0023] S203 inputs the set of interactive points from the box-level detection training dataset into the cue encoder of the SAM model to obtain point features.

[0024] S2031 pads the point set of each sample to a fixed length, constructs a point mask, assigns a labeled / unlabeled type to each point, and performs preprocessing synchronously with the image data; S2032 normalizes the pixel coordinates of each point according to the current valid area of ​​the image and inputs them into the position encoding MLP to obtain the point position features; S2033 adds the corresponding target category and static / dynamic state type embedding according to the type of the point, adds the point location features and type embedding to obtain the point features, and retains the point mask.

[0025] In this step, point mask preservation is used to mask invalid points in the subsequent attention mechanism. This step completes the resolution-independent point coordinate representation and semantic injection of prompts, and uses a masking mechanism to achieve batch processing robustness.

[0026] S3 fuses the intermediate image features with the point features based on the point features to obtain enhanced fused features; S301 inputs the intermediate image features and point features into the GSAF module to calculate the spatial gating weights corresponding to the point features. The calculation formula is as follows: , in, For the Sigmoid function, It is a convolutional layer. For intermediate image features, Point features; S302 multiplies the spatial gating weights with the intermediate image features and then fuses them to obtain enhanced image features; S303 fuses enhanced image features with point features to obtain enhanced fused features.

[0027] In this step, the GSAF (Gated Spatial Adaptive Fusion) module introduces a spatial gating mechanism. This gating acts like a "smart switch" or "dimmer," adaptively and selectively determining which parts of the image features should be enhanced and which should be suppressed based on the location and semantics of each point's cues before fusion.

[0028] S4 inputs the enhanced fusion features into the regression branch of the detector head, and performs regression classification through the multilayer perceptron of the detector head to obtain the prediction box corresponding to the moving target.

[0029] The S401 detector head's multilayer perceptron performs moving target bounding box regression and classification prediction on the enhanced fusion features, and outputs the bounding box confidence and class confidence of the corresponding candidate targets; S402 combines the bounding box parameters output by the regression branch with a mapping strategy to restore the bounding box in the original image pixel coordinates; S403 performs threshold filtering on the output category confidence and box confidence to eliminate low-confidence target candidate boxes; S4031 weights the category confidence and bounding box confidence according to the actual detection scenario to obtain the joint confidence; S4032 sets a filtering threshold based on the target region, retains target candidate boxes with a joint confidence level greater than or equal to the filtering threshold, and discards the remaining target candidate boxes.

[0030] In this step, to reduce the risk of false detections in autonomous driving systems and prevent background clutter from being identified as targets, a joint confidence score filtering mechanism is designed. Category confidence score refers to the probability score by which the model predicts that the target within the detection box belongs to a specific category (such as vehicle, pedestrian, traffic sign, etc.). Box confidence score refers to the probability score by which the model predicts whether a target actually exists at that location; it can also be understood as: whether the box contains an actual object. Autonomous driving requires extremely low false negatives and negatives. Certain backgrounds or complex environments can easily cause the model to mistakenly identify the background as a target, resulting in high category confidence scores even when there is no actual target. If only category confidence score is considered, the false detection rate will increase. Therefore, box confidence score must be used to supplement and filter out high-scoring boxes that do not contain objects. Thus, a joint filtering of these two confidence scores is necessary.

[0031] Some scenarios are more vulnerable to missed detections (such as detecting pedestrians on highways), while others are more vulnerable to false detections (such as detecting foreign objects in static structural areas). In certain scenarios (such as rainy days or nighttime, when lighting is extremely low), the category confidence score is easily affected by interference, but if we only focus on whether there is an object in front, we need to give more weight to the bounding box confidence score. Therefore, we need to obtain a joint confidence score by weighting according to the actual detection scenario.

[0032] S404 uses the Hungarian algorithm to match the filtered candidate boxes, removes redundant and duplicate boxes, and obtains the final set of predicted target boxes.

[0033] S4041 For each retained candidate target, calculate its matching cost with all real targets or historically tracked targets, and construct a matching cost matrix; S4042 inputs the cost matrix into the Hungarian algorithm, performs a global one-to-one optimal matching, and outputs the optimal matching relationship; S4043 performs a set of predicted target boxes on all candidate boxes. Based on the matching results, it retains the box that is uniquely and successfully assigned, and removes the remaining unassigned and / or redundant boxes that highly overlap with the assigned boxes, thus obtaining the predicted target box set.

[0034] In this step, deep features are used to extract the location and category probabilities of the target through a multilayer perceptron, removing false or low-confidence targets and improving the accuracy and reliability of the detection results. The Hungarian algorithm is used to assign each detection result to the corresponding real target end-to-end, resolving issues such as duplication, overlap, and redundancy in one go. This ensures that each target is assigned only a unique high-confidence prediction, achieving accurate detection and labeling of moving targets.

[0035] S405 decodes the predicted bounding box into the coordinates of four corner points, which are used as a new set of point cues. The new set of point cues is then input into the cue encoder again, and secondary feature fusion is performed with the image features to predict the bounding box. S406 performs a weighted fusion of the two prediction results to obtain the final prediction box.

[0036] In this step, the first predicted bounding boxes (S401-S402) may have a deviation of a few pixels on the boundary due to the initial cue position, feature fusion quality, or model tolerance. Iterative optimization provides the model with more accurate contextual information by decoding the initial predicted bounding boxes into richer corner cues (4 vs. 1 original), enabling fine-tuning and resulting in the final bounding boxes with extremely high edge fit.

[0037] Furthermore, occlusion, truncation, and small targets are typical challenging scenarios for autonomous driving, and initial predictions may be incomplete or inaccurate. Iterative optimization mechanisms can self-correct. For example, for an occluded car, the model may only predict the visible part on the first attempt; during iteration, the model infers the overall structure based on the already predicted parts, thus outputting a more complete bounding box.

[0038] S5 calculates the IOU loss between the predicted bounding box and the manual bounding box, and continuously updates the model parameters through the backpropagation algorithm until the model converges on the training dataset.

[0039] In this step, during object detection, IOU represents the ratio of the area of ​​intersection between the predicted bounding box and the ground truth bounding box to the area of ​​their union. It is used to measure the degree of overlap between the predicted bounding box and the ground truth bounding box. The trained model can automatically label moving objects on the image.

[0040] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0041] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0042] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0043] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0044] It should be noted that any reference signs placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The words first, second, and third, etc., do not indicate any order. These words can be interpreted as names.

[0045] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.

[0046] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

[0047] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0048] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms should not be construed as necessarily referring to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

Claims

1. An interactive object detection bounding box annotation method based on a large SAM model, characterized in that, include: Acquire image data of autonomous driving scenarios, manually annotate moving targets with bounding boxes, and construct a bounding box detection training dataset through random point sampling and label generation, including: Image data is input into an uncertainty prediction network to generate a corresponding uncertainty heatmap. An adaptive sampling probability map is generated by overlaying an uncertainty heatmap with a priori risk area. On each adaptive sampling probability map, n pixels are randomly selected as the interaction point set. If a point falls into any manually labeled box, it is marked as a marked point; otherwise, it is marked as an unmarked point. For each marked point, a corresponding image point set supervision label is generated based on the coordinate difference between the center point of its bounding box and the point, the bounding box scale parameter, the target category, and the static and dynamic state. The box-level detection training dataset is input into the modified SAM model backbone to obtain intermediate image features and point features, including: Remove the mask decoder from the SAM model and remap the point type semantics to labeled / unlabeled to obtain the modified SAM model backbone; The image data from the box-level detection training dataset is input into the image encoder of the SAM model to obtain intermediate image features; Input the set of interactive points from the box-level detection training dataset into the cue encoder of the SAM model to obtain point features; The intermediate image features and point features are fused based on the point features to obtain enhanced fused features, including: The intermediate image features and point features are input into the GSAF module to calculate the spatial gating weights corresponding to the point features. The calculation formula is as follows: , in, For the Sigmoid function, It is a convolutional layer. For intermediate image features, Point features; The spatial gating weights are multiplied with the intermediate image features and then fused to obtain the enhanced image features; The enhanced image features are fused with the point features to obtain the enhanced fused features; The enhanced fusion features are input into the regression branch of the detector head, and the multilayer perceptron of the detector head is used for regression classification to obtain the predicted box corresponding to the moving target. Calculate the IOU loss between the predicted bounding box and the manual bounding box, and continuously update the model parameters through the backpropagation algorithm until the model converges on the training dataset.

2. The interactive object detection bounding box annotation method based on the SAM large model according to claim 1, characterized in that, The step of inputting image data from the box-level detection training dataset into the image encoder of the SAM model to obtain intermediate image features includes: Normalize and align the image data in the box-level detection training dataset, and record the effective region and the original size; Normalized and size-aligned image data is input into an image encoder to obtain downsampled grid features, while generating matching two-dimensional position codes and key target region image masks. The downsampled grid features, two-dimensional position codes, and key target region image masks are flattened into a one-dimensional sequence and aligned according to the channel dimension to obtain intermediate image features, flattened position codes, and masks.

3. The interactive object detection bounding box annotation method based on the SAM large model according to claim 1, characterized in that, The step of inputting the set of interactive points from the box-level detection training dataset into the cue encoder of the SAM model to obtain point features includes: For each sample, the point set is padded to a fixed length, a point mask is constructed, each point is assigned a labeled / unlabeled type, and preprocessing is performed synchronously with the image data; The pixel coordinates of each point are normalized according to the current valid area of ​​the image and input into the position encoding MLP to obtain the point position features; Add the corresponding target category and static / dynamic state type embedding according to the type of the point. Add the point location features and type embedding to obtain the point features and retain the point mask.

4. The interactive object detection bounding box annotation method based on the SAM large model according to claim 1, characterized in that, The step of inputting the enhanced fusion features into the regression branch of the detection head, and performing regression classification through the multilayer perceptron of the detection head to obtain the prediction box corresponding to the moving target includes: The multilayer perceptron of the detection head performs moving target bounding box regression and classification prediction on the enhanced fusion features, and outputs the bounding box confidence and class confidence of the corresponding candidate targets; The bounding box parameters output by the regression branch are combined with a mapping strategy to restore the bounding box in the original image pixel coordinates; Threshold filtering is applied to the output category confidence and bounding box confidence to remove low-confidence target candidate boxes; The filtered candidate boxes are matched against the target using the Hungarian algorithm to remove redundant and duplicate boxes, resulting in the final set of predicted target boxes. The predicted bounding box is decoded into the coordinates of four corner points, which are used as a new set of point cues. The new set of point cues is then input into the cue encoder again, and secondary feature fusion is performed with the image features to predict the bounding box. The two prediction results are weighted and fused to obtain the final prediction box.

5. The interactive object detection bounding box annotation method based on the SAM large model according to claim 4, characterized in that, The threshold filtering of the output category confidence and bounding box confidence to eliminate low-confidence target candidate boxes includes: The joint confidence score is obtained by weighting the category confidence score and the bounding box confidence score according to the actual detection scenario. Based on the target region, a filtering threshold is set to retain target candidate boxes with a joint confidence level greater than or equal to the filtering threshold, while other target candidate boxes are discarded.

6. The interactive object detection bounding box annotation method based on the SAM large model according to claim 4, characterized in that, The process of matching the filtered candidate boxes with targets using the Hungarian algorithm to remove redundant and duplicate boxes, resulting in the final set of predicted target boxes, includes: For each retained candidate target, calculate its matching cost with all real targets or historically tracked targets, and construct a matching cost matrix; Input the cost matrix into the Hungarian algorithm, perform a global one-to-one optimal matching, and output the optimal matching relationship; For all candidate boxes, based on the matching results, the boxes that are uniquely and successfully assigned are retained, and the remaining unassigned and / or redundant boxes that highly overlap with the assigned boxes are removed, resulting in the final set of predicted target boxes.

Citation Information

Patent Citations

  • Target detection model training method and device and target detection method

    CN115131655A

  • Vehicle scene image data automatic labeling method and system based on Co-Deet model and SAM large model

    CN119579951A