2d interpolation label optimization method and device, equipment and computer readable storage medium
Patent Information
- Application Number
- CN202610965235.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-15
Smart Images

Figure CN122760992A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision applications, specifically to a 2D interpolation label optimization method, apparatus, device, and computer-readable storage medium. Background Technology
[0002] With the continuous development of computer vision applications such as intelligent driving, vehicle-road cooperation, and intelligent security, high-quality 2D object detection annotation data is a necessary foundation for training high-performance perception models. Since video sequences typically have high frame rates, the human and time costs of full manual annotation are extremely high. Therefore, the industry urgently needs an annotation data generation solution that can balance annotation costs and data quality, placing high demands on the accuracy, completeness, and generation efficiency of interpolated annotation labels.
[0003] In related technologies, the strategy of "sparse annotation + interpolation completion" is generally adopted to generate 2D annotation data for video sequences. That is, after manually annotating the target boxes of a small number of key frames, the annotation results of intermediate frames are generated through interpolation algorithms. Some solutions improve the interpolation accuracy by optimizing the interpolation algorithm itself, while others correct errors in the annotated data by manual review.
[0004] However, due to factors such as nonlinear target motion, scene occlusion, and sudden appearance / disappearance of targets, labels generated by traditional interpolation methods are prone to positional offsets and inaccurate dimensions. Furthermore, the interpolation method itself cannot solve the defects of missed target detection and false target detection in intermediate frames, resulting in insufficient accuracy and completeness of the labeled data. In addition, existing optimization schemes are either computationally complex and difficult to implement in engineering, or still rely on a large amount of manual review, failing to automate and efficiently complete the correction and error correction of interpolated labels. Summary of the Invention
[0005] This application provides a 2D interpolation label optimization method, apparatus, device, and computer-readable storage medium, which can solve the technical problems in related technologies such as insufficient accuracy of interpolation label position and size, inability to automatically handle missed detection and false detection, and difficulty in balancing engineering feasibility and label generation efficiency.
[0006] In a first aspect, embodiments of this application provide a 2D interpolation label optimization method, the method comprising: Obtain the 2D interpolated target set and the 2D target detection result set corresponding to the same video frame; wherein, the 2D interpolated target set is generated by interpolation of the annotations of the preceding and following keyframes, and the 2D target detection result set is obtained by inference from the target detection model; Based on category consistency constraints and location overlap matching rules, a one-to-one optimal matching is performed on the two target sets to obtain successfully matched target pairs, unmatched interpolated targets, and unmatched detected targets. For successfully matched target pairs, the bounding box information of the interpolated target and the detected target are fused to optimize and correct the position and size of the interpolated target; For unmatched interpolation targets, perform retention or deletion operations based on interpolation confidence; for unmatched detection targets, perform addition or ignore operations based on detection confidence. Summarize all processed targets, generate and output an optimized set of 2D interpolation labels.
[0007] In conjunction with the first aspect, in one implementation, the step of performing a one-to-one optimal matching of the two target sets based on category consistency constraints and position overlap matching rules includes: The interpolation target and the detection target are filtered by category, and only the target pairs with the same category label are selected as candidate matching pairs; Calculate the positional overlap between the two bounding boxes in each candidate matching pair, and construct the matching matrix; The optimal one-to-one matching result is obtained by solving the matching matrix, and the target pairs that are successfully matched, the interpolated targets that are not matched, and the detected targets that are not matched are divided.
[0008] In conjunction with the first aspect, in one implementation, the step of solving for the one-to-one optimal matching result based on the matching matrix includes: The overlap threshold for effective matching is dynamically adjusted based on the number of targets to be matched within the current frame. When the number of targets is small, the overlap threshold is reduced to improve the matching coverage. When the number of targets is large, maintain or increase the overlap threshold to ensure matching accuracy. Only matching pairs with a positional overlap degree not lower than the current overlap degree threshold are included in the optimal matching solution range as valid matching pairs.
[0009] In conjunction with the first aspect, in one implementation, the matching algorithm used to solve for the one-to-one optimal matching result based on the matching matrix includes any one of the Hungarian algorithm, the KM algorithm, and the greedy matching algorithm.
[0010] In conjunction with the first aspect, in one implementation, the step of fusing the bounding box information of the interpolated target and the detected target for the successfully matched target pair, and optimizing and correcting the position and size of the interpolated target, includes: The center point position parameter and the width and height dimension parameters of the target box are weighted and fused separately to obtain the optimized target box parameters; The fusion weights are dynamically adjusted based on the confidence level of the corresponding detection target. The higher the confidence level of the detection target, the greater the proportion of the fusion weights corresponding to the detection target bounding box parameters.
[0011] In conjunction with the first aspect, in one implementation, the step of performing a retain or deletion operation based on interpolation confidence for unmatched interpolation targets, and performing an add or ignore operation based on detection confidence for unmatched detection targets, includes: For each unmatched interpolation target, if its interpolation confidence is lower than the preset deletion threshold, it is determined to be a false detection target and deleted from the label set; if its interpolation confidence is not lower than the preset deletion threshold, the original interpolation target is retained and no correction is made. For each unmatched detection target, if its detection confidence is higher than the preset addition threshold, it is determined to be a missed detection target and added to the interpolation label set; if its detection confidence is not higher than the preset addition threshold, the detection target is ignored.
[0012] In conjunction with the first aspect, in one embodiment, the method further includes: For each target in the optimized 2D interpolation label set, the evaluation is carried out from three dimensions: source data credibility, optimization effectiveness, and time series consistency. The final confidence score corresponding to each target is calculated, and the final confidence score is associated with the label information of the corresponding target.
[0013] Secondly, embodiments of this application provide a 2D interpolation label optimization device, the 2D interpolation label optimization device comprising: The data acquisition module is used to acquire the 2D interpolated target set and the 2D target detection result set corresponding to the same video frame; wherein, the 2D interpolated target set is generated by interpolation of the annotations of the preceding and following keyframes, and the 2D target detection result set is obtained by inference from the target detection model; The target matching module is used to perform one-to-one optimal matching of two target sets based on category consistency constraints and location overlap matching rules, and to divide them into successfully matched target pairs, unmatched interpolated targets, and unmatched detected targets. The classification processing module is used to perform label optimization and error correction based on the matching results, including optimizing the frame of successfully matched target pairs and determining the confidence level and adding / deleting unmatched targets. The results output module is used to summarize all processed targets, generate and output an optimized set of 2D interpolation labels.
[0014] Thirdly, embodiments of this application provide a 2D interpolation label optimization device, which includes a processor, a memory, and a 2D interpolation label optimization program stored in the memory and executable by the processor. When the 2D interpolation label optimization program is executed by the processor, it implements the steps of the 2D interpolation label optimization method as described in some of the above embodiments.
[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing a 2D interpolation label optimization program, wherein when the 2D interpolation label optimization program is executed by a processor, it implements the steps of the 2D interpolation label optimization method as described in some of the above embodiments.
[0016] The beneficial effects of the technical solutions provided in this application include: Using the 2D label set generated by interpolation and the inference result set of the 2D object detection model as input, the two types of objects naturally correspond to the same video frame, without the need to perform additional cross-frame alignment operations. By constructing a closed-loop processing logic of matching-correction-addition and deletion, the position offset and size deviation of the interpolated labels can be systematically corrected. At the same time, it can make up for the defects of missed detection and false detection that the interpolation method itself cannot handle, improve the integrity and accuracy of the labeled data, and support the batch optimization of interpolated labels for large-scale continuous video frames without relying on a large amount of manual review. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating an embodiment of the 2D interpolation label optimization method of this application; Figure 2 This is a schematic diagram of the hardware structure of the 2D interpolation label optimization device involved in the embodiments of this application. Detailed Implementation
[0018] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0019] First, some of the technical terms used in this application will be explained to help those skilled in the art understand this application.
[0020] Keyframe: A frame in a video sequence that has been manually annotated with a complete 2D target. It serves as the reference frame for interpolating and generating intermediate frame labels.
[0021] 2D interpolation labels: These are the 2D object detection bounding boxes generated for intermediate non-keyframes after inter-frame interpolation of sparsely labeled keyframes in a video sequence. They include the object's position, size, and category information.
[0022] 2D interpolation target set: refers to the set of 2D interpolation labels generated by interpolation of keyframes before and after within the same video frame.
[0023] 2D object detection model: refers to a pre-trained computer vision model that can identify objects in an input image and output corresponding 2D detection boxes, category labels and confidence scores. It is the data source for the detection results of this solution. Specifically, it can include various types such as single-stage detection models (such as YOLO series, SSD) and detection models based on the Transformer architecture (such as DETR, DINO).
[0024] 2D object detection result set: refers to the set of all detected objects output by the model after inputting a single frame video image into a pre-trained 2D object detection model. Each detected object includes the bounding box information, category label, and detection confidence.
[0025] Interpolation confidence: A quantitative parameter used to characterize the credibility of 2D labels generated by interpolation, and is the core basis for determining whether to retain or delete unmatched interpolation targets.
[0026] Detection confidence: A quantitative parameter output by the target detection model that characterizes the credibility of the detection results. It is the core basis for determining whether to add or ignore unmatched targets.
[0027] Category consistency constraint: refers to the pre-constraint condition when matching interpolation targets with detection targets. Only target pairs with consistent category labels can enter the subsequent matching judgment stage.
[0028] Location overlap: refers to the ratio of the intersection area to the union area of two 2D target boxes, used to quantify the degree of spatial overlap between the two target boxes.
[0029] Matching matrix: A two-dimensional matrix used to carry the positional overlap values of all candidate matching pairs. It is the basis for solving the optimal one-to-one matching result.
[0030] One-to-one optimal matching: refers to the matching method in which each target participates in the matching at most once when pairing two types of targets, and the total matching score of the overall matching result reaches the optimal.
[0031] Bipartite Graph Maximum Weight Matching: A graph theory optimization algorithm for finding the set of edge matching with the maximum total weight in a bipartite graph structure, ensuring that each node participates in matching at most once. In this application, it is used to solve for the optimal pairing result between the interpolation target and the detection target.
[0032] The Hungarian algorithm is a classic algorithm for finding the maximum weight matching in a bipartite graph and is one of the alternative implementations of the one-to-one optimal matching in this application.
[0033] KM Algorithm: A bipartite graph optimal matching algorithm based on an extension of the Hungarian algorithm, suitable for solving perfect matching problems in weighted bipartite graphs, and is one of the optional implementation methods for one-to-one optimal matching in this application.
[0034] Greedy matching algorithm: A matching algorithm based on a local optimum strategy, which completes pairing in descending order of matching score, and is one of the optional implementation methods of one-to-one optimal matching in this application.
[0035] Adaptive overlap threshold: refers to the matching effective threshold that is dynamically adjusted according to the number of targets to be matched in the current frame; when the number of targets is small, the threshold is lowered to improve the matching coverage, and when the number of targets is large, the threshold is maintained or increased to ensure matching accuracy.
[0036] Weighted fusion: In this scheme, the position and size parameters of the interpolated target box and the detected target box are weighted and calculated to generate a corrected target box after optimization.
[0037] Final confidence score: A quantitative score obtained by comprehensively evaluating the optimized 2D labels from multiple dimensions, used to characterize the overall quality of the optimized labels.
[0038] Temporal smoothing: Applying inter-frame smoothing constraints to the label results of multiple consecutive frames to ensure the continuous and stable motion trajectory of the target box in post-processing.
[0039] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0040] In a first aspect, embodiments of this application provide a 2D interpolation label optimization method.
[0041] In one embodiment, reference is made to Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the 2D interpolation label optimization method of this application. Figure 1 As shown, the 2D interpolation label optimization method includes: S100: Obtain the 2D interpolated target set and the 2D target detection result set corresponding to the same video frame; wherein, the 2D interpolated target set is generated by interpolation of the annotations of the preceding and following keyframes, and the 2D target detection result set is obtained by inference from the target detection model; S200: Based on category consistency constraints and location overlap matching rules, one-to-one optimal matching is performed on two target sets to divide them into successfully matched target pairs, unmatched interpolated targets, and unmatched detected targets. S300: For successfully matched target pairs, fuse the bounding box information of the interpolated target and the detected target, and optimize and correct the position and size of the interpolated target; S400: For unmatched interpolation targets, perform a retain or delete operation based on the interpolation confidence; for unmatched detection targets, perform an add or ignore operation based on the detection confidence. S500: Summarizes all processed targets, generates and outputs an optimized set of 2D interpolation labels.
[0042] In this embodiment, to address the shortcomings of traditional interpolation methods, such as the susceptibility to positional and dimensional deviations due to nonlinear target motion and scene occlusion, and the inherent inability to handle missed and false detections, the single-frame inference results of a high-precision target detection model are used as the correction basis and fused with the interpolation labels within the same frame. An automated processing loop of "matching-correction-addition / deletion-output" is constructed, enabling the entire process of interpolation label optimization without manual review. Since the two types of target data naturally correspond to the same video frame, no additional cross-frame alignment is required; only the correspondence between targets within the set needs to be matched, ensuring correction accuracy while reducing computational complexity. The final optimized labels retain the temporal continuity of the interpolation labels while possessing the single-frame accuracy of the detection model, effectively improving the overall quality of the labeled data and adapting to the batch labeling and optimization needs of large-scale video sequences.
[0043] Furthermore, in one embodiment, the 2D interpolation target set is denoted as set. It includes m interpolation targets, each interpolation target Including the center point coordinates of the target bounding box Width and height dimensions Category tags and interpolation confidence The set of 2D target detection results is denoted as set. It includes n detection targets, each detection target Including target bounding box information Category tags and model detection confidence .
[0044] In this embodiment, the 2D interpolation target set is extracted from the output of the interpolation module and generated through interpolation operations based on manually annotated keyframes before and after the interpolation. It serves as the foundation for optimization processing. The 2D target detection result set is generated by the pre-trained target detection model through inference on the current frame image and serves as the reference for optimization correction. Clearly defining the field composition and data format of these two sets ensures data consistency and computability in subsequent matching calculations, optimization corrections, and confidence determination.
[0045] Furthermore, in one embodiment, the 2D object detection result set is obtained by inferring the current frame image from a pre-trained 2D object detection model. The 2D object detection model adopts the YOLO series model; it can also adopt any one of the DETR model, DINO model, and SSD model, or the inference results of multiple detection models can be fused together as the input of the detection result set.
[0046] In this embodiment, the target detection model can be flexibly selected based on the accuracy requirements and computing resources of the actual application scenario: the YOLO series models have the characteristics of fast inference speed and low deployment difficulty, and are suitable for batch processing of large-scale data; the DETR, DINO and other Transformer architecture models have higher detection accuracy and are suitable for scenarios with higher requirements for annotation quality; multi-model fusion can further improve the robustness of detection results. The flexible model selection mechanism can ensure that the solution is adapted to different engineering implementation conditions and accuracy requirements.
[0047] Furthermore, in one embodiment, step S200 includes the following steps: S201: Perform category filtering on the interpolation target and the detection target, and only select target pairs with the same category label as candidate matching pairs; S202: Calculate the positional overlap of the two target boxes in each candidate matching pair and construct the matching matrix; S203: Based on the matching matrix, solve for the one-to-one optimal matching result and divide the target pairs that are successfully matched, the interpolated targets that are not matched, and the detected targets that are not matched.
[0048] In this embodiment, a hierarchical matching logic is used to establish the correspondence between the interpolation target and the detection target: first, a first-round screening is completed through category constraints to eliminate invalid cross-category matching pairs; then, quantitative matching is performed through positional overlap; and finally, the globally optimal one-to-one pairing result is obtained. This hierarchical logic can significantly reduce the amount of invalid overlap calculations and avoid cross-category mismatch problems at the root, ensuring the computational efficiency of the matching process and the accuracy of the matching results.
[0049] Furthermore, in one embodiment, in S202, the positional overlap is quantified using the intersection-over-union ratio (IoU) for interpolation targets that satisfy class consistency. With the detection target Its IoU calculation formula is:
[0050] in The area function representing the bounding box.
[0051] Based on the IoU calculation results of all candidate matching pairs, a structure of size is constructed. The matching matrix M; the elements in the matrix The rule for determining the value is: if the interpolation target With the detection target If the categories are the same, then If the two categories are inconsistent, then The corresponding target pair is directly judged as a mismatch.
[0052] In this embodiment, the intersection-union ratio (IUGR) is used as a quantitative indicator of the degree of overlap in the target bounding box space. It is a common standard in the field of target detection for measuring the degree of box matching, with simple calculation logic and clear physical meaning. Setting the scores of targets with inconsistent categories to 0 directly embeds category constraints into the matching matrix, eliminating the need for additional category judgment in subsequent optimal matching solutions. This simplifies the algorithm logic and provides a standardized computational basis for finding the global optimal matching.
[0053] Furthermore, in one embodiment, step S203 includes the following steps: S203-1: Dynamically adjust the overlap threshold for effective matching based on the number of targets to be matched within the current frame; S203-2: When the number of targets is small, reduce the overlap threshold to improve the matching coverage. S203-3: When the number of targets is large, maintain or increase the overlap threshold to ensure matching accuracy; S203-4: Only matching pairs with a positional overlap degree not lower than the current overlap degree threshold are included in the optimal matching solution range as valid matching pairs.
[0054] In this embodiment, the overlap threshold for effective matching is denoted as . Its value is dynamically adjusted based on the number of targets in the current frame: When the interpolation target number And the number of targets to be detected When the target quantity is determined to be small, the overlap threshold is calculated using the following formula:
[0055] in The basic matching threshold is typically set to 0.5. The threshold adjustment amount is adaptively reduced, and is usually set to 0.2~0.3.
[0056] When the interpolation target number Or the number of targets to be detected When the number of targets is deemed large, the overlap threshold is set. In high-density scenarios, it can be further improved to .
[0057] The core logic behind using adaptive thresholds is as follows: In scenarios with sparse targets, the probability of overlapping bounding boxes of different targets is extremely low. Lowering the matching threshold will not cause the risk of false matching, but can effectively cover correct target pairs whose IoU is slightly lower than the fixed threshold due to the overall offset of the interpolation box caused by the z-axis deviation in the interpolation operation, thus improving the matching coverage rate. In scenarios with dense targets, the probability of overlapping bounding boxes between targets increases. Maintaining or increasing the threshold can reduce false matching and ensure matching accuracy. By dynamically adjusting the threshold according to the scenario, the matching success rate in sparse scenarios is significantly improved without sacrificing matching accuracy, enabling the solution to perform well even in road scenarios with low traffic volume.
[0058] Furthermore, in one embodiment, in S203, the Hungarian Algorithm is used in the matching matrix M to obtain the set of optimal one-to-one matching pairs between the interpolation set I and the detection set D, with the goal of maximizing the total matching score, while ensuring that each target is matched at most once. After performing the optimal pairing, the interpolation targets and detection targets are divided into the following three subsets:
[0059] In this embodiment, the Hungarian algorithm is a classic algorithm for finding the maximum weighted matching in a bipartite graph. It performs global optimization with the goal of maximizing the total matching score and strictly adheres to the one-to-one matching constraint to avoid duplicate pairing of a single target. This algorithm yields globally optimal matching results, achieving higher matching accuracy compared to locally optimal matching strategies. The explicit division of the target subsets into three categories through the aforementioned set definition provides a clear classification basis for subsequent differentiated processing, making the logical boundaries for optimizing matching targets and adding / removing unmatched targets clear and highly executable.
[0060] Furthermore, in one embodiment, the matching algorithm used to solve for the one-to-one optimal matching result based on the matching matrix includes any one of the Hungarian algorithm, the KM algorithm, and the greedy matching algorithm.
[0061] In this embodiment, the matching algorithm can be flexibly selected based on actual computing power and accuracy requirements: greedy matching is faster and suitable for batch processing scenarios with limited computing power; the KM algorithm is more suitable for perfect matching solutions of weighted bipartite graphs, with better matching accuracy. Diversifying the threshold adjustment methods can further enhance the scenario adaptability of the solution and meet the matching performance requirements of different scenarios.
[0062] Furthermore, in one embodiment, step S300 includes the following steps: S301: Perform weighted fusion calculations on the center point position parameters and width and height parameters of the target box to obtain the optimized target box parameters; The fusion weights are dynamically adjusted based on the confidence level of the corresponding detection target. The higher the confidence level of the detection target, the greater the proportion of the fusion weights corresponding to the detection target bounding box parameters.
[0063] In this embodiment, the interpolation target box is set as follows: ,in The coordinates of the center point, The width and height are given; the corresponding detection bounding box is given. ,in The coordinates of the center point, For width and height.
[0064] Formula for sizing optimization: ,
[0065] in This is the interpolation size weighting factor, which typically ranges from 0.2 to 0.8.
[0066] The formula for calculating location optimization is as follows: ,
[0067] in The interpolation position weighting factor is set in a way similar to... Consistent.
[0068] The optimized target bounding box is denoted as .
[0069] The core logic of using a weighted fusion approach for bounding box optimization lies in the fact that interpolated bounding boxes exhibit smooth inter-frame motion and good temporal consistency, but are prone to single-frame position and size deviations due to nonlinear target motion. Detection models output high single-frame positioning accuracy for their bounding boxes, but experience slight fluctuations between frames. By weighting and fusing the position and size parameters of both types of bounding boxes, the advantages of both can be combined. This corrects the deviation of the interpolated bounding boxes while preserving the temporal smoothness of the interpolation labels, avoiding the jump problems caused by single-frame detection. The fusion weights are dynamically adjusted based on the detection confidence level because a higher detection confidence level indicates stronger reliability of the detected bounding boxes, thus assigning a higher weight percentage. This allows the optimization results to better reflect the true position and size of the target, further improving optimization accuracy.
[0070] Furthermore, in one embodiment, in S300, when optimizing and correcting the position and size of the interpolation target, in addition to the weighted fusion method, the detection target box can be directly used to completely replace the interpolation target box, or the interpolation target box can be used as a reference and the detection target box can be slightly adjusted; Kalman filtering or extended Kalman filtering can also be used to perform temporal fusion optimization of the interpolation box and the detection box in combination with historical information from previous frames.
[0071] In this embodiment, the method of directly replacing the detection box is suitable for scenarios with extremely high detection model accuracy and extremely high requirements for single-frame positioning accuracy; the small-scale fine-tuning method can preserve the temporal characteristics of the interpolation label to the greatest extent and only correct obvious deviations; temporal fusion methods such as Kalman filtering can introduce historical motion information from previous frames to further smooth and optimize the results, which is suitable for video tracking and annotation scenarios with higher requirements for temporal continuity. Multiple optimization implementation methods can adapt to different accuracy and temporal requirements, improving the scenario adaptability of the solution.
[0072] Furthermore, in one embodiment, step S400 includes the following steps: S401: For each unmatched interpolation target, if its interpolation confidence is lower than the preset deletion threshold, it is determined to be a false detection target and deleted from the label set; if its interpolation confidence is not lower than the preset deletion threshold, the original interpolation target is retained and no correction is made. S402: For each unmatched detection target, if its detection confidence is higher than the preset addition threshold, it is determined to be a missed detection target and added to the interpolation label set; if its detection confidence is not higher than the preset addition threshold, the detection target is ignored.
[0073] In this embodiment, the deletion threshold is denoted as... The value is usually set to 0.35; adding a threshold is denoted as The value is usually 0.65.
[0074] The logic behind using confidence level determination for unmatched targets is as follows: unmatched interpolated targets have two possibilities: either they are false targets generated by interpolation (false detections), or they are real targets not detected by the detection model. Targets with low confidence and no supporting detection results are highly likely to be false targets, and deleting them can effectively eliminate false detections. Unmatched detected targets also have two possibilities: either they are real targets not covered by the interpolation process (missed detections, such as targets that suddenly appear between keyframes or reappear after being occluded), or they are false alarms from the detection model. Detected targets with high confidence are highly likely to be real targets, and adding them can effectively supplement targets missed by interpolation, solving the inherent limitation of interpolation methods in handling missed detections. The preset deletion and addition thresholds can be calibrated through experimental data based on specific application scenarios such as urban roads and highways, ensuring the rationality of the thresholds and scenario adaptability. This mechanism can systematically improve the completeness and accuracy of labeled data.
[0075] Furthermore, in one embodiment, in S500, each target in the generated optimized 2D interpolation label set includes complete target bounding box position and size information, tracking ID, category label, optimized confidence score, and operation history marker; the operation history marker includes four corresponding types: match correction, high confidence retention, low confidence deletion, and model supplementation.
[0076] In this embodiment, all processed targets are aggregated to generate a final label set. Complete attribute information and operation history markers are attached to each target, enabling full-chain traceability of the optimization process for each label. The operation history markers clearly distinguish the optimization source of the target, facilitating subsequent usability assessment, quality evaluation, and tiered screening of the labeled data, thus supporting the full lifecycle management of labeled data.
[0077] Furthermore, in one embodiment, the method further includes: S600: For each target in the optimized 2D interpolation label set, evaluate it from three dimensions: source data credibility, optimization effectiveness, and time series consistency. Calculate the final confidence score for each target and associate the final confidence score with the label information of the corresponding target.
[0078] In this embodiment, the final confidence score is recorded as follows: Its calculation formula is:
[0079] in For source data confidence: For successfully matched targets, the interpolation confidence and detection confidence are combined; for retained unmatched interpolated targets, the interpolation confidence is the primary factor in the calculation; for supplementary detection targets, the detection confidence is the primary factor in the calculation.
[0080] To optimize the effectiveness score: a comprehensive evaluation is conducted based on the magnitude of the change in the target bounding box before and after optimization, as well as the rationality of the weighting factors in the optimization formula.
[0081] For temporal consistency score: a smooth evaluation is performed by combining the consistency of optimization results of adjacent frames.
[0082] , , Let be the weighting coefficient, satisfying The value is usually taken as .
[0083] The logic behind using a multi-dimensional weighted calculation to determine the final confidence score is that a single source data confidence score cannot fully reflect the overall quality of the optimized labels. It is necessary to comprehensively evaluate three dimensions: source data credibility, the rationality of the optimization operation, and temporal consistency, corresponding to the input quality, optimization quality, and sequence quality levels, respectively. The quantified scoring results provide a unified quality assessment benchmark for labeled data, facilitating the selection of labeled data of different quality levels based on model training needs, and also providing a quantifiable basis for the quality control of batch data.
[0084] Furthermore, in one embodiment, the method further includes: S700: After performing optimization processing on each frame in the video sequence, the optimized interpolated label sequence of the entire video is subjected to cross-frame temporal smoothing processing to ensure the continuity of label motion between frames.
[0085] In this embodiment, the aforementioned optimization steps are repeated for each frame in the video sequence to obtain a complete optimized 2D interpolated label sequence. After frame-by-frame optimization, cross-frame temporal smoothing is added to eliminate small label jumps that may be introduced by single-frame optimization, further ensuring the inter-frame continuity of the target box motion trajectory. Since the overall optimization is based on keyframe interpolation, the correction operations are all completed within the same frame, naturally maintaining consistency with the logic of manual annotation. Combined with cross-frame smoothing, the final label sequence has both higher single-frame accuracy and good temporal smoothness, making it suitable for both single-frame 2D target detection training and video multi-target tracking training, thus expanding the applicability of the labeled data.
[0086] Secondly, embodiments of this application also provide a 2D interpolation label optimization device, which includes: a data acquisition module, used to acquire a 2D interpolation target set and a 2D target detection result set corresponding to the same video frame; wherein, the 2D interpolation target set is generated by interpolation of keyframe annotations before and after, and the 2D target detection result set is obtained by inference from a target detection model; a target matching module, used to perform one-to-one optimal matching of the two target sets based on category consistency constraints and position overlap matching rules, and to divide them into successfully matched target pairs, unmatched interpolated targets, and unmatched detected targets; a classification processing module, used to perform label optimization and error correction according to the matching results, including frame optimization for successfully matched target pairs and confidence determination and addition / deletion operations for unmatched targets; and a result output module, used to summarize all processed targets, generate and output the optimized 2D interpolation label set. The functions of each module in the above-mentioned 2D interpolation label optimization device correspond to the steps in the above-mentioned 2D interpolation label optimization method embodiment, and their functions and implementation processes will not be described in detail here.
[0087] Thirdly, embodiments of this application provide a 2D interpolation label optimization device, which can be a personal computer (PC), laptop computer, server, or other device with data processing capabilities.
[0088] Reference Figure 2 , Figure 2 This is a schematic diagram of the hardware structure of the 2D interpolation label optimization device involved in the embodiments of this application. In the embodiments of this application, the 2D interpolation label optimization device may include a processor, a memory, a communication interface, and a communication bus.
[0089] The communication bus can be of any type and is used to interconnect the processor, memory, and communication interface.
[0090] The communication interface includes input / output (I / O) interfaces, physical interfaces, and logical interfaces used for interconnecting devices within the 2D interpolation label optimization device, as well as interfaces used for interconnecting the 2D interpolation label optimization device with other devices (such as other computing devices or user equipment). Physical interfaces can be Ethernet interfaces, fiber optic interfaces, ATM interfaces, etc.; user equipment can be displays, keyboards, etc.
[0091] Memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.
[0092] The processor can be a general-purpose processor, which can call the 2D interpolation label optimization program stored in memory and execute the 2D interpolation label optimization method provided in the embodiments of this application. For example, the general-purpose processor can be a central processing unit (CPU). The method executed when the 2D interpolation label optimization program is called can be referred to in the various embodiments of the 2D interpolation label optimization method of this application, and will not be repeated here.
[0093] Those skilled in the art will understand that Figure 2 The hardware structure shown does not constitute a limitation of this application and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0094] Fourthly, embodiments of this application also provide a readable storage medium.
[0095] The present application has a 2D interpolation label optimization program stored on a readable storage medium, wherein when the 2D interpolation label optimization program is executed by a processor, it implements the steps of the 2D interpolation label optimization method as described above.
[0096] The method implemented when the 2D interpolation label optimization procedure is executed can be referred to in various embodiments of the 2D interpolation label optimization method of this application, and will not be repeated here.
[0097] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0098] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover a non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus. The terms "first," "second," and "third," etc., are used to distinguish different objects, etc., and do not indicate a sequence, nor do they limit "first," "second," and "third" to different types.
[0099] In the description of the embodiments of this application, terms such as "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a concrete manner.
[0100] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.
[0101] In some processes described in the embodiments of this application, multiple operations or steps are included in a specific order. However, it should be understood that these operations or steps may not be executed in the order they appear in the embodiments of this application, or they may be executed in parallel. The sequence number of the operation is only used to distinguish different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed sequentially or in parallel, and these operations or steps may be combined.
[0102] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device to execute the methods described in the various embodiments of this application.
[0103] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A 2D interpolation label optimization method, characterized in that, The method includes: Obtain the 2D interpolated target set and the 2D target detection result set corresponding to the same video frame; wherein, the 2D interpolated target set is generated by interpolation of the annotations of the preceding and following keyframes, and the 2D target detection result set is obtained by inference from the target detection model; Based on category consistency constraints and location overlap matching rules, a one-to-one optimal matching is performed on the two target sets to obtain successfully matched target pairs, unmatched interpolated targets, and unmatched detected targets. For successfully matched target pairs, the bounding box information of the interpolated target and the detected target are fused to optimize and correct the position and size of the interpolated target; For unmatched interpolation targets, perform retention or deletion operations based on interpolation confidence; for unmatched detection targets, perform addition or ignore operations based on detection confidence. Summarize all processed targets, generate and output an optimized set of 2D interpolation labels.
2. The 2D interpolation label optimization method as described in claim 1, characterized in that, The method of performing a one-to-one optimal matching of two target sets based on category consistency constraints and position overlap matching rules includes: The interpolation target and the detection target are filtered by category, and only the target pairs with the same category label are selected as candidate matching pairs; Calculate the positional overlap between the two bounding boxes in each candidate matching pair, and construct the matching matrix; The optimal one-to-one matching result is obtained by solving the matching matrix, and the target pairs that are successfully matched, the interpolated targets that are not matched, and the detected targets that are not matched are divided.
3. The 2D interpolation label optimization method as described in claim 2, characterized in that, The process of finding the optimal one-to-one matching result based on the matching matrix includes: The overlap threshold for effective matching is dynamically adjusted based on the number of targets to be matched within the current frame. When the number of targets is small, the overlap threshold is reduced to improve the matching coverage. When the number of targets is large, maintain or increase the overlap threshold to ensure matching accuracy. Only matching pairs with a positional overlap degree not lower than the current overlap degree threshold are included in the optimal matching solution range as valid matching pairs.
4. The 2D interpolation label optimization method as described in claim 2, characterized in that, The matching algorithm used to solve the one-to-one optimal matching result based on the matching matrix includes any one of the Hungarian algorithm, KM algorithm, and greedy matching algorithm.
5. The 2D interpolation label optimization method as described in claim 1, characterized in that, For successfully matched target pairs, the frame information of the interpolated target and the detected target is fused to optimize and correct the position and size of the interpolated target, including: The center point position parameter and the width and height dimension parameters of the target box are weighted and fused separately to obtain the optimized target box parameters; The fusion weights are dynamically adjusted based on the confidence level of the corresponding detection target. The higher the confidence level of the detection target, the greater the proportion of the fusion weights corresponding to the detection target bounding box parameters.
6. The 2D interpolation label optimization method as described in claim 1, characterized in that, The steps of performing retention or deletion operations based on interpolation confidence for unmatched interpolation targets, and adding or ignoring operations based on detection confidence for unmatched detection targets, include: For each unmatched interpolation target, if its interpolation confidence is lower than the preset deletion threshold, it is determined to be a false detection target and deleted from the label set; if its interpolation confidence is not lower than the preset deletion threshold, the original interpolation target is retained and no correction is made. For each unmatched detection target, if its detection confidence is higher than the preset addition threshold, it is determined to be a missed detection target and added to the interpolation label set; if its detection confidence is not higher than the preset addition threshold, the detection target is ignored.
7. The 2D interpolation label optimization method of claim 1, wherein, The method further includes: For each target in the optimized 2D interpolation label set, the evaluation is carried out from three dimensions: source data credibility, optimization effectiveness, and time series consistency. The final confidence score corresponding to each target is calculated, and the final confidence score is associated with the label information of the corresponding target.
8. A 2D interpolation label optimization apparatus, characterized by, The 2D interpolation label optimization device includes: The data acquisition module is used to acquire the 2D interpolated target set and the 2D target detection result set corresponding to the same video frame; wherein, the 2D interpolated target set is generated by interpolation of the annotations of the preceding and following keyframes, and the 2D target detection result set is obtained by inference from the target detection model; The target matching module is used to perform one-to-one optimal matching of two target sets based on category consistency constraints and location overlap matching rules, and to divide them into successfully matched target pairs, unmatched interpolated targets, and unmatched detected targets. The classification processing module is used to perform label optimization and error correction based on the matching results, including optimizing the frame of successfully matched target pairs and determining the confidence level and adding / deleting unmatched targets. The results output module is used to summarize all processed targets, generate and output an optimized set of 2D interpolation labels.
9. A 2D interpolation label optimization device, characterized in that, The 2D interpolation label optimization device includes a processor, a memory, and a 2D interpolation label optimization program stored in the memory and executable by the processor, wherein when the 2D interpolation label optimization program is executed by the processor, it implements the steps of the 2D interpolation label optimization method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a 2D interpolation label optimization program, wherein when the 2D interpolation label optimization program is executed by a processor, it implements the steps of the 2D interpolation label optimization method as described in any one of claims 1 to 7.