A two-dimensional target annotation method and system for multi-scenario autonomous driving image data

CN122510852BActive Publication Date: 2026-09-29JIANGSU DALUOTOU ZHIJIA TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610993332.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-29
Estimated Expiration
2046-07-06

AI Technical Summary

Technical Problem

标注执行过程中仍以人工目视和主观经验作为主要判定路径,导致同一目标在不同批次的标注结果之间一致性差,难以满足感知模型对标注精度的要求

Benefits of technology

本发明通过构建包含数据完整性校验与稀疏光流运动估计的有效帧筛选步骤,将镜头炫光、长时间静止等无效帧提前剔除,使得后续标注处理仅针对具备实质感知价值的图像帧进行,减少了因低质量帧引入的误标和算力浪费。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122510852B_ABST
    Figure CN122510852B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data processing technology and discloses a two-dimensional target annotation method and system for multi-scenario autonomous driving image data. The method includes: filtering a set of valid frames; constructing a five-dimensional annotation vector; jointly determining the occlusion area ratio and key point visibility by category to obtain annotation decisions; generating the final two-dimensional bounding box; and exporting the label file and delivery log. Compared with the existing technology where annotation specifications rely on qualitative descriptions, especially in dense traffic flow where targets occlude each other and have diverse categories, it is impossible to formulate separate occlusion ratio thresholds and key part visibility constraints for different categories to solve the technical problems of annotation inconsistency and omission / mislabeling caused by subjective judgment. Because this application sets category-differentiated occlusion area percentage and key part visibility rules, and cascades filtering, vector construction, joint determination, filtering, and export, it achieves automated and reproducible annotation for multiple categories, improving annotation consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a two-dimensional target annotation method and system for multi-scenario autonomous driving image data. Background Technology

[0002] Currently, in the data annotation stage of autonomous driving perception systems for mining transportation and port operations, large-scale two-dimensional bounding box annotations are required for various targets such as mining trucks, trailers, and pedestrians to train and validate target detection models. These confined scenarios involve a wide range of vehicle sizes, dense traffic flow, and frequent occlusion between targets. Existing technologies typically rely on annotators making judgments based on simplified qualitative rules, such as "no annotation for most occluded targets" or "no annotation for truncated targets." The annotation specifications themselves lack quantifiable thresholds based on category and viewpoint, and do not set operable detection conditions for the visibility of key components such as wheels, heads, and connecting parts. During the annotation process, manual visual inspection and subjective experience remain the primary decision-making path, resulting in poor consistency between annotation results for the same target across different batches, making it difficult to meet the accuracy requirements of perception models.

[0003] For example, when mining trucks are densely packed and waiting to be loaded, the area obscured by the following truck may exceed 60%, but the front or rear wheels of the preceding truck may still be clearly visible. According to existing qualitative rules, such targets would be discarded entirely, resulting in the loss of valuable training samples. Similarly, when container trailers are traveling in a port, the trailers and tractor units exhibit drastically different occlusion patterns. If annotators rely solely on subjective judgment, they are prone to incorrectly removing targets with large areas of obscuration but visible gooseneck joints (the articulated connection between the trailer and tractor unit) or wheels at the bottom of the trailer, or retaining targets with truncated front edges but complete outlines. This leads to conflicting judgment standards among different annotators or between different batches. Such omissions and mislabelings, caused by a lack of quantitative criteria and key component inspection mechanisms, introduce significant annotation noise into the training data, exacerbating the degradation of model feature learning and reducing the robustness of the detection model in densely occluded scenarios. Existing qualitative rules cannot adequately meet the needs for high-quality annotation in densely occluded, multi-category scenarios.

[0004] Therefore, there is an urgent need for a labeling method that can still formulate quantifiable thresholds for the proportion of occlusion area and visibility constraints for key parts such as wheels, heads, and gooseneck joints for different categories in dense traffic flow where targets occlude each other and are of various types. This method can eliminate human subjective bias through an automated joint judgment mechanism, thereby improving the reproducibility, consistency and delivery quality of cross-batch labeling and ensuring that the labeling results can provide reliable supervision signals for downstream perception models. Summary of the Invention

[0005] To address the aforementioned technical shortcomings, the present invention aims to propose a two-dimensional target annotation method for multi-scenario autonomous driving image data. This method addresses the shortcomings of existing annotation standards, which rely on qualitative descriptions for determining occlusion and truncation, and fail to provide quantitative thresholds and visibility conditions for key parts based on category and viewpoint. In particular, when targets in dense traffic flow occlude each other and have diverse categories, it is impossible to formulate separate occlusion ratio thresholds and visibility constraints for key parts for different categories. This method aims to solve the technical problems of inconsistent annotation and omissions or errors caused by subjective judgment.

[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: The present invention provides a two-dimensional target annotation method for multi-scenario autonomous driving image data.

[0007] The two-dimensional target annotation method for multi-scenario autonomous driving image data includes: Step S10: Obtain the image sequence from the vehicle-mounted camera. Based on the image sequence from the vehicle-mounted camera, perform a valid frame filtering task using data integrity verification and Lucas-Kanade sparse optical flow motion estimation mechanism, and output a set of valid frames. Step S20: Based on the set of valid frames, a pixel coordinate system and a five-dimensional annotation vector construction mechanism are used to perform the annotation vector generation task and output the five-dimensional annotation vector; Step S30: Based on the five-dimensional annotation vector, the annotation generation and judgment task is performed using a category-adaptive occlusion area ratio and key point joint judgment mechanism, and the annotation generation and judgment result is output. Step S40: Based on the annotation generation judgment result, the bounding box generation and filtering task is performed using the scale-adaptive white space and directional gradient histogram and local binary pattern discriminability verification mechanism, and the final two-dimensional bounding box is output. Step S50: Export the data in a standardized format and verify the number of files based on the final two-dimensional bounding box, and output the tag file and delivery log.

[0008] Preferably, step S10, which involves acquiring an image sequence from the vehicle-mounted camera, performing a valid frame filtering task based on the image sequence using data integrity verification and a Lucas-Kanade sparse optical flow motion estimation mechanism, and outputting a set of valid frames, specifically includes: Step S101: Perform data integrity verification frame by frame on the image sequence of the vehicle camera, read the magic number, image encoding format, resolution field and data block integrity flag in the header of each image file, and mark the corresponding image frame as invalid and discard it when any image file cannot be parsed, verification fails or the resolution field is inconsistent with the preset camera acquisition specifications. Step S102: For image frames that have passed data integrity verification, read their timestamp field and query the distortion correction completion flag in the metadata tag of the corresponding image frame; when the distortion correction completion flag indicates that the intrinsic parameter correction has not been completed, remove the corresponding image frame; when the distortion correction completion flag indicates that the intrinsic parameter correction has been completed, convert the corresponding image frame from the original color space to the HSV color space, and calculate the area of ​​overly bright and overly dark regions based on the luminance component V in the HSV color space; wherein, the connected regions of pixels with luminance component V greater than a preset overly bright threshold are determined as overly bright regions, and the connected regions of pixels with luminance component V less than a preset overly dark threshold are determined as overly dark regions, the area of ​​the overly bright region is the sum of the number of pixels contained in all overly bright regions, and the area of ​​the overly dark region is the sum of the number of pixels contained in all overly dark regions; if the proportion of the overly bright region area to the total area of ​​the image frame exceeds 80%, or the proportion of the overly dark region area to the total area of ​​the image frame exceeds 20%, then the image frame is determined to be an unrecoverable occluded frame or a strong light frame and is removed. Step S103: Perform still frame determination on the retained image frames, use the ORB feature point detection algorithm to extract feature points in adjacent image frames, and use the Lucas-Kanade sparse optical flow method to track the inter-frame displacement of the feature points, and calculate the average displacement amplitude of all successfully tracked feature points; at the same time, read the vehicle speed, and when the average displacement amplitude is less than 0.5 pixels per frame and the vehicle speed is less than 0.1 meters per second in three consecutive frames, mark the three consecutive frames as still invalid frames and discard them, and form the set of valid frames with the image frames that are not discarded; wherein, the average displacement amplitude is used to characterize the degree of change in the field of view between adjacent image frames that can be used for target annotation.

[0009] Preferably, step S20, which involves using a pixel coordinate system and a five-dimensional annotation vector construction mechanism to perform the annotation vector generation task based on the effective frame set and outputting the five-dimensional annotation vector, specifically includes: Step S201: Establish a pixel coordinate system for each valid frame in the set of valid frames, with the upper left corner of the image as the origin of the pixel coordinate system, the horizontal direction to the right as the positive x-axis, and the vertical direction downward as the positive y-axis, and determine the visible area boundary of the target to be labeled under the pixel coordinate system. Step S202: Represent the annotation information of the target to be annotated as a five-dimensional annotation vector V, wherein the five-dimensional annotation vector V satisfies:

[0010] Where label_type is the category identifier string of the target to be labeled, x and y are the horizontal and vertical coordinates of the center point of the two-dimensional bounding box in the pixel coordinate system, and w and h are the pixel width and pixel height of the two-dimensional bounding box, respectively.

[0011] Preferably, in step S202, the category identifier string label_type of the target to be labeled includes vehicle class, truck class, pedestrian class, trailer class and construction machinery class; wherein, the same category identifier string is consistent in the set of valid frames, the five-dimensional annotation vector, the annotation generation judgment result, the final two-dimensional bounding box and the label file.

[0012] Preferably, step S30, which involves performing the annotation generation and determination task based on the five-dimensional annotation vector using a category-adaptive occlusion area ratio and keypoint joint determination mechanism, and outputting the annotation generation and determination result, specifically includes: Step S301: Determine the visible contour mask of the target to be annotated in the current frame based on the five-dimensional annotation vector. and complete contour mask Calculate the shading area ratio according to the following formula. :

[0013] in, This represents the percentage of the occlusion area of ​​the target to be labeled. The pixel position in the mask. pixel position The value in the visible contour mask pixel position The value in the complete contour mask Indicates the area of ​​the visible region. Indicates the area of ​​the complete target region; Step S302: Determine the intersection of the complete bounding box of the target to be labeled with the image boundary. When the target is only truncated by the image boundary in the horizontal direction, calculate the truncation ratio C using the width truncation ratio method. When the target is truncated only along the vertical direction by the image boundary, the truncation ratio C is calculated using the height truncation ratio method; When the target is truncated by the image boundary along both the horizontal and vertical directions, the truncation ratio C is calculated using the area truncation ratio method. Step S303: Set the occlusion ratio threshold based on the category identifier string of the target to be labeled. and cutoff ratio threshold Specifically, the occlusion threshold for vehicles is 0.60 and the truncation threshold is 0.50; the occlusion threshold for pedestrians is 0.50 and the truncation threshold is 0.40; and the occlusion threshold for trailers is... The threshold value is 0.70 and the cutoff ratio is 0.60. If the proportion of the occluded area Greater than the occlusion ratio threshold Or the cutoff ratio C is greater than the cutoff ratio threshold. If the condition is met, then it is determined that no annotation will be generated; otherwise, it is determined that an annotation will be generated, and the result of the determination of whether or not an annotation is generated will be used as the annotation generation determination result.

[0014] Preferably, in step S40, based on the annotation generation judgment result, the bounding box generation and filtering task is performed using a scale-adaptive whitespace and directional gradient histogram and local binary pattern discriminability verification mechanism to output the final two-dimensional bounding box, specifically including: Step S401: For the target whose annotation was generated in step S30, obtain its visible area A, and calculate the adaptive whitespace spacing g according to the following formula:

[0015] in, For the target to be labeled in the valid frame Image location in This is an image boundary constraint function; when the expanded boundary does not exceed the image boundary, it follows... Expand the bounding box on all four sides; when the expanded boundary exceeds the image boundary, keep the edge that exceeds the image boundary at the image boundary, and do not expand the bounding box on the truncated edge. Step S402: For targets with a visible area A less than 225 pixels squared, extract the histogram of directional gradients and local binary pattern texture features of the visible region; wherein, the histogram of directional gradients is used to characterize the distribution of the target edge direction, and the local binary pattern texture features are used to characterize the local grayscale texture structure of the target. Step S403: Calculate the similarity between the histogram of oriented gradients features and the histogram of oriented gradients reference features of untruncated targets of the same category to obtain the histogram of oriented gradients similarity. The local binary pattern texture features are compared with the local binary pattern reference features of the same category of untruncated targets to obtain the local binary pattern similarity. If the similarity of the directional gradient histogram is less than 0.35 and the similarity of the local binary pattern is less than 0.40, the corresponding target will be filtered as an unlabeled target; otherwise, the expanded bounding box will be retained as the final two-dimensional bounding box.

[0016] Preferably, in step S50, the delivery log includes the total number of scenarios, the number of image files in each scenario, the number of tag files, the consistency verification results of the number of image files and tag files, and a summary table of the number of labeled instances for each category.

[0017] This invention also provides a two-dimensional target annotation system for multi-scenario autonomous driving image data, comprising: The effective frame filtering module is used to acquire the image sequence of the vehicle camera, and perform the effective frame filtering task based on the image sequence of the vehicle camera using data integrity verification and Lucas-Kanade sparse optical flow motion estimation mechanism, and output the effective frame set. The annotation vector generation module is used to perform the annotation vector generation task based on the effective frame set, using a pixel coordinate system and a five-dimensional annotation vector construction mechanism, and output a five-dimensional annotation vector. The annotation decision module is used to perform the annotation generation decision task based on the five-dimensional annotation vector, using a category-adaptive occlusion area ratio and key point joint decision mechanism, and output the annotation generation decision result. The bounding box generation and filtering module is used to generate a judgment result based on the annotation, and to perform the bounding box generation and filtering tasks by using scale-adaptive white space, directional gradient histogram and local binary pattern discernibility verification mechanism, and output the final two-dimensional bounding box. The export log module is used to export the logs in a standardized format and verify the number of files based on the final two-dimensional bounding box, and output the tag file and delivery log.

[0018] The present invention also provides a two-dimensional target annotation device for multi-scene autonomous driving image data. The two-dimensional target annotation device for multi-scene autonomous driving image data includes: a memory, a processor, and a two-dimensional target annotation program for multi-scene autonomous driving image data stored in the memory and executable on the processor. When the two-dimensional target annotation program for multi-scene autonomous driving image data is executed by the processor, it implements the above-described method.

[0019] The present invention also provides a computer program product, the computer program product including a two-dimensional target annotation program for multi-scene autonomous driving image data, the two-dimensional target annotation program for multi-scene autonomous driving image data implementing the above method when executed by a processor.

[0020] The beneficial effects of this invention are as follows: This invention constructs an effective frame screening step that includes data integrity verification and sparse optical flow motion estimation. Invalid frames such as lens flare and long periods of stillness are removed in advance, so that subsequent annotation processing is only performed on image frames with substantial perceptual value, reducing mislabeling and wasted computing power caused by low-quality frames.

[0021] This invention employs a category-adaptive occlusion area ratio and key point joint judgment mechanism to set differentiated occlusion ratio thresholds and truncation ratio thresholds for vehicle, pedestrian, and trailer categories, respectively. It also uses the visibility of key parts such as wheels, heads, and gooseneck joints as veto conditions. Without increasing human intervention, the annotation decision is transformed from qualitative experience into a reproducible mathematical judgment, making the judgment results between different batches and different annotators tend to be consistent. It is particularly suitable for operation scenarios with dense occlusion of multiple types of targets, such as mines and ports. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating the first embodiment of a two-dimensional target annotation method for multi-scenario autonomous driving image data according to the present invention. Detailed Implementation

[0023] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0024] Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Example 1: As Figure 1 The diagram shown is a flowchart of the first embodiment of a two-dimensional target annotation method for multi-scenario autonomous driving image data according to the present invention. The first embodiment of the two-dimensional target annotation method for multi-scenario autonomous driving image data according to the present invention is presented.

[0026] In the first embodiment, the two-dimensional target annotation method for multi-scenario autonomous driving image data includes: Step S10: Obtain the image sequence from the vehicle-mounted camera. Based on the image sequence from the vehicle-mounted camera, perform a valid frame filtering task using data integrity verification and Lucas-Kanade sparse optical flow motion estimation mechanism, and output a set of valid frames. The camera image sequences in this step are sourced from the front-view, side-view, and surround-view cameras mounted on autonomous mining trucks, port container trucks, or test vehicles. The raw bitstream is decoded into a continuous frame set, which serves as input for data integrity verification and optical flow motion estimation mechanisms. Data integrity verification reads the magic number, data block integrity markers, and checksums from the image file header frame by frame. This metadata may be corrupted during storage or writing due to vibration, power outages, or transmission errors; verification identifies and removes damaged frames. For frames that pass verification, the distortion correction completion flag in the metadata is further queried; frames that have not completed intrinsic parameter correction are discarded. Then, the image is converted to the HSV color space. Regions with luminance components exceeding a preset glare threshold and a connected area exceeding 8% of the total image area are detected to identify local overexposure caused by direct headlights from oncoming vehicles or strong surface reflections. Simultaneously, local contrast analysis is performed on the image. The ratio of the area with contrast below a preset haze threshold to the total area is used as a measure of the degree of haze obstruction. When the area below the threshold exceeds 20% of the total area, the frame is determined to be an irrecoverable obstruction or glare frame and is excluded. After the above screening, the retained frames are judged as stationary frames: the ORB feature points extracted from the image are tracked using the Lucas-Kanade sparse optical flow method, and the average displacement amplitude of all feature points is calculated. At the same time, the speed amplitude is obtained through the vehicle's CAN bus odometer or GPS / IMU. When the average displacement amplitude of three consecutive frames is less than 0.5 pixels per frame and the vehicle speed is less than 0.1 meters per second, it is considered a redundant frame caused by a stationary scene and is discarded. The final output set of valid frames contains only image frames with complete data, corrected intrinsic parameters, no large-area glare or haze, and effective scene motion information, providing a clean and interpretable input source for the annotation vector generation in the subsequent step S20.

[0027] After this step, invalid frames generated by hardware failures, ambient light interference, or static scenes in the original camera image sequence are systematically removed, leaving each frame with a resolvable data structure and sufficient visual information. Glare and haze detection in the HSV color space prevents annotators from working on frames where target outlines are completely indistinguishable, while the joint static determination based on optical flow and vehicle motion accurately distinguishes between natural stillness such as queuing and parking and camera freeze failures, ensuring that only frames with annotation value are delivered. The formation of the effective frame set significantly reduces the number of frames required for subsequent annotation vector construction and significantly improves content effectiveness, thus improving the utilization efficiency of annotation resources. In particular, when these clear and effective frames enter the category-adaptive occlusion and keypoint joint determination step S30, the calculation of the occlusion area ratio and truncation ratio will be based on real pixel information unaffected by glare contamination and haze attenuation. This avoids misjudging normal targets as occluded or incorrectly evaluating real occluded areas due to image quality degradation, thus providing a reliable basis for quantification.

[0028] Traditional image annotation preprocessing typically only checks whether a file can be opened or extracts frames at fixed time intervals, neglecting to address specific data degradation caused by factors such as mine dust, port vehicle headlight glare, and vehicle queues. In mining operations, storage corruption due to heavy vehicle vibration, sudden drops in image contrast caused by dust, and duplicate frames resulting from prolonged parking are common occurrences. If these low-quality frames are mixed into the annotation process, annotators often cannot accurately judge the degree of occlusion when faced with glare-covered vehicle outlines or fog-masked target edges, relying solely on subjective judgment. This further amplifies the already existing qualitative occlusion judgment error, resulting in poorer annotation consistency between different personnel and batches. This step, at the source of the process, identifies and removes these systematically defective frames through integrity checks, glare / fog area ratio thresholds, and optical flow checks coupled with vehicle motion. It eliminates the inducement for occlusion truncation misjudgment caused by unreadable images or severe information loss at the data input level, ensuring that subsequent steps always operate on valid visual information and providing a fair comparison environment for quantitative occlusion judgment.

[0029] For example, during continuous image acquisition at a mine transport lane, direct high beams from oncoming mining trucks caused large areas of overexposure in multiple frames. Simultaneously, dust kicked up by passing vehicles resulted in extremely low local contrast in some frames. Without filtering, annotators would find the truck's outline or wheels obscured by light spots or dust, making it impossible to define visible boundaries. They would only be able to roughly estimate the occlusion ratio or even abandon annotation altogether. This step addresses this by reading the file header verification information to exclude frames with bad storage blocks due to vibration. Next, the image is converted to HSV space, and areas exceeding the glare threshold and having a connected area greater than 8% of the image are identified. Frames where the truck's front area is saturated due to direct headlights are detected and removed. Contrast analysis of dust-affected frames reveals that areas with contrast below the haze threshold account for more than 20% of the total area; these frames are also removed. Subsequently, sparse optical flow is used to track ORB feature points, combined with GPS / IMU velocity, to remove duplicate frames where the mining truck remains stationary for extended periods in front of the loading point. The final dataset has a reasonable volume, and each frame of the image is clearly distinguishable. Based on this, the annotators can accurately delineate the visible outline and key parts of the occluded mining card, laying the foundation for image quality for subsequent quantitative occlusion determination.

[0030] Step S20: Based on the set of valid frames, a pixel coordinate system and a five-dimensional annotation vector construction mechanism are used to perform the annotation vector generation task and output the five-dimensional annotation vector; This step, based on the valid frame set output in step S10, constructs a unified five-dimensional annotation vector for each target to be labeled in each frame in a pixel coordinate system. The pixel coordinate system has the top-left corner of the image as the origin, with the x-axis increasing horizontally to the right and the y-axis increasing vertically downwards, standardizing all spatial metrics into discrete pixel values. For targets that are only partially visible due to occlusion by other objects or truncation at the image edge, this step does not directly use the boundary of the visible contour as the values ​​of w and h. Instead, it obtains the visible contour of the target in the current frame and infers the complete spatial range of the target by retrieving the complete projection when it was not occluded in historical adjacent frames or by completing it based on a prior morphological skeleton model of similar targets. Then, the inferred complete width and height are assigned to the vector. This output vector carries category information and the completed true scale and is directly passed to step S30, providing accurate reference geometric primitives for calculating the occlusion area ratio and truncation ratio.

[0031] By constructing a unified pixel coordinate system and structured five-dimensional annotation vectors, this step forces the target representations from images from different cameras and resolutions to align to the same metric space, eliminating coordinate ambiguity caused by diverse image sources. The bounding box uses a representation of center point coordinates plus width and height, which naturally corresponds to the anchor point regression mechanism of the target detection algorithm, eliminating the need for repeated corner point format conversion in subsequent processing. More importantly, for occluded or truncated targets, this step completes the outline by combining historical frame projection or morphological models, ensuring that the final written w and h values ​​reflect the true spatial size of the target in the unoccluded state, rather than the incomplete scale of the currently visible portion.

[0032] In traditional manual annotation, annotators typically draw bounding boxes based solely on the visual boundaries of the visible portion of the target in the current frame, without considering the target's full size under unobstructed conditions. When the degree of occlusion of the same vehicle changes across consecutive frames, the width and height of its annotation box may fluctuate several times, sometimes only selecting a small area revealing a corner of the headlights, and sometimes encompassing a larger area, lacking a stable scale reference. This behavior leads to occlusion judgments being entirely based on the subjective visible area, with drastic jumps in the size features of the annotation vector over time, failing to provide consistent geometric constraints for depth estimation and target tracking. This scale inconsistency problem is particularly pronounced for targets with extremely large depth spans, such as large mining trucks in mines and port trailers. This step introduces a joint inference of historical frame projection and morphological models to restore the reference full bounding box of the occluded target, ensuring that the values ​​of w and h maintain a proportional relationship with the target structure under occlusion conditions. Thus, when comparing the visible area and the full area in step S30, the occlusion and truncation ratio can be calculated based on a unified and physically accurate reference standard, significantly reducing the uncertainty introduced by scale drift.

[0033] For example, in a scenario where container trailers are crisscrossing at a port, a trailer consists of a tractor unit and a three-axle trailer, with a total length of approximately 860 pixels and a width of approximately 320 pixels. When the side of this trailer is severely obscured by another trailer in the adjacent lane, only a 200-pixel segment of the trailer's rear is visible. Traditional annotation methods might simply define a narrow rectangle based on this 200-pixel segment, resulting in severe distortion of the target's w and h values. Applying this step creates a five-dimensional annotation vector for this target. Set it to 'trailer'. By querying the historical complete bounding box records of this trailer in earlier unoccluded frames and combining them with the prior aspect ratio morphology model of the trailer category, it is inferred that its complete width and height in the current frame, if unoccluded, should be approximately 320×860 pixels. The final vector x,y,w,h represent the center position and inferred size of this complete bounding box.

[0034] Step S30: Based on the five-dimensional annotation vector, the annotation generation and judgment task is performed using a category-adaptive occlusion area ratio and key point joint judgment mechanism, and the annotation generation and judgment result is output. After adopting this category-adaptive joint decision-making mechanism, the annotation decision-making process changes from fuzzy qualitative judgment based on personal experience to deterministic calculation based on occlusion area ratio, truncation ratio, and visibility of key parts. The system no longer uses a single threshold for all targets, but instead... The system automatically retrieves matching occlusion and truncation thresholds and initiates visibility detection for specific key areas within that category. Traditional annotation standards often use qualitative terminology for occlusion and truncation determinations, such as "do not annotate if mostly occluded" or "do not annotate if severely truncated," lacking clear percentage figures and visibility conditions for key areas. When mining trucks are densely packed in queues, a truck in front may have more than 60% of its body obscured by a truck behind, but its front or rear wheels may still be fully visible. Based on qualitative rules, different annotators may make diametrically opposed annotation decisions; some may discard it because it "mostly obscures," while others may retain it because "the wheels are visible." Similarly, when port trailers are largely obscured, annotators cannot determine whether the gooseneck joint or bottom wheels are visible, relying solely on overall appearance impressions. This can easily lead to missing samples with remaining key areas or mistakenly including targets that have completely lost their key areas in the dataset. This step sets quantifiable thresholds for occlusion and truncation ratios for different categories, and uses the visibility of key parts such as wheels, heads, and gooseneck joints as a veto condition. It replaces subjective judgment with reproducible mathematical judgment, ensuring that the labeling decision for the same target remains consistent regardless of which operator or batch it is executed by. This maintains data utilization while reducing the mislabeling and omission rates caused by subjective bias.

[0035] Step S40: Based on the annotation generation judgment result, the bounding box generation and filtering task is performed using the scale-adaptive white space and directional gradient histogram and local binary pattern discriminability verification mechanism, and the final two-dimensional bounding box is output. The scale-adaptive whitespace processing in this step maintains a reasonable background spacing between the bounding box edge and the target outline. This avoids the box cutting off the target's edge outline features while also preventing the inclusion of excessively large areas of useless background, thus achieving a balance between preserving target integrity and controlling noise. For extremely small targets, HOG-LBP dual discriminability verification is initiated, using edge direction distribution and local texture patterns to semantically filter candidate regions. If a small region's similarity to any template of the same type is below a threshold, it indicates that the region's outline direction and texture brightness pattern do not possess the typical characteristics of the category, and it is likely a false positive region caused by environmental clutter. Through this filtering step, seemingly similar but actually false targets caused by shadows, water stains, gravel, etc., in occluded scenes are eliminated. The final retained extended bounding box not only passes the occlusion quantification judgment but also the texture discriminability verification, ensuring that the area covered by the annotation box delivered to step S50 truly contains the semantically discriminative target appearance, which helps improve the internal purity of the annotation data.

[0036] In traditional annotation workflows, annotators typically select targets by closely fitting the bounding box along the visible contour, without intentionally leaving white space or verifying the texture features of the selected content. In far-field scenes with dust and water mist, such as mines and ports, small bright spots or dark patches often appear on the image, such as reflections from road surface water stains, edges of gravel piles, and dust clumps. These areas may resemble pedestrians or small vehicle components in shape. Due to a lack of texture analysis tools, annotators sometimes mislabel these as pedestrians or vehicles, resulting in false positives mixed into the dataset. During model training, these clutters are fitted to specific categories, affecting detection accuracy. This step adds an embedded discriminability check at the end of the traditional workflow through adaptive white space and HOG-LBP joint filtering. It can identify and reject target candidate areas whose contour features and texture patterns do not match. With almost no change to the annotators' operating habits, it reduces mislabeling caused by environmental artifacts and ensures the accuracy of the final label file.

[0037] Step S50: Export the data in a standardized format and verify the number of files based on the final two-dimensional bounding box, and output the tag file and delivery log.

[0038] File quantity consistency verification provides an automated quality checkpoint at the end of the delivery process: if individual tag files are missing due to intermediate processing anomalies or write failures, the verification log can clearly report the quantity mismatch on a scenario-by-scenario basis, allowing data managers to quickly locate the problem without relying on additional scripts for manual troubleshooting. The labeled instance count table summarized by category intuitively displays the sample distribution of each target type, helping algorithm engineers assess whether the sample size of each category is balanced, and promptly identify situations where the number of instances in a certain category is too low due to dense occlusion or excessively small size, providing a reference for subsequent model training strategy adjustments.

[0039] In traditional image annotation delivery, label files are typically manually compiled, packaged, and exported by annotators. This often results in discrepancies between the number of image and label files, missing annotations for some images without generated blank label files, and inconsistent category naming formats. This issue is particularly pronounced when multiple annotation teams work in parallel: different personnel have varying handling habits for frames without target data—some skip them without creating files, while others create empty files arbitrarily. This causes the training framework to mistakenly interpret the frames as missing data and report errors or skip those frames. Furthermore, deliverables lack structured quality reports, requiring recipients to write their own checks to verify data integrity. This step, by enforcing standardized directory structures, unifying lowercase category naming, automatically generating blank label files, and embedding file quantity consistency checks and category count summaries, ensures that the delivered annotation data meets automated quality control standards in both form and content completeness. This effectively reduces integration errors caused by non-standard manual delivery and lowers the initial processing costs for subsequent data users.

[0040] For example, in a port multi-camera joint annotation project, there are two scenes, daytime and nighttime, with a total of 8 camera positions. After processing in steps S10 to S40, the final 2D bounding boxes are distributed across the processing pipelines of each camera position. After this step starts, it iterates through all image files under scene / camera / images, generating a corresponding label file for each image with annotated targets, for example... Although the image 00342.jpg in the directory was selected and retained, all candidate targets on it were filtered out due to severe occlusion or unidentifiable textures, resulting in the creation of an empty file 00342.txt in the labels directory. Subsequent scans for quantity consistency verification revealed... The number of tag files was 3 fewer than the number of image files, and this was immediately flagged as an anomaly in the log. Simultaneously, the daytime scene counts were summarized as 12,865 truck instances, 3,409 pedestrian instances, and 7,822 trailer instances; the corresponding counts for nighttime scenes were 4,021, 968, and 3,120 respectively. The log included confirmation that "the number of image files and tag files for daytime scenes are both 25,000, and the quantities match," and provided a category count table. Based on this, the receiving party's quality control personnel quickly identified the issue. The issue of missing data was identified, and it was confirmed that the sample size for the trailer category was significantly lower at night. Therefore, it was decided to supplement the collection of nighttime trailer operation data to balance the distribution of the training set.

[0041] Example 2: Furthermore, the present invention provides a two-dimensional target annotation system for multi-scenario autonomous driving image data, employing a two-dimensional target annotation method for multi-scenario autonomous driving image data from the above embodiments, which can solve the technical problem of two-dimensional target annotation for multi-scenario autonomous driving image data. The beneficial effects of the two-dimensional target annotation system for multi-scenario autonomous driving image data provided by the present invention are the same as those of the two-dimensional target annotation method for multi-scenario autonomous driving image data provided in the above embodiments, and other technical features of the two-dimensional target annotation system for multi-scenario autonomous driving image data are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0042] Example 3: This invention provides a two-dimensional target annotation device for multi-scene autonomous driving image data. The device includes at least one processor and a memory communicatively connected to the processor. The memory stores instructions executable by the processor, which are then executed to enable the processor to perform the two-dimensional target annotation method for multi-scene autonomous driving image data described in Example 1. The two-dimensional target annotation device for multi-scene autonomous driving image data in this invention may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. This two-dimensional target annotation device for multi-scene autonomous driving image data is merely an example and should not limit the functionality or scope of the invention. The device may also include a processing unit (e.g., a central processing unit, a graphics processing unit), which can perform various appropriate actions and processes based on a program stored in a read-only memory or a program loaded from a storage device into a random access memory. The random access memory also stores various programs and data required for the operation of a two-dimensional target annotation device for multi-scene autonomous driving image data. The processing unit, read-only memory, and random access memory are interconnected via a bus. The I / O interface is also connected to the bus. Typically, the following systems can be connected to the I / O interface: input devices including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. The communication device allows the two-dimensional target annotation device for multi-scene autonomous driving image data to communicate wirelessly or wiredly with other devices to exchange data. While a two-dimensional target annotation device for multi-scene autonomous driving image data with various systems has been described, it should be understood that it is not required to implement or possess all the systems described. Alternatively, more or fewer systems can be implemented.

[0043] Example 4: This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method for two-dimensional target annotation of multi-scene autonomous driving image data. The computer program product provided by this invention can solve the technical problem of two-dimensional target annotation of multi-scene autonomous driving image data. Compared with the prior art, the beneficial effects of the computer program product provided by this invention are the same as those of the two-dimensional target annotation method for multi-scene autonomous driving image data provided in the above embodiments, and will not be repeated here.

[0044] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a read-only memory. When the computer program is executed by a processing device, it performs the functions defined in the methods of the embodiments disclosed in this invention.

[0045] It should be understood that the various parts disclosed in this invention can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0046] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the present invention and its equivalents, the present invention also intends to include these modifications and variations.

Claims

1. A two-dimensional target annotation method for multi-scenario autonomous driving image data, characterized in that, The methods include: Step S10: Obtain the image sequence from the vehicle-mounted camera. Based on the image sequence from the vehicle-mounted camera, perform a valid frame filtering task using data integrity verification and Lucas-Kanade sparse optical flow motion estimation mechanism, and output a set of valid frames. Step S20: Based on the set of valid frames, a pixel coordinate system and a five-dimensional annotation vector construction mechanism are used to perform the annotation vector generation task, and the five-dimensional annotation vector is output. Step S20, specifically, includes the following steps: Step S201: Establish a pixel coordinate system for each valid frame in the set of valid frames, with the upper left corner of the image as the origin of the pixel coordinate system, the horizontal direction to the right as the positive x-axis, and the vertical direction downward as the positive y-axis, and determine the visible area boundary of the target to be labeled under the pixel coordinate system. Step S202: Represent the annotation information of the target to be annotated as a five-dimensional annotation vector V, wherein the five-dimensional annotation vector V satisfies: Where label_type is the category identifier string of the target to be labeled, x and y are the horizontal and vertical coordinates of the center point of the two-dimensional bounding box in the pixel coordinate system, and w and h are the pixel width and pixel height of the two-dimensional bounding box, respectively. Step S30: Based on the five-dimensional annotation vector, a joint determination mechanism of category adaptive occlusion area ratio and keypoints is used to perform the annotation generation determination task, and the annotation generation determination result is output. Step S30, specifically, includes the following steps: Step S301: Determine the visible contour mask of the target to be annotated in the current frame based on the five-dimensional annotation vector. and complete contour mask Calculate the shading area ratio according to the following formula. : in, This represents the percentage of the occlusion area of ​​the target to be labeled. The pixel position in the mask. pixel position The value in the visible contour mask pixel position The value in the complete contour mask Indicates the area of ​​the visible region. Indicates the area of ​​the complete target region; Step S302: Determine the intersection of the complete bounding box of the target to be labeled with the image boundary. When the target is only truncated by the image boundary in the horizontal direction, calculate the truncation ratio C using the width truncation ratio method. When the target is truncated only along the vertical direction by the image boundary, the truncation ratio C is calculated using the height truncation ratio method; When the target is truncated by the image boundary along both the horizontal and vertical directions, the truncation ratio C is calculated using the area truncation ratio method. Step S303: Set the occlusion ratio threshold based on the category identifier string of the target to be labeled. and cutoff ratio threshold Specifically, the occlusion threshold for vehicles is 0.60 and the truncation threshold is 0.50; the occlusion threshold for pedestrians is 0.50 and the truncation threshold is 0.40; and the occlusion threshold for trailers is... The threshold value is 0.70 and the cutoff ratio is 0.

60. If the proportion of the occluded area Greater than the occlusion ratio threshold Or the cutoff ratio C is greater than the cutoff ratio threshold. If the condition is met, it is determined that no annotation will be generated; otherwise, it is determined that an annotation will be generated, and the result of the determination of whether or not an annotation is generated will be used as the annotation generation determination result. Step S40: Based on the annotation generation judgment result, the bounding box generation and filtering task is performed using the scale-adaptive white space and directional gradient histogram and local binary pattern discriminability verification mechanism, and the final two-dimensional bounding box is output. Step S50: Export the data in a standardized format and verify the number of files based on the final two-dimensional bounding box, and output the tag file and delivery log.

2. The two-dimensional target annotation method for multi-scenario autonomous driving image data as described in claim 1, characterized in that, Step S10 involves acquiring an image sequence from the vehicle-mounted camera, performing a valid frame filtering task based on the image sequence using data integrity verification and a Lucas-Kanade sparse optical flow motion estimation mechanism, and outputting a set of valid frames. Specifically, this includes: Step S101: Perform data integrity verification frame by frame on the image sequence of the vehicle camera, read the magic number, image encoding format, resolution field and data block integrity flag in the header of each image file, and mark the corresponding image frame as invalid and discard it when any image file cannot be parsed, verification fails or the resolution field is inconsistent with the preset camera acquisition specifications. Step S102: For image frames that have passed data integrity verification, read their timestamp field and query the distortion correction completion flag in the metadata tag of the corresponding image frame; when the distortion correction completion flag indicates that the intrinsic parameter correction has not been completed, remove the corresponding image frame; when the distortion correction completion flag indicates that the intrinsic parameter correction has been completed, convert the corresponding image frame from the original color space to the HSV color space, and calculate the area of ​​overly bright and overly dark regions based on the luminance component V in the HSV color space; wherein, the connected regions of pixels with luminance component V greater than a preset overly bright threshold are determined as overly bright regions, and the connected regions of pixels with luminance component V less than a preset overly dark threshold are determined as overly dark regions, the area of ​​the overly bright region is the sum of the number of pixels contained in all overly bright regions, and the area of ​​the overly dark region is the sum of the number of pixels contained in all overly dark regions; if the proportion of the overly bright region area to the total area of ​​the image frame exceeds 80%, or the proportion of the overly dark region area to the total area of ​​the image frame exceeds 20%, then the image frame is determined to be an unrecoverable occluded frame or a strong light frame and is removed. Step S103: Perform still frame determination on the retained image frames, use the ORB feature point detection algorithm to extract feature points in adjacent image frames, and use the Lucas-Kanade sparse optical flow method to track the inter-frame displacement of the feature points, and calculate the average displacement amplitude of all successfully tracked feature points; at the same time, read the vehicle speed, and when the average displacement amplitude is less than 0.5 pixels per frame and the vehicle speed is less than 0.1 meters per second in three consecutive frames, mark the three consecutive frames as still invalid frames and discard them, and form the set of valid frames with the image frames that are not discarded; wherein, the average displacement amplitude is used to characterize the degree of change in the field of view between adjacent image frames that can be used for target annotation.

3. The two-dimensional target annotation method for multi-scenario autonomous driving image data as described in claim 1, characterized in that, In step S202, the category identifier string label_type of the target to be labeled includes vehicle class, truck class, pedestrian class, trailer class and construction machinery class; wherein, the same category identifier string is consistent in the set of valid frames, the five-dimensional annotation vector, the annotation generation judgment result, the final two-dimensional bounding box and the label file.

4. The two-dimensional target annotation method for multi-scenario autonomous driving image data as described in claim 1, characterized in that, In step S40, based on the annotation generation judgment result, the bounding box generation and filtering task is performed using a scale-adaptive whitespace and directional gradient histogram and local binary pattern discriminability verification mechanism to output the final two-dimensional bounding box. This step specifically includes: Step S401: For the target whose annotation was generated in step S30, obtain its visible area A, and calculate the adaptive whitespace spacing g according to the following formula: in, For the target to be labeled in the valid frame Image location in This is an image boundary constraint function; when the expanded boundary does not exceed the image boundary, it follows... Expand the bounding box on all four sides; when the expanded boundary exceeds the image boundary, keep the edge that exceeds the image boundary at the image boundary, and do not expand the bounding box on the truncated edge. Step S402: For targets with a visible area A less than 225 pixels squared, extract the histogram of directional gradients and local binary pattern texture features of the visible region; wherein, the histogram of directional gradients is used to characterize the distribution of the target edge direction, and the local binary pattern texture features are used to characterize the local grayscale texture structure of the target. Step S403: Calculate the similarity between the histogram of oriented gradients features and the histogram of oriented gradients reference features of untruncated targets of the same category to obtain the histogram of oriented gradients similarity. The local binary pattern texture features are compared with the local binary pattern reference features of the same category of untruncated targets to obtain the local binary pattern similarity. If the similarity of the directional gradient histogram is less than 0.35 and the similarity of the local binary pattern is less than 0.40, the corresponding target will be filtered as an unlabeled target; otherwise, the expanded bounding box will be retained as the final two-dimensional bounding box.

5. The two-dimensional target annotation method for multi-scenario autonomous driving image data as described in claim 1, characterized in that, In step S50, the delivery log includes the total number of scenarios, the number of image files in each scenario, the number of tag files, the consistency verification results of the number of image files and tag files, and a summary table of the number of labeled instances for each category.

6. A two-dimensional target annotation system for multi-scene autonomous driving image data, applied to the two-dimensional target annotation method for multi-scene autonomous driving image data according to any one of claims 1 to 5, characterized in that, The two-dimensional target annotation system includes: The effective frame filtering module is used to acquire the image sequence of the vehicle camera, and perform the effective frame filtering task based on the image sequence of the vehicle camera using data integrity verification and Lucas-Kanade sparse optical flow motion estimation mechanism, and output the effective frame set. The annotation vector generation module is used to perform the annotation vector generation task based on the effective frame set, using a pixel coordinate system and a five-dimensional annotation vector construction mechanism, and output a five-dimensional annotation vector. The annotation decision module is used to perform the annotation generation decision task based on the five-dimensional annotation vector, using a category-adaptive occlusion area ratio and key point joint decision mechanism, and output the annotation generation decision result. The bounding box generation and filtering module is used to generate a judgment result based on the annotation, and to perform the bounding box generation and filtering tasks by using scale-adaptive white space, directional gradient histogram and local binary pattern discernibility verification mechanism, and output the final two-dimensional bounding box. The export log module is used to export the logs in a standardized format and verify the number of files based on the final two-dimensional bounding box, and output the tag file and delivery log.

7. A two-dimensional target annotation device for multi-scenario autonomous driving image data, characterized in that, The two-dimensional target annotation device for multi-scenario autonomous driving image data includes: a memory, a processor, and a two-dimensional target annotation program for multi-scenario autonomous driving image data stored in the memory and executable on the processor. When the two-dimensional target annotation program for multi-scenario autonomous driving image data is executed by the processor, it implements a two-dimensional target annotation method for multi-scenario autonomous driving image data according to any one of claims 1 to 5.

8. A computer program product, characterized in that, The computer program product includes a two-dimensional target annotation program for multi-scenario autonomous driving image data. When the two-dimensional target annotation program for multi-scenario autonomous driving image data is executed by the processor, it implements a two-dimensional target annotation method for multi-scenario autonomous driving image data according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Urban traffic perception data generation and analysis method based on virtual simulation

    CN121725313A

  • Parking detection method and device based on monitoring video

    WO2018130016A1