Automatic labeling method and system for small target detection data, storage medium and electronic equipment
By using the automatic labeling method of wide and narrow field of view camera image registration and open vocabulary object detection model in small object detection data, the problems of missing and inaccurate small object labeling caused by manual labeling are solved, and high-precision small object detection data labeling and detection performance improvements are achieved.
Patent Information
- Application Number
- CN202510279404.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The existing small target annotation methods rely on manual annotation, resulting in missing or inaccurate small target annotation in the dataset, affecting the performance of small target detectors.
The automatic labeling method of small object detection data based on wide and narrow field of view camera image registration and open vocabulary object detection model is adopted. The narrow field of view image is detected through the object detection model, image registration is performed, the position of the wide field of view small object bounding box is adjusted, and the object detection model is finally trained.
Improve the labeling accuracy of small object detection data, reduce labeling missing, and improve the performance of small object detectors.
Smart Images

Figure CN120219818A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to an automatic annotation method, system, storage medium and electronic device for small target detection data. Background Art
[0002] Object detection is an important task in computer vision, aiming to identify and locate objects in images or videos. With the rapid development of deep learning technology, object detection has made remarkable progress and is widely used in multiple fields such as monitoring, autonomous driving, drone analysis, and robot vision. In the high-altitude security monitoring scenario, target detection is usually directly carried out by a wide-field-of-view camera. Due to the high installation height of the camera, the ground resolution of the wide-field-of-view camera image is low, so the size of the targets in the image is extremely small, and accurately locating and annotating small targets remains a challenge.
[0003] The existing means for small target annotation still adopt manual annotation, and the data sets obtained by manual annotation often have problems such as missing small target annotation or inaccurate annotation, which directly affects the performance of small target detectors.
[0004] In terms of large models for open-vocabulary object detection, a data intelligent annotation method and device are disclosed in the patent with the publication number CN118298250A. In this solution, an open-vocabulary object detection large model is used to annotate targets, but this method has better annotation effects only for large targets in images with high ground resolution, and it is difficult to annotate small targets in images with low ground resolution.
[0005] Therefore, it is necessary to provide an annotation scheme for small target detection data to improve problems such as missing small target annotation or inaccurate annotation existing in the existing small target detection data annotation scheme. Summary of the Invention
[0006] The main purpose of the present invention is to provide an automatic annotation method, system, storage medium and electronic device for small target detection data based on the image registration of narrow and wide field-of-view cameras and an open-vocabulary object detection model, so as to improve the annotation accuracy of small target detection data and reduce missing annotations.
[0007] To achieve the foregoing invention purpose, the technical solutions adopted by the present invention include: An automatic annotation method for small target detection data, comprising:
[0008] S1, using an object detection model to perform object detection on the collected narrow field-of-view camera image to obtain a large target detection bounding box of the narrow field-of-view image;
[0009] S2. Perform image registration on the wide and narrow field-of-view camera images to obtain the homography matrix for transforming the narrow field-of-view image to the wide field-of-view image, the registered narrow field-of-view camera image, and the bounding box of the small target in the wide field-of-view image before adjustment;
[0010] S3. Adjust the position of the bounding box of the small target in the wide field-of-view image to obtain an accurately labeled bounding box of the small target;
[0011] S4. Use the accurately labeled bounding box of the small target to train the target detection model to obtain a trained small target detection model.
[0012] In a preferred embodiment, in S1, the target detection model is an open-vocabulary target detection model, and / or, S1 includes:
[0013] S11. Extract word vectors from the vocabulary using the BERT model;
[0014] S12. Input the word vectors and the narrow field-of-view camera image into the target detection model for feature fusion;
[0015] S13. The target detection model infers to obtain the bounding box of the large target in the narrow field-of-view image.
[0016] In a preferred embodiment, in S2, use an image registration method to perform image registration on the wide and narrow field-of-view camera images, and / or, S2 includes:
[0017] S21. Extract feature points of the wide and narrow field-of-view camera images;
[0018] S22. Match the feature points to obtain multiple matching pairs;
[0019] S23. Eliminate mis-matched points and calculate the homography matrix for transforming the narrow field-of-view image to the wide field-of-view image;
[0020] S24. Use the homography matrix to transform the narrow field-of-view camera image onto the wide field-of-view camera image to obtain the registered narrow field-of-view camera image;
[0021] S25. Use the homography matrix to transform the bounding box of the large target in the narrow field-of-view image onto the wide field-of-view camera image to obtain the bounding box of the small target in the wide field-of-view image.
[0022] In a preferred embodiment, in S21, use SuperPoint to extract the feature points of the narrow field-of-view camera image, and / or, in S22, use LightGlue to match the feature points, and / or, in S23, use the random sample consensus method to eliminate mis-matched points.
[0023] In a preferred embodiment, in step S3, using the target tracking method, with the registered narrow-field camera image, wide-field camera image, and the wide-field small target bounding box obtained through image registration in step S2 as inputs, by iteratively adjusting the position of the wide-field small target bounding box, an accurately labeled small target bounding box is generated, and / or step S3 includes:
[0024] S31, respectively extract features from the registered narrow-field camera image and the wide-field small target bounding box;
[0025] S32, according to the kernel correlation filtering method in the target tracking method, calculate the response map of the extracted features;
[0026] S33, take the part with the largest value in the response map as the position adjustment result of the wide-field small target bounding box.
[0027] In a preferred embodiment, in step S31, first, respectively magnify the registered narrow-field camera image and the wide-field small target bounding box by a multiple, and respectively perform HOG feature extraction, and / or in step S32, map the features extracted in step S31 to a high-dimensional space through a kernel method to train a kernel correlation filter, and use the trained kernel correlation filter to perform a fast Fourier convolution on the extracted features to obtain the response map, and / or in step S33, the part with the largest value in the response map is the center point position of the adjusted wide-field small target bounding box.
[0028] On the other hand, the technical solution adopted by the present invention includes: an automatic annotation system for small target detection data, including:
[0029] A large target detection bounding box obtaining module, configured to perform target detection on the collected narrow-field camera image using a target detection model to obtain a large target detection bounding box of the narrow-field image;
[0030] A registration module, configured to perform image registration on the narrow and wide-field camera images to obtain a homography matrix for transforming the narrow-field image to the wide-field image, the registered narrow-field camera image, and the wide-field small target bounding box before adjustment;
[0031] A bounding box position adjustment module, configured to adjust the position of the wide-field small target bounding box to obtain an accurately labeled small target bounding box;
[0032] A model training module, configured to train the target detection model using the accurately labeled small target bounding box to obtain a trained small target detection model.
[0033] In a preferred embodiment, the large target detection bounding box obtaining module performs target detection using an open vocabulary target detection model, and the registration module performs image registration using an image registration method.
[0034] On the other hand, the technical solution adopted by the present invention includes: a readable storage medium, characterized in that: a computer program is stored in the readable storage medium, and when the computer program is run, the steps in the above-mentioned automatic annotation method for small target detection data are executed.
[0035] On the other hand, the technical solution adopted by the present invention includes: an electronic device, including a memory and a processor, a computer program is stored in the memory, and when the computer program is run by the processor, the steps in the above-mentioned automatic annotation method for small target detection data are executed.
[0036] Compared with the prior art, the beneficial effects of the present invention are at least as follows:
[0037] The present invention provides an automatic annotation solution for small target detection data based on the image registration of wide and narrow field-of-view cameras and an open-vocabulary object detection large model. The large objects in the narrow field-of-view image are annotated by the open-vocabulary object detection large model, and then the large objects are inversely mapped to the small objects in the wide field-of-view through image registration, and finally the automatic annotation of small target detection data is realized. The automatic annotation solution for small target detection data provided by the present invention improves the annotation accuracy of small target detection data and reduces annotation omission. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic flow chart of an automatic annotation method for small target detection data provided by the present invention;
[0040] Figure 2 It is a schematic diagram of using an open-vocabulary object detection model to detect the narrow field-of-view camera image of the present invention;
[0041] Figure 3 It is a schematic diagram of the registration result of the wide and narrow field-of-view images and their registration bounding boxes of the present invention;
[0042] Figure 4 It is a schematic diagram of the response map during the correction process of the wide and narrow field-of-view target positions of the present invention;
[0043] Figure 5 It is a schematic diagram of the bounding boxes before and after the correction of the wide and narrow field-of-view target positions of the present invention;
[0044] Figure 6 Schematic diagram of the small target detection data finally obtained by the present invention. Detailed implementation manners
[0045] The present invention will be more fully understood by the following detailed implementation manners which should be read in conjunction with the accompanying drawings. Specific embodiments of the present invention are disclosed herein; however, it should be understood that the disclosed embodiments are merely exemplary of the present invention, and the present invention can be embodied in various forms. Therefore, the specific functional details disclosed herein should not be construed as limiting, but merely as a basis for the claims and as a representative basis for teaching those skilled in the art to employ the present invention in any appropriate detailed embodiment in different ways.
[0046] As Figure 1 shown, an automatic annotation method for small target detection data disclosed by the present invention specifically includes the following steps:
[0047] S1. Use a target detection model to perform target detection on the narrow field of view camera images collected, and obtain the large target detection bounding boxes of the narrow field of view images.
[0048] Specifically, first collect narrow field of view camera images from the real environment through a narrow field of view camera, and then use an open vocabulary target detection large model, specifically the Grounding-Dino model (i.e., an open set detection model) in the open vocabulary target detection large model, to perform target detection on the collected narrow field of view camera images, and obtain the large target detection bounding boxes of the narrow field of view images, denoted as B n , as Figure 2 shown, Figure 2 Schematic diagram of the present invention using an open vocabulary target detection model to detect narrow field of view camera images; the left side is the process of narrow field of view target detection, the upper left corner of the right side figure is the effect of the open vocabulary target detection model detecting the wide field of view image failing, and the lower left corner of the right side figure is the effect of the open vocabulary target detection model detecting the narrow field of view image successfully. It shows that the success rate of narrow field of view image target annotation is higher than that of wide field of view images. Since the ground resolution of narrow field of view camera images is high and the target size is large, the detection success rate is high and the annotation effect is good. In addition, the wide field of view camera and the narrow field of view camera are synchronously collected, and the wide field of view camera images collected by the wide field of view camera are used for subsequent registration annotation.
[0049] The process of obtaining the large target detection bounding boxes of the narrow field of view images specifically includes the following steps:
[0050] S11. Use the BERT model (i.e., Bidirectional Encoder Representations from Transformers) to extract word vectors from the vocabulary.
[0051] Specifically, the vocabulary here is the words input by the user, such as "person.car.", etc., which are text data. Word vectors are extracted from the vocabulary input by the user (such as text data like "person", "car", etc.). The BERT model is a pre-trained language model based on the Transformer architecture, which can capture the semantic information of vocabulary through bidirectional context encoding. The specific extraction process is as follows: First, the vocabulary input by the user is segmented into Tokens (such as "person" and "car"), and special tokens (such as [CLS] and [SEP]) are added; then, the Tokens are subjected to bidirectional context encoding through the multi-layer Transformer encoder of BERT to generate the semantic vectors of each Token; finally, these vectors are extracted as the semantic representation of the vocabulary.
[0052] S12, Input the word vector and the narrow field of view camera image into the object detection model for feature fusion.
[0053] Specifically, first, map the word vector and the image features to the same feature space to ensure consistent dimensions; then, through the cross-attention mechanism, use the text features as Query, the image features as Key and Value, calculate the attention weights of the text to the image, and inject the semantic information into the image features; then, fuse the enhanced image features with the original features by weighted summation or concatenation to generate a multi-modal feature representation; finally, input the fused features into the object detection model.
[0054] S13, The object detection model infers to obtain the large object detection bounding box Bn of the narrow field of view image.
[0055] Specifically, the multi-modal features after feature fusion are input into the object detection model, and the detection bounding box Bn of the large object in the narrow field of view image is inferred through the detection head. The detection head is the core component of the object detection model, usually composed of a classification branch and a regression branch: the classification branch is used to predict the category of the object (such as "person", "car"), and the regression branch is used to predict the coordinates of the bounding box of the object. Based on the image features fused with text semantic information, the detection head decodes the features through convolutional or fully connected layers and directly outputs the category probability and bounding box position information of the object.
[0056] S2, Perform image registration on the wide and narrow field of view camera images to obtain the homography matrix for transforming the narrow field of view image to the wide field of view image, the registered narrow field of view camera image, and the small object bounding box of the wide field of view before adjustment.
[0057] The wide and narrow field of view camera images here are the wide field of view camera image and the narrow field of view camera image collected above. Specifically, use the image registration method to register the wide and narrow field of view camera images to obtain the homography matrix for transforming the narrow field of view image to the wide field of view image, denoted as Hn2w , according to the homography matrix H n2w , the registered narrow field of view camera image is obtained, denoted as Ir, and the bounding box of the wide field of view small target before adjustment, denoted as B w , such as Figure 3 shown Figure 3 is a schematic diagram of the registration result of the wide and narrow field of view images and its registration bounding box of the present invention; the part with obvious color difference from the background in the figure is the registered area, and the red marked box is the registration bounding box
[0058] The process of performing image registration on the wide and narrow field of view camera images to obtain the homography matrix for transforming the narrow field of view image to the wide field of view image, the registered narrow field of view camera image, and the bounding box of the wide field of view small target before adjustment includes the following steps:
[0059] S21, extract the feature points of the wide and narrow field of view camera images
[0060] In implementation, specifically use SuperPoint (SuperPoint: Self-Supervised Interest Point Detection and Description, a self-supervised learning superpoint feature extraction algorithm) to extract the feature points of the narrow field of view camera image. SuperPoint is a feature point detector based on deep learning. It extracts significant feature points from the image through a convolutional neural network (CNN) and simultaneously generates a descriptor for each feature point. These descriptors are used to characterize the local visual information of the feature points and provide a basis for subsequent matching
[0061] S22, match the extracted feature points to obtain multiple matching pairs
[0062] In implementation, specifically use LightGlue (LightGlue: Local Feature Matching at Light Speed, a lightweight and fast local feature matching algorithm) to match the above feature points. LightGlue is based on Transformer and graph neural networks. By optimizing the cross-attention layer, introducing position encoding and matching prediction module, it dynamically adjusts the computational complexity to achieve efficient and accurate feature matching
[0063] S23, eliminate the mis-matched points and calculate the homography matrix H for transforming the narrow field of view image to the wide field of view image n2w .
[0064] During implementation, the Random Sample Consensus (RANSAC) method is specifically used to eliminate mismatched points. RANSAC randomly samples a set of matching point pairs, calculates the candidate homography matrix, and counts the number of matching points that conform to this matrix (i.e., inliers). This process is repeated, and finally, the matrix with the most inliers is selected as the optimal solution, while the matching points that do not conform to this matrix (i.e., outliers, mismatched points) are eliminated. RANSAC can effectively handle noise and outliers, improving the robustness of the matching. The homography matrix is a 3×3 matrix used to describe the perspective transformation relationship from a narrow field of view image to a wide field of view image. After eliminating the mismatched points, the remaining matching point pairs are used to solve the homography matrix through the least squares method or other optimization methods. Specifically, the homography matrix satisfies the following relationship:
[0065]
[0066] where (x n , y n ) are the points in the narrow field of view image, and (x w , y w ) are the corresponding points in the wide field of view image. By solving this system of linear equations, the parameters of the homography matrix can be obtained.
[0067] S24. Using this homography matrix H n2w , transform the narrow field of view camera image onto the wide field of view camera image to obtain the registered narrow field of view camera image I r .
[0068] S25. Using this homography matrix H n2w , transform the large target detection bounding box B n of the narrow field of view image onto the wide field of view camera image to obtain the small target bounding box B w of the wide field of view.
[0069] S3. Adjust the position of the small target bounding box B w of the wide field of view to obtain the accurately labeled small target bounding box.
[0070] Due to the frame asynchrony and registration error phenomena between the narrow and wide field of view camera images, the small target bounding box B w obtained in the above step S2 has a position error on the wide field of view camera image, so its position needs to be adjusted. Specifically, using the target tracking method, taking the above registered narrow field of view camera image I r , the wide field of view camera image, and the small target bounding box B w obtained through image registration in the above S2 as inputs, and iteratively adjusting the initial small target bounding box B wGenerate a precisely labeled small target bounding box at the position of
[0071] The process of adjusting the position of the wide - field small target bounding box B w to obtain a precisely labeled small target bounding box specifically includes the following steps:
[0072] S31, respectively extract features from the registered narrow - field camera image I r and the wide - field small target bounding box B w When implemented, first, respectively magnify the registered narrow - field camera image I
[0073] and the wide - field small target bounding box B r by a certain multiple, such as 1.5 times magnification. The main purpose of this operation is to expand the search range, so as to provide more sufficient context information for the KCF (Kernelized Correlation Filters) tracker, and avoid target loss or tracking failure caused by the initial mapping deviation. KCF itself enhances the ability to capture the target by expanding the search range. After magnifying the image and the bounding box, it can more effectively extract the HOG features in Ir and Bw, improving the robustness and accuracy of feature matching. Preferably, the magnification factor is controlled between 1.2 and 1.8 times, which can not only ensure that the search range is large enough but also avoid introducing too much irrelevant background information due to excessive magnification, affecting the tracking effect. After magnification, then extract the HOG (Histogram of Oriented Gradients) features in the registered narrow - field camera image I w and the HOG features in the wide - field small target bounding box B r w w w w
[0074] S32, according to the kernel correlation filtering method in the target tracking method, calculate the response map of the extracted features.
[0075] When implemented, use the HOG features extracted in step S31 above to train a correlation filter, and perform a fast Fourier convolution on the extracted HOG features using the trained correlation filter to obtain the response map of the extracted HOG features. The response map is as Figure 4 shown, Figure 4 is a schematic diagram of the response map in the process of correcting the target position of the wide - and narrow - field of the present invention; this response map only contains the target bounding box area, and the larger the pixel value, the more likely the corrected target is at this position. And it can be observed that Figure 4 the position of the maximum pixel value indicated by the arrow in Figure 5 corresponds to the correction direction of the bounding box in
[0076] S33. Use the part with the largest value in the response map as the wide-field small target bounding box B w 's position adjustment result to obtain an accurately labeled small target bounding box.
[0077] During implementation, the part with the largest pixel value in the above response map is the center point position of the adjusted wide-field small target bounding box B w . The adjusted wide-field small target bounding box B w As Figure 5 shown, Figure 5 is a schematic diagram of the bounding boxes of the narrow and wide-field target positions before and after correction in the present invention; the left side in the figure shows the effect before the bounding box correction, and the right side in the figure shows the effect after the bounding box correction. It can be seen that the effectiveness of the target position correction.
[0078] S4. Use the above accurately labeled small target bounding box to train the target detection model to obtain a trained small target detection model. As Figure 6 shown, Figure 6 is a schematic diagram of the small target detection data finally obtained in the present invention. The top row of images is the large target obtained by annotation, the middle row of images is the medium target obtained by annotation, and the bottom row of images is the small target obtained by annotation. This indicates that the algorithm can annotate a large amount of data.
[0079] Correspondingly, the present invention also discloses an automatic annotation system for small target detection data, including:
[0080] A large target detection bounding box acquisition module, which is used to perform target detection on the narrow-field camera images collected by using the target detection model to obtain the large target detection bounding box of the narrow-field images.
[0081] A registration module, which is used to perform image registration on the narrow and wide-field camera images to obtain the homography matrix for transforming the narrow-field image to the wide-field image, the registered narrow-field camera image, and the wide-field small target bounding box before adjustment.
[0082] A bounding box position adjustment module, which is used to adjust the position of the wide-field small target bounding box to obtain an accurately labeled small target bounding box.
[0083] A model training module, which is used to train the target detection model by using the accurately labeled small target bounding box to obtain a trained small target detection model.
[0084] Among them, the working principle of each of the above modules can be respectively referred to the description of each step in the above method, and will not be elaborated here.
[0085] On the other hand, the present invention also provides a readable storage medium, on which a computer program is stored, and when the program is run, it implements the steps in the automatic annotation method for small target detection data provided in the above embodiment.
[0086] In another aspect, the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory. When the computer program is run by the processor, it executes the steps in the above-mentioned automatic annotation method for small target detection data.
[0087] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0088] More specific examples (non-exhaustive list) of readable storage media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the readable storage media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0089] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0090] The present invention has the following advantages: The present invention provides a small target detection data automatic annotation scheme based on the image registration of wide and narrow field cameras and an open vocabulary object detection large model. The large objects in the narrow field image are annotated by the open vocabulary object detection large model, and then the large objects are inversely mapped to the small targets in the wide field through image registration, finally realizing the automatic annotation of small target detection data. The small target detection data automatic annotation scheme provided by the present invention improves the annotation accuracy of small target detection data and reduces annotation omission.
[0091] All aspects, embodiments, features and examples of the present invention should be considered illustrative in all respects and are not intended to limit the present invention, the scope of which is only defined by the claims. Without departing from the spirit and scope of the claimed invention, those skilled in the art will appreciate other embodiments, modifications and uses.
[0092] The use of headings and sections in the present invention does not imply a limitation of the present invention; each section can be applied to any aspect, embodiment or feature of the present invention.
Claims
1. A method for automatic labeling of small target detection data, characterized in that: The method comprises: S1, using the target detection model to perform target detection on the captured narrow field of view camera image to obtain a large target detection bounding box of the narrow field of view image; S2, performing image registration on the wide and narrow field of view camera images to obtain the homography matrix of the narrow field of view image transformed into the wide field of view image, the registered narrow field of view camera image, and the wide field of view small target bounding box before adjustment; S3, adjusting the position of the wide field of view small object bounding box to obtain a precisely marked small object bounding box; S4, training the target detection model using the accurately labeled small target bounding box to obtain a trained small target detection model.
2. The automatic labeling method for small target detection data according to claim 1, characterized in that: In S1, the object detection model is an open vocabulary object detection model, and / or S1 includes: S11, extract word vectors from vocabulary using BERT model; S12, inputting the word vector and the narrow field of view camera image into a target detection model for feature fusion; S13, the object detection model infers the large object detection bounding box of the narrow field of view image.
3. The automatic labeling method for small target detection data according to claim 1, characterized in that: In S2, the wide and narrow field of view camera images are registered using an image registration method, and / or S2 includes: S21, extracting feature points of the wide and narrow field of view camera image; S22, matching the feature points to obtain a plurality of matching pairs; S23, eliminating mismatched points and calculating the homography matrix for transforming the narrow field of view image into the wide field of view image; S24, using the homography matrix, transforming the narrow field of view camera image to the wide field of view camera image to obtain the registered narrow field of view camera image; S25, using the homography matrix, transforming the large object detection bounding box of the narrow field of view image to the wide field of view camera image to obtain the wide field of view small object bounding box.
4. The automatic labeling method for small target detection data according to claim 3, characterized in that: In the S21, the feature points of the narrow field of view camera image are extracted using SuperPoint, and / or, in the S22, the feature points are matched using LightGlue, and / or, in the S23, the mismatched points are eliminated using a random sampling consistency method.
5. The automatic labeling method for small target detection data according to claim 1, characterized in that: In S3, a target tracking method is used to take the aligned narrow field of view camera image, the wide field of view camera image and the wide field of view small target bounding box obtained by image registration in S2 as input, and the position of the wide field of view small target bounding box is iteratively adjusted to generate a precisely labeled small target bounding box, and / or S3 includes: S31, extracting features from the registered narrow field of view camera image and the wide field of view small target bounding box respectively; S32, calculating a response map of the extracted features according to a kernel correlation filtering method in a target tracking method; S33, taking the part with the largest value in the response graph as the position adjustment result of the wide field of view small target bounding box.
6. The automatic labeling method for small target detection data according to claim 5, characterized in that: In the S31, the aligned narrow field of view camera image and wide field of view small target bounding box are first magnified by multiples, and HOG feature extraction is performed respectively, and / or, in the S32, the features extracted in step S31 are mapped to a high-dimensional space by a kernel method to train a kernel correlation filter, and the extracted features are fast Fourier convolved using the trained kernel correlation filter to obtain the response map, and / or, in the S33, the part with the largest value in the response map is the center point position of the adjusted wide field of view small target bounding box.
7. An automatic labeling system for small target detection data, characterized in that: The system comprises: A large target detection bounding box acquisition module is used to perform target detection on the acquired narrow field of view camera image using the target detection model to obtain a large target detection bounding box of the narrow field of view image; A registration module is used to perform image registration on wide and narrow field of view camera images, obtain a homography matrix of the narrow field of view image transformed into a wide field of view image, the registered narrow field of view camera image, and a wide field of view small target bounding box before adjustment; A bounding box position adjustment module, used to adjust the position of the wide field of view small object bounding box to obtain an accurately marked small object bounding box; The model training module is used to train the target detection model using the accurately labeled small target bounding box to obtain a trained small target detection model.
8. The automatic labeling system for small target detection data according to claim 7, characterized in that: The large object detection bounding box acquisition module performs object detection using an open vocabulary object detection model, and the registration module performs image registration using an image registration method.
9. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed, the steps of the automatic labeling method for small target detection data according to any one of claims 1 to 6 are executed.
10. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the automatic labeling method for small target detection data according to any one of claims 1 to 6 are executed.
Citation Information
Patent Citations
Intelligent data labeling method and device
CN118298250A
Cited By
Image registration method and device based on multi-frame fusion
CN120833362A
Image positioning calibration method and device, electronic equipment and computer readable medium
CN121937535A