Large model-based data enhancement method, apparatus and device, and storage medium

Through the combination of Grounded-SAM detection algorithm and camera calibration parameters, data augmentation for flexible acquisition of scarce targets is achieved, and the problem of single data acquisition form in the data enhancement method is solved, and the robustness and detection capabilities of the autonomous driving model are improved.

CN120298677AActive Publication Date: 2025-07-11ZHIZI AUTOMOTIVE TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510782891.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The data acquisition form in the existing data augmentation method is single, and scarce targets cannot be flexibly acquired, resulting in insufficient robustness of the autonomous driving model in the long-tail data problem.

Method used

The Grounded-SAM detection algorithm is used to detect images, and the target frame is merged by judging the intersection ratio and loss matrix of the target frame, and the target outline is accurately pasted with the camera calibration parameters. The generalization ability of the large model is used to obtain scarce targets and expand data.

Benefits of technology

It improves the quality of data enhancement and the robustness of the autonomous driving model, enhances the ability to identify complex interactive targets, reduces background interference, and improves the accuracy and generalization performance of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298677A_ABST
    Figure CN120298677A_ABST
Patent Text Reader

Abstract

The invention discloses a data enhancement method, device and equipment based on a large model and a storage medium, relates to the technical field of automatic driving data enhancement, and can solve the problem of single data acquisition form in a data enhancement method in the prior art. The method specifically comprises the steps of determining a target category and generating corresponding prompt information; detecting and segmenting the detection image by adopting a Ground-SAM detection algorithm according to the prompt information to obtain a target contour and a corresponding target frame; when two target frames corresponding to different target categories respectively intersect in the target contour, combining the target frames of the two target frames according to a preset rule to form a new target contour; and pasting the target contour to the original image to obtain an enhanced image. According to the method, the large model is utilized to predict the to-be-enhanced target, by means of the segmentation and generalization ability of the large model, the target is accurately positioned, the data enhancement quality is improved, scarce target data is expanded, the long tail problem is solved, and the robustness and the detection ability of the automatic driving model are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving data augmentation, and in particular to a data augmentation method, device, equipment and storage medium based on a large model. Background Art

[0002] Object detection is a core technology for an autonomous driving system to perceive the surrounding environment. An autonomous driving vehicle collects environmental data through sensors (such as cameras, lidars, radars, etc.), and uses object detection technology to identify and classify objects such as pedestrians, vehicles, traffic signs, and obstacles on the road. This enables the autonomous driving system to perceive the road conditions in real time and make corresponding decisions. However, there has always been a pain point in object detection in the field of autonomous driving, that is, the long-tail data problem. Long-tail data refers to the situation where the class distribution is unbalanced, where common objects of certain classes occupy most of the samples, while rare or uncommon object samples are less. However, these few-sample data are equally important for autonomous driving. If the detection effect is not good, it may lead to major accidents. Therefore, how to handle rare classes has become an important challenge.

[0003] The paper "Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation" in CVPR 2021 proposed the "copy-paste" data augmentation idea. However, the current data augmentation methods using this idea have certain limitations and always rely on the form and morphology of the existing data set, and cannot obtain the desired data more flexibly. For example, the data augmentation method mentioned in WO2020CN113998 fuses the object detection frame of an image with other background images to achieve the purpose of data augmentation. However, since the detection frame contains not only the object but also a part of the background image, it will increase the possibility of false detection and does not achieve pure data augmentation; moreover, the form of data acquisition in this method is single and cannot flexibly obtain the desired data. Summary of the Invention

[0004] The present invention provides a data augmentation method, device, equipment and storage medium based on a large model, which can solve the problem of single form of data acquisition in the existing data augmentation methods. The technical solution is as follows:

[0005] According to the first aspect of the present invention, there is provided a data augmentation method based on a large model, the method comprising:

[0006] Determine at least one target class and generate corresponding prompt information, where the target class is the class to be augmented, including but not limited to pedestrians, transportation means, traffic signs, obstacles;

[0007] Detect and segment the detection image using the Grounded-SAM detection algorithm according to at least one piece of the prompt information, and obtain at least one target contour and the corresponding target box;

[0008] Determine whether there are two intersecting target boxes corresponding to different target categories among the at least one target contour. If so, merge the target boxes of the two according to the preset rules to form a new target contour;

[0009] Paste the at least one target contour onto the original image to obtain at least one enhanced image.

[0010] The data enhancement method based on the large model provided by the present invention first determines at least one target category and generates the corresponding prompt information. The target category is the category to be enhanced, including but not limited to pedestrians, vehicles, traffic signs, and obstacles. Then, detect and segment the detection image using the Grounded-SAM detection algorithm according to at least one piece of the prompt information to obtain at least one target contour and the corresponding target box. Next, determine whether there are two intersecting target boxes corresponding to different target categories among the at least one target contour. If so, merge the target boxes of the two according to the preset rules to form a new target contour. Finally, paste the at least one target contour onto the original image to obtain at least one enhanced image. The present invention uses the Grounded-SAM large model to predict the target to be enhanced, can efficiently extract the detection results without additional training, and accurately locates the target using the segmentation ability of the large model, effectively reducing background interference and improving the quality of data enhancement. In addition, with the generalization ability of the large model, scarce targets can be flexibly obtained from any image for data augmentation, solving the long-tail data problem and enhancing the robustness and detection ability of the autonomous driving model.

[0011] As a further scheme of the present invention: determining whether there are two intersecting target boxes corresponding to different target categories among the at least one target contour includes:

[0012] Judge by calculating the first intersection over union of the target box of the first target category and the target box of the second target category. If the first intersection over union is non-zero, it is judged that there is an intersection;

[0013] If the first intersection over union is zero, it is judged that there is no intersection;

[0014] The first intersection over union is calculated by the first formula, and the first formula includes:

[0015]

[0016] Among them, is the first intersection over union (IoU), area(P) is the area of the bounding box of the first target category, and area(T) is the area of the bounding box of the second target category

[0017] The method of the present invention calculates the first intersection over union (IoU) of the bounding box of the first target category and the bounding box of the second target category, and quickly and accurately determines whether there is an intersection between the two. If the IoU is non-zero, there is an intersection; if it is zero, there is no intersection. The method of the present invention is simple and efficient, and can improve the accuracy of object detection and analysis

[0018] As a further aspect of the present invention: the merging of the bounding boxes of the two according to the preset rules to form a new target contour includes:

[0019] Calculating a loss matrix according to the preset rules and the coordinate difference between the centers of the bounding box of the first target category and the bounding box of the second target category

[0020] Performing a minimum-cost bipartite matching based on the loss matrix to obtain two corresponding target categories

[0021] Merging the contours of the two matched target categories to form the new target contour

[0022] The method of the present invention calculates the loss matrix through the coordinate difference of the center points, and uses the minimum-cost bipartite matching to achieve precise pairing of two target categories. After merging the contours, a new target is generated, effectively improving the robustness of object detection in occlusion scenarios and enhancing the recognition ability of autonomous driving for complex interactive objects

[0023] As a further aspect of the present invention: the loss matrix is calculated by a second formula, and the second formula includes:

[0024]

[0025] where represents the loss matrix of the bounding box of the first target category and the bounding box of the second target category; (x i , y i ) and (x j , y j ) respectively represent the center point coordinates of the bounding box of the first target category and the bounding box of the second target category

[0026] The method of the present invention calculates the Euclidean distance of the center point coordinates of the bounding box of the first target category and the bounding box of the second target category, quantifies the position difference between the two, provides a precise pairing cost for bipartite matching, improves the accuracy of object association, optimizes the contour merging effect in occlusion scenarios, and enhances the robustness of multi-object tracking

[0027] As a further aspect of the present invention: The step of pasting the at least one target contour onto the original image to obtain at least one enhanced image includes:

[0028] Obtain the internal and external parameters of the camera through the Zhang-Zhengyou calibration method;

[0029] Combining the internal and external parameters of the camera, as well as the actual size and pixel size of the at least one target contour, determine the longitudinal position of the at least one target contour on the original image;

[0030] Determine the horizontal position according to the central symmetry position of the at least one target contour and the existing targets on the original image;

[0031] According to the longitudinal position and the horizontal position, paste the at least one target contour onto the original image to obtain at least one enhanced image.

[0032] The method of the present invention utilizes the relationship between the camera calibration parameters and the target contour size, accurately calculates the longitudinal position of the target in the image, and determines the horizontal position through central symmetry, ensuring that the enhanced target is consistent with the scene perspective relationship and improving the authenticity of data enhancement.

[0033] As a further aspect of the present invention: The step of determining the horizontal position according to the central symmetry position of the at least one target contour and the existing targets on the original image includes:

[0034] Establish a set X of existing targets at the position symmetric to the center point of the existing targets on the original image, and establish a set Y of target contours according to the at least one target contour;

[0035] According to the center point position of the X j th target in the set X of existing targets, select a target subset Z in the set Y of target contours that meets this position condition;

[0036] Paste the Z i target to the center point position of X j , and calculate the second intersection-over-union ratio between this pasting position and other targets;

[0037] If the second intersection-over-union ratio is zero, then paste the Z i target to the center point position of X j , and remove this position from the set X of existing targets;

[0038] If the second intersection-over-union ratio is not zero or Z iIf the vertical position of the target is not at the center point of the existing target, then according to the determined vertical position, starting from the center point of the horizontal axis of the image, with half of the target box width as the step size, gradually search for suitable positions on both sides. The suitable position is the position where the second intersection over union with other targets is zero;

[0039] If no suitable position is still found, then abandon the pasting of Z i and paste Z i+1 the target to X j position and calculate its second intersection over union again;

[0040] This process continues to loop until a position that meets the conditions is found or all target contours are traversed;

[0041] If there is no pasting target that meets the conditions at a certain X j position, then no pasting is performed at this position.

[0042] The method of the present invention ensures the reasonable integration of the newly added target contour and the original scene through symmetric position matching and dynamic adjustment strategies. Using the intersection over union verification to avoid target overlap, and the step size search mechanism to improve the position adaptability, enhances the authenticity and spatial rationality of image synthesis, and effectively improves the data augmentation quality.

[0043] As a further aspect of the present invention: The method further includes:

[0044] Performing edge smoothing processing on the enhanced image using an edge blur algorithm.

[0045] The method of the present invention can effectively eliminate the hard boundary of the synthesized target through the edge blur algorithm, make it naturally blend with the background, improve the visual realism of the enhanced image, help reduce the overfitting of the deep learning model to artificial traces, and improve the generalization performance in real scenes.

[0046] According to a second aspect of the present invention, there is provided a data augmentation device based on a large model, including: a determination module, a detection module, a processing module, and a pasting module;

[0047] The determination module is used to determine at least one target category and generate corresponding prompt information. The target category is the category to be augmented, including but not limited to pedestrians, vehicles, traffic signs, obstacles;

[0048] The detection module is used to detect and segment the detection image using the Grounded-SAM detection algorithm according to at least one of the prompt information, and obtain at least one target contour and the corresponding target box;

[0049] The processing module is configured to determine whether there are two target bounding boxes corresponding to different target categories that intersect among the at least one target contour. If so, the target bounding boxes of the two are merged according to a preset rule to form a new target contour.

[0050] The pasting module is configured to paste the at least one target contour onto the original image to obtain at least one enhanced image.

[0051] The data enhancement device based on a large model provided by the present invention includes a determination module, a detection module, a processing module, and a pasting module. The determination module determines at least one target category and generates corresponding prompt information. The target category is the category to be enhanced, including but not limited to pedestrians, vehicles, traffic signs, and obstacles. The detection module uses the Grounded-SAM detection algorithm to detect and segment the detection image according to the at least one prompt information to obtain at least one target contour and the corresponding target bounding box. The processing module determines whether there are two target bounding boxes corresponding to different target categories that intersect among the at least one target contour. If so, the target bounding boxes of the two are merged according to a preset rule to form a new target contour. The pasting module pastes the at least one target contour onto the original image to obtain at least one enhanced image. The present invention uses the Grounded-SAM large model to predict the target to be enhanced, can efficiently extract the detection results without additional training, and accurately locates the target according to the segmentation ability of the large model, reduces background interference, and improves the quality of data enhancement. In addition, by virtue of the generalization ability of the large model, scarce targets can be flexibly obtained from any image for data augmentation, solving the long-tail data problem and enhancing the robustness and detection ability of the autonomous driving model.

[0052] According to a third aspect of the present invention, there is provided a data enhancement device based on a large model. The data enhancement device based on a large model includes a processor and a memory. At least one computer instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the steps performed in any one of the above-mentioned data enhancement methods based on a large model.

[0053] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium. At least one computer instruction is stored in the storage medium, and the instruction is loaded and executed by a processor to implement the steps performed in any one of the above-mentioned data enhancement methods based on a large model.

[0054] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Description of the Drawings

[0055] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0056] Figure 1 is a flowchart of a data augmentation method based on a large model provided by an embodiment of the present invention;

[0057] Figure 2 is a perspective schematic diagram of a camera in an embodiment of the present invention;

[0058] Figure 3 is a structural diagram of a data augmentation device based on a large model provided by an embodiment of the present invention. Detailed Embodiments

[0059] Exemplary embodiments will be described in detail here, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention.

[0060] An embodiment of the present invention provides a data augmentation method based on a large model, as Figure 1 shown, the data augmentation method based on a large model includes the following steps:

[0061] Step 101, determine at least one target category and generate corresponding prompt information;

[0062] In this embodiment, the target category is the category to be augmented, including but not limited to pedestrians, vehicles, traffic signs, and obstacles. In the object detection of autonomous driving, vehicles generally include several types such as buses, cars, trucks, pedestrians, bicycles, tricycles, and motorcycles; traffic signs include warning signs, prohibition signs, indication signs, guide signs, road construction safety signs, etc.; obstacles include but not limited to fallen trees, dropped goods, animals, illegally parked vehicles, etc.

[0063] Step 102, use the Grounded-SAM detection algorithm to detect and segment the detection image according to at least one piece of prompt information to obtain at least one target contour and corresponding target box;

[0064] In this embodiment, the Grounded-SAM model is used to detect the required targets, and the target contour information is obtained. Specifically, from the perspective of the set-based model, the Grounded-SAM model integrates the open-set detector model Grounding DINO and the promptable segmentation model SAM, enabling the model to accurately detect and segment text according to user input in both regular and long-tail scenarios. Among them, Grounding DINO is an open-object detector that can detect any object with any free-form text prompt. The model has been trained on more than 10 million images, including detection data, visual grounding data, and image-text pairs. It has strong zero-shot learning detection performance. The Segment Anything Model (SAM) is an open-world segmentation model that can "cut out" any object in any image according to appropriate prompts (such as points, boxes, or text). The model has been trained on more than 11 million images and 1.1 billion masks, and the zero-shot learning performance of the model is very strong. Therefore, given an input image and a text prompt, first, Grounding DINO uses the text information as a condition to generate accurate boxes for objects or regions in the image. Subsequently, the annotation boxes obtained by Grounding DINO are used as boxes to prompt SAM to generate accurate mask annotations. By leveraging the capabilities of these two models, the open-set detection and segmentation tasks can be completed more easily.

[0065] Step 103: Determine whether there are two intersecting target boxes in at least one target contour that correspond to different target categories. If so, merge the target boxes of the two according to a preset rule to form a new target contour.

[0066] In this embodiment, the spatial positions of different detection targets are analyzed with the detection boxes, and they are reasonably matched according to the correlation between the two. For example, when the detected transportation vehicles (such as bicycles, motorcycles, etc.) and pedestrians are highly coincident or close in space, they can be merged into a new detection category, namely "cyclist and pedestrian", through the matching strategy. This strategy enables the detection results to not only reflect individual targets but also capture more complex scene combinations, achieving a higher level of semantic understanding.

[0067] In one embodiment, determining whether there are two intersecting target boxes in at least one target contour that correspond to different target categories includes:

[0068] Judge by calculating the first intersection over union of the target box of the first target category and the target box of the second target category. If the first intersection over union is non-zero, it is determined that there is an intersection;

[0069] If the first intersection over union (IoU) is zero, it is determined that there is no intersection.

[0070] The first IoU is calculated by the first formula, and the first formula includes:

[0071]

[0072] Where, is the first IoU, area(P) is the area of the bounding box of the first target category, and area(T) is the area of the bounding box of the second target category.

[0073] The method of the present invention calculates the first IoU of the bounding box of the first target category and the bounding box of the second target category, and quickly and accurately determines whether there is an intersection between the two. If the IoU is non-zero, there is an intersection; if it is zero, there is no intersection. The method of the present invention is simple and efficient, and can improve the accuracy of target detection and analysis.

[0074] In one embodiment, the bounding boxes of the two are merged according to a preset rule to form a new target contour, including:

[0075] Calculate the loss matrix according to the preset rule and the coordinate difference between the centers of the bounding box of the first target category and the bounding box of the second target category;

[0076] Perform the minimum-cost matching of the bipartite graph based on the loss matrix to obtain two corresponding target categories;

[0077] Merge the contours of the two matched target categories to form a new target contour.

[0078] The method of the present invention calculates the loss matrix through the coordinate difference of the center points, and uses the minimum-cost matching of the bipartite graph to achieve the precise pairing of two target categories. After merging the contours, a new target is generated, effectively improving the robustness of target detection in occlusion scenarios and enhancing the ability of autonomous driving to recognize complex interactive targets.

[0079] In one embodiment, the loss matrix is calculated by the second formula, and the second formula includes:

[0080]

[0081] Where, represents the loss matrix of the bounding box of the first target category and the bounding box of the second target category; (x i , y i ) and (x j , y j ) respectively represent the center point coordinates of the bounding box of the first target category and the bounding box of the second target category.

[0082] The method of the present invention quantifies the positional difference between the center point coordinates of the bounding boxes of the first target category and the second target category by calculating the Euclidean distance, provides an accurate pairing cost for bipartite graph matching, improves the accuracy of target association, optimizes the contour merging effect in occlusion scenarios, and enhances the robustness of multi-target tracking.

[0083] In summary, the matching strategy of the present invention effectively solves the limitation of the single detection type existing in the detection process of large models. The object detection of large models is mostly the recognition of a single category and cannot fully represent complex combinations in actual scenarios (such as cyclists, people pushing carts, etc.). However, through the matching method proposed by the present invention, multiple associated targets can be combined into a new category, making the detection results more in line with the actual business needs. In the field of autonomous driving, the detection results in different scenarios may affect the accuracy and safety of driving decisions. For example, the detection of cyclists helps to anticipate dangerous situations in advance. This matching strategy not only improves the flexibility of detection but also ensures that the categories of detected targets are more adaptable to the scenario, improving the practicality and reliability of the model in actual applications.

[0084] Step 104: Paste at least one target contour onto the original image to obtain at least one enhanced image.

[0085] In one embodiment, pasting at least one target contour onto the original image to obtain at least one enhanced image includes:

[0086] Obtain the internal and external parameters of the camera through Zhang Zhengyou calibration method;

[0087] Combining the internal and external parameters of the camera, as well as the actual size and pixel size of at least one target contour, determine the longitudinal position of at least one target contour on the original image;

[0088] Determine the horizontal position according to the symmetric position of the center point of at least one target contour and the existing targets on the original image;

[0089] According to the longitudinal position and the horizontal position, paste at least one target contour onto the original image to obtain at least one enhanced image.

[0090] In this embodiment, the internal and external parameters of the camera are obtained through Zhang Zhengyou calibration method. Combining the internal and external parameters of the camera and the actual size (length, width, height) of the object to be pasted, the approximate pasting position can be determined according to the bounding box information of the target in the pixel space.

[0091] Specifically, Zhang Zhengyou calibration method fixes the world coordinate system on the checkerboard, so the physical coordinate of any point on the checkerboard is W = 0. Since the world coordinate system of the calibration board is defined in advance artificially and the size of each grid on the calibration board is known, the physical coordinates (U, V, W = 0) of each corner point in the world coordinate system can be calculated. These information will be used: the pixel coordinates (u, v) of each corner point, the physical coordinates ((U, V, W = 0) of each corner point in the world coordinate system, to calibrate the camera and obtain the internal and external parameter matrices and distortion parameters of the camera. The perspective principle of the camera is as Figure 2 shown.

[0092] As shown in the above formula, according to the principle of similar triangles and the internal and external parameters of the camera, the depth distance of the target can be obtained. Where y1 is the pixel ordinate of the target, f is the camera focal length, H is the camera height, and Z1 is the depth distance of the target. According to the depth distance and the actual size and pixel size of the target to be pasted, the longitudinal position of the target to be pasted on the image can be estimated.

[0093] The method of the present invention uses the relationship between the camera calibration parameters and the target contour size to accurately calculate the longitudinal position of the target in the image, and determines the horizontal position through central symmetry to ensure that the enhanced target is consistent with the scene perspective relationship, improving the authenticity of data enhancement.

[0094] In one embodiment, determining the horizontal position according to the symmetric position of at least one target contour and the center point of the existing target on the original image includes:

[0095] Establish a position symmetric to the center point of the existing target on the original image as the existing target set X, and establish a target contour set Y according to at least one target contour;

[0096] According to the center point position of the X j th target in the existing target set X, select a target subset Z in the target contour set Y that meets the position condition;

[0097] Paste the Z i target in the target subset Z to the center point position of X j and calculate the second intersection over union ratio of this paste position with other targets;

[0098] If the second intersection over union ratio is zero, paste the Z i target to the center point position of X j and remove this position from the existing target set X;

[0099] If the second intersection over union ratio is not zero or Z iIf the vertical position of the target is not at the center point of the existing target, then according to the determined vertical position, starting from the center point of the image horizontal axis, with half of the target box width as the step size, gradually search for suitable positions on both sides. The suitable position is the position where the second intersection over union with other targets is zero;

[0100] If no suitable position is still found, then give up the Z i pasting of the target, and for the Z in the target subset Z i+1 paste the target to X j position, and calculate its second intersection over union again;

[0101] This process continues to loop until a qualified position is found or all target contours are traversed;

[0102] If there is no qualified pasting target at a certain X j position, then no pasting is performed at this position.

[0103] The method of the present invention ensures the reasonable integration of the newly added target contour and the original scene through symmetric position matching and dynamic adjustment strategies. Using the intersection over union verification to avoid target overlap, and the step size search mechanism to improve the position adaptability, enhancing the authenticity and spatial rationality of image synthesis, and effectively improving the quality of data augmentation.

[0104] In actual use, after determining the approximate vertical position, it is necessary to reasonably determine the horizontal position. In the field of autonomous driving, more attention is paid to the targets on the road, and the most attention is paid to the targets directly in front of the road. Since the boundary of the road cannot be determined and to ensure the rationality of pasting, this design preferably pastes the target to be pasted at the symmetric position of the center point of the existing target. This pasting strategy ensures the rationality and accuracy of target pasting, effectively avoiding the overlap problem between the newly pasted target and the existing target, and at the same time maintaining the authenticity and consistency of the overall image through reasonable size and position adjustment.

[0105] In summary, the pasting strategy of the present invention fully considers the imaging principle of images and the actual requirements of the autonomous driving scenario. In the autonomous driving environment, most detection targets are usually concentrated on the road, and the objects in the middle area of the road are often more important, and the recognition performance is also higher than that of the areas on both sides of the road. Therefore, in order to avoid the pasted target appearing in the area that does not need to be focused on, thus affecting the data augmentation effect, the strategy preferentially selects to paste the target to be pasted to the symmetric position of the existing target or the central area of the image, so as to ensure that the position of the pasted target is reasonable and has a high detection value in the scenario. This method effectively reduces the randomness of the target pasting position and improves the pertinence of data augmentation. It can be seen that the pasting strategy of the present invention not only comprehensively considers whether the target to be pasted will cover or interfere with the existing targets in the original image, but also carefully evaluates whether the size of the target to be pasted conforms to the image imaging principle. This process ensures the coordination between the pasted target and the original image, so as to optimize the data augmentation effect. At the same time, the size and position of the target also conform to the natural vision law, avoiding the situation of false detection or category jump caused by unreasonable target size. For example, if an overly large target is pasted in the distance of the image or an overly small target is pasted nearby, it may lead to incorrect identification or confusion of the category by the model. Through this pasting strategy, it is ensured that the pasting process of each target conforms to the requirements of the actual scenario, providing high-quality and real-world-like augmented data for the model.

[0106] In one embodiment, before pasting at least one target contour onto the original image to obtain at least one augmented image, the above method further includes:

[0107] Screen at least one target contour to obtain at least one valid data;

[0108] Correspondingly, pasting at least one target contour onto the original image to obtain at least one augmented image includes:

[0109] Paste at least one valid data onto the original image to obtain at least one augmented image.

[0110] In actual use, before pasting, it is also necessary to conduct a sampling inspection on all the sorted data to screen out valid target contours and avoid images with incomplete or incorrectly segmented targets by the large model to ensure the correctness of the data.

[0111] In one embodiment, the above method further includes:

[0112] Use an edge blurring algorithm to perform edge smoothing processing on the augmented image.

[0113] Specifically, in order to make the pasting of the target to be pasted on the picture more natural, the edge area of the image is processed by Gaussian blur to make the edge transition between the pasting target and the background smoother.

[0114] The method of the present invention can effectively eliminate the rigid boundary of the synthesized target through the edge blur algorithm, make it naturally blend with the background, enhance the visual realism of the enhanced image, help reduce the overfitting of the deep learning model to artificial traces, and improve the generalization performance in real scenarios.

[0115] The data augmentation method based on a large model provided by the present invention first determines at least one target category and generates corresponding prompt information. The target category is the category to be augmented, including but not limited to pedestrians, vehicles, traffic signs, and obstacles. Then, according to at least one piece of prompt information, the Grounded-SAM detection algorithm is used to detect and segment the detection image to obtain at least one target contour and the corresponding target box. Then, it is judged whether there are two intersecting target boxes in at least one target contour that respectively correspond to different target categories. If so, the target boxes of the two are merged according to the preset rules to form a new target contour. Finally, at least one target contour is pasted onto the original image to obtain at least one enhanced image. The present invention uses the Grounded-SAM large model to predict the target to be augmented, can efficiently extract the detection results without additional training, and accurately locates the target according to the segmentation ability of the large model, reduces background interference, and improves the quality of data augmentation. In addition, with the generalization ability of the large model, scarce targets can be flexibly obtained from any image for data expansion, solving the long-tail data problem and improving the robustness and detection ability of the autonomous driving model.

[0116] Based on the above Figure 1 In the corresponding embodiments, the data augmentation method based on a large model described above, the following is an embodiment of the device of the present invention, which can be used to execute the method embodiment of the present invention.

[0117] An embodiment of the present invention provides a data augmentation device based on a large model, as Figure 3 shown, the device includes: a determination module 201, a detection module 202, a processing module 203, and a pasting module 204;

[0118] The determination module 201 is used to determine at least one target category and generate corresponding prompt information. The target category is the category to be augmented, including but not limited to pedestrians, vehicles, traffic signs, and obstacles;

[0119] The detection module 202 is used to detect and segment the detection image according to at least one piece of prompt information by using the Grounded-SAM detection algorithm to obtain at least one target contour and the corresponding target box;

[0120] The processing module 203 is configured to determine whether there are two intersecting target bounding boxes corresponding to different target categories in at least one target contour. If so, the target bounding boxes of the two are merged according to a preset rule to form a new target contour.

[0121] The pasting module 204 is configured to paste at least one target contour onto the original image to obtain at least one enhanced image.

[0122] The data enhancement device based on a large model provided by the present invention includes a determination module 201, a detection module 202, a processing module 203, and a pasting module 204. The determination module 201 determines at least one target category and generates corresponding prompt information. The target category is the category to be enhanced, including but not limited to pedestrians, vehicles, traffic signs, and obstacles. The detection module 202 uses the Grounded-SAM detection algorithm to detect and segment the detection image according to at least one piece of prompt information to obtain at least one target contour and corresponding target bounding boxes. The processing module 203 determines whether there are two intersecting target bounding boxes corresponding to different target categories in at least one target contour. If so, the target bounding boxes of the two are merged according to a preset rule to form a new target contour. The pasting module 204 pastes at least one target contour onto the original image to obtain at least one enhanced image. The present invention uses the Grounded-SAM large model to predict the target to be enhanced, can efficiently extract the detection results without additional training, and accurately locates the target according to the segmentation ability of the large model, reducing background interference and improving the quality of data enhancement. In addition, by virtue of the generalization ability of the large model, scarce targets can be flexibly obtained from any image for data expansion, solving the long-tail data problem and improving the robustness and detection ability of the autonomous driving model.

[0123] Based on the above Figure 1 corresponding embodiment of the data enhancement method based on a large model, another embodiment of the present invention further provides a data enhancement device based on a large model. The data enhancement device based on a large model includes a processor and a memory. At least one computer instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the above Figure 1 corresponding embodiment of the data enhancement method based on a large model described.

[0124] Based on the above Figure 1For the data augmentation method based on large models described in the corresponding embodiments, the embodiments of the present invention also provide a computer-readable storage medium. For example, a non-transitory computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc. At least one computer instruction is stored on this storage medium for executing the above-mentioned Figure 1 data augmentation method based on large models described in the corresponding embodiments, which will not be elaborated here.

[0125] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed by the present invention. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the claims.

[0126] It should be understood that the present invention is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A data augmentation method based on large models, characterized in that, The method includes: Determine at least one target category and generate corresponding prompt information, where the target category is the category to be enhanced, including but not limited to pedestrians, vehicles, traffic signs, and obstacles; Detect and segment the detection image using the Grounded-SAM detection algorithm according to at least one piece of the prompt information to obtain at least one target contour and corresponding target box; Judge whether there are two intersecting target boxes in the at least one target contour that correspond to different target categories respectively. If so, merge the target boxes of the two according to a preset rule to form a new target contour; Paste the at least one target contour onto the original image to obtain at least one enhanced image.

2. The data augmentation method based on a large model according to claim 1, wherein The judgment of whether there are two intersecting target boxes in the at least one target contour that correspond to different target categories respectively includes: Judge by calculating the first intersection over union of the target box of the first target category and the target box of the second target category. If the first intersection over union is non-zero, it is judged that there is an intersection; If the first intersection over union is zero, it is judged that there is no intersection; The first intersection over union is calculated by a first formula, and the first formula includes: Among them, IOU is the first intersection over union, area(P) is the area of the bounding box of the first target category, and area(T) is the area of the bounding box of the second target category.

3. The data augmentation method based on a large model according to claim 2, wherein The merging of the target boxes of the two according to a preset rule to form a new target contour includes: Calculate the loss matrix according to a preset rule and the coordinate difference between the center points of the target box of the first target category and the target box of the second target category; Perform bipartite graph minimum cost matching based on the loss matrix to obtain two target categories in one-to-one correspondence; Merge the contours of the two matched target categories to form the new target contour.

4. The data augmentation method based on a large model according to claim 3, wherein The loss matrix is calculated by a second formula, and the second formula includes: Among them, cost i,j represents the loss matrix of the target boxes of the first target category and the target boxes of the second target category; (x i , y i ) and (x j , y j ) respectively represent the center point coordinates of the target boxes of the first target category and the target boxes of the second target category.

5. The data augmentation method based on a large model according to claim 1, wherein The pasting of the at least one target contour onto the original image to obtain at least one enhanced image includes: Obtain the internal and external parameters of the camera through the Zhang Zhengyou calibration method; Combined with the internal and external parameters of the camera, as well as the actual size and pixel size of the at least one target contour, determine the longitudinal position of the at least one target contour on the original image; Determine the horizontal position according to the symmetric position of the center point of the at least one target contour and the existing target on the original image; According to the longitudinal position and the horizontal position, paste the at least one target contour onto the original image to obtain at least one enhanced image.

6. The data enhancement method based on a large model according to claim 5, wherein The determination of the horizontal position according to the symmetric position of the center point of the at least one target contour and the existing target on the original image includes: Establish a set X of existing targets at the symmetric position of the center point of the existing target on the original image, and establish a set Y of target contours according to the at least one target contour; According to the center point position of the Xth target in the existing target set X j select the target subset Z in the target contour set Y that meets the position condition; Paste Z i to the center point position of X j and calculate the second intersection over union of this paste position and other targets; If the second intersection-over-union ratio is zero, then paste Z i the target to the center point position of X j and remove this position from the existing target set X; If the second intersection over union is not zero or Z i If the vertical position of the target is not at the center point position of the existing targets, then according to the determined vertical position, starting from the center point of the horizontal axis of the image, with half of the target box width as the step size, search step by step along both sides to find a suitable position, where the suitable position is the position where the second intersection over union with other targets is zero; If a suitable position is still not found, then give up pasting the Z i target, and paste the Z i+1 target at the X j position, and calculate its second intersection over union again; This process continues to loop until a qualified position is found or all target contours are traversed; If a certain X j has no eligible paste target at the position, no paste is performed at that position.

7. The data augmentation method based on a large model according to claim 1, wherein The method further includes: Perform edge smoothing processing on the enhanced image using an edge blurring algorithm.

8. A data augmentation device based on a large model, characterized in that, It includes: A determination module, a detection module, a processing module, and a pasting module; The determination module is used to determine at least one target category and generate corresponding prompt information, where the target category is the category to be enhanced, including but not limited to pedestrians, vehicles, traffic signs, and obstacles; The detection module is used to detect and segment a detection image by using the Grounded-SAM detection algorithm according to at least one of the prompt messages, so as to obtain at least one target contour and a corresponding target box; The processing module is used to judge whether there are two target boxes corresponding to different target categories that intersect among the at least one target contour. If so, the target boxes of the two are merged according to a preset rule to form a new target contour; The pasting module is used to paste the at least one target contour onto the original image to obtain at least one enhanced image.

9. A data augmentation device based on a large model, characterized in that, The large model-based data augmentation device includes a processor and a memory. At least one computer instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the steps executed in the large model-based data augmentation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, At least one computer instruction is stored in the storage medium, and the instruction is loaded and executed by the processor to implement the steps executed in the large model-based data augmentation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Power distribution cabinet on-off state identification method based on binocular vision

    CN111260788A

  • Automatic data enhancement and expansion method, recognition method and system for deep learning

    CN111881720A

  • Data enhancement method for traffic sign target detection field

    CN113793279A

  • Domestic garbage data set generation method based on data enhancement

    CN114429573A

  • Image augmentation for machine learning based defect examination

    US20240095903A1