Picture synthesis method for target detection

Through a target detection-oriented image synthesis method, the problems of insufficient automation, poor stability and manual annotation in the prior art are solved, and the rapid generation of synthetic data and automatic annotation that meets the actual scenarios are realized, and the performance of the detection model is improved.

CN119941527AActive Publication Date: 2025-05-06CHONGQING YUDONG EXPRESSWAY CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510017343.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The data synthesis method in the field of target detection has problems such as inappropriate size, inappropriate size, inadequate synthesis, inadequate automation, instability, and manual labeling of pictures after synthesis, resulting in high data acquisition cost and low efficiency.

Method used

A method of image synthesis for object detection is proposed, including a base map image acquisition module, a synthetic object acquisition module, a synthetic object parameter determination module on the base map and a synthetic object fusion module. Through these modules, large batches of diverse synthetic data can be quickly generated and labels of synthetic pictures can be automatically generated to avoid subsequent manual processing.

Benefits of technology

It realizes the rapid generation of synthetic data that conforms to the actual scenario, and automatically generates the label of synthetic pictures, which improves the performance of the detection model and reduces the cost of data acquisition and labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941527A_ABST
    Figure CN119941527A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection-oriented picture synthesis method in the field of computer vision, which comprises the following steps of: firstly, manually making a detection model training set, training to obtain a preliminary detection model, and obtaining an object foreground as a synthesis object according to a marked detection coordinate frame by using an SAM segmentation model; detecting each frame of the video of the real scene by using the trained detection model, storing the detection information of each frame, including object category, object size and object position information, and making a base map through the information; and finally, traversing the stored detection information of each frame of video, selecting the same kind of foreground according to the size information or the center distance information, and fusing the foreground to the center position of the detected object. According to the method, the object size, the position and the category of the synthetic object on the base map are ensured to accord with the actual scene condition, the label of the synthetic picture can be automatically generated, subsequent manual labeling is not needed, and the performance of a detection model can be effectively improved by the synthetic picture data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a method for synthesizing images for target detection. Background Art

[0002] Providing image data to a model built with a neural network for training is an important means of target detection in the field of computer vision. This method is highly dependent on the quality of data. Data acquisition and production are crucial to model training. However, in practical applications, data acquisition has high labor costs and long-tail data is difficult to obtain. Therefore, it is of great significance to produce effective training data through data synthesis. However, the current methods of synthesizing data on the market have problems such as the size of objects synthesized into the target background image is not appropriate, the synthesis size is not appropriate, the synthesis is not automated enough, the synthesis is not stable enough, and the synthesized image needs to be labeled separately.

[0003] In the "Automatic Labeling Method and Device for Traffic Images Based on Difficult Sample Mining" disclosed in patent CN108492343B, the method of automatically obtaining the bounding box of objects in the image using the k-means algorithm is inaccurate, and the position of the synthesized object on the source image is directly copied to the base map without considering the base map situation, resulting in the synthesized image not being consistent with the actual image situation.

[0004] In the "Methods, devices, equipment and media for image synthesis and image synthesis model training" disclosed in patent CN109472764B, a detection network is first used to obtain object attributes, and then a generative adversarial network is trained to modify the objects in the image. This method has high requirements on usage conditions, requires a lot of preliminary detection and labeling and generative adversarial network training, and the rendering effect is unstable.

[0005] Patent CN113762422B discloses "A method and system for synthesizing image training sets", which trains a semantic segmentation model for a scene, transforms the obtained foreground objects and then fuses them back into the source image. The data synthesized by this method cannot be automatically annotated and further manual annotation of the synthesized image is required before it can be used.

[0006] Patent CN118379570A discloses a method and device for synthesizing image training data. Although it uses segmentation technology to obtain synthesized objects and then synthesizes the synthesized objects onto a base image, the patent does not have an accurate judgment on the size and position of the synthesized objects onto the base image. It only uses a rough clustering combination method. This method results in large deviations in the size and position of the synthesized objects in the actual synthesized images.

[0007] The methods described in the existing data synthesis patents do not yet have a simple, fast, and low-cost data synthesis method that can automatically generate large quantities of relatively realistic synthetic data, and does not require subsequent manual annotation of synthetic images, and can be directly used for training. In order to expand the number of images, manual data collection is often used, but this method will consume huge manpower and material resources, and it is difficult to collect data that rarely exists in reality. For example, there is a situation where a certain category of objects rarely appear in a certain area, so there is insufficient training data. If the detection model is used to detect objects of a certain category in the corresponding area, the detection accuracy will be greatly reduced.

[0008] To address the above issues, there are currently two commonly used methods on the market: 1. Use Photoshop and other photo editing software to artificially create images. This method has high labor costs and cannot quickly generate large quantities of training data; 2. Create a data synthesis system to automatically synthesize images. However, the images generated by the data synthesis methods in the prior art have serious problems with distorted synthesis effects, such as the inappropriate position, category and size of the synthesized objects on the base map. In addition, most of the existing methods do not have a simple and clear process to support the mass generation of synthetic images and automatically form annotations for the synthetic images. The synthetic images cannot be directly used for model training, and subsequent manual processing is required.

[0009] Therefore, the present invention provides a method for image synthesis for target detection to solve the problems raised in the above background technology. Summary of the invention

[0010] The purpose of the present invention is to provide a method for image synthesis for target detection to solve the problems raised in the above background technology.

[0011] To achieve the above object, the present invention provides the following technical solutions:

[0012] A method for image synthesis for target detection, comprising a base image acquisition module, a synthetic object acquisition module, a synthetic object parameter determination module on the base image, and a synthetic object fusion module;

[0013] The base map acquisition module has the following three ways to obtain the base map:

[0014] a) It can be obtained from the model training data (the annotation file of the training data can directly show whether it is a pure background image. If it is a pure background image, it will be used as the base image);

[0015] b) Use the trained detection model to detect new data (non-labeled training data), and list the pictures without detection output as suspicious pure background pictures for subsequent manual screening as base pictures;

[0016] c) Apply the gmic erasing algorithm to automatically erase according to the object detection frame, manually screen the erased images, and use the images without objects as the base image;

[0017] The synthetic object acquisition module locates the object through the marked detection frame coordinates, and calls the advanced segmentation algorithm on the market to segment all objects to save the area within the object outline in the detection frame as a mask. The mask can be used to locate the precise pixels of the object as a sample library for future synthetic objects;

[0018] The synthetic object parameter determination module on the base map refers to the determination of the position, scale, etc. of the synthetic object synthesized on the base map from the position and size of the object of the same category in the object source scene, by calculating the distance between the center of the synthetic object itself and the center of the object of the same category in the object source scene or the area difference between the two. If it is less than a certain threshold, the position and size of the object of the same category in the object source scene are used to apply to the parameters of the synthetic object synthesized on the base map;

[0019] The synthetic object fusion module extracts the object from the source image through the synthetic object mask and fuses it into the base image. In order to ensure the consistency of the connection between the synthetic object and the base image, the mask needs to be processed by median filtering to make the edge of the synthetic object smooth, and the erosion and dilation operation needs to fill the holes in the mask. Then, the mask is used as the image fusion weight to fuse the synthetic object into the base image.

[0020] As a further solution of the present invention: specifically comprising the following steps:

[0021] Step S1: When capturing video frames from real-scene traffic videos as detection model training data, the captured video frames are named as the concatenation of the video name and the image serial number identifier, specifically "video name_image serial number identifier.jpg". For example, there are two types of objects in the training data, namely car and person. Then, the objects in the data are labeled and trained to obtain a detection model;

[0022] Step S2: Get the source video of the image through the name of the data image marked in step S1, collect and organize all the source videos, save each frame of the video, and save the directory name as the video name and the image name as the frame number;

[0023] Step S3: Use the trained detection model to detect each image in each directory in step S2, and save the detection results (the saved results are json files with the same name as the image name, such as image name.json). The saved directory name is consistent with that in step S1;

[0024] Step S4: Count the detection results of each frame of each video, save the pictures that are judged as having no objects as candidate background images, and select gmic (automatic erasure code library) to automatically erase the pictures detected as having objects according to the detection coordinate frame; if the detection results of all frames of a video are pictures without objects, it is necessary to sample some pictures and use gmic to automatically erase the objects according to the object detection frame obtained by the detection as candidate background images;

[0025] Step S5: Manually screen the candidate background images, store the images without any objects in the background image directory, and traverse the training set to find pure background images and store them in the background image directory. Finally, the mapping relationship between the candidate background image and the video source needs to be saved;

[0026] Step S6: The source of the synthesized object is the object in the annotated data training set. The SAM algorithm is used to segment the object with the annotated bounding box information in the annotated image to obtain the object mask. The mask is named as the image name + bboxcords + bounding box information + object category name (note that the specific video from which the mask corresponding image comes can be obtained from S1), such as video1_bboxcords_1816_202_1916_274_car.png. The source image name, object coordinates and object category can be located by the name of the mask, and the pixel position of the object can be obtained by the mask to separate the object.

[0027] Step S7: Further process the object detection information of each frame of each video obtained in step S2 to obtain a set of center points and corresponding sizes of all objects of each category in each video, and save the results as a json file with a name format of classes_size_results_+video frame name.json, and the content is {car:[[x_center,y_center,w,h],...], person:[[x_center,y_center,w,h],...]}, where x_center, y_center, w, h correspond to the horizontal coordinate of the center point of the object, the vertical coordinate of the center point, the width of the object, and the height of the object, respectively.

[0028] Step S8: Further process the object detection information of each frame of each video obtained in step S2 to obtain the set of object categories and corresponding center points of each frame of each video, and save the results as a json file named after the video name, such as video1.json, with the specific content as follows:

[0029] {0.jpg:{car:[[x_center,y_center,w,h],...]},person:[[x_center,y_center,w,h],...]}};

[0030] Step S9: Randomly sample the category and center point of each object in a frame of the image from the information saved in step S7 as the category and center point of the object synthesized into the base map. At the same time, randomly select a corresponding base map according to the record in step S1. The category, center point and size of the synthesized object obtained in the above steps will be used directly as the annotation information of the synthesized image for training. Traverse the objects of the same category recorded by the mask name in step S6. If the synthesized object and the base map are from the same video (or from the same camera), calculate the distance between the center points of the two. If the distance is less than a certain threshold, it is adopted as a synthesized object. If the synthesized object and the base map are not from the same video (or not from the same camera), calculate the difference between the areas of the two. If the difference is less than a certain threshold, it is adopted as a synthesized object. This operation is performed each time. The traversal order of the mask in S6 will be disrupted to ensure the diversity of synthesis. If each synthetic object category and center point are manually set, the distance between the center point and the center point set of the corresponding category object in step S6 will be calculated first, and the size corresponding to the center point of step S6 with the minimum distance will be selected as the size parameter of synthesis. The category, center point and size of the synthetic object obtained in the above steps will be used directly as the annotation information of the synthetic image for training. Then, objects of the same category recorded by the mask name in step S6 will be traversed. If the synthetic object and the base map are from the same video, the center point will be calculated as a distance. If the distance is less than a certain threshold, it will be adopted as a synthetic object. Otherwise, the area of ​​the two will be calculated as a difference. If the difference is less than a certain threshold, it will be adopted as a synthetic object. Each time this operation is performed, the traversal order of the mask in S6 will be disrupted.

[0031] Step S10: According to the target center position and target size information of the object to be synthesized on the base map obtained in step S9, the mask is scaled proportionally in combination with the ratio of the original size of the synthesized object to the target size recorded in the mask name generated in step S6 so that the size of the object synthesized on the base map is consistent with the target size. Then, the mask is operated multiple times by median filtering to smooth the mask edge, and a dilation operation is added to fill the holes in the mask. Finally, the object is synthesized into the base map with the mask as the weight. The synthesis formula is as follows:

[0032] new_image=image1×mask+image2×(1-mask)

[0033] Among them, new_image is the final synthesized image, image1 is the source image of the synthesized object, and mask is the foreground of the synthesized object segmented using the SAM algorithm.

[0034] As a further solution of the present invention: the base map in the base map image acquisition module refers to a pure background image without any detected object.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] The present invention can quickly generate large quantities of diverse data on the base map, ensure that the object size and object category of the synthesized object on the base map conform to the actual scene conditions, and can automatically generate annotations for the synthesized pictures without subsequent manual annotation. The synthesized picture data can effectively improve the performance of the detection model, and the method of automatically fusing objects on the labeled detection frame picture into other pictures can ensure the appropriateness of the object size and the authenticity of the fused object, realize the method of synthesizing objects in pictures collected from the same camera onto the background picture, and the method of synthesizing objects in different scenes into another scene, which can simply and quickly complete the synthesis of images close to the real scene, greatly promoting the model training effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is the overall flow chart of the system of the present invention;

[0038] Figure 2 It is the base map image acquisition module of the present invention;

[0039] Figure 3 A schematic diagram of a synthetic object acquisition module of the present invention;

[0040] Figure 4 It is a schematic diagram of a module for determining parameters of a synthetic object on a base map of the present invention;

[0041] Figure 5 It is a schematic diagram of a synthetic object fusion module of the present invention;

[0042] Figure 6 A view of a synthetic base map in the present invention;

[0043] Figure 7 and Figure 8 For the present invention Figure 5 a view of the corresponding composite image;

[0044] Fig. 9 For the present invention Figure 6 View of the corresponding real image. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] See also Figures 2 to 5 ,In an embodiment of the present invention, an image synthesis method for target detection includes a base image acquisition module, a synthetic object acquisition module, a synthetic object parameter determination module on the base image, and a synthetic object fusion module;

[0047] The base map acquisition module has the following three ways to obtain the base map:

[0048] a) It can be obtained from the model training data (the annotation file of the training data can directly show whether it is a pure background image);

[0049] b) Use the trained detection model to detect new data (non-labeled training data), and list the images without detection output as suspicious pure background images for subsequent manual screening;

[0050] c) Apply the gmic erasing algorithm to automatically erase according to the object detection frame, and manually screen the erased image to use the erased image as the base image;

[0051] The synthetic object acquisition module locates the object through the marked detection frame coordinates, and calls the advanced segmentation algorithm on the market to segment all objects to save the area within the object outline in the detection frame as a mask. The mask can be used to locate the precise pixels of the object as a sample library for future synthetic objects;

[0052] The module for determining parameters of the synthesized object on the base map refers to determining the position, scale, etc. of the synthesized object on the base map from the position and size of the object of the same category in the object source scene, calculating the distance between the center of the synthesized object itself and the center of the object of the same category in the object source scene or the area difference between the two, and if it is less than a certain threshold, using the position and size of the object of the same category in the object source scene to apply to the parameters of the synthesized object on the base map;

[0053] The synthetic object fusion module extracts the object from the source image through the synthetic object mask and fuses it into the base image. In order to ensure the consistency of the connection between the synthetic object and the base image, the mask needs to be processed by median filtering to make the edge of the synthetic object smooth, and the erosion and dilation operation needs to fill the holes in the mask. Then, the mask is used as the image fusion weight to fuse the synthetic object into the base image.

[0054] By adopting the above technical solution, a large amount of diverse data can be quickly generated on the base map to ensure that the object size and object category of the synthesized object on the base map are consistent with the actual scene conditions, and the annotations of the synthesized images can be automatically generated without the need for subsequent manual annotation. The synthesized image data can effectively improve the performance of the detection model, and the method of automatically fusing objects on the labeled detection frame image into other images can ensure the appropriateness of the object size and the authenticity of the fused object. It realizes the method of synthesizing objects in images collected by the same camera onto the background image, and the method of synthesizing objects from different scenes into another scene. The synthesis of images close to the real scene can be completed simply and quickly, greatly promoting the model training effect.

[0055] In this embodiment, the following steps are specifically included:

[0056] Step S1: When capturing video frames from real-scene traffic videos as detection model training data, the captured video frames are named as the concatenation of the video name and the image serial number identifier, specifically "video name_image serial number identifier.jpg". For example, there are two types of objects in the training data, namely car and person. Then, the objects in the data are labeled and trained to obtain a detection model;

[0057] Step S2: Get the source video of the image through the name of the data image marked in step S1, collect and organize all the source videos, save each frame of the video, and save the directory name as the video name and the image name as the frame number;

[0058] Step S3: Use the trained detection model to detect each image in each directory in step S3, save the detection results (the saved results are json files with the same name as the image name, such as image name.json), and keep the saved directory name consistent with that in step S1;

[0059] Step S4: Count the detection results of each frame of each video, save the pictures that are judged as having no objects as candidate background images, and select gmic (automatic erasure code library) to automatically erase the pictures detected as having objects according to the detection coordinate frame; if the detection results of all frames of a video are pictures without objects, it is necessary to sample some pictures and use gmic to automatically erase the objects according to the object detection frame obtained by the detection as candidate background images;

[0060] Step S5: Manually screen the candidate background images, store the images without any objects in the background image directory, and traverse the training set to find pure background images and store them in the background image directory. Finally, the mapping relationship between the candidate background image and the video source needs to be saved;

[0061] Step S6: The source of the synthesized object is the object in the labeled data training set. The object in each picture is segmented to obtain the foreground mask of the object. The object is segmented by the marked bounding box information in the labeled picture through the SAM algorithm to obtain the object mask. The mask is named as picture name + bboxcords + bounding box information + object category name (note that the specific video from which the mask corresponding picture comes can be obtained from S1), such as video1_bboxcords_1816_202_1916_274_car.png. At that time, the source picture name, object coordinates and object category can be located through the name of the mask, and the pixel position of the object can be obtained through the mask to separate the object;

[0062] Step S7: Further process the object detection information of each frame of each video obtained in step S2 to obtain a set of center points and corresponding sizes of all objects of each category in each video, and save the results as a json file with a name format of classes_size_results_+video frame name.json, and the content is {car:[[x_center,y_center,w,h],...], person:[[x_center,y_center,w,h],...]}, where x_center, y_center, w, h correspond to the horizontal coordinate of the center point of the object, the vertical coordinate of the center point, the width of the object, and the height of the object, respectively.

[0063] Step S8: Further process the object detection information of each frame of each video obtained in step S2 to obtain the set of object categories and corresponding center points of each frame of each video, and save the results as a json file named after the video name, such as video1.json, with the specific content as follows:

[0064] {0.jpg:{car:[[x_center,y_center,w,h],...]},person:[[x_center,y_center,w,h],...]}};

[0065] Step S9: Randomly sample the category and center point of each object in a frame of the picture from the information saved in step S7 as the category and center point of the object synthesized into the base map, and also randomly select a corresponding base map according to the record of step S1. The category, center point and size of the synthetic object obtained in the above steps will be used directly as the annotation information of the synthetic picture for training, and traverse the objects of the same category recorded by the mask name in step S6. If the synthetic object and the base map are from the same video (or from the same camera), the distance between the center points of the two is calculated. If the distance is less than a certain threshold, it is adopted as a synthetic object. If the synthetic object and the base map are not from the same video (or from the same camera), the difference between the areas of the two is calculated. If the difference is less than a certain threshold, it is adopted as a synthetic object. Each time this operation is performed, The traversal order of the mask in S6 is disrupted once to ensure the diversity of synthesis. If each synthetic object category and center point are manually set, the distance between the center point and the center point set of the corresponding category object in step S6 will be calculated first, and the size corresponding to the center point of step S6 with the minimum distance will be selected as the size parameter of synthesis. The category, center point and size of the synthetic object obtained in the above steps will be used directly as the annotation information of the synthetic image for training. Then, objects of the same category recorded by the mask name in step S6 are traversed. If the synthetic object and the base map are from the same video, the center point is calculated as a distance. If the distance is less than a certain threshold, it is adopted as a synthetic object. Otherwise, the area of ​​the two is calculated as a difference. If the difference is less than a certain threshold, it is adopted as a synthetic object. Each time this operation is performed, the traversal order of the mask in S6 will be disrupted once;

[0066] Step S10: According to the target center position and target size information of the object to be synthesized on the base map obtained in step S8, the mask is scaled proportionally in combination with the ratio of the original size of the synthesized object to the target size recorded in the mask name generated in step S6 so that the size of the object synthesized on the base map is consistent with the target size. Then, the mask is operated multiple times by median filtering to smooth the mask edge, and a dilation operation is added to fill the holes in the mask. Finally, each unit of the mask is used as a weight to synthesize the object into the base map. The synthesis formula is as follows:

[0067] new_image=image1×mask+image2×(1-mask)

[0068] Where new_image is the final synthesized image, image1 is the source image of the synthesized object, and mask is the foreground of the synthesized object segmented using the SAM algorithm;

[0069] In order to intuitively demonstrate the synthesis effect of this method, this patent also compares the synthesized image and the real image, image2 is the base image; the synthesized base image is as follows Figure 6 As shown, the corresponding composite images are as shown in 7 and Figure 8 As shown, Picture 9 is a real picture.

[0070] In this embodiment, the base image in the base image acquisition module refers to a pure background image without any detected object.

[0071] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A method for image synthesis for target detection, characterized in that: It includes a base map image acquisition module, a synthetic object acquisition module, a synthetic object parameter determination module on the base map, and a synthetic object fusion module; The base map acquisition module has the following three ways to obtain the base map: a) Images without objects can be obtained from the model training data as base maps; b) Use the trained detection model to detect new data, and list the pictures without detection output as suspicious pure background pictures for subsequent manual screening to determine the pictures as base pictures; c) Apply the gmic erasing algorithm to automatically erase objects according to the object detection frame, and then perform manual screening to use the picture of the erased object as the base map; The synthetic object acquisition module locates the object through the marked detection frame coordinates, and calls the advanced segmentation algorithm on the market to segment all objects to save the area within the object outline in the detection frame as a mask. The mask can be used to locate the precise pixels of the object as a sample library for future synthetic objects; The synthetic object parameter determination module on the base map refers to the determination of the position, scale, etc. of the synthetic object synthesized on the base map from the position and size of the object of the same category in the object source scene, by calculating the distance between the center of the synthetic object itself and the center of the object of the same category in the object source scene or the area difference between the two. If it is less than a certain threshold, the position and size of the object of the same category in the object source scene are used to apply to the parameters of the synthetic object synthesized on the base map; The synthetic object fusion module extracts the object from the source image through the synthetic object mask and fuses it into the base image. In order to ensure the consistency of the connection between the synthetic object and the base image, the mask needs to be processed by median filtering to make the edge of the synthetic object smooth, and the corrosion and expansion operation needs to be performed to fill the holes in the mask. Then, the mask is used as the image fusion weight to fuse the synthetic object into the base image.

2. The method for image synthesis for target detection according to claim 1, characterized in that: The specific steps include: Step S1: When capturing video frames from real-scene traffic videos as detection model training data, the captured video frames are named as the concatenation of the video name and the image serial number, and then the objects in the data are labeled and trained to obtain a detection model; Step S2: Get the source video of the image through the name of the data image marked in step S1, collect and organize all the source videos, save each frame of the video, and save the directory name as the video name and the image name as the frame number; Step S3: Use the trained detection model to detect each image in each directory in step S3, save the detection results, and keep the directory name consistent with that in step S1; Step S4: Count the detection results of each frame of each video, save the pictures that are determined to have no objects as candidate background images, or use the gmic automatic erasure code library to automatically erase the pictures that are detected as having objects according to the detection coordinate frame as candidate background images; Step S5: Manually screen the candidate background images, store the images without any objects in the background image directory, traverse the training set to find pure background images and store them in the background image directory, and finally save the mapping relationship between the candidate background images and the video source; Step S6: Segment the objects in each image to obtain the foreground mask of the object. Segment the objects using the marked bounding box information in the labeled image using the SAM algorithm to obtain the mask of the object. The source image name, the coordinates of the object and the object category can be located by the name of the mask. The pixel position of the object can be obtained by the mask to separate the object. Step S7: further process the object detection information of each frame of each video obtained in step S2 to obtain the center points and corresponding size sets of all objects of each category in each video, and save the results as a json file with the name format of classes_size_results_+video frame name.json, and the content is {car:[[x_center,y_center,w,h],...], person:[[x_center,y_center,w,h],...]}, where x_center, y_center, w, h correspond to the horizontal coordinate of the center point of the object, the vertical coordinate of the center point, the width of the object, and the height of the object respectively; Step S8: further process the object detection information of each frame of each video obtained in step S2 to obtain a set of object categories and corresponding center points of each frame of each video, and save the result as a json file named as the video name, such as video1.json, and the specific content is {0.jpg:{car:[[x_center,y_center,w,h],...]},person:[[x_center,y_center,w,h],...]}}; Step S9: Randomly sample the category and center point of each object in a frame of the image from the information saved in step S7 as the category and center point of the object synthesized into the base map. At the same time, randomly select a corresponding base map according to the record of step S1. The category, center point and size of the synthetic object obtained in the above steps will be used directly for training as the annotation information of the synthetic image. Traverse the objects of the same category recorded by the mask name in step S6. If the synthetic object and the base map are from the same video or from the same camera, calculate the distance between the center points of the two. If the distance is less than a certain threshold, it is adopted as a synthetic object. If the synthetic object and the base map are not from the same video or from the same camera, calculate the difference between the areas of the two. If the difference is less than a certain threshold, it is adopted as a synthetic object. Each time this operation is performed, it will be disrupted. The mask traversal order in S6 is performed once to ensure the diversity of synthesis. If each synthetic object category and center point are manually set, the distance between the center point and the center point set of the corresponding category object in step S6 will be calculated first, and the size corresponding to the center point of step S6 with the minimum distance will be selected as the size parameter of synthesis. The category, center point and size of the synthetic object obtained in the above steps will be used directly as the annotation information of the synthetic image for training. Then, objects of the same category recorded by the mask name in step S6 are traversed. If the synthetic object and the base map are from the same video, the center point is calculated as a distance. If the distance is less than a certain threshold, it is adopted as a synthetic object. Otherwise, the area of ​​the two is calculated as a difference. If the difference is less than a certain threshold, it is adopted as a synthetic object. Each time this operation is performed, the traversal order of the mask in S6 will be disrupted. Step S10: According to the target center position and target size information of the object to be synthesized on the base map obtained in step S8, the mask is scaled proportionally in combination with the ratio of the original size of the synthesized object to the target size recorded in the mask name generated in step S6 so that the size of the object synthesized on the base map is consistent with the target size. Then, the mask is operated multiple times by median filtering to smooth the mask edge, and a dilation operation is added to fill the holes in the mask. Finally, each unit of the mask is used as a weight to synthesize the object into the base map. The synthesis formula is as follows: new_image=image1×mask+image2×(1-mask) Among them, new_image is the final synthesized image, image1 is the source image of the synthesized object, and mask is the foreground of the synthesized object segmented using the SAM algorithm.

3. The method for image synthesis for target detection according to claim 2, characterized in that: The base image in the base image acquisition module refers to a pure background image without any detected objects.

4. The method for image synthesis for target detection according to claim 2, characterized in that: In step S4, when the detection results of all frames of a video are pictures without objects, it is necessary to sample some pictures and use gmic to automatically erase the objects according to the object detection frames obtained by detection as candidate background pictures.

5. The method for image synthesis for target detection according to claim 2, characterized in that: The source of synthetic objects is the objects in the labeled data training set.

Citation Information

Patent Citations

  • An image synthesis method for augmenting training data for target recognition

    CN108492343B

  • Methods, apparatus, equipment and media for image synthesis and training of image synthesis models

    CN109472764B

  • A method and system for image training ensemble assembly

    CN113762422B

  • Semi-automatic labeling method and segmentation positioning system for stacking planar target objects

    CN113420839A

  • Image training set synthesis method and system

    CN113762422A