Sample data generation method, object detection method, and device
Generating diverse sample data through image synthesis and image generation, solving the problem of insufficient generation of negative sample data in the prior art and improving the accuracy of the object detection model.
Patent Information
- Application Number
- PCT/CN2024/137398
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-26
AI Technical Summary
The prior art is difficult to effectively generate diverse negative sample data, resulting in high error detection rate and insufficient accuracy when facing interferers with complex diversity.
By acquiring the interference source image and the original sample image, a foreground image containing the interference is extracted, and a first sample image is synthesized with the original sample image. At the same time, using the trained image generation model, a second sample image is generated based on the original sample image, interferer characteristics and generated position information, and the sample data for object detection is generated in combination.
By generating diverse sample data, we can effectively deal with the diversity of interferers, reduce the error detection rate, and improve the accuracy of the target detection model.
Smart Images

Figure CN2024137398_26062025_PF_FP_ABST
Abstract
Description
A sample data generation method, target detection method and device
[0001] This application claims priority to Chinese patent application No. 202311790221.9 filed on December 22, 2023, entitled “A sample data generation method, target detection method and device”. The entire contents of the above Chinese patent application are incorporated into this application by reference. Technical Field
[0002] The present application relates to the field of image processing technology, and in particular to a sample data generation method, a target detection method and a device. Background Art
[0003] Currently, machine vision technology using deep learning methods is widely used in various fields. In the field of machine vision-based object detection, simply by collecting sufficient positive and negative sample data, selecting an appropriate deep learning model, and training it, a well-performing object detection model can be obtained. Therefore, sample data is of great value in improving the detection accuracy of object detection models.
[0004] However, with the increasing complexity and diversity of interference objects in target detection scenarios, collecting a large number of difficult samples that match interference scenarios is very costly. Therefore, existing techniques generally use manual image editing operations based on positive samples from non-interference scenes to generate negative samples, and then use these generated samples to conduct targeted training on the target detection model. Due to the limited operating space of image editing technology and its reliance on manual labor, this approach cannot generate a large number of diverse negative samples. Furthermore, since the generated negative samples are still significantly different from the negative samples in the actual scene, it still cannot effectively solve the problem of false detection of interference objects and effectively improve the accuracy of target detection. Summary of the Invention
[0005] In order to overcome the above-mentioned defects, the sample data generation method, target detection method and equipment proposed in this application can solve the problem of false detection caused by interference in specific scenarios and effectively improve the accuracy of target detection.
[0006] In a first aspect, the present application provides a method for generating sample data, comprising:
[0007] Obtain interference source image and original sample image;
[0008] Extracting a foreground image containing an interference object based on the interference source image, and synthesizing a first sample image based on the foreground image and the original sample image;
[0009] Based on the original sample image, the information including the characteristics of the interference object, and the information including the generation location of the interference object, using the trained image generation model to generate a second sample image;
[0010] Sample data for target detection is obtained based on the first sample image and the second sample image.
[0011] In one technical solution, extracting a foreground image containing an interference object based on the interference source image and synthesizing a first sample image based on the foreground image and the original sample image specifically include: performing foreground target segmentation on the interference source image to obtain an interference object mask image; taking the original sample image as the background image, fusing the interference object mask image and the background image to synthesize the first sample image.
[0012] Furthermore, the fusion processing of the interference mask image and the background image to synthesize the first sample image is specifically: determining an area in the background image where no detection target exists as a mapping area, and mapping the interference mask image to the mapping area in the background image to synthesize the first sample image.
[0013] Furthermore, before synthesizing to obtain the first sample image, the method further includes: performing harmonization processing on the fused image to output the first sample image.
[0014] In one technical solution, the information containing the characteristics of the interference object is specifically a description text of the interference object, and the information containing the generation location of the interference object is specifically a description text of the generation location; the input of the image generation model includes text input and picture input;
[0015] The method of generating the second sample image based on the original sample image, the information containing the characteristics of the interference object, and the information containing the generation location of the interference object using the trained image generation model is specifically as follows: the original sample image is used as a picture input, the interference object description text and the generation location description text are used as text input, and the second sample image is generated using the image generation model.
[0016] In one technical solution, the information containing the characteristics of the interference object is specifically the foreground image containing the interference object, and the information containing the generation location of the interference object is specifically the generation location description text; the input of the image generation model includes text input and picture input;
[0017] The method of generating the second sample image based on the original sample image, the information containing the features of the interference object, and the information containing the generation location of the interference object using the trained image generation model is specifically as follows: the original sample image and the foreground image containing the interference object are input as pictures, the generation location description text is input as text, and the second sample image is generated using the image generation model.
[0018] In one technical solution, obtaining the interference source image specifically includes obtaining the interference source image containing the interference object from network data.
[0019] In a second aspect, the present application provides a target detection method, the method comprising:
[0020] Obtain the image to be detected;
[0021] Inputting the image to be detected into a pre-trained target detection model, and obtaining a target detection result through the output of the target detection model;
[0022] The target detection model is trained using sample data generated by the above-mentioned sample data generation method.
[0023] In a third aspect, the present application provides a smart device, comprising: at least one processor; and a memory communicatively connected to the at least one processor;
[0024] The memory stores a computer program, and when the computer program is executed by the at least one processor, the target detection method is implemented.
[0025] In a fourth aspect, the present application provides a computer-readable storage medium storing a plurality of program codes, wherein the program codes are suitable for being loaded and run by a processor to execute the method described in any one of the technical solutions of the above-mentioned sample data generation method or target detection method.
[0026] One or more of the above-mentioned technical solutions of the present application have at least one or more of the following beneficial effects: in the sample data generation method provided by the technical solution of the present application, a first sample image is synthesized based on the foreground image containing the interference object and the original sample image, and a second sample image is generated using a trained image generation model based on the original sample image, information containing the characteristics of the interference object, and information containing the generation position of the interference object; and sample data for target detection is obtained based on the first sample image and the second sample image. By combining the two methods of image synthesis and image generation, sample data with diversity is obtained to address the problem of false detection caused by the diversity of interference objects, and the training of the target detection model is completed based on the diverse sample data, thereby improving the accuracy of target detection performed using the target detection model. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The disclosure of this application will be more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application. Furthermore, similar numbers in the figures represent similar components, where:
[0028] FIG1 is a schematic diagram showing the principle of a sample data generation method capable of generating diverse negative sample data according to an embodiment of the present application;
[0029] FIG2 is a flow chart showing the main steps of a method for generating sample data according to an embodiment of the present application;
[0030] FIG3 is a flow chart of a method for generating sample data for intelligent cockpit target detection according to an embodiment of the present application;
[0031] FIG4 is a schematic diagram showing the effect of obtaining a composite image by synthesizing a doll image and a cockpit image according to an embodiment of the present application;
[0032] FIG5 is a schematic diagram showing the effect of obtaining an image through an image generation method based on a doll image and a cockpit image in an embodiment of the present application.
[0033] FIG6 is a structural block diagram of a sample data generating device according to an embodiment of the present application. DETAILED DESCRIPTION
[0034] Some embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.
[0035] In the description of this application, "module" and "processor" may include hardware, software, or a combination of both. A module may include hardware circuitry, various suitable sensors, communication ports, and memory. It may also include software components, such as program code, or a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. A processor has data and / or signal processing capabilities. A processor may be implemented in software, hardware, or a combination of both. Non-transitory computer-readable storage media include any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" refers to all possible combinations of A and B, such as only A, only B, or both A and B. The terms "at least one of A or B" or "at least one of A and B" have similar meanings to "A and / or B" and may include only A, only B, or both A and B. The singular forms "a" and "the" may also include the plural forms.
[0036] The following describes the embodiments of the present application in conjunction with the drawings in the specification. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. In addition, the embodiments and features in the embodiments of the present application can be combined with each other if there is no conflict.
[0037] In visual detection tasks, due to the interference objects that are easily confused with the target objects in specific scenes, the detection accuracy of the target detection model will be greatly reduced. Therefore, if a large number of interference scene images can be obtained as negative sample data, the target detection model can be trained using these negative sample data to improve the detection accuracy of the model. However, in actual applications, the interference objects in the application scenes are diverse. Therefore, in order to obtain diverse negative sample data, the present application provides a sample data generation method that can generate diverse negative sample data. The principle of this method is shown in Figure 1. Based on the interference source image and the original sample image (such as the collected application scene background image), diverse negative sample data are generated through image synthesis and image generation. The target detection model is trained based on the diverse negative sample data thus generated, which can effectively improve the target detection accuracy.
[0038] FIG2 shows a method for generating sample data according to an embodiment of the present application. The method mainly includes the following steps:
[0039] Step S11: Acquire the interference source image and the original sample image;
[0040] Specifically in this embodiment, for the diverse interference sources that appear in actual application scenarios, a large number of various interference source images can be obtained from the network data. Therefore, a specific implementation method for obtaining the interference source image is, including but not limited to, obtaining the interference source image containing the interference object from the network data by crawling.
[0041] The original sample images obtained are specifically images of actual application scenes acquired by image acquisition devices such as sensors, for example, images of a smart cockpit acquired by an infrared camera.
[0042] Step S12: extracting a foreground image containing the interference object based on the interference source image, and synthesizing a first sample image based on the foreground image and the original sample image;
[0043] Specifically in this embodiment, this step mainly includes the following steps S121 and S122:
[0044] Step S121: performing foreground object segmentation on the interference source image to obtain an interference object mask image.
[0045] Specifically, in this embodiment, the interference source image is segmented into foreground objects based on an existing object segmentation algorithm, such as a trained large language segmentation model, to obtain an interference mask image. For example, the large language segmentation model may specifically adopt a GroundSAM model.
[0046] Step S122: using the original sample image as a background image, fusing the interference mask image and the background image to synthesize the first sample image.
[0047] Specifically, in this embodiment, the original sample image is used as the background image, an area of the background image where the detection target is not present is determined as a mapping area, and the interference mask image is mapped onto the mapping area of the background image to synthesize the first sample image. For example, in an application scenario of in-cabin person detection, the detection target is a person, and the interference object is a decorative object such as a doll inside the vehicle.
[0048] Specifically, an image processing toolkit (such as OpenCV) may be used to implement mapping the interference mask image to the mapping area in the background image to synthesize the first sample image.
[0049] Furthermore, for the problem of inconsistent foreground and background styles, the foreground style can be made consistent with the background by performing image harmonization processing. Accordingly, before synthesizing the first sample image, the above-mentioned process may also include: performing harmonization processing on the fused image to obtain the first sample image.
[0050] Specifically, the existing image harmonization model can be used to implement the harmonization processing of the image.
[0051] Step S13: Based on the original sample image, the information including the characteristics of the interference object, and the information including the generation location of the interference object, a second sample image is generated using the trained image generation model;
[0052] The image generation model in this embodiment can be obtained by using a neural network model structure and model training method for generating images in the prior art. For example, a pre-trained neural network model is trained based on a sample training set to obtain a trained image generation model. This embodiment does not involve improvements to the neural network model. It mainly obtains the model output image by designing the model input, and then obtains the second sample image. Specifically, the input of the image generation model in this embodiment can include at least text input and / or image input; for example, the image generation model can adopt the stable diffusion model (Stable Diffusion) in the prior art, or a diffusion model combined with a control network (ControlNet).
[0053] In this embodiment, the method of using the trained image generation model to generate the second sample image includes but is not limited to the following two specific implementations:
[0054] First, the information containing the characteristics of the interference object and the information containing the generation location of the interference object are both customizable text content, for example, the information containing the characteristics of the interference object is specifically the interference object description text, and the information containing the generation location of the interference object is specifically the generation location description text;
[0055] The method of generating the second sample image based on the original sample image, the information containing the characteristics of the interference object, and the information containing the generation location of the interference object using the trained image generation model is specifically as follows: the original sample image is used as a picture input, the interference object description text and the generation location description text are used as text input, and the second sample image is generated using the image generation model.
[0056] Secondly, the information containing the characteristics of the interference object is specifically the foreground image containing the interference object, and the information containing the generation location of the interference object is specifically the generation location description text; the input of the image generation model includes text input and picture input;
[0057] The method of generating the second sample image based on the original sample image, the information containing the features of the interference object, and the information containing the generation location of the interference object using the trained image generation model is specifically as follows: the original sample image and the foreground image containing the interference object are input as pictures, the generation location description text is input as text, and the second sample image is generated using the image generation model.
[0058] It is understandable that the original sample image can also be described with customized text content. Accordingly, this step can also be designed as: the original sample image description text, the interference object description text and the generation location description text are used as text inputs of the image generation model, and the second sample image is obtained using the output of the image generation model.
[0059] Step S14: obtaining sample data for target detection based on the first sample image and the second sample image.
[0060] Optionally, target detection is generally implemented using a target detection model. The sample data used for target detection in this step specifically refers to sample data used to train the target detection model. The sample data generally includes positive sample data and negative sample data. For example, in an application scenario of detecting people in a cab, the target is a person, and the interference objects are decorative objects such as dolls in the car. The positive sample data is an image containing the target, and the negative sample data is an image containing only the interference object or an image containing both the target and the interference object.
[0061] One implementation of this step is to use the first sample image and the second sample image as negative sample data for training the target detection model.
[0062] Another implementation of this step is: fusing the first sample image and the second sample image to obtain a third sample image; and using at least one sample image among the first sample image, the second sample image, and the third sample image as a negative sample image for training the target detection model.
[0063] The sample data generated by the implementation method of the present application provides diverse negative sample data for the training of the target detection model, thereby solving the problem of easy false detection caused by the diversity of interferers, and can effectively improve the detection accuracy of the target detection model.
[0064] In a specific application scenario of target detection, such as smart cockpit visual detection, smart cockpit target detection usually uses various sensors (color cameras, infrared cameras) to collect image data and uses target detection algorithms to detect people in the car. However, in practice, due to human factors, there are often interior decorations such as dolls and pillows in the car. Some of these decorations are similar to the human body, causing the target detection model to mistakenly detect doll decorations as people, which brings challenges to the smart cockpit visual detection task. As shown in Figure 3, the embodiment of the present application provides a sample data generation method process specifically applied to smart cockpit target detection based on the sample data generation principle shown in Figure 1.
[0065] In this embodiment, the interference source image is a puppet image crawled from online data, and the original sample image is a cockpit image captured by a sensor. The image synthesis process based on the puppet and cockpit images includes foreground extraction, foreground and background fusion, and image harmonization. The image generation process based on the puppet and cockpit images includes designing a prompt, designing a generation location, and generating an image based on the prompt or a reference image. Through image synthesis and generation, synthetic training data is obtained for target detection.
[0066] It can be understood that in the above-mentioned image synthesis process, foreground extraction is performed from the doll image through target segmentation to extract a doll mask image, and the extracted doll mask image is used as the foreground and the cockpit image is used as the background to perform foreground and background fusion. In order to achieve the effect of consistency in the style of the foreground and background, image harmonization processing is further used to obtain a synthesized image.
[0067] In the embodiment of the present application, a schematic diagram of an exemplary effect of obtaining a composite image by image synthesis based on the doll image and the cockpit image is shown in FIG4 .
[0068] It is understood that the above-described image generation process can be directly implemented using existing image generation models (such as Stable Diffusion). By designing the prompt (the model's text input) and the generation location, a variety of description information and generation location information for the doll image can be obtained. One image generation method implemented based on the image generation model is to use the cabin image as input and utilize the image generation model to output negative sample data based on the prompt and generation location information. Another image generation method implemented based on the image generation model is to use the cabin image as input and utilize the image generation model to output negative sample data based on a reference image and generation location information. The reference image can be a doll foreground image extracted from an image containing the doll. In practical applications, any coordinate information within the seat position area in the cabin image is generally designed as the generation location information input to the image generation model.
[0069] In the embodiment of the present application, a schematic diagram of an exemplary effect of generating an image based on a doll image and a cockpit image through an image generation method is shown in FIG5 .
[0070] This embodiment of the present application provides an effective method for addressing the issue of false detection of puppets in cockpit object detection. This method comprehensively and effectively considers various scenarios and situations within the cockpit. While maintaining the cabin scene distribution, it uses various image data processing methods to generate negative sample data containing puppets in diverse locations and types, thereby reducing false detections caused by puppet data and improving the accuracy of in-cabin object detection.
[0071] In summary, the sample data generation method provided in this application can be adapted to various target detection application scenarios where diverse interferences are likely to occur, and has good versatility.
[0072] Another aspect of the present application provides a sample data generating device.
[0073] 6 , which is a block diagram of the main structure of a sample data generating device 500 according to an embodiment of the present application. The device mainly includes:
[0074] Image acquisition module 501: used to acquire interference source images and original sample images.
[0075] The interference source image may be an interference source image containing interference objects that is crawled from network data by a crawler; and the original sample image may be an image collected and acquired by a sensor in an actual application scenario.
[0076] Image synthesis module 502: configured to extract a foreground image containing an interference object based on the interference source image, and synthesize a first sample image based on the foreground image and the original sample image.
[0077] Image generation module 503: used to generate a second sample image based on the original sample image, information containing the characteristics of the interference object, and information containing the generation position of the interference object using the trained image generation model.
[0078] The sample data generating module 504 is configured to obtain sample data for target detection based on the first sample image and the second sample image. For example, the first sample image and the second sample image are simultaneously used as generated sample data.
[0079] For ease of explanation, the introduction to the above-mentioned sample data generating device only shows the part related to the embodiment of the present application. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application.
[0080] It should be understood that since the configuration of each module is merely for the purpose of illustrating the functional units of this application, the physical devices corresponding to these modules may be the processor itself, or a portion of the software in the processor, a portion of the hardware, or a combination of software and hardware. Therefore, the number of modules in the figure is merely illustrative.
[0081] Those skilled in the art will appreciate that the various modules in the system can be adaptively split or merged. Such splitting or merging of specific modules will not cause the technical solution to deviate from the principles of this application. Therefore, the technical solutions after splitting or merging will fall within the scope of protection of this application.
[0082] Another aspect of the present application further provides a target detection method. First, based on the above-mentioned sample data generation device, sample data can be obtained. The obtained sample data is used as negative sample data to train a target detection model. Then, target detection is performed based on the trained target detection model. The target detection method provided in this embodiment mainly includes the following steps:
[0083] Step S21: Acquire the image to be detected;
[0084] Step S22: inputting the image to be detected into a pre-trained target detection model, and obtaining a target detection result through the output of the target detection model;
[0085] It should be noted that the target detection model is trained based on the sample data generated by the sample data generation method provided in the embodiment of the method of this application.
[0086] Another aspect of the present application provides an intelligent device, which may include at least one processor and a memory communicatively connected to the at least one processor. The memory stores a computer program that, when executed by the at least one processor, implements the target detection method described in the above embodiment. The intelligent device described in the present application may include a driving device, a smart car, a robot, or other devices.
[0087] In some embodiments of the present application, the smart device further includes at least one sensor for sensing information. The sensor is communicatively connected to any type of processor mentioned in the present application. Optionally, the smart device further includes a smart cockpit system, which is an integrated in-vehicle digital platform integrating multiple IT and artificial intelligence technologies to provide the driver with an intelligent experience. The processor communicates with the sensor and / or the smart cockpit system to perform the target detection method described in the above embodiment.
[0088] Furthermore, the present application also provides a computer device.
[0089] In an embodiment of a computer device according to the present application, the computer device primarily includes a processor and a storage device. The storage device may be configured to store a program for executing the sample data generation method of the above-described method embodiment, and the processor may be configured to execute the program in the storage device, including but not limited to a program for executing the sample data generation method of the above-described method embodiment. For ease of illustration, only the portions relevant to the embodiments of the present application are shown. For specific technical details not disclosed, please refer to the method section of the embodiments of the present application.
[0090] In the embodiment of the present application, the computer device may be a control device device formed by various electronic devices. In some possible implementations, the computer device may include multiple storage devices and multiple processors. The program for executing the sample data generation method of the above method embodiment can be divided into multiple subroutines, and each subroutine can be loaded and run by the processor to execute different steps of the sample data generation method of the above method embodiment. Specifically, each subroutine can be stored in different storage devices respectively, and each processor can be configured to execute the program in one or more storage devices to jointly implement the sample data generation method of the above method embodiment, that is, each processor executes different steps of the sample data generation method of the above method embodiment respectively to jointly implement the sample data generation method of the above method embodiment.
[0091] The aforementioned multiple processors may be processors deployed on the same device. For example, the aforementioned computer device may be a high-performance device composed of multiple processors, and the aforementioned multiple processors may be processors configured on the high-performance device. Furthermore, the aforementioned multiple processors may also be processors deployed on different devices. For example, the aforementioned computer device may be a server cluster, and the aforementioned multiple processors may be processors on different servers in the server cluster.
[0092] Furthermore, the present application also provides a computer-readable storage medium.
[0093] In a computer-readable storage medium embodiment according to the present application, the computer-readable storage medium can be configured to store a program for executing the sample data generation method or target detection method of the above-mentioned method embodiment, and the program can be loaded and run by the processor to implement the above-mentioned sample data generation method or target detection method. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method section of the embodiment of the present application. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiment of the present application is a non-transitory computer-readable storage medium.
[0094] It will be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment of the present application can also be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium, etc. that can carry the computer program code.
[0095] The relevant user personal information that may be involved in the various embodiments of this application is strictly in accordance with the requirements of laws and regulations, following the principles of legality, legitimacy and necessity, and based on the reasonable purposes of business scenarios, to process the personal information that users actively provide during the use of products / services or generated due to the use of products / services, as well as the personal information obtained with the user's authorization.
[0096] The user personal information processed in this application will vary depending on the specific product / service scenario and will be based on the specific scenario in which the user uses the product / service. This may involve the user's account information, device information, driving information, vehicle information, or other related information. The applicant will treat the user's personal information and its processing with a high degree of diligence.
[0097] This application attaches great importance to the security of user personal information and has taken reasonable and feasible security protection measures that comply with industry standards to protect user information and prevent personal information from being accessed, disclosed, used, modified, damaged or lost without authorization.
[0098] Thus far, the technical solutions of the present application have been described in conjunction with the embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of the present application is obviously not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present application.
Claims
1. A method for generating sample data, characterized in that: The method comprises: Obtain interference source image and original sample image; Extracting a foreground image containing interference objects based on the interference source image, and synthesizing a first sample image based on the foreground image and the original sample image; Based on the original sample image, the information containing the characteristics of the interferer, and the information containing the generation position of the interferer, using the trained image generation model to generate a second sample image; Sample data for target detection is obtained based on the first sample image and the second sample image.
2. The method according to claim 1, characterized in that: The extracting of a foreground image containing interference objects based on the interference source image, and synthesizing a first sample image based on the foreground image and the original sample image specifically includes: Performing foreground target segmentation on the interference source image to obtain an interference object mask image; The original sample image is used as a background image, and the interference mask image and the background image are fused to synthesize the first sample image.
3. The method according to claim 2, characterized in that The step of fusing the interference mask image and the background image to synthesize the first sample image is specifically as follows: An area in the background image where no detection target exists is determined as a mapping area, and the interference mask image is mapped to the mapping area in the background image to synthesize the first sample image.
4. The method according to claim 2, characterized in that: Before synthesizing to obtain the first sample image, the method further includes: performing harmonization processing on the fused image to output the first sample image.
5. The method according to claim 1, characterized in that The information containing the characteristics of the interference object is specifically the interference object description text, and the information containing the generation location of the interference object is specifically the generation location description text; the input of the image generation model includes text input and picture input; The method of generating the second sample image based on the original sample image, the information containing the characteristics of the interference object and the information containing the generation location of the interference object by using the trained image generation model is specifically as follows: the original sample image is used as a picture input, the interference object description text and the generation location description text are used as text input, and the second sample image is generated by using the image generation model.
6. The method according to claim 1, characterized in that The information containing the characteristics of the interference object is specifically the foreground image containing the interference object, and the information containing the generation location of the interference object is specifically the generation location description text; the input of the image generation model includes text input and picture input; The method of generating the second sample image based on the original sample image, the information containing the features of the interference object and the information containing the generation location of the interference object by using the trained image generation model is specifically as follows: the original sample image and the foreground image containing the interference object are used as picture inputs, the generation location description text is used as text input, and the second sample image is generated by using the image generation model.
7. The method according to claim 1, characterized in that The obtaining of the interference source image specifically includes: obtaining the interference source image containing the interference object from the network data.
8. A target detection method, characterized in that: The method comprises: Acquire the image to be detected; Inputting the image to be detected into a pre-trained target detection model, and obtaining a target detection result through the output of the target detection model; The target detection model is trained by using sample data generated by the sample data generation method described in any one of claims 1 to 7.
9. A smart device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein, a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the target detection method according to claim 8 is implemented.
10. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and executed by a processor to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Image detection method and device, electronic equipment and medium
CN113065607A
Official seal positive and negative sample generation method, official seal verification method and terminal
CN113505841A
Physical world confrontation sample generation method and device, electronic equipment and storage medium
CN114005168A
Sample data generation method, target detection method and equipment
CN117726904A
Data generation device, detection device, and program
JP2022017098A
Cited By
Food material data generation method and device based on image processing, storage medium and electronic device
CN121482526A
Training data generation method and system for image defect detection
CN121788532A