Object Detection Method, System and Device Based on Image Generation Model

By generating a diverse training image and fine-tuning the model, the problem of lack of training images is solved, and the performance and accuracy of the object detection model is improved, especially in detection capabilities in occlusions and complex scenarios.

CN119904625BActive Publication Date: 2025-07-25HANGZHOU EZVIZ SOFTWARE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510377170.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-25
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The prior art is difficult to acquire sufficient training images containing dangerous animals, resulting in the impact of the performance and accuracy of the target detection model.

Method used

By generating sample images of the same object in different scenarios based on the image generation model, fine-tuning the current image generation model, using the masked area to generate training images, combining the segmentation model to automatically generate contour labels and area position labels, and training the object detection model.

Benefits of technology

The training image samples are enriched, and the performance and accuracy of the object detection model are improved, especially in detection capabilities with occlusions and complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904625B_ABST
    Figure CN119904625B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a target detection method, system and device based on an image generation model. In this embodiment, by adopting the method of dynamically optimizing the image generation model, it can ensure that the finally obtained image generation model is optimal. Furthermore, it can ensure that the training images generated by the optimal image generation model for training the target detection model are optimal, so as to realize training the target detection model with the optimal training images generated by the optimal image generation model. This enriches the training images for training the detection model and avoids affecting the performance of the detection model and the accuracy of target detection due to the lack of training images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to a target detection method, system, and device based on an image generation model. Background Art

[0002] In specific applications, a detection model is often used to detect an image to detect a target. For example, a dangerous animal detection model is used to detect whether there are dangerous animals such as poisonous snakes and tigers in the image.

[0003] However, in specific implementations, due to the particularity of the target, such as a dangerous animal, it is often difficult to directly collect many training images containing dangerous animals in real scenarios. For example, it is difficult to collect images such as a tiger in a garden or a bear in a gas station. This leads to a lack of training images for training the detection model. The lack of training images will affect the performance of the finally trained detection model and also affect the accuracy of target detection. Summary of the Invention

[0004] In view of this, embodiments of this application provide a target detection method, system, and device based on an image generation model to dynamically generate training images and avoid affecting the performance of the detection model, the accuracy of target detection, etc. due to the difficulty of collecting training images containing targets such as dangerous animals in real scenarios.

[0005] Embodiments of this application provide a target detection method based on an image generation model. The method includes:

[0006] Generating an object model corresponding to the object based on sample images of the same object in different scenarios generated by using different image generation methods and the current image generation model;

[0007] Fine-tuning the current image generation model based on object models corresponding to different objects;

[0008] Determining whether the current image generation model meets the model iteration requirement. If not, return to the step of generating the object model corresponding to the object; if so, then:

[0009] Generating training images for training a target detection model based on the current image generation model and by means of mask regions corresponding to image regions in each training scene image; the mask region is used to indicate the region where the object is located;

[0010] Automatically generating a contour label of the object and a region position label of the region where the object is located by using the mask regions corresponding to the object regions in each training image and in combination with a segmentation model;

[0011] Training a target detection model by using each training image and the contour label of the object and the region position label of the region where the object is located in the training image;

[0012] Use the trained object detection model to detect objects in an image.

[0013] An embodiment of the present application further provides an object detection device, which includes:

[0014] A first generation module, configured to generate an object model corresponding to the object based on sample images of the same object in different scenarios generated by using different image generation methods and the current image generation model;

[0015] A fine-tuning module, configured to fine-tune the current image generation model based on object models corresponding to different objects;

[0016] A second generation module, configured to determine whether the current image generation model meets the model iteration requirement. If not, return to the step of generating the object model corresponding to the object; if so, then:

[0017] Generate a training image for training the object detection model based on the current image generation model and by means of mask regions corresponding to image regions in each training scenario image; the mask region is used to indicate the region where the object is located;

[0018] A third generation module, configured to automatically generate a contour label of the object and a region position label where the object is located by using the mask regions corresponding to the object regions in each training image and in combination with a segmentation model;

[0019] A training module, configured to train the object detection model by using each training image and the contour label of the object and the region position label where the object is located in the training image; the trained object detection model is used to detect objects in an image.

[0020] An embodiment of the present application further provides an object detection system, which includes:

[0021] An image acquisition device, configured to acquire a scene image;

[0022] A model device end, which is a processing device for running an image generation model set locally or in the cloud, obtains the scene image acquired by the image acquisition device, and executes the steps in the above method based on the scene image.

[0023] An embodiment of the present application further provides an electronic device, including: a processor and a memory for storing computer program instructions, and the computer program instructions, when run by the processor, cause the processor to execute the steps in the above method.

[0024] An embodiment of the present application further provides a machine-readable storage medium, which stores computer program instructions, and when the computer program instructions are executed, the steps in the above method can be implemented.

[0025] As can be seen from the above technical solutions, in this embodiment, first, with the sample images of the same object generated by different image generation methods in different scenarios and the current image generation model, an object model corresponding to the object is generated. Then, based on the object models corresponding to different objects, the current image generation model is fine-tuned to optimize the current image generation model. If the optimized current image generation model does not meet the model iteration requirements, then the step of generating the object model corresponding to the object is returned; if the current image generation model meets the model iteration requirements, then based on the current image generation model and with the help of the mask regions corresponding to the image regions in each training scenario image, training images for training the target detection model are generated. This way of dynamically optimizing the image generation model can ensure that the finally obtained image generation model is optimal, and further ensure that the training images generated by the optimal image generation model for training the target detection model are optimal, so as to realize training the target detection model with the optimal training images generated by the optimal image generation model. This enriches the training images for training the detection model and avoids affecting the performance of the detection model and the accuracy of target detection due to the lack of training images. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a schematic flowchart of the method provided by the embodiment of the present application;

[0027] Figure 2a is a schematic diagram of generating image samples provided by the embodiment of the present application;

[0028] Figure 2b is a schematic diagram of generating image samples provided by the embodiment of the present application;

[0029] Figure 2c is a schematic diagram of generating image samples provided by the embodiment of the present application;

[0030] Figure 2d is a schematic diagram of generating image samples provided by the embodiment of the present application;

[0031] Figure 2e is a schematic diagram of generating image samples provided by the embodiment of the present application;

[0032] Figure 3 is a schematic diagram of the device structure provided by the embodiment of the present application;

[0033] Figure 4 is a schematic diagram of the system structure provided by the embodiment of the present application;

[0034] Figure 5 is a schematic diagram of the electronic device structure provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] To enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application and to make the above-mentioned objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0036] It should be noted that the object detection method provided by the embodiments of the present application is applicable to multiple fields such as the detection of dangerous animals and the detection of target faces. For the sake of convenience of elaboration, the method provided by the present application will be described by taking the detection of dangerous animals as an example.

[0037] See Figure 1 , Figure 1 the schematic flowchart of the method provided by the embodiments of the present application. This method is applied to an electronic device, which can be a server, and the present application does not specifically limit it.

[0038] As Figure 1 shown, this process may include the following steps:

[0039] S101, based on the sample images of the same object in different scenarios generated by using different image generation methods and the current image generation model, generate an object model corresponding to the object.

[0040] Here, the object is a designated dangerous animal, such as a wild boar, a venomous snake, etc., or the object is a designated occluder, such as a table, a chair, etc. The sample images of the same object in different scenarios generated by using different image generation methods can effectively enrich the sample images and ensure the diversity of the sample images, so as to optimize the current image generation model in a better direction.

[0041] In this embodiment, the current image generation model is an initial image generation model or an image generation model obtained after fine-tuning the previous time by using each object model. As an embodiment, the initial image generation model is obtained by training a neural network model (such as a Diffusion model) with the original images including dangerous animals directly collected by an image acquisition device (such as a camera) in different scenarios.

[0042] In this embodiment, the specific implementation manners of the sample images of the same object in different scenarios generated by using different image generation methods and the generation of the object model corresponding to the object will be described later and will not be elaborated here.

[0043] S102, based on the object models corresponding to different objects, fine-tune the current image generation model.

[0044] In this embodiment, the specific implementation manner of this step S102 will also be described later and will not be elaborated here.

[0045] S103. Determine whether the current image generation model meets the model iteration requirements.

[0046] If the result of executing step S103 is negative, return to step S101 above. If the result of executing step S103 is positive, execute the following step S104.

[0047] Here, the current image generation model in step S103 is different from the current image generation models mentioned in steps S101 and S102, and is fine-tuned compared to the previous current image generation model.

[0048] As for the specific implementation manner of this step S103, it will be elaborated later and will not be repeated here.

[0049] S104. Based on the current image generation model and by means of the mask regions corresponding to the image regions in each training scenario image, generate training images for training the object detection model.

[0050] In this embodiment, the mask region is obtained by marking the image region in the scenario image or smearing the image region, and the mask region is used to indicate the region where the object is located. The current image generation model here is an image generation model that meets the model iteration requirements.

[0051] As for the specific implementation manner of this step S104, it will be elaborated later and will not be repeated here.

[0052] S105. Utilize the mask regions corresponding to the object regions in each training image, and in combination with the segmentation model, automatically generate the contour labels of the object and the region position labels where the object is located.

[0053] In this embodiment, input the training image into the trained segmentation model to automatically generate the contour labels of the object and the region position labels where the object is located. This way of automatically generating contour labels and region position labels can greatly reduce the workload and cost of manual annotation and improve the generation efficiency of the subsequent object detection model.

[0054] S106. Utilize each training image and the contour labels of the object and the region position labels where the object is located in the training image to train the object detection model.

[0055] In this embodiment, a large number of training images are generated according to the current image generation model that meets the model iteration requirements. Each training image is input into a neural network model, such as a CNN convolutional model, to obtain the predicted contour and predicted position information output by the model. The loss values are calculated by using the predicted contour and predicted position information and the contour label and position label of the training sample respectively. The parameters of the neural network model are adjusted according to the loss values until the neural network model meets the stop iteration training condition. The neural network model that meets the stop iteration training condition is determined as the target detection model. The trained target detection model is used to detect objects in images.

[0056] So far, the Figure 1 shown process is completed.

[0057] Through Figure 1 the shown process, it can be seen that first, sample images of the same object generated by different image generation methods in different scenarios and the current image generation model are used to generate an object model corresponding to the object. Then, the current image generation model is fine-tuned based on the object models corresponding to different objects to optimize the current image generation model. If the optimized current image generation model does not meet the model iteration requirements, then return to the step of generating the object model corresponding to the object above; if the current image generation model meets the model iteration requirements, then based on the current image generation model and by means of the mask regions corresponding to the image regions in each training scenario image, training images for training the target detection model are generated. This way of dynamically optimizing the image generation model can ensure that the finally obtained image generation model is optimal, and further ensure that the training images generated by the optimal image generation model for training the target detection model are optimal, so as to realize training the target detection model by using the optimal training images generated by the optimal image generation model. This enriches the training images for training the detection model and avoids affecting the performance of the detection model and the accuracy of target detection due to the lack of training images.

[0058] Next, in combination with Figures 2a to 2e an example description is given of the sample images of the same object generated by different image generation methods in different scenarios in step S101 above.

[0059] The first sample image generation method is:

[0060] For any object, the scene image in any scene and the text description information of the object are input into the current image generation model, so that the current image generation model can obtain the sample image of the object in that scene. Here, the scene image can be a real scene image collected by an image acquisition device in different real scenes, such as real scenes like a courtyard, a playground, a desert, etc. The text description information of the object is used to describe generating the object within the mask region corresponding to the image region in the scene image.

[0061] For example, as Figure 2a shown, the object is a wild boar, the scene image is an image under a courtyard scene, and the text description information of this object is Figure 2a "It is necessary to generate a wild boar in the masked area in the figure" shown in. Inputting this scene image and the text description information of this object into the current image generation model, a sample image of "There is a wild boar in the garden" is obtained.

[0062] Another example, as Figure 2a shown, the object is a viper, the scene image is an image under a courtyard scene, and the text description information of this object is Figure 2a "It is necessary to generate a viper in the masked area in the figure" shown in. Inputting this scene image and the text description information of this object into the current image generation model, a sample image of "There is a viper in the garden" is obtained.

[0063] Another example, as Figure 2b shown, the object is a table, the scene image is an image under a courtyard scene, and the text description information of this object is Figure 2b "It is necessary to generate a table in the masked area in the figure" shown in. Inputting this scene image and the text description information of this object into the current image generation model, a sample image of "There is a table in the garden" is obtained.

[0064] Another example, as Figure 2b shown, the object is a chair, the scene image is an image under a courtyard scene, and the text description information of this object is Figure 2b "It is necessary to generate a chair in the masked area in the figure" shown in. Inputting this scene image and the text description information of this object into the current image generation model, a sample image of "There is a chair in the garden" is obtained.

[0065] In this way of generating sample images, for any object, by combining the scene images under different real scenes with this object, sample images of this object under different scenes can be obtained, especially images that are difficult to be directly collected (for example, images such as "There is a jaguar in the garden" that are difficult to be collected), thus greatly enriching the diversity of image samples.

[0066] The second way of generating sample images is:

[0067] For any object, input the scene image under any scene and the obtained object image into the current image generation model, so that the current image generation model fuses the object image into the masked area corresponding to the image area in the scene image, and a sample image of this object under this scene is obtained. Here, the object image can be a local cropped image obtained from other images.

[0068] For example, as Figure 2cAs shown, the object image is a "wild boar picture", for example, a partial cut-out of a "wild boar" cut from other images, and the scene image is an image in a courtyard scene. The scene image and the "wild boar picture" are input into the current image generation model to obtain a sample image of "a wild boar in the garden".

[0069] In this sample image generation method, the scene images in different real scenes can also be combined with the object, thus greatly enriching the diversity of image samples.

[0070] The third sample image generation method is:

[0071] Fuse the first image and the second image. The first image and the second image are images in the same scene. The objects in the first image and the second image are different, and the image regions where the different objects are located have an intersection. Preferably, the categories of the objects in the first image and the second image are different, one is a dangerous animal and the other is an occluder. Use the fusion result to obtain the sample image in this scene. In the obtained sample image, there is an occlusion relationship between the object originally in the first image and the object originally in the second image.

[0072] For example, as Figure 2d shown, fuse the first image and the second image obtained in the courtyard scene. The first image has a "wild boar", and the second image has a "table". The image regions of the "wild boar" and the "table" have an intersection. After fusion, a sample image of "the table occludes the wild boar in the courtyard" is obtained.

[0073] Optionally, in specific implementation, the above-mentioned first image and the above-mentioned second image can be generated by the above-mentioned first sample image generation method or the second sample image generation method, and the same scene image is used in the generation process.

[0074] In this sample image generation method, a sample image with a specific occlusion relationship between different objects is constructed, further enriching the diversity of sample images.

[0075] The fourth sample image generation method is:

[0076] Input the depth image and the scene guidance generation condition into the current image generation model, so that the current image generation model generates a sample image in the scene indicated by the scene guidance generation condition based on the contour information and depth information of two objects with an occlusion relationship in the depth image and the scene guidance generation condition; there are two objects that conform to the occlusion relationship in the sample image.

[0077] Here, the depth image can be extracted from the sample image obtained by using the third sample image generation method, or can be extracted from an image directly collected by an image acquisition device that includes two objects with an occlusion relationship.

[0078] It should be noted that the occlusion relationship between the two objects in the generated sample images above can be the same as or different from the occlusion relationship in the previously input depth image.

[0079] For example, as Figure 2e shown, the depth image and the scene guidance generation condition of "the scene where the generated object is located is a courtyard" are input into the current image generation model to generate 3 sample images as shown in Figure 2e . The "table" and "wild boar" in these 3 sample images present 3 different occlusion relationships.

[0080] In this way of generating sample images, the two objects in the generated sample images present more than one occlusion relationship, enriching the diversity of the sample images. Subsequently, using the obtained sample images to fine-tune the image generation model can improve the generation ability of the image generation model.

[0081] The above has introduced in detail the sample images of the same object generated by different image generation methods in different scenes.

[0082] The following elaborates in detail on the generation of the object model corresponding to the object in the above step S101:

[0083] After generating the sample images corresponding to the object using at least one of the above 4 image generation methods, for any object, on the premise that the parameter values corresponding to at least one parameter attribute matched with the object in the current image generation model remain unchanged, use the sample images of the object in different scenes to train the other variable parameter values in the current image generation model, and take the trained current image generation model as the object model.

[0084] For example, if the object is a venomous snake, after obtaining the sample images of the venomous snake in different scenes, the current image generation model has a total of M parameter attributes, and N parameter attributes (N is less than M) are selected from them. The parameter values corresponding to these N parameter attributes remain unchanged, and the sample images of the venomous snake in different scenes are used to train the current image generation model. During the iterative training process, the parameter values corresponding to the remaining M - N parameters change respectively. Iterative training is performed until the stop iteration training condition is met (for example, the number of iterations reaches a set threshold such as 20), and a venomous snake model, denoted as venomous snake lora, is obtained. In the above manner, object models of different objects can be obtained, such as Chinese wild boar lora, Chinese snake lora, North American grizzly bear lora, North American gray wolf lora, Australian kangaroo lora, chair lora, table lora, sundries lora, trash can, and various occluders lora, etc.

[0085] The generation of the object model corresponding to the object in the above step S101 is elaborated in detail above.

[0086] After obtaining the object models of different objects, based on the object models corresponding to different objects, the current image generation model is fine-tuned. The following elaborates on the above step S102 in detail.

[0087] In the specific implementation of step S102, as an embodiment, for any parameter attribute, find the parameter values belonging to the parameter attribute from each object model, and perform a specified operation (here, the specified operation can be operations such as weighted average) on the found parameter values to obtain the target parameter value corresponding to the parameter attribute. Depending on the target parameter value corresponding to the parameter attribute, adjust the current parameter value under the parameter attribute in the current image generation model. For example, replace the current parameter value under the parameter attribute with the target parameter value corresponding to the parameter attribute.

[0088] Through the above fine-tuning, the generation accuracy of the current image generation model when generating each object can be improved, thereby optimizing the accuracy of the image generated by the current graphics generation model.

[0089] After fine-tuning the current image generation model in the above manner, it is necessary to determine whether the current image generation model (that is, the image generation model obtained after executing the above step S102) meets the model iteration requirements. The following elaborates in detail on how to determine whether the current image generation model meets the model iteration requirements:

[0090] In the specific implementation, as an embodiment, when generating sample images of the same object in different scenarios using different image generation methods in the above step S101, the method further includes: for any object, based on the description information of the sample image of the object in each scenario, and based on the current image generation model, generate a test image corresponding to the object by the current image generation model according to the description information. Here, the description information is used to describe the object and the environment where the object is located in the sample image.

[0091] For example, as shown in, Figure 2a the description information corresponding to the first image sample in is "There is a wild boar in the courtyard". Output the description information of the sample image to the current image generation model, so that the current image generation model generates an image according to the scenario and object described by the description information, and determine the generated image as the test image corresponding to the object.

[0092] After obtaining the test images corresponding to each object, the specific implementation method for determining whether the current image generation model meets the model iteration requirements is as follows: Based on the test images corresponding to each object, test whether the current image generation model meets the accuracy requirements. If not, it is determined that the current image generation model does not meet the model iteration requirements. If so, it is determined that the current image generation model meets the model iteration requirements.

[0093] The specific implementation method for testing whether the current image generation model meets the accuracy requirements based on the test images corresponding to each object can be as follows: For each sample image, determine whether the difference degree between the sample image and the test image corresponding to the sample image is less than or equal to the set difference degree threshold. If so, it is determined that the difference degree condition is met; otherwise, it is determined that the difference degree condition is not met. If the number of images that meet the above difference degree condition among the sample images of each object obtained is greater than the set number value, it is determined that the current image generation model meets the accuracy requirements; otherwise, it is determined that the current image generation model does not meet the accuracy requirements.

[0094] The above has elaborated in detail on how to determine whether the current image generation model meets the model iteration requirements.

[0095] After determining that the current image generation model meets the model iteration requirements, training images for training the object detection model will be generated based on the current image generation model and by means of the mask regions corresponding to the image regions in each training scenario image.

[0096] The following elaborates in detail on generating training images for training the object detection model based on the current image generation model and by means of the mask regions corresponding to the image regions in each training scenario image in the above step S103:

[0097] The specific method for generating training samples is at least one of the following methods:

[0098] The first training image generation method is:

[0099] For any training object, input the training scenario image in any training scenario and the text description information of the training object into the current image generation model, so that the current image generation model obtains the training image of the training object in the training scenario. The text description information of the training object is used to describe generating the training object within the mask region corresponding to the image region in the training scenario image.

[0100] The second training image generation method is:

[0101] For any training object, input the training scene image in any training scene and the obtained training object image into the current image generation model, so that the current image generation model fuses the training object image into the mask area corresponding to the image area in the training scene image, and obtains the training image of the training object in the training scene.

[0102] The third training image generation method is as follows:

[0103] Fuse the third image and the fourth image; the third image and the fourth image are images in the same training scene, the training objects in the third image and the fourth image are different, and the image areas where the different training objects are located have an intersection;

[0104] Use the fusion result to obtain the training image in the training scene; among them, in the training image, there is an occlusion relationship between the training object originally existing in the third image and the training object originally existing in the fourth image.

[0105] The fourth training image generation method is as follows:

[0106] Input the depth image and the scene-guided generation condition into the current image generation model, so that the current image generation model generates a sample image in the scene indicated by the scene-guided generation condition based on the contour information and depth information of two objects with an occlusion relationship in the depth image and the scene-guided generation condition; there are two objects with an occlusion relationship in the sample image.

[0107] The above four training image generation methods are similar to the four sample image generation methods mentioned before, and will not be elaborated here.

[0108] Through the above method, diverse training samples can be generated, greatly enriching the training images of the object detection model, and effectively improving the accuracy of the object detection model. And fully considering the influence of occluders and different scenes, using these training samples to train the object detection model enables the object detection model to better learn how to accurately detect dangerous animals in the presence of occluders and complex scenes, thereby improving the detection ability of the object detection model in the presence of occluders and complex scenes, and further improving the detection accuracy of the object detection model.

[0109] The above has elaborated in detail on the generation of training images for training the object detection model.

[0110] After obtaining the training samples through the above method, it is necessary to execute the above steps S105 and S106 to obtain the object detection model.

[0111] After obtaining the target detection model, the method further includes: obtaining the effect evaluation parameters of the target detection model in a specified scenario. If the effect evaluation parameters do not meet the requirements of the set evaluation parameters, then generate optimized images of different training objects in the specified scenario by using the current image generation model when the model iteration requirements are met (the generation method refers to the above four sample image generation methods and will not be elaborated here). Use each optimized image to optimize the target detection model.

[0112] In this embodiment, collect the effect evaluation parameters of the target detection model in the specified scenario, and use the effect evaluation parameters to guide the generation direction of subsequent training images, so as to continuously obtain high-quality training images, and then optimize the target detection model in the direction of improving the detection accuracy of the specified scenario, and improve the adaptability of the target detection model.

[0113] The method provided by the embodiments of the present application has been described above. Next, the device provided by the embodiments of the present application will be described:

[0114] See Figure 3 , Figure 3 which is the structure diagram of the device provided by the embodiments of the present application. This device is applied to an electronic device. As Figure 3 shown, the device may include: a first generation module 301, a fine-tuning module 302, and a second generation module 303, a third generation module 304, and a training module 305.

[0115] The first generation module 301 is used to generate an object model corresponding to the object based on the sample images of the same object in different scenarios generated by using different image generation methods and the current image generation model;

[0116] The fine-tuning module 302 is used to fine-tune the current image generation model based on the object models corresponding to different objects;

[0117] The second generation module 303 is used to determine whether the current image generation model meets the model iteration requirements. If not, return to the step of generating the object model corresponding to the object. If so, then:

[0118] Generate training images for training the target detection model based on the current image generation model and by means of the mask regions corresponding to the image regions in each training scenario image; the mask region is used to indicate the region where the object is located;

[0119] The third generation module 304 is used to automatically generate the contour label of the object and the region position label where the object is located by using the mask regions corresponding to the object regions in each training image and in combination with the segmentation model;

[0120] The training module 305 is used to train an object detection model by using each training image, as well as the contour label of the object in the training image and the region position label of the region where the object is located. The trained object detection model is used to detect the object in the image.

[0121] As an embodiment, generating sample images of the same object in different scenarios by using different image generation methods includes:

[0122] Generating sample images of the object in different scenarios according to at least one of the following image generation methods:

[0123] The first image generation method is:

[0124] For any object, input the scene image in any scenario and the text description information of the object into the current image generation model, so that the current image generation model obtains the sample image of the object in this scenario; the text description information of the object is used to describe generating the object within the mask region corresponding to the image region in the scene image.

[0125] The second image generation method is:

[0126] For any object, input the scene image in any scenario and the obtained object image into the current image generation model, so that the current image generation model fuses the object image into the mask region corresponding to the image region in the scene image to obtain the sample image of the object in this scenario.

[0127] The third image generation method is:

[0128] Fuse the first image and the second image; the first image and the second image are images in the same scenario, the objects in the first image and the second image are different, and the image regions where the different objects are located have an intersection.

[0129] Obtain the sample image of this scenario by using the fusion result; among them, in the sample image, there is an occlusion relationship between the object originally existing in the first image and the object originally existing in the second image.

[0130] The fourth image generation method is:

[0131] Input the depth image and the scene-guided generation condition into the current image generation model, so that the current image generation model generates the sample image of the scenario indicated by the scene-guided generation condition based on the contour information and depth information of two objects with an occlusion relationship in the depth image, and the scene-guided generation condition; there are two objects with an occlusion relationship in the sample image.

[0132] As an example, generating an object model corresponding to an object based on sample images of the same object generated by different image generation methods in different scenarios and the current image generation model includes:

[0133] For any object, on the premise that the parameter values corresponding to at least one parameter attribute matching the object in the current image generation model remain unchanged, use the sample images of the object in different scenarios to train other variable parameter values in the current image generation model, and use the trained current image generation model as the object model.

[0134] As an example, fine-tuning the current image generation model based on object models corresponding to different objects includes:

[0135] For any parameter attribute, find the parameter values belonging to the parameter attribute from each object model, and perform a specified operation on the found parameter values to obtain the target parameter value corresponding to the parameter attribute;

[0136] Depending on the target parameter value corresponding to the parameter attribute, adjust the current parameter value under the parameter attribute in the current image generation model.

[0137] As an example, generating training images for training an object detection model based on the current image generation model and with the aid of mask regions corresponding to image regions in each training scenario image includes:

[0138] For any training object, input the training scenario image in any training scenario and the text description information of the training object into the current image generation model, so that the current image generation model obtains the training image of the training object in the training scenario; the text description information of the training object is used to describe generating the training object within the mask region corresponding to the image region in the training scenario image;

[0139] And / or,

[0140] For any training object, input the training scenario image in any training scenario and the obtained training object image into the current image generation model, so that the current image generation model fuses the training object image into the mask region corresponding to the image region in the training scenario image to obtain the training image of the training object in the training scenario;

[0141] And / or,

[0142] Fuse the third image and the fourth image; the third image and the fourth image are images in the same training scenario, the training objects in the third image and the fourth image are different, and there is an intersection in the image regions where the different training objects are located;

[0143] Obtain a training image in this training scenario using the fusion result; among them, in the training image, there is an occlusion relationship between the training object originally existing in the third image and the training object originally existing in the fourth image;

[0144] and / or,

[0145] Input the depth image and the training scenario guidance generation condition into the current image generation model, so that the current image generation model generates a training image in the training scenario indicated by the training scenario guidance generation condition based on the contour information and depth information of two training objects with an occlusion relationship in the depth image, and the training scenario guidance generation condition; there are two training objects that meet the occlusion relationship in the training image.

[0146] As an embodiment, the object is a designated dangerous animal or a designated occluder.

[0147] As an embodiment, after obtaining the target detection model, the training module is further used for:

[0148] Obtain the effect evaluation parameters of the target detection model in the designated scenario;

[0149] If the effect evaluation parameters do not meet the set evaluation parameter requirements, then generate optimized images of different training objects in this designated scenario by means of the current image generation model when the model iteration requirements are met;

[0150] Optimize the target detection model using each optimized image.

[0151] As an embodiment, generating sample images of the same object in different scenarios using different image generation methods further includes:

[0152] For any object, based on the description information of the sample image of the object in each scenario, and based on the current image generation model, so that the current image generation model generates a test image corresponding to the object according to the description information; the description information is used to describe the object and the environment where the object is located in the sample image.

[0153] As an embodiment, determining whether the current image generation model meets the model iteration requirements includes:

[0154] Test whether the current image generation model meets the accuracy requirements based on the test images corresponding to each object: if not, determine that the current image generation model does not meet the model iteration requirements; if so, determine that the current image generation model meets the model iteration requirements.

[0155] Thus, the Figure 3 structural description of the device shown is completed.

[0156] Next, the system provided by the embodiments of the present application will be described:

[0157] See Figure 4 , Figure 4 which is the system structure diagram provided by the embodiment of the present application. The system includes:

[0158] An image acquisition device for acquiring a scene image;

[0159] A model device end, which is a processing device for running an image generation model set locally or in the cloud, obtains the scene image acquired by the image acquisition device, and executes the steps in the above method based on the scene image.

[0160] Optionally, there is a communication connection between the image acquisition device and the model device end;

[0161] The model device end obtaining the scene image acquired by the image acquisition device includes: receiving the scene image acquired by the image acquisition device sent by the image acquisition device through the communication connection.

[0162] Optionally, the image acquisition device is deployed on the model device end.

[0163] Optionally, the system further includes: a client device;

[0164] The client device is communicatively connected to the model device end, and is used to send the sample image of the object in each scene and the corresponding description information to the model device end through the communication connection, so that the model device end generates a test image corresponding to the object based on the description information of the sample image of the object in each scene by using the current image generation model; the description information is used to describe the object in the sample image and the environment where the object is located; the test image is used to test whether the current image generation model meets the model iteration requirements.

[0165] Optionally, the client device is further used to send the text description information of the object to the model device end, so that the model device end inputs the scene image in any scene and the text description information of the object into the current image generation model, so that the current image generation model obtains the sample image of the object in this scene; the text description information of the object is used to describe that the object is generated in the mask area corresponding to the image area in the scene image; and / or,

[0166] The client device is further used to send the object image to the model device end, so that the model device end inputs the scene image in any scene and the obtained object image into the current image generation model, so that the current image generation model fuses the object image into the mask area corresponding to the image area in the scene image to obtain the sample image of the object in this scene; and / or,

[0167] The image acquisition device is further configured to acquire a depth image;

[0168] The client is further configured to send scene guidance generation conditions to the model device side, so that the model device side inputs the depth image and the scene guidance generation conditions into the current image generation model, and the current image generation model generates a sample image in the scene indicated by the scene guidance generation conditions based on the contour information and depth information of two objects with an occlusion relationship in the depth image and the scene guidance generation conditions; there are two objects with an occlusion relationship in the sample image.

[0169] Next, the hardware structure of the device provided in the embodiments of the present application will be described: Figure 3 as shown in the following:

[0170] Please refer to Figure 5 , Figure 5 , which is the structural diagram of the electronic device provided in the embodiments of the present application. As Figure 5 shown, the hardware structure may include: a processor and a machine-readable storage medium, and the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is configured to execute the machine-executable instructions to implement the method disclosed in the above examples of the present application.

[0171] Based on the same inventive concept as the above method, the embodiments of the present application further provide a machine-readable storage medium, on which a number of computer instructions are stored, and when the computer instructions are executed by a processor, the method disclosed in the above examples of the present application can be implemented.

[0172] Exemplarily, the above machine-readable storage medium can be any electronic, magnetic, optical or other physical storage device, which can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.

[0173] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A target detection method based on an image generation model, characterized in that, The method includes: Generating an object model corresponding to the object based on sample images of the same object in different scenarios generated by different image generation methods and the current image generation model; wherein, one of the different image generation methods is: inputting a depth image and a scene-guided generation condition into the current image generation model, so that the current image generation model generates a sample image in the scene indicated by the scene-guided generation condition based on the contour information and depth information of two objects with an occlusion relationship in the depth image and the scene-guided generation condition; there are two objects conforming to the occlusion relationship in the sample image; Fine-tuning the current image generation model based on the object models corresponding to different objects; Determining whether the current image generation model meets the model iteration requirement, if not, then returning to the step of generating the object model corresponding to the object; if so, then: Generating training images for training the object detection model based on the current image generation model and the mask regions corresponding to the image regions in each training scene image, by means of different training image generation methods; the mask region is used to indicate the region where the object is located; wherein, one of the different training image generation methods includes: inputting a depth image and a training scene-guided generation condition into the current image generation model, so that the current image generation model generates a training image in the training scene indicated by the training scene-guided generation condition based on the contour information and depth information of two training objects with an occlusion relationship in the depth image and the training scene-guided generation condition; there are two objects conforming to the occlusion relationship in the training image; Using the mask regions corresponding to the object regions in each training image and combining with the segmentation model to automatically generate the contour label of the object and the region position label where the object is located; Training the object detection model using each training image and the contour label of the object and the region position label where the object is located in the training image; Using the trained object detection model to detect the object in the image.

2. The method according to claim 1, wherein The different image generation methods further include at least one of the following image generation methods: The first image generation method is: For any object, inputting the scene image in any scene and the text description information of the object into the current image generation model, so that the current image generation model obtains a sample image of the object in the scene; the text description information of the object is used to describe generating the object within the mask region corresponding to the image region in the scene image; The second image generation method is: For any object, inputting the scene image in any scene and the obtained object image into the current image generation model, so that the current image generation model fuses the object image into the mask region corresponding to the image region in the scene image to obtain a sample image of the object in the scene; The third image generation method is: Fusing a first image and a second image; the first image and the second image are images in the same scene, the objects in the first image and the second image are different, and the image regions where the different objects are located have an intersection; Obtain a sample image in this scenario using the fusion result; wherein, in the sample image, there is an occlusion relationship between the object originally existing in the first image and the object originally existing in the second image.

3. The method according to claim 1, characterized in that, Generating the object model corresponding to the object based on the sample images of the same object in different scenarios generated by different image generation methods and the current image generation model includes: For any object, on the premise that the parameter values corresponding to at least one parameter attribute matched with the object in the current image generation model remain unchanged, use the sample images of the object in different scenarios to train the other variable parameter values in the current image generation model, and use the trained current image generation model as the object model.

4. The method according to claim 1 or 3, characterized in that Fine-tuning the current image generation model based on the object models corresponding to different objects includes: For any parameter attribute, find the parameter values belonging to this parameter attribute from each object model, and perform a specified operation on the found parameter values to obtain the target parameter value corresponding to this parameter attribute; Depending on the target parameter value corresponding to this parameter attribute, adjust the current parameter value under this parameter attribute in the current image generation model.

5. The method according to claim 1, wherein The different training image generation methods further include: For any training object, input the training scene image in any training scene and the text description information of the training object into the current image generation model, so that the current image generation model obtains the training image of the training object in this training scene; the text description information of the training object is used to describe generating the training object within the mask region corresponding to the image region in this training scene image; and / or For any training object, input the training scene image in any training scene and the obtained training object image into the current image generation model, so that the current image generation model fuses the training object image into the mask region corresponding to the image region in this training scene image to obtain the training image of the training object in this training scene; and / or Fuse the third image and the fourth image; the third image and the fourth image are images in the same training scene, the training objects in the third image and the fourth image are different, and there is an intersection in the image regions where the different training objects are located; Obtain the training image in this training scene using the fusion result; wherein, in the training image, there is an occlusion relationship between the training object originally existing in the third image and the training object originally existing in the fourth image.

6. The method according to claim 1, wherein The object is a designated dangerous animal or a designated occluder.

7. The method according to claim 1, characterized in that, After obtaining the target detection model, the method further includes: Obtain the effect evaluation parameters of the target detection model in the designated scenario; If the effect evaluation parameters do not meet the requirements of the set evaluation parameters, then generate optimized images of different training objects in this designated scenario by means of the current image generation model when the model iteration requirements are met; Use each optimized image to optimize the target detection model.

8. The method according to claim 1, wherein Generating sample images of the same object in different scenarios using different image generation methods further includes: For any object, based on the description information of the sample image of the object in each scenario, and using the current image generation model to generate a test image corresponding to the object according to the description information; the description information is used to describe the object and the environment where the object is located in the sample image.

9. The method according to claim 1 or 8, characterized in that, The determination of whether the current image generation model meets the model iteration requirements includes: Testing whether the current image generation model meets the accuracy requirements based on the test images corresponding to each object: if not, it is determined that the current image generation model does not meet the model iteration requirements; if so, it is determined that the current image generation model meets the model iteration requirements.

10. A target detection device, characterized in that, The device includes: A first generation module, configured to generate an object model corresponding to the object based on the sample images of the same object in different scenarios generated by using different image generation methods and the current image generation model; wherein, one of the different image generation methods is: inputting a depth image and a scene-guided generation condition into the current image generation model, so that the current image generation model generates a sample image in the scene indicated by the scene-guided generation condition based on the contour information and depth information of two objects with an occlusion relationship in the depth image and the scene-guided generation condition; there are two objects that conform to the occlusion relationship in the sample image. A fine-tuning module, configured to fine-tune the current image generation model based on the object models corresponding to different objects. A second generation module, configured to determine whether the current image generation model meets the model iteration requirements. If not, it returns to the step of generating the object model corresponding to the object; if so, then: Based on the current image generation model and the mask regions corresponding to the image regions in each training scene image, by means of different training image generation methods, generate training images for training the object detection model; the mask regions are used to indicate the regions where the objects are located; wherein, one of the different training image generation methods includes: inputting a depth image and a training scene-guided generation condition into the current image generation model, so that the current image generation model generates a training image in the training scene indicated by the training scene-guided generation condition based on the contour information and depth information of two training objects with an occlusion relationship in the depth image and the training scene-guided generation condition; there are two objects that conform to the occlusion relationship in the training image. A third generation module, configured to automatically generate contour labels of the object and region position labels where the object is located by using the mask regions corresponding to the object regions in each training image and combining with a segmentation model. A training module, configured to train the object detection model by using each training image and the contour labels of the objects and the region position labels where the objects are located in the training images; the trained object detection model is used to detect objects in images.

11. The device according to claim 10, characterized in that, The different image generation methods further include at least one of the following image generation methods: The first image generation method is: For any object, input the scene image in any scene and the text description information of the object into the current image generation model, so that the current image generation model can obtain the sample image of the object in this scene; the text description information of the object is used to describe generating the object within the mask region corresponding to the image region in the scene image. The second image generation method is as follows: For any object, input the scene image in any scene and the obtained object image into the current image generation model, so that the current image generation model can fuse the object image into the mask region corresponding to the image region in the scene image to obtain the sample image of the object in this scene. The third image generation method is as follows: Fuse the first image and the second image; the first image and the second image are images in the same scene, the objects in the first image and the second image are different, and there is an intersection in the image regions where the different objects are located. Use the fusion result to obtain the sample image in this scene; wherein, in the sample image, there is an occlusion relationship between the object originally existing in the first image and the object originally existing in the second image.

12. The device according to claim 10, characterized in that, Generating the object model corresponding to the object based on the sample images of the same object in different scenes generated by different image generation methods and the current image generation model includes: For any object, on the premise that the parameter values corresponding to at least one parameter attribute matching the object in the current image generation model remain unchanged, use the sample images of the object in different scenes to train other variable parameter values in the current image generation model, and use the trained current image generation model as the object model.

13. The device according to claim 10 or 12, characterized in that, Fine-tuning the current image generation model based on the object models corresponding to different objects includes: For any parameter attribute, find the parameter values belonging to this parameter attribute from each object model, and perform a specified operation on the found parameter values to obtain the target parameter value corresponding to this parameter attribute. Depending on the target parameter value corresponding to this parameter attribute, adjust the current parameter value under this parameter attribute in the current image generation model.

14. The device according to claim 10, characterized in that, The different training image generation methods also include: For any training object, input the training scene image in any training scene and the text description information of the training object into the current image generation model, so that the current image generation model can obtain the training image of the training object in this training scene; the text description information of the training object is used to describe generating the training object within the mask region corresponding to the image region in the training scene image. and / or For any training object, input the training scene image in any training scene and the obtained training object image into the current image generation model, so that the current image generation model can fuse the training object image into the mask region corresponding to the image region in the training scene image to obtain the training image of the training object in this training scene. and / or Fuse the third image and the fourth image; the third image and the fourth image are images under the same training scenario, the training objects in the third image and the fourth image are different, and there is an intersection in the image areas where the different training objects are located; Obtain a training image under this training scenario using the fusion result; wherein, in the training image, there is an occlusion relationship between the training object originally existing in the third image and the training object originally existing in the fourth image.

15. The device according to claim 10, characterized in that, The object is a designated dangerous animal or a designated occluder.

16. The device according to claim 10, characterized in that, After obtaining the target detection model, the training module is further configured to: Obtain the effect evaluation parameters of the target detection model in the designated scenario; If the effect evaluation parameters do not meet the requirements of the set evaluation parameters, then generate optimized images of different training objects in the designated scenario by means of the current image generation model when the model iteration requirements are met; Optimize the target detection model using each optimized image.

17. The device according to claim 10, characterized in that, Generating sample images of the same object in different scenarios using different image generation methods further includes: For any object, based on the description information of the sample image of the object in each scenario, and based on the current image generation model, the current image generation model generates a test image corresponding to the object according to the description information; the description information is used to describe the object and the environment where the object is located in the sample image.

18. The device according to claim 10 or 17, characterized in that, The determination of whether the current image generation model meets the model iteration requirements includes: Testing whether the current image generation model meets the accuracy requirements based on the test images corresponding to each object: if not, it is determined that the current image generation model does not meet the model iteration requirements; if so, it is determined that the current image generation model meets the model iteration requirements.

19. A target detection system, characterized in that, The system includes: An image acquisition device for acquiring scene images; A model device end, which is a processing device for running an image generation model set locally or in the cloud, obtains the scene images acquired by the image acquisition device, and executes the steps in the method according to any one of claims 1 to 9 based on the scene images.

20. The object detection system according to claim 19, wherein There is a communication connection between the image acquisition device and the model device end; The model device end obtaining the scene images acquired by the image acquisition device includes: receiving the scene images acquired by the image acquisition device sent by the image acquisition device through the communication connection.

21. The object detection system according to claim 19, wherein, The image acquisition device is deployed on the model device end.

22. The object detection system according to claim 20 or 21, characterized in that, The system further includes: a client device; The client device is communicatively connected to the model device end, and is configured to send the sample images of the object in each scenario and the corresponding description information to the model device end through the communication connection, so that the model device end, based on the description information of the sample images of the object in each scenario, uses the current image generation model to generate a test image corresponding to the object according to the description information; the description information is used to describe the object and the environment where the object is located in the sample image; the test image is used to test whether the current image generation model meets the model iteration requirements.

23. The object detection system according to claim 22, wherein The client device is further configured to send the text description information of the object to the model device side, so that the model device side inputs the scene image in any scene and the text description information of the object into the current image generation model, and the current image generation model obtains the sample image of the object in this scene; the text description information of the object is used to describe generating the object within the mask area corresponding to the image area in the scene image; and / or, The client device is further configured to send the object image to the model device side, so that the model device side inputs the scene image in any scene and the obtained object image into the current image generation model, and the current image generation model fuses the object image into the mask area corresponding to the image area in the scene image to obtain the sample image of the object in this scene; and / or, The image acquisition device is further configured to acquire a depth image; The client device is further configured to send the scene guidance generation condition to the model device side, so that the model device side inputs the depth image and the scene guidance generation condition into the current image generation model, and the current image generation model generates the sample image in the scene indicated by the scene guidance generation condition based on the contour information and depth information of two objects with an occlusion relationship in the depth image and the scene guidance generation condition; there are two objects with an occlusion relationship in the sample image.

24. An electronic device, characterized in that, The electronic device includes: a processor; and a computer-readable storage medium, in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 9.

25. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are run by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Video generation method and device, equipment and storage medium

    CN113313790A

  • Self-adaptive defect detection method based on small sample and terminal equipment

    CN117094986A

  • Image generation method and device, electronic equipment and storage medium

    CN117953091A