Object detection based on image generation model

WO2026200493A1PCT designated stage Publication Date: 2026-10-01HANGZHOU EZVIZ SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/082108
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-09
Publication Date
2026-10-01

Smart Images

  • Figure CN2026082108_01102026_PF_FP_ABST
    Figure CN2026082108_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to object detection based on an image generation model. By dynamically optimizing an image generation model, a finally obtained image generation model can be ensured to be optimal, thereby ensuring that training images generated by the optimal image generation model and used for training an object detection model are optimal; accordingly, the optimal training images generated by the optimal image generation model are used to train the object detection model, thereby enriching the training images used for training the detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Object detection based on image generation model Technical Field

[0001] This application relates to the field of image processing technology, and in particular to object detection based on image generation models. Background Technology

[0002] In practical applications, detection models are often used to detect targets in images. For example, dangerous animal detection models are used to detect whether there are dangerous animals such as venomous snakes or tigers in an image.

[0003] However, in practice, due to the specific nature of targets such as dangerous animals, it is often difficult to directly collect many training images of real-world scenes containing dangerous animals. For example, images of tigers in a garden or bears in a gas station are hard to capture. This leads to a lack of training images for training the detection model. The lack of training images, in turn, affects the performance of the final trained detection model and the accuracy of object detection. Summary of the Invention

[0004] In view of this, embodiments of this application provide a target detection method, system, and apparatus based on an image generation model.

[0005] This application provides a target detection method based on an image generation model. The method includes: for each of a plurality of first objects, training a current image generation model based on multiple sample images of the first object generated using different image generation methods in different scenes to generate an object model corresponding to the first object. The different image generation methods include: generating a first sample image including the first object in the first scene based on mask information and a first scene image corresponding to the first scene, wherein the mask information is used to indicate the position information of the first object in the first scene image; combining the first image including the first object, which corresponds to a second scene, with a second image including a second object different from the first object, to obtain a second sample image in the second scene where the first object and the second object have an occlusion relationship, wherein the second scene is the same as or different from the first scene; and / Or, based on a depth image including the first object and a third object different from the first object, generate a third sample image in a third scene indicated by scene-guided generation conditions, wherein the first object and the third object have an occlusion relationship in the third sample image, the third scene is the same as or different from the first scene, and the third scene is the same as or different from the second scene; fine-tune the current image generation model based on multiple object models corresponding to the multiple first objects; determine whether the current image generation model meets the model iteration requirements, if not, return to the step of generating the object model; if yes, then: take one of the multiple first objects as the target object, and based on the current image generation model and with the help of the different image generation methods, generate multiple training images including the target object corresponding to different training scenes; use the multiple training images to train an object detection model for detecting the target object.

[0006] This application embodiment also provides a target detection device, which includes: a target detection device, comprising: a first generation module, configured to, for each of a plurality of first objects, train a current image generation model based on multiple sample images of the first object generated using different image generation methods in different scenes, to generate an object model corresponding to the first object, wherein the different image generation methods include: generating a first sample image including the first object in the first scene based on mask information and a first scene image corresponding to the first scene, wherein the mask information is used to indicate the position information of the first object in the first scene image; combining the first image including the first object that corresponds to a second scene with a second image including a second object different from the first object to obtain a second sample image in the second scene in which the first object and the second object have an occlusion relationship, wherein the second scene is the same as or different from the first scene; and / or based on The system includes depth images of the first object and a third object different from the first object, and generates a third sample image in a third scene indicated by scene-guided generation conditions. The first object and the third object have an occlusion relationship in the third sample image. The third scene may be the same as or different from the first scene, and the third scene may be the same as or different from the second scene. A fine-tuning module is used to fine-tune the current image generation model based on multiple object models corresponding to the multiple first objects. A second generation module is used to determine whether the current image generation model meets the model iteration requirements. If not, it returns to the step of generating the object model; if yes, it uses one of the multiple first objects as the target object, and generates multiple training images corresponding to different training scenes, including the target object, based on the current image generation model and using the different image generation methods. A training module is used to train a target detection model for detecting the target object using the multiple training images.

[0007] This application embodiment also provides a target detection system, which includes: an image acquisition device for acquiring scene images; and a model device, which is a processing device set locally or in the cloud for running an image generation model, to obtain the scene images acquired by the image acquisition device and to perform the steps in the above method based on the scene images.

[0008] This application also provides an electronic device, including: a processor and a memory for storing computer program instructions, which, when executed by the processor, cause the processor to perform the steps of the method described above.

[0009] This application also provides a machine-readable storage medium storing computer program instructions that, when executed, enable the implementation of the steps described above. Attached Figure Description

[0010] Figure 1 is a flowchart illustrating the method provided in an embodiment of this application.

[0011] Figure 2a is a schematic diagram of the generated image sample provided in an embodiment of this application.

[0012] Figure 2b is a schematic diagram of the generated image sample provided in an embodiment of this application.

[0013] Figure 2c is a schematic diagram of the generated image sample provided in an embodiment of this application.

[0014] Figure 2d is a schematic diagram of the generated image sample provided in the embodiment of this application.

[0015] Figure 2e is a schematic diagram of the generated image sample provided in an embodiment of this application.

[0016] Figure 3 is a schematic diagram of the device structure provided in the embodiment of this application.

[0017] Figure 4 is a schematic diagram of the system structure provided in the embodiment of this application.

[0018] Figure 5 is a schematic diagram of the electronic device structure provided in an embodiment of this application. Detailed Implementation

[0019] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, and to make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, the technical solutions in the embodiments of this application will be further described in detail below with reference to the accompanying drawings.

[0020] It should be noted that the target detection method provided in this application is applicable to multiple fields such as the detection of dangerous animals and the detection of target faces. For ease of explanation, this application uses the detection of dangerous animals as an example to illustrate the method provided in this application.

[0021] Referring to Figure 1, Figure 1 is a schematic flowchart of the target detection method based on an image generation model provided in an embodiment of this application. This method can be applied to electronic devices, which can be servers; however, this application does not specifically limit the application.

[0022] As shown in Figure 1, the process may include the following steps S101 to S106.

[0023] In step S101, based on sample images of the same object generated using different image generation methods in different scenes, and the current image generation model, an object model corresponding to the object is generated.

[0024] Here, the object can be a designated dangerous animal, such as a wild boar or a venomous snake; or, the object can be a designated occlusion, such as a table or a chair. Generating sample images of the same object in different scenes using different image generation methods effectively enriches the sample image library, ensures the diversity of sample images, and enables the current image generation model to be optimized in a better direction.

[0025] In this embodiment, the current image generation model is either the initial image generation model or the image generation model obtained after fine-tuning the previous model using various object models. As an example, the initial image generation model is obtained by training a neural network model (e.g., a Diffusion model) with raw images of dangerous animals directly captured by an image acquisition device (e.g., a camera) in different scenes.

[0026] In this embodiment, the sample images of the same object generated using different image generation methods in different scenes, as well as the specific implementation methods for generating the object model corresponding to the object, will be described later and will not be repeated here.

[0027] In step S102, the current image generation model is fine-tuned based on the object model corresponding to different objects.

[0028] The specific implementation of step S102 in this embodiment will be described later, and will not be repeated here.

[0029] In step S103, it is determined whether the current image generation model, after being fine-tuned in step S102, meets the model iteration requirements.

[0030] If the result of step S103 is negative, return to step S101 above; if the result of step S103 is positive, execute step S104 below.

[0031] Here, the current image generation model in step S103 is different from the current image generation model mentioned in steps S101 and S102. The former is fine-tuned compared to the latter.

[0032] The specific implementation of step S103 will be explained later and will not be repeated here.

[0033] In step S104, training images for training the target detection model are generated based on the current image generation model and the mask regions in each training scene image.

[0034] In this embodiment, the mask region is an image region labeled in the training scene image, or obtained by masking an image region in the training scene image. The mask region in the training scene image is used to indicate the region where the training object is located. The current image generation model here is an image generation model that meets the model iteration requirements.

[0035] The specific implementation of step S104 will be explained later and will not be repeated here.

[0036] In step S105, for each training image, using the mask region corresponding to the object region in the training image and combining it with the segmentation model, the contour label of the training object in the training image and the location label of the region where the training object is located in the training image are automatically generated.

[0037] In this embodiment, the training image is input into a trained segmentation model (e.g., U-Net, Mask R-CNN, or DeepLab) to automatically generate contour labels of the training objects in the training image and location labels of the regions where the training objects are located in the training image. This method of automatically generating contour labels and location labels can greatly reduce the workload and cost of manual annotation and improve the generation efficiency of subsequent object detection models.

[0038] In step S106, the target detection model is trained using each training image, the contour labels of the training objects in each training image, and the location labels of the regions where the training objects are located in each training image.

[0039] In this embodiment, a large number of training images are generated based on the current image generation model that meets the model iteration requirements. Each training image is input into a neural network model, such as a CNN (Convolutional Neural Network) model, to obtain the predicted contour and predicted location information output by the model. The predicted contour and predicted location information are used to calculate the loss value based on the contour label and location label of the training object in the training image, respectively. The parameters of the neural network model are adjusted according to the loss value until the neural network model meets the stopping iteration training condition. The neural network model that meets the stopping iteration training condition is determined as the object detection model.

[0040] In some embodiments, images captured in real-time by video surveillance cameras installed at gas stations are transmitted to the video surveillance camera's server (which is configured with a trained object detection model). The server can monitor the images in real-time for the presence of dangerous animals. In still other embodiments, images captured in real-time by video surveillance cameras are transmitted to the video surveillance camera's server (which is configured with a trained object detection model). The server can monitor the images in real-time for the presence of target faces. The embodiments of this application do not limit the application scenarios of the trained object detection model.

[0041] As shown in Figure 1, the process first generates sample images of the same object in different scenes using different image generation methods, along with the current image generation model, to create an object model corresponding to the object. Then, the current image generation model is fine-tuned based on the object models corresponding to different objects to optimize it. If the optimized current image generation model does not meet the model iteration requirements, the process returns to the previous step of generating the object model corresponding to the object. If the current image generation model meets the model iteration requirements, training images for training the object detection model are generated based on the current image generation model and the mask regions in each training scene image. This dynamic optimization of the image generation model ensures that the final image generation model is relatively optimal, and consequently, the training images generated by the image generation model for training the object detection model are also relatively optimal. This enriches the training images used to train the object detection model and avoids the performance and accuracy of the object detection model being affected by a lack of training images.

[0042] The following examples, using Figures 2a to 2e, illustrate sample images of the same object generated in step S101 using different image generation methods in different scenes. The first sample image generation method...

[0043] For a given object, a scene image and its text description are input into the current image generation model to generate a sample image of the object within that scene. Here, the scene image can be captured by an image acquisition device in different real-world scenarios, such as a courtyard, a playground, or a desert. The text description of the object is used to describe the masked region in the scene image to generate the image of that object.

[0044] For example, as shown in the upper part of Figure 2a, the object is a wild boar, the scene image is a courtyard scene, and the text description of the object is "Wild boar needs to be generated in the masked area of ​​the image". Inputting the scene image and the text description of the object into the current image generation model yields a sample image of "wild boar in the courtyard".

[0045] For example, as shown in the lower part of Figure 2a, the object is a venomous snake, the scene image is a courtyard scene, and the text description information of the object is "a venomous snake needs to be generated in the masked area of ​​the image". Inputting the scene image and the text description information of the object into the current image generation model yields a sample image of "a venomous snake in the courtyard".

[0046] For example, as shown at the top of Figure 2b, the object is a table, the scene image is a courtyard scene, and the text description of the object is "a table needs to be generated in the masked area of ​​the image". Inputting the scene image and the text description of the object into the current image generation model yields a sample image of "a table in the courtyard".

[0047] For example, as shown in the lower part of Figure 2b, the object is a chair, the scene image is a courtyard scene, and the text description of the object is "A chair needs to be generated in the masked area of ​​the image". Inputting the scene image and the object's text description into the current image generation model yields a sample image of "a chair in the courtyard".

[0048] In this sample image generation method, for any object, by combining scene images from different real-world scenarios with the object itself, sample images of that object in different scenarios can be obtained. In particular, it can obtain images that are difficult to directly acquire (e.g., an image of a "jaguar in a yard"), thus greatly enriching the diversity of image samples. The second sample image generation method...

[0049] For a specific object, a scene image and an already obtained object image are input into the current image generation model. The current image generation model then fuses the object image into a masked region within the scene image, resulting in a sample image of the object within that scene. Here, the object image can be a local matted image of the object obtained from other images.

[0050] For example, as shown in Figure 2c, the object image is a "wild boar picture," such as a partial cutout of a "wild boar" from another image, and the scene image is a courtyard scene. Inputting this scene image and the "wild boar picture" into the current image generation model yields a sample image of "wild boar in a courtyard."

[0051] This method of generating sample images also allows for the combination of scene images from different real-world scenarios with the object, thus greatly enriching the diversity of image samples. The third method of generating sample images...

[0052] The first image and the second image are fused. The first image and the second image are images from the same scene, containing different objects, and the image regions containing these different objects have at least partial overlap in the image coordinate system (hereinafter referred to as having intersection). In some embodiments, the different objects in the first image and the second image have different categories; for example, one is a dangerous animal, and the other is an occluder. A sample image of the scene is obtained through fusion, and in this sample image, there is an occlusion relationship between the objects originally present in the first image and the objects originally present in the second image.

[0053] For example, as shown at the top of Figure 2d, the first and second images of the courtyard scene are fused. The first image contains a "wild boar," and the second image contains a "table." The images of the "wild boar" and the "table" overlap. After fusion, a sample image of "the table in the courtyard obscuring the wild boar" is obtained.

[0054] For example, as shown in the lower part of Figure 2d, the first and second images of the courtyard scene are fused. The first image contains a "venomous snake," and the second image contains a "chair." The images of the "venomous snake" and the "chair" overlap. After fusion, a sample image of "a chair obscuring a venomous snake in the courtyard" is obtained.

[0055] Optionally, in a specific implementation, the first image and the second image may be generated by either the first or the second sample image generation method, using the same scene image in the generation process.

[0056] This sample image generation method constructs sample images with occlusion relationships between different objects, further enriching the diversity of the sample images. The fourth sample image generation method...

[0057] The depth image and scene-guided generation conditions are input into the current image generation model, so that the current image generation model generates a sample image of the scene indicated by the scene-guided generation conditions based on the contour information and depth information of two objects with occlusion relationship in the depth image; the sample image contains the two objects with occlusion relationship in the depth image.

[0058] Here, the depth image can be extracted from a sample image obtained using the third sample image generation method, or it can be extracted from an image directly acquired by an image acquisition device that includes two objects with an occlusion relationship.

[0059] It should be noted that the occlusion relationship between the two objects in the generated sample image can be the same as or different from the occlusion relationship in the previously input depth image.

[0060] For example, as shown in Figure 2e, the depth image and the scene-guided generation condition "the scene where the object is located is a courtyard" are input into the current image generation model to generate three sample images as shown in Figure 2e. The "table" and "wild boar" in these three sample images present three different occlusion relationships.

[0061] In this sample image generation method, the two objects in the generated sample image present more than one occlusion relationship, which enriches the diversity of the sample image. Subsequently, the obtained sample images can be used to fine-tune the image generation model, which can improve the generation capability of the image generation model.

[0062] The above details sample images of the same object generated in different scenes using different image generation methods.

[0063] The following section provides a detailed explanation of how the object model corresponding to the object is generated in step S101.

[0064] After generating sample images corresponding to the object using at least one of the four image generation methods mentioned above, while keeping the parameter values ​​of at least one parameter matching the object in the current image generation model unchanged (here, each parameter represents an attribute of the object, and the parameter value corresponds to the specific situation of the attribute it represents. For example, for the parameter Danger representing the attribute of danger, if the parameter value of Danger is 10, it means that the object is extremely dangerous; if the parameter value of Danger is 5, it means that the object is generally dangerous; if the parameter value of Danger is 0, it means that the object is not dangerous), other variable parameter values ​​in the current image generation model are trained using sample images of the object in different scenarios, and the trained current image generation model is used as the object model.

[0065] For example, if the object is a venomous snake, after obtaining sample images of the snake in different scenes, assuming there are M parameters matching the snake in the current image generation model, N parameters (N less than M) are selected. The values ​​of these N parameters remain unchanged. The current image generation model is then trained using the sample images of the snake in different scenes. Specifically, during iterative training, the values ​​of the remaining M parameters change until the stop-training condition is met (for example, the number of iterations reaches a set threshold, such as 20), resulting in a snake model, denoted as the snake lora. Following this method, object models for different objects can be obtained, such as Chinese wild boar lora, Chinese snake lora, North American grizzly bear lora, North American gray wolf lora, Australian kangaroo lora, chair lora, table lora, clutter lora, trash can lora, and various occlusion lora, etc.

[0066] The above provides a detailed explanation of how the object model corresponding to the object is generated in step S101.

[0067] After obtaining the object models of different objects, the current image generation model is fine-tuned based on the object models of different objects. The above step S102 is explained in detail below.

[0068] In a specific implementation of step S102, as an example, for a certain parameter, the parameter value of the parameter is found from each object model, and a specified operation (here, the specified operation can be a weighted average or other operation) is performed on each found parameter value to obtain the target parameter value corresponding to the parameter. Based on the target parameter value corresponding to the parameter, the current parameter value of the parameter in the current image generation model is adjusted, for example, the current parameter value of the parameter is replaced with the target parameter value corresponding to the parameter.

[0069] By making the above fine-tuning, the generation accuracy of the current image generation model when generating each object can be improved, thereby optimizing the accuracy of the images generated by the current image generation model.

[0070] After fine-tuning the current image generation model using the methods described above, it is necessary to determine whether the fine-tuned current image generation model (i.e., the image generation model obtained after performing step S102) meets the model iteration requirements. The following section elaborates on how to determine whether the current image generation model meets the model iteration requirements.

[0071] In a specific implementation, as an example, when generating sample images of the same object in different scenes using different image generation methods in step S101 above, the method further includes: for any object, based on the description information of the sample images of the object in each scene, the current image generation model generates a test image corresponding to the object according to the description information. Here, the description information is used to describe the object in the sample image and the environment in which the object is located.

[0072] For example, as shown in Figure 2a, the description information corresponding to the first sample image in Figure 2a is "there is a wild boar in the courtyard". The description information of the sample image is output to the current image generation model, so that the current image generation model can generate an image based on the scene and object described by the description information, and the generated image is determined as the test image corresponding to the object.

[0073] After obtaining the test images corresponding to each object, the specific implementation method for determining whether the current image generation model meets the model iteration requirements is as follows: test whether the current image generation model meets the accuracy requirements based on the test images corresponding to each object; if the test result is "no", that is, the current image generation model does not meet the accuracy requirements, then it is determined that the current image generation model does not meet the model iteration requirements; if the test result is "yes", that is, the current image generation model meets the accuracy requirements, then it is determined that the current image generation model meets the model iteration requirements.

[0074] The specific implementation method for testing whether the current image generation model meets the accuracy requirements based on the test images corresponding to each object can be as follows: For each sample image, determine whether the difference between the sample image and the corresponding test image is less than or equal to the set difference threshold. If the determination result is "yes", then the sample image is determined to meet the difference condition; otherwise, the sample image is determined not to meet the difference condition. Then, if the number of images that meet the above difference condition in the sample images of each object is greater than the set number value, then the current image generation model is determined to meet the accuracy requirements; otherwise, the current image generation model is determined not to meet the accuracy requirements.

[0075] The above provides a detailed explanation of how to determine whether the current image generation model meets the model iteration requirements.

[0076] After determining that the current image generation model meets the model iteration requirements, training images for training the object detection model will be generated based on the current image generation model and the mask regions in each training scene image.

[0077] The following section elaborates on step S103 above, which involves generating training images for training the target detection model based on the current image generation model and the mask regions in each training scene image.

[0078] The specific method for generating training images is at least one of the following methods. The first method for generating training images...

[0079] For a specific training object, a training scene image and its text description are input into the current image generation model to generate a training image of the training object within that training scene. The text description describes how to generate an image of the training object within a masked region of the training scene image. (Second training image generation method)

[0080] For a specific training object, the training scene image and the obtained training object image are input into the current image generation model. The current image generation model then fuses the training object image into the mask region of the training scene image to obtain the training image of the training object in that training scene. The training object image can be a patch of an image containing the object obtained from other images. (Third method for generating training images)

[0081] The third and fourth images are fused. The third and fourth images are from the same training scene, containing different training objects whose regions overlap. The fused images yield a training image for the same training scene, where the training objects originally in the third image and those originally in the fourth image have an occlusion relationship. A fourth method for generating training images.

[0082] The depth image and training scene guidance generation conditions are input into the current image generation model, so that the current image generation model generates training sample images under the training scene indicated by the training scene guidance generation conditions based on the contour information and depth information of two training objects with occlusion relationship in the depth image; the training sample images contain the two training objects with occlusion relationship in the depth image.

[0083] The generation methods for the four training images mentioned above are similar to those for the four sample images mentioned earlier, and will not be repeated here.

[0084] The above methods generate diverse training samples, greatly enriching the training image resources for the object detection model and effectively improving its accuracy. Furthermore, the effects of occlusion and interference from different scenes are fully considered. Using these training samples to train the object detection model allows it to better learn how to accurately detect various targets, including dangerous animals, in occluded and complex scenes. This improves the model's detection capabilities in occluded and complex environments, further enhancing its accuracy.

[0085] The above provides a detailed explanation of how to generate training sample images for training the object detection model.

[0086] After obtaining the training samples in the above manner, steps S105 and S106 need to be performed to obtain the object detection model.

[0087] After obtaining the object detection model, the method further includes: obtaining performance evaluation parameters of the object detection model in a specified scene. If the performance evaluation parameters do not meet the requirements of the set evaluation parameters, then using the current image generation model that meets the model iteration requirements, optimized images of different objects in the specified scene are generated (the generation method refers to the generation method of the four sample images mentioned above, and will not be repeated here). The object detection model is optimized using each optimized image.

[0088] In this embodiment, performance evaluation parameters of the target detection model under specified scenarios are collected, and these parameters are used to guide the generation direction of subsequent training images, thereby continuously obtaining high-quality training images. This enables the target detection model to be optimized in the direction of improving the detection accuracy of the specified scenarios, thus improving the adaptability of the target detection model.

[0089] The methods provided in the embodiments of this application have been described above. The apparatus provided in the embodiments of this application is described below:

[0090] Referring to Figure 3, which is a structural diagram of the device provided in an embodiment of this application, the device is applied to an electronic device. As shown in Figure 3, the device may include: a first generation module 301, a fine-tuning module 302, a second generation module 303, a third generation module 304, and a training module 305.

[0091] The first generation module 301 is used to generate an object model corresponding to the object based on sample images of the same object generated using different image generation methods in different scenes, and the current image generation model.

[0092] The fine-tuning module 302 is used to fine-tune the current image generation model based on the object model corresponding to different objects.

[0093] The second generation module 303 is used to determine whether the current image generation model after being fine-tuned by the fine-tuning module 302 meets the model iteration requirements. If not, it returns the step of generating the object model corresponding to the object. If yes, it generates training images for training the object detection model based on the current image generation model and the mask regions in each training scene image. The mask regions are used to indicate the area where the object is located.

[0094] The third generation module 304 is used to automatically generate the contour label of the training object and the location label of the region where the training object is located by utilizing the mask region corresponding to the training object region in each training image and combining it with the segmentation model.

[0095] The training module 305 is used to train the target detection model for each training image, using the training image and the contour labels of the training objects in the training image and the location labels of the regions where the training objects are located.

[0096] As an example, generating sample images of the same object in different scenes using different image generation methods includes: generating sample images of the object in different scenes according to at least one of the following image generation methods.

[0097] The first image generation method is:

[0098] For a given object, a scene image and the object's text description are input into the current image generation model to obtain a sample image of the object in that scene. The object's text description is used to describe the generation of the object within the masked area corresponding to the image region in the scene image.

[0099] The second method of image generation is:

[0100] For a given object, the scene image and the obtained object image are input into the current image generation model, so that the current image generation model can fuse the object image into the mask region of the scene image to obtain a sample image of the object in that scene.

[0101] The third method of image generation is:

[0102] The first image and the second image are fused. The first image and the second image are images from the same scene, and the first image and the second image contain different objects. Furthermore, the image regions containing these different objects have at least partial overlap in their positions within the image coordinate system (hereinafter referred to as having intersection).

[0103] Sample images of the scene are obtained by fusion; in the sample images, there is an occlusion relationship between the objects that originally existed in the first image and the objects that originally existed in the second image.

[0104] The fourth method of image generation is:

[0105] The depth image and scene-guided generation conditions are input into the current image generation model, so that the current image generation model generates a sample image of the scene indicated by the scene-guided generation conditions based on the contour information and depth information of two objects with occlusion relationship in the depth image; the sample image contains the two objects with occlusion relationship in the depth image.

[0106] As an example, based on sample images of the same object generated using different image generation methods in different scenes, and the current image generation model, the object model corresponding to the generated object includes: for any object, under the premise that the parameter value corresponding to at least one parameter attribute that matches the object in the current image generation model remains unchanged, other variable parameter values ​​in the current image generation model are trained using sample images of the object in different scenes, and the trained current image generation model is used as the object model.

[0107] As an example, fine-tuning the current image generation model based on the object models corresponding to different objects includes: for any parameter attribute, finding the parameter value belonging to the parameter attribute from each object model, and performing specified operations on each found parameter value to obtain the target parameter value corresponding to the parameter attribute; and adjusting the current parameter value under the parameter attribute in the current image generation model based on the target parameter value corresponding to the parameter attribute.

[0108] As an example, generating training images for training the target detection model based on the current image generation model and using the mask regions corresponding to image regions in each training scene image includes: for any training object, inputting the training scene image in any training scene and the text description information of the training object into the current image generation model, so that the current image generation model can obtain the training image of the training object in the training scene; the text description information of the training object is used to describe the generation of the training object within the mask region corresponding to the image region in the training scene image;

[0109] And / or, for any training object, input the training scene image and the obtained training object image in any training scenario into the current image generation model, so that the current image generation model can fuse the training object image into the mask region corresponding to the image region in the training scene image to obtain the training image of the training object in the training scene.

[0110] And / or, fuse the third image and the fourth image; the third image and the fourth image are images in the same training scene, the training objects in the third image and the fourth image are different, and the image regions where the different training objects are located have an intersection; use the fusion result to obtain the training image in the training scene; wherein, in the training image, the training objects originally existing in the third image and the training objects originally existing in the fourth image have an occlusion relationship.

[0111] And / or, input the depth image and training scene-guided generation conditions into the current image generation model, so that the current image generation model generates a training image in the training scene indicated by the training scene-guided generation conditions based on the contour information and depth information of two training objects with occlusion relationship in the depth image and the training scene-guided generation conditions; the training image contains two training objects that meet the occlusion relationship.

[0112] As one example, the object is a designated dangerous animal or a designated obstruction.

[0113] As an example, after obtaining the object detection model, the training module is further used to: obtain the performance evaluation parameters of the object detection model in a specified scene; if the performance evaluation parameters do not meet the requirements of the set evaluation parameters, then use the current image that meets the model iteration requirements to train the model to generate optimized images of different training objects in the specified scene; and use each optimized image to optimize the object detection model.

[0114] As one embodiment, generating sample images of the same object in different scenes using different image generation methods further includes: for any object, based on the description information of the sample images of the object in each scene, and based on the current image generation model, generating a test image corresponding to the object by the current image generation model according to the description information; the description information is used to describe the object in the sample image and the environment in which the object is located.

[0115] As an example, determining whether the current image generation model meets the model iteration requirements includes: testing whether the current image generation model meets the accuracy requirements based on the test images corresponding to each object; if not, it is determined that the current image generation model does not meet the model iteration requirements; if yes, it is determined that the current image generation model meets the model iteration requirements.

[0116] This completes the structural description of the device shown in Figure 3.

[0117] The system provided in the embodiments of this application is described below:

[0118] Referring to Figure 4, which is a system structure diagram provided in an embodiment of this application, the system includes: an image acquisition device 410 for acquiring scene images; and a model device terminal 420, which is a processing device set locally or in the cloud for running an image generation model, to obtain the scene images acquired by the image acquisition device 410 and to perform the steps in the above method based on the scene images.

[0119] Optionally, there is a communication connection between the image acquisition device 410 and the model device terminal 420; the model device terminal 420 obtains the scene image acquired by the image acquisition device 410 by receiving the scene image acquired by the image acquisition device 410 sent by the image acquisition device 410 through the communication connection.

[0120] Optionally, the image acquisition device 410 is deployed on the model device end 420.

[0121] Optionally, the system further includes: a client device 430; the client device 430 is communicatively connected to the model device 420, and is used to send sample images of the object in each scene and corresponding descriptive information to the model device 420 through the communication connection, so that the model device 420 can generate a test image corresponding to the object based on the descriptive information of the sample images of the object in each scene using the current image generation model according to the descriptive information; the descriptive information is used to describe the object in the sample image and the environment in which the object is located; the test image is used to test whether the current image generation model meets the model iteration requirements.

[0122] Optionally, the client device 430 is further configured to send text description information of an object to the model device 420, so that the model device 420 inputs a scene image of any scene and the text description information of the object into the current image generation model, so that the current image generation model obtains a sample image of the object in that scene; the text description information of the object is used to describe the generation of the object within a mask area corresponding to the image region in the scene image; and / or,

[0123] The client device 430 is further configured to send the object image to the model device 420, so that the model device 420 inputs a scene image and the obtained object image into the current image generation model, so that the current image generation model fuses the object image into the mask region corresponding to the image region in the scene image to obtain a sample image of the object in that scene; and / or,

[0124] The image acquisition device 410 is also used to acquire depth images;

[0125] The client is also used to send scene guidance generation conditions to the model device terminal 420, so that the model device terminal 420 inputs the depth image and scene guidance generation conditions into the current image generation model, so that the current image generation model generates a sample image of the scene indicated by the scene guidance generation conditions based on the contour information and depth information of two objects with occlusion relationship in the depth image and the scene guidance generation conditions; the sample image contains two objects with occlusion relationship.

[0126] The hardware structure of the device shown in Figure 3 provided in the embodiments of this application is described below:

[0127] Please refer to Figure 5, which is a structural diagram of an electronic device provided in an embodiment of this application. As shown in Figure 5, the hardware structure may include: a processor 510 and a machine-readable storage medium 520, the machine-readable storage medium 520 storing machine-executable instructions that can be executed by the processor; the processor 510 is used to execute the machine-executable instructions to implement the method disclosed in the above example of this application.

[0128] Based on the same application concept as the above method, this application embodiment also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the method disclosed in the above examples of this application.

[0129] For example, the aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as executable instructions, data, etc. For instance, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.

[0130] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A target detection method based on an image generation model, comprising: For each of a plurality of first objects, the current image generation model is trained based on multiple sample images of the first object generated using different image generation methods in different scenes, thereby generating an object model corresponding to the first object. The different image generation methods include: Based on the mask information and the first scene image corresponding to the first scene, a first sample image including the first object in the first scene is generated, wherein the mask information is used to indicate the position information of the first object in the first scene image; A first image, which includes the first object and corresponds to the second scene, is combined with a second image, which includes a second object different from the first object, to obtain a second sample image in the second scene in which the first object and the second object have an occlusion relationship, wherein the second scene is the same as or different from the first scene; and / or Based on the depth image including the first object and a third object different from the first object, a third sample image is generated in a third scene indicated by the scene guidance generation conditions, wherein the first object and the third object have an occlusion relationship in the third sample image, the third scene is the same as or different from the first scene, and the third scene is the same as or different from the second scene. Based on the multiple object models corresponding to the multiple first objects, the current image generation model is fine-tuned; Determine whether the current image generation model meets the model iteration requirements. If not, return to the step of generating the object model; If so: take one of the plurality of first objects as the target object, generate a plurality of training images including the target object corresponding to different training scenarios based on the current image generation model and with the aid of the different image generation methods; use the plurality of training images to train an object detection model for detecting the target object.

2. The method according to claim 1, characterized in that, The step of generating a first sample image including the first object in the first scene based on mask information and a first scene image corresponding to the first scene includes: The first scene image and text description information are input into the current image generation model to obtain the first sample image; the text description information is used to describe the generation of an image of the first object within the area indicated by the mask information in the first scene image; or The first scene image and the obtained object image including the first object are input into the current image generation model, so that the current image generation model can fuse the object image into the area indicated by the mask information in the first scene image to obtain the first sample image.

3. The method according to claim 1, characterized in that, The step of combining a first image, which corresponds to the second scene and includes the first object, with a second image, which includes a second object different from the first object, includes: In the same image coordinate system, the first image and the second image are fused, wherein the first position information of the first object in the first image and the second position information of the second object in the second image have at least partial overlap.

4. The method according to claim 1, characterized in that, The step of generating a third sample image in a third scene indicated by scene-guided generation conditions based on a depth image including the first object and a third object different from the first object includes: The depth image and the scene-guided generation conditions are input into the current image generation model, so that the current image generation model generates the third sample image in the third scene indicated by the scene-guided generation conditions based on the contour information and depth information of the first object and the third object in the depth image. The occlusion relationship between the first object and the third object in the third sample image is the same as or different from that in the depth image.

5. The method according to claim 1, characterized in that, The step of training the current image generation model based on multiple sample images of the first object generated using different image generation methods in different scenes to generate an object model corresponding to the first object includes: While keeping the parameter values ​​of at least one parameter in the current image generation model that matches the first object unchanged, other attribute parameters with variable parameter values ​​in the current image generation model are trained using the multiple sample images. The trained current image generation model is used as the object model corresponding to the first object.

6. The method according to claim 1 or 3, characterized in that, The fine-tuning of the current image generation model based on the multiple object models corresponding to the multiple first objects includes: For each of the multiple parameters of the current image generation model, Find the parameter value of the parameter from each of the multiple object models; Perform specified operations on the found multiple parameter values ​​to obtain the target parameter value; and Based on the target parameter value of this parameter, the current parameter value of this parameter in the current image generation model is adjusted.

7. The method according to claim 1, characterized in that, The first object is a designated dangerous animal, and the second or third object is a designated obstruction.

8. The method according to claim 1, characterized in that, After obtaining the target detection model, the method further includes: Obtain the performance evaluation parameters of the target detection model in the specified scenario; If the performance evaluation parameters do not meet the set evaluation parameter requirements, then return to the step of generating the multiple training images, or return to the step of generating the object model; and The target detection model is optimized using the newly generated training images.

9. The method according to claim 1, characterized in that, Determining whether the current image generation model meets the model iteration requirements includes: For each of the plurality of first objects, Based on sample description information in multiple sample images associated with the first object, multiple test images corresponding to the first object are generated using the current image generation model, wherein the sample description information is used to describe the first object in the sample images and the environment in which the first object is located; Based on the multiple test images corresponding to each of the multiple first objects, the current image generation model is tested to see if it meets the accuracy requirements. If not, then it is determined that the current image generation model does not meet the model iteration requirements; If so, then the current image generation model is determined to meet the model iteration requirements.

10. The method according to claim 1, characterized in that, After generating multiple training images including the target object corresponding to different training scenarios, the method further includes: For each training image, using a segmentation model, based on the mask information corresponding to the target object in the training image, a contour label and a location label for the target object are generated; and The step of training a target detection model for detecting the target object using the plurality of training images includes: The target detection model is trained using the multiple training images and the contour and location labels of the target objects in each training image.

11. The method according to claim 1, characterized in that, Also includes: The trained target detection model is used to detect the target object in the image in real time.

12. A target detection device, characterized in that, The device includes: A first generation module is configured to, for each of a plurality of first objects, train a current image generation model based on multiple sample images of the first object generated using different image generation methods in different scenes, thereby generating an object model corresponding to the first object. The different image generation methods include: Based on the mask information and the first scene image corresponding to the first scene, a first sample image including the first object in the first scene is generated, wherein the mask information is used to indicate the position information of the first object in the first scene image; A first image, which includes the first object and corresponds to the second scene, is combined with a second image, which includes a second object different from the first object, to obtain a second sample image in the second scene in which the first object and the second object have an occlusion relationship, wherein the second scene is the same as or different from the first scene; and / or Based on the depth image including the first object and a third object different from the first object, a third sample image is generated in a third scene indicated by the scene guidance generation conditions, wherein the first object and the third object have an occlusion relationship in the third sample image, the third scene is the same as or different from the first scene, and the third scene is the same as or different from the second scene. The fine-tuning module is used to fine-tune the current image generation model based on the multiple object models corresponding to the multiple first objects; The second generation module is used to determine whether the current image generation model meets the model iteration requirements. If not, it returns to the step of generating the object model; if yes, then: Using one of the plurality of first objects as the target object, and based on the current image generation model and with the aid of the different image generation methods, generate a plurality of training images, including the target object, corresponding to different training scenarios; The training module is used to train an object detection model for detecting the target object using the plurality of training images.

13. The apparatus according to claim 12, characterized in that, Based on the mask information and the first scene image corresponding to the first scene, a first sample image including the first object in the first scene is generated, including: The first scene image and text description information are input into the current image generation model to obtain the first sample image; the text description information is used to describe the generation of an image of the first object within the area indicated by the mask information in the first scene image; or The first scene image and the obtained object image including the first object are input into the current image generation model, so that the current image generation model can fuse the object image into the area indicated by the mask information in the first scene image to obtain the first sample image.

14. The apparatus according to claim 12, characterized in that, The first image, which corresponds to the second scene and includes the first object, is combined with a second image, which includes a second object different from the first object. In the same image coordinate system, the first image and the second image are fused, wherein the first position information of the first object in the first image and the second position information of the second object in the second image have at least partial overlap; The step of generating a third sample image in a third scene indicated by scene-guided generation conditions based on a depth image including the first object and a third object different from the first object includes: The depth image and the scene-guided generation conditions are input into the current image generation model, so that the current image generation model generates the third sample image in the third scene indicated by the scene-guided generation conditions based on the contour information and depth information of the first object and the third object in the depth image. The occlusion relationship between the first object and the third object in the third sample image is the same as or different from that in the depth image. And / or, The step of training the current image generation model based on multiple sample images of the first object generated using different image generation methods in different scenes to generate an object model corresponding to the first object includes: Under the premise that the parameter value of at least one parameter that matches the first object in the current image generation model remains unchanged, other attribute parameters with variable parameter values ​​in the current image generation model are trained using the multiple sample images, and the trained current image generation model is used as the object model corresponding to the first object. And / or, The fine-tuning of the current image generation model based on the multiple object models corresponding to the multiple first objects includes: For each parameter in the current image generation model, find the parameter value from the multiple object models; perform a specified operation on the found parameter values ​​to obtain the target parameter value; and Based on the target parameter value of this parameter, adjust the current parameter value of this parameter in the current image generation model; And / or, The first object is a designated dangerous animal, and the second or third object is a designated obstruction; And / or, After obtaining the target detection model, the method further includes: Obtain the performance evaluation parameters of the target detection model in the specified scenario; If the performance evaluation parameters do not meet the set evaluation parameter requirements, then return to the step of generating the multiple training images, or return to the step of generating the object model; and The target detection model is optimized using the newly generated training images; And / or, Determining whether the current image generation model meets the model iteration requirements includes: For each of the plurality of first objects, Based on sample description information in multiple sample images associated with the first object, multiple test images corresponding to the first object are generated using the current image generation model, wherein the sample description information is used to describe the first object in the sample images and the environment in which the first object is located; Based on the multiple test images corresponding to each of the multiple first objects, the current image generation model is tested to see if it meets the accuracy requirements; if not, it is determined that the current image generation model does not meet the model iteration requirements; if yes, it is determined that the current image generation model meets the model iteration requirements.

15. A target detection system, characterized in that, The system includes: Image acquisition equipment, used to acquire scene images; The model device is a processing device set up locally or in the cloud for running an image generation model, which obtains scene images captured by the image acquisition device and performs the steps of the method as described in any one of claims 1 to 10 based on the scene images.

16. The target detection system according to claim 15, characterized in that, There is a communication connection between the image acquisition device and the model device. The model device obtains the scene image acquired by the image acquisition device by receiving the scene image acquired by the image acquisition device and sending it through the communication connection.

17. The target detection system according to claim 15, characterized in that, The image acquisition device is deployed on the model device.

18. The target detection system according to claim 16 or 17, characterized in that, The system also includes: client devices; The client device is connected to the model device and is used to send sample images of the object in each scene and corresponding descriptive information to the model device through the communication connection. The model device then uses the current image generation model to generate a test image of the object based on the descriptive information of the sample images of the object in each scene. The descriptive information describes the object in the sample image and the environment in which the object is located. The test image is used to test whether the current image generation model meets the model iteration requirements.

19. The target detection system according to claim 18, characterized in that, The client device is further configured to send text description information of an object to the model device, so that the model device inputs a scene image of any scenario and the text description information of the object into the current image generation model, so that the current image generation model obtains a sample image of the object in that scenario; the text description information of the object is used to describe the generation of the object within a mask area corresponding to the image region in the scene image; and / or, The client device is also configured to send the object image to the model device, so that the model device can input the scene image and the obtained object image in any scene into the current image generation model, so that the current image generation model can fuse the object image into the mask region corresponding to the image region in the scene image to obtain a sample image of the object in that scene; and / or, The image acquisition device is also used to acquire depth images; The client is also used to send scene-guided generation conditions to the model device, so that the model device can input the depth image and scene-guided generation conditions into the current image generation model, so that the current image generation model can generate a sample image of the scene indicated by the scene-guided generation conditions based on the contour information and depth information of two objects with occlusion relationship in the depth image and the scene-guided generation conditions; the sample image contains two objects with occlusion relationship.

20. An electronic device, comprising: processor; as well as A computer-readable storage medium storing computer program instructions that, when executed by the processor, cause the processor to perform the steps of the method as described in any one of claims 1 to 11.

21. A computer-readable storage medium storing computer program instructions that, when executed by a processor, cause the processor to perform the steps of the method as described in any one of claims 1 to 11.