Synthetic data for chemical industry

By generating synthetic object images for training object detection models, the method addresses the challenge of detecting single objects in complex scenes, enhancing model performance and production efficiency in chemical industries.

WO2025162768A1PCT designated stage Publication Date: 2025-08-07BASF SE
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/051417
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-29
Filing Date
2025-01-21
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Reliable detection of single objects with varying appearances in multi-object scenes remains challenging in image recognition, particularly in industrial applications like agriculture and chemical production.

Method used

A computer-implemented method generates synthetic object images by combining target objects with backgrounds using object and background generators, trained on historical data, to create training images for object detection models, enhancing model performance and accuracy.

Benefits of technology

The method enables resource-efficient and time-effective object detection tailored to specific use cases, improving production efficiency and reducing time-to-market by generating high-quality training data for chemical product development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000033_0000
    Figure 00000033_0000
  • Figure 00000034_0000
    Figure 00000034_0000
  • Figure 00000035_0000
    Figure 00000035_0000
Patent Text Reader

Abstract

A method, in particular a computer-implemented method, for detecting one or more target object(s), the method comprising: providing an image scene including one or more target object(s), providing the image scene to an object detection model for determining an indication on the presence of the one or more target object(s) within the image scene, wherein the object detection model is parametrized based on one or more training image(s) and one or more corresponding indication(s) on the presence of the one or more target object(s) within the one or more training image(s), and wherein at least a part of the one or more training image(s) are synthetic images obtained by providing one or more synthetic object image(s) comprising one or more target object(s) and object background, and placing the target object extracted from the one or more synthetic object image(s) on one or more background(s), providing the indication on the presence of the target object within the image scene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SYNTHETIC DATA FOR CHEMICAL INDUSTRY

[0002] TECHNICAL FIELD

[0003] The invention relates to a computer-implemented method for detecting an object, a method for training an object detection model, a method for generating training data, use of an object detection model, use of a training image, use of a indication on a presence of an object within an image, an apparatus for detecting an object, an apparatus for training an object detection model, an apparatus for generating training data.

[0004] TECHNICAL BACKGROUND

[0005] Image recognition based on convolutional neural networks has been a growing field since years. Image recognition is used in multiple technical applications such as agriculture or industrial production. However, recognition of single objects with different appearance in multi object scenes remain difficult to reliably detect.

[0006] US1 1475246B2 describes a system and method for training a model using a training dataset. The training dataset can be made up of only real data, only synthetic data, or any combination of synthetic data and real data. The images are segmented to define objects with known labels. The object is pasted onto backgrounds to generated synthetic datasets. The various aspects of the invention include generation of data that is used to supplement or augment real data. Labels or attributes can be automatically added to the data as it is generated. The data can be generated using seed data. The data can be generated using synthetic data. The data can be generated from any source, including the user's thoughts or memory. Using the training dataset, various domain adaptation models can be trained.

[0007] US20210089845A1 describes a method and apparatus for joint image and per-pixel annotation synthesis with a generative adversarial network (GAN) are provided. The method includes: by inputting data to a generative adversarial network (GAN), obtaining a first image from the GAN; inputting, to a decoder, a first feature value that is obtained from at least one intermediate layer of the GAN according to the inputting of the data to the GAN; and obtaining a first semantic segmentation mask from the decoder according to the inputting of the first feature value to the decoder.

[0008] WO2023242236A1 relates to image processing. The computer-implemented method may be used to improve the computer vision technique for the application in the technical field of agriculture and in production environment. SUMMARY

[0009] In another aspect, the disclosure relates to a computer-implemented method, in particular a computer-implemented method, for detecting one or more target object(s), the method comprising: providing an image scene including one or more target object(s), providing the image scene to an object detection model for determining an indication on the presence of the one or more target object(s) within the image scene, wherein the object detection model is parametrized based on one or more training image(s) and one or more corresponding indication(s) on the presence of the one or more target object(s) within the one or more training image(s), and wherein at least a part of the one or more training image(s) are synthetic images obtained by providing one or more synthetic object image(s) comprising one or more target object(s) and object background, and placing the target object extracted from the one or more synthetic object image(s) on one or more background(s), providing the indication on the presence of the target object within the image scene.

[0010] In another aspect, it relates to a method, in particular a computer-implemented method, for generating one or more synthetic object image(s) comprising one or more target object(s) and object background, wherein the synthetic object image(s) are used for generating one or more training image(s) for training an object detection model, the method comprising: providing a request for generating one or more synthetic object image(s) to an object image generator, wherein the object image generator is parametrized to generate synthetic object images in response to receiving requests for generating one or more synthetic object image(s), generating the one or more synthetic object image(s) by the object image generator, providing the one or more synthetic object image(s).

[0011] In another aspect, it relates to a computer-implemented method for generating training data comprising a training image suitable for training an object detection model, the method comprising: receiving a request for generating a training image comprising a target object and first background, generating a first object image comprising the target object and a target object background based on an object image generator, wherein the object image generator is parametrized and / or trained based on a first historical image comprising the target object and first historical object background, receiving a one or more background(s) comprising the first background, combining the one or more background(s) and the first object image to obtain the training image, providing the training image.

[0012] In another aspect, it relates to use of a training image as obtained by any one of the methods as described herein for training an object detection model. In another aspect, it relates to use of an indication on the presence of the object within the image as obtained by any one of the methods as described herein for controlling and / or monitoring a production of a chemical product and / or an application of a chemical product.

[0013] In another aspect, it relates to a computer-implemented method for training an object detection model suitable for detecting an object within an image, the method comprising receiving a training image as generated by any one of any one of methods described herein, receiving an indication on the presence of the object within the training image, training the object detection model based on the training image and the indication on the presence of the object within the training image.

[0014] In another aspect, it relates to use of a training image as generated by any one of the methods as described herein for training an object detection model.

[0015] In another aspect, it relates to an apparatus comprising: a processor; and a memory storing instructions that, when executed by the processor, configure the apparatus to perform the any one of the methods described herein.

[0016] In another aspect, it relates to a computer-implemented method for detecting a at least a part of a crop in an image, the method comprising: receiving an image showing at least the part of the crop, detecting at least the part of the crop within the image by providing the image to an image detection model, and; receiving an indication on the presence of at least the part of the crop within the image, wherein the object detection model has been trained based on a plurality of training images, and, wherein the plurality of training images are generated by : receiving one or more historical images comprising at least one or more parts of the crop, generating an object image based on an object image generator, wherein the object image comprises at least one of the one or more parts of the crop and background, wherein the object image generator is parametrized and / or trained based on the one or more historical images to generate a plurality of object images, receiving a one or more background(s), combining the one or more background(s) and the object image to obtain the training image, and; providing the indication on the presence of at least the part of the crop within the image.

[0017] EMBODIMENTS

[0018] Any disclosure, embodiments and examples described herein relate to the methods, the systems, apparatuses, uses, models and training data sets lined out above and below. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples. For producing reliable and high-quality products object detection plays a crucial rule from raw material supply, monitoring of growth to quality control of the furnished product. Hence, it is of high importance to enable use case specific and reliable object detection.

[0019] The herein presented invention allows for object detection that can be tailored to the needs of the use case by training the object detection model based on a first object image comprising the target object and a target object background. The so-generated synthetic images provide realistic image features and hence, object detection models trained based on the so-generated synthetic images can provide improved performances. Further, full control over generating synthetic data for training an object detection model is provided. The realistic-looking synthetic target objects can be pasted in a tailored manner to allow for realistic overlapping object on diverse backgrounds. Further, already available workflows for evaluating the production and the performance of chemical products can be improved, thereby saving energy and time needed for new developments. This enables resource-efficient application and production of chemical products while reducing the time-to-market when developing chemical products.

[0020] Hence, the invention allows for a time-efficient and barrier-free solution to the above-specified problem.

[0021] In an embodiment, background may comprise a scenery and / or one or more object independent of the target object(s). Background may be defined dependent on the one or more target object(s). Background may comprise at least a part of a surrounding of the one or more target object(s).

[0022] In an embodiment, background generator may refer to a generative model. The background generator may be configured for generating one or more background(s), preferably a plurality of different backgrounds. Examples for a generative model may include and / or may be based on a generative adversarial network, a variational autoencoder, a diffusion model or the like. The background generator may be parametrized and / or trained based on one or more historical background(s), in particular a plurality of different historical background(s). The historical background(s) may be different from the one or more background(s). Preferably, the one or more background(s) generated by the background generator may be synthetic background(s). The background generator may receive a trigger for generating the one or more background(s). The trigger for generating the one or more background(s) may initiate the generating of the one or more background(s). The trigger for generating the one or more background(s) may be included in the request for generating a training image and / or may be generated based on the request for generating the training image. The trigger for generating the one or more background(s) may comprise text, audio, noise, in particular a noise image and / or a combination thereof. A noise image may refer to data representing noise, in particular associated with the same format as the one or more background(s). The trigger for generating the one or more background(s) may be included in and / or may be generated in response to the request for generating the one or more training image(s). The background generator may comprise at least one machine-learning architecture and model parameters. Parametrizing the background generator may comprise initializing and / or defining the model parameters, e.g. by assigning random numbers to the model parameters. Training the background generator may comprise updating the model parameters. For example, the machine-learning architecture may be or may comprise one or more of: linear regression, logistic regression, random forest, piecewise linear, nonlinear classifiers, support vector machines, naive Bayes classifications, nearest neighbours, neural networks, convolutional neural networks, generative adversarial networks, diffusion model, support vector machines, or gradient boosting algorithms or the like. In the case of a neural network, the model can be a multi-scale neural network or a recurrent neural network (RNN) such as, but not limited to, a gated recurrent unit (GRU) recurrent neural network, vision transformer or a long short-term memory (LSTM) recurrent neural network.

[0023] In an embodiment, target object mask may be indicative of a contour of the one or more target object(s). Further, the target object mask may be indicative of the one or more parts of an image scene and / or training image associated with the one or more target object(s). For example, a target object mask may indicate the part of the image scene and / or training image associated with the one or more target object(s) by a numerical value, and / or a Boolean value. Hence, the target object mask may be associated with an equal image format and / or an equal number of pixel values as the one or more target object(s) and / or the one or more synthetic object image(s). The target object mask may be further indicative on the contour of the one or more background(s). For example, the target object mask may indicate the contour of the one or more target object(s), in particular within the one or more synthetic object image(s), by a positive Boolean value and / or a numerical value such as a 1 or a numerical value > or equal to 0.5. The target object mask may indicate the background, in particular the object background by a negative Boolean value and / or a numerical value such as 0 or a numerical values < 0.5.

[0024] In an embodiment, image augmentation technique may refer to a technique suitable for changing an orientation of the one or more target object(s), changing a scaling associated with the one or more target object(s) and / or changing one or more pixel value(s) associated with the one or more target object(s). For example, image augmentation technique may comprise at least one of scaling, cutting, rotating, blurring, warping, shearing, resizing, folding, changing the contrast, changing the brightness, adding noise, multiply at least a part of the pixel values, drop out, adjusting colors, applying a convolution, embossing, sharpening, flipping, averaging pixel values or a combination thereof. See advanced Graphics Programming Using OpenGL - A volume in The Morgan Kaufmann Series in Computer Graphics by TOM McREYNOLDS and DAVID BLYTHE (2005) ISBN 9781558606593, https: / / doi.org / 10.1016 / B978-1-55860-659-3.50030-5 for a non-exhaustive list of image augmentation techniques.

[0025] In an embodiment, mask generator may be parametrized and / or trained based on one or more mask training image(s) associated with the one or more target object(s) and an indication on the contour of the one or more target object(s). Mask generator may be a generative model. The mask generator may be configured for generating a target object mask, preferably for generating a plurality of target object masks, preferably a plurality of different target object masks. Examples for a generative model may include and / or may be based on a generative adversarial network, a variational autoencoder, a diffusion model or the like. The mask generator may receive a trigger for generating the target object mask. The trigger for generating the target object mask may initiate the generating of the target object mask. The trigger for generating the target object mask may be included in the request for generating one or more training image(s) and / or may be generated based on the request for generating the training image. The trigger for generating the target object mask may comprise text, audio, noise, in particular a noise image and / or a combination thereof. A noise image may refer to data representing noise, in particular associated with the same format as the target object mask.

[0026] In an embodiment, the mask generator may be further parametrized on a representation associated with the one or more synthetic object image(s) and / or the request for generating the one or more synthetic object image(s). The representation of the one or more synthetic object image(s) may be a numerical representation of the one or more synthetic object image(s), in particular a vector representation of the one or more synthetic object image(s). The representation of the one or more synthetic object image(s) may be obtained by the object image generator. The mask generator may be parametrized to generate the target object mask based on the representation associated with the one or more synthetic object image(s) and / or the request for generating the one or more synthetic object image(s).

[0027] In an embodiment, memory may be a physical system memory which may be volatile, non-volatile, or a combination thereof. The memory may include non-volatile mass storage such as physical storage media. The memory may be a computer-readable storage media such as RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, non-magnetic disk storage such as solid-state disk or any other physical and tangible storage medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by the computing system. Moreover, the memory may be a computer-readable media that carries computer- executable instructions (also called transmission media). Further, upon reaching various computing system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to storage media (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a “NIC”), and then eventually transferred to computing system RAM and / or to less volatile storage media at a computing system. Thus, it should be understood that storage media can be included in computing components that also (or even primarily) utilize transmission media.

[0028] In an embodiment, object may refer to a thing, a part of a thing, a part of a living organism such as a body part and / or a living organism. Object may be characterized by a shape associated with the object and / or parts of the object. The object may be a crop and / or may comprise an attribute of the crop. The target object may be an object to be detected, in particular by the object detection model. Target object may be associated with one or more type(s) of object(s), in particular one or more different type(s) of object(s). A type of an object may be associated with a group of object having a common denotation.

[0029] In an embodiment, object background may refer to background associated with the object. Preferably, object background may be background in an image showing an object. Object background may be background surrounding the object at least partially. Object background may be a target object background and / or a second object background. In particular, object background may be background within a predefined distance from the one or more target object(s) in one or more synthetic object image(s).

[0030] In an embodiment, object detection model may be parametrized and / or trained to receive an image scene associated with the one or more target object(s) and providing an indication on the presence of the one or more target object(s) within the image scene. The object detection model may be parametrized and / or trained based on one or more training image(s). The object detection model may be a classification model trained and / or parametrized for classifying the image according to the presence of the one or more target object(s). The classification model may comprise at least one machine-learning architecture and model parameters.

[0031] In an embodiment, object image may comprise the object and object background. Object image may be a synthetic object image. The synthetic object image may be generated by an object image generator, in particular an object image generator model. Object image may be associated with an image area showing the object and the object background. For example, the image area may be associated with a plurality of pixel values. At least 1 % of the image area may show and / or at least 1 % of the pixel values of the image may be associated with the object background, preferably at least 2 %, more preferably at least 3 %, more preferably at least 4 %, more preferably at least 5 %, more preferably at least 7 %, most preferably at least 10 %. Taking some background into account when generating the object image for training an object image generator increases the likelihood of the object images generated by the object image generator to be classified as real images. Adding some background to the object images helps the object image generator to learn to correctly generate the boundaries of the object and to mimic the relation between the object and the background. Hence, this increases quality of the object images which in turn increases the accuracy of an object detection model trained based on the object images. Ultimately, this increases the yield obtained by treating crops according to detected diseases, increases the performance of a product line otherwise falsely provided to consumers of the products or reduces the amount of resources used for producing products by correctly identifying faulty supply. The object image may comprise the one or more target object(s) and object background at the edges. The object image may be a cutout of the one or more target object(s) from an image scene, in particular a historical image scene. Further, the object image may comprise remaining background resulting from the cutout of the one or more target object(s). In an embodiment, the object image generator may be an object image generator model.

[0032] In an embodiment, the object localization model may be parametrized to provide and / or generate indications on the location of one or more object(s) including the one or more target object(s) within image scenes, preferably two or more different object(s) including the one or more target object(s). The two or more different object(s) may include two or more target object(s) of two or more different types or one or more target object(s) and one or more object(s) different from the one or more target object(s). The object localization model may be parametrized based on localization training images including the one or more target object(s), optionally including two or more objects including the one or more target object(s), and corresponding indications on the location of the one or more target object(s), optionally including two or more objects including the one or more target object(s) within the localization training images. Object localization model and / or object image generator may comprise at least one machinelearning architecture and model parameters. Parametrizing the object localization model may comprise initializing and / or defining the model parameters, e.g. by assigning random numbers to the model parameters. Training the object localization model may comprise updating the model parameters. For example, the machine-learning architecture may be or may comprise one or more of: linear regression, logistic regression, random forest, piecewise linear, nonlinear classifiers, support vector machines, naive Bayes classifications, nearest neighbours, neural networks, convolutional neural networks, generative adversarial networks, diffusion model, support vector machines, or gradient boosting algorithms or the like. In the case of a neural network, the model can be a multiscale neural network or a recurrent neural network (RNN) such as, but not limited to, a gated recurrent unit (GRU) recurrent neural network, vision transformer or a long short-term memory (LSTM) recurrent neural network.

[0033] In an embodiment, processor may refer to an arbitrary logic circuitry configured to perform basic operations of a computer or system, and / or, generally, to a device which is configured for performing calculations or logic operations. In particular, the processor, or computer processor may be configured for processing basic instructions that drive the computer or system. It may be a semi-conductor based processor, a quantum processor, or any other type of processor configures for processing instructions. As an example, the processor may be or may comprise a Central Processing Unit ("CPU"). The processor may be a (“GPU”) graphics processing unit, (“TPU”) tensor processing unit, ("CISC") Complex Instruction Set Computing microprocessor, Reduced Instruction Set Computing ("RISC") microprocessor, Very Long Instruction Word ("VLIW") microprocessor, or a processor implementing other instruction sets or processors implementing a combination of instruction sets. The processing means may also be one or more special-purpose processing devices such as an Application-Specific Integrated Circuit ("ASIC"), a Field Programmable Gate Array ("FPGA"), a Complex Programmable Logic Device ("CPLD"), a Digital Signal Processor ("DSP"), a network processor, or the like. The methods, systems and devices described herein may be implemented as software in a DSP, in a micro-controller, or in any other side-processor or as hardware circuit within an ASIC, CPLD, or FPGA. It is to be understood that the term processor may also refer to one or more processing devices, such as a distributed system of processing devices located across multiple computer systems (e.g., cloud computing), and is not limited to a single device unless otherwise specified. The processor may also be an interface to a remote computer system such as a cloud service. The processor may include or may be a secure enclave processor (SEP). An SEP may be a secure circuit configured for processing the spectra. A "secure circuit" is a circuit that protects an isolated, internal resource from being directly accessed by an external circuit. The processor may be an image signal processor (ISP) and may include circuitry suitable for processing images, in particular images with personal and / or confidential information.

[0034] In an embodiment, providing the one or more synthetic object image(s) may comprise generating the one or more synthetic object image(s) by an object image generator. The object image generator may be parametrized to generate one or more synthetic object image(s) in response to receiving a request for generating one or more synthetic image(s). The object image generator may be parametrized and / or trained based a plurality of historical object images. The request may be provided via a user interface. The request may trigger and / or may comprise providing one or more synthetic object image(s). Any one of the methods may further include providing a request for generating the one or more training image(s) associated with a target object and one or more background(s). The request may comprise a random tensor, in particular a random vector. The object image generator may be parametrized and / or trained to receive the random tensor, in particular the random vector, and generate, in particular by removing noise from the random tensor and / or random vector, one or more synthetic object image(s) from the random tensor, in particular random vector. Removing noise from the random tensor may include passing the random tensor through one or more layer(s) of a neural network. The random tensor may be associated with a distribution of values following a random distribution of values. The random tensor may be associated with, in particular represent, a random image scene showing a random distribution of pixel values. The one or more layer(s) of the neural network may be configured to increase a difference to the distribution of values following a random distribution of values. The object image generator may be parametrized and / or trained to generate a representation, in particular a multidimensional representation, associated with the one or more synthetic image(s) and / or the request for generating one or more synthetic image(s). Further, the object image generator may be parametrized and / or trained to generate the one or more synthetic image(s) from the representation, in particular the multidimensional representation.

[0035] In an embodiment, placing the one or more synthetic object image(s) on one or more background(s), may include providing the one or more background(s) and placing the one or more synthetic object image(s) on the one or more provided background(s). Further, placing the one or more synthetic object image(s) on the one or more background(s) may include determining if the object background matches the one or more background(s), in particular by comparing the object background and the one or more background(s), and placing the one or more synthetic object image(s) in response to determining that the object background matches the one or more background(s). Comparing the object background and the one or more background(s) may include determining if at least a part of the pixel values associated with the object background and at least a part of the pixel values associated with the one or more background(s) correspond. At least the part of the pixel values associated with the one or more background(s) correspond to at least the part of the pixel values associated with the object background if at least the part of the pixel values associated with the one or more background(s) is within a tolerance range of 15% associated with the at least one pixel value of the pixel values associated with the object background. Placing the one or more synthetic object image(s) may include placing the one or more synthetic object image(s) on the one or more background(s) and filling the one or more background(s) prior or subsequently to the placing of the one or more synthetic object image(s) on the one or more background(s).

[0036] In an embodiment, providing the one or more background(s) may comprise selecting the one or more background(s) from a plurality of backgrounds randomly or according to the one or more target object(s). The plurality of backgrounds may be asscociated with a plurality of different colors, textures, patterns or the like. Thereby, the diversity of the training data can be increased.

[0037] In an embodiment, the object detection model may be parametrized based on one or more training image(s) and one or more corresponding indication(s) on the presence of the one or more target object(s) within the one or more training image(s).

[0038] In an embodiment, any one of the methods as described herein may further providing an indication on a location of the one or more target object(s) within the image scene to the object detection model. The object detection model may be further parametrized based on indications on the location of the one or more target object(s). Any one of the methods may further comprise providing the one or more training image(s) and / or the image scene to the object localization model for generating the indication on the location of the one or more target object(s) within the image scene and / or the one or more training image(s). The indication on the location of the one or more target object(s) may be provided and / or generated by an object localization model. The indication on the location of the one or more object(s) including the one or more target object(s) within image scenes may be location data indicative of a part of the image scene associated with the one or more object(s) including the one or more target object(s) within image scenes. For example, the indication on the location of the one or more object(s) including the one or more target object(s) within image scenes may be a bounding box. Followingly, the indication on the location, in particular the part of the image scene associated with the one or more object(s) including the one or more target object(s), may comprise the one or more object(s) including the one or more target object(s) and at least a part of the background associated with the image scene. This is advantageous since object localization models may be readily available and no training may further training may be required. Such models may be trained to localize a plurality of objects. Thus, this makes efficient use already available capabilities. Hence, resources otherwise used for further training, and time, otherwise used for establishing a training data set, can be saved. In an embodiment, the object detection model may be further parametrized and / or trained to generate an indication on the location of the object within the image. The object detection model may be further trained and / or parametrized based on one or more training image(s) associated with, in particular comprising, the object and object background, and an indication on a location of the object within the one or more training image(s), optionally providing the indication on the location of the object within the image. Hence, the object detection model may be trained and / or parametrized for identifying the location of the object within the image. The object detection model may be an object detection an localization model. The object detection and localization model may be further configured for providing the indication on the location of the one or more target object(s) within the image scene. This is advantageous since a model configured for localizing and detecting the object usually performs better than using a first model to localize the object and providing the identified location to a second model for detecting the object. Therefore, an increased performance for detecting an object correctly can be achieved by training a model for localizing and detecting the object within the image.

[0039] In an example, the indication on the presence of the object within the image may be represented as a label for training and / or parametrizing the object detection model. The indication on the presence of the object within the image scene and / or the one or more training image(s) may be associated with a Boolean value and / or a numerical value indicative of true and / or false value. In an example, the indication on may be a Boolean value, a yes or nor and / or a numerical value, e.g. a value between 0 and 1 . In particular, the numerical values may indicate a confidence score indicative of whether the object may be present within the image scene and / or training image. In an example, the indication on the location of the object within the image scene and / or one or more training image(s) may be represented as one or more labels associated with at least a part of the image. Preferably, the indication on the location of the object may be assigned to a part of the image scene and / or one or more training image(s). The indication on the location may indicate whether the part of the image may be associated with the object. A plurality of indications on locations of the one or more target object(s) and / or one or more object(s) different from the one or more target object(s) may be associated with, in particular assigned to, the one or more training image(s) and / or the image scene. This may be known as semantic segmentation. For example, the indication on the location of the object may comprise one or more values associated with one or more parts of the image. The values associated with the one or more parts of the image may be for example numerical values, one or more boolean values and / or one or more yes or no's. Numerical values may be any one of the values between 0 and 1 for example, wherein the numerical values may indicate a confidence score indicative of whether the part of the image scene and / or the one or more training image(s) associated with the numerical values may comprise at least a part of the object and / or may show at least a part of the object.

[0040] In an embodiment, any one of the methods as described herein may further comprise providing the image scene showing the object, preferably via a user interface and / or providing the indication on the presence of the object within the image, preferably via a user interface. The request for generating one or more synthetic image(s) may be received via a user interface and / or the one or more training image(s) may be provided via a user interface. This allows the user to interact and tailor the training data to be generated to current needs. Hence, this feature enables a precise adaption of the training data to the use case resulting in an improved accuracy of the object detection model.

[0041] In an embodiment, the request for generating the first object image may comprise the first historical image. This allows the user to interact and tailor the training data to be generated to examples available. Hence, this feature enables a precise adaption of the training data to current use cases resulting in an improved accuracy of the object detection model.

[0042] In an embodiment, any one of the methods as described herein may further comprise training the object image generator based on the first historical image.

[0043] In an embodiment, generating a first object image based on an object image generator may comprise providing a trigger for generating the first object image to the object image generator and receiving the first object image from the object image generator. The trigger allows to precisely define the training data needed. For example, conditional GANs can be instructed to generate certain synthetic data. This is especially necessary in situations where specific objects need to be detected and / or where the generated real world data does not represent the certain configurations of the object sufficient in quantity to train a robust object detection model. For example, where a specific product is usually produced with a high accuracy and precision, synthetic data will reflect this. Followingly, influencing what kind of data is exactly generated comes with the benefit of generating the desired data for each specific use case. Thus, the accuracy of the detection trained based on the training data will be increased.

[0044] In an embodiment, the request for generating the training image and / or the image may be provided, in particular via the user interface and / or the training image and / or the indication on the presence of the object within the image may be received, in particular via the user interface. This allows the user to interact and tailor the training data to be generated to examples available. Hence, this feature enables a precise adaption of the training data to current use cases resulting in an improved accuracy of the object detection model.

[0045] In an embodiment, the one or more background(s) may be provided and / or generated by a background generator. The background generator may be parametrized and / ortrained to generate backgrounds. The background generator may be parametrized to generate backgrounds based on one or more historical background(s). Furthermore, one or more background(s) may be selected from a plurality of background(s) according to the synthetic object image and / or a selection of the one or more background(s). Any one of the methods may further comprise providing a selection of the one or more background(s), preferably via a user interface. This is advantageous since this allows to tailor the background according to the specific use case. Hence, this feature enables a precise adaption of the training data to current use cases resulting in an improved accuracy of the object detection model. Additionally or alternatively, any one of the methods may further comprise applying one or more image augmentation techniques to the background image, in particular prior to combining the first object image and / or the second object image with the background image. The one or more background(s) may be one or more background image(s) and / or background specification data associated with instructions for obtaining the one or more background(s). The instructions for obtaining the one or more background(s) may indicate one or more hue(s) and one or more corresponding location(s) within the one or more background(s).

[0046] In an embodiment, placing the one or more synthetic object image(s) on one or more background(s) may include extracting the target object from the one or more synthetic object image(s) and placing the one or more target object(s) on the one or more background(s). This feature allows for introducing the target object to the background, thereby defining the relation between the target object and the background. Any one of the methods may further comprise applying one or more image augmentation technique(s) to the extracted one or more target object(s), in particular prior to placing the one or more target object(s) on the one or more background(s). The one or more image augmentation technique(s) may change an orientation of the one or more target object(s), change a scaling associated with the one or more target object(s) and / or change one or more pixel value(s) associated with the one or more target object(s). Changing the one or more pixel value(s) associated with the one or more target object(s) may result in changing one or more hue(s) associated with the one or more target object(s). Examples for changing the orientation of the one or more target object(s) may include rotating or mirroring the one or more target object(s), in particular as compared to the one or more synthetic object image(s). Examples for changing the scaling associated with the one or more target object(s) may include increasing and / or decreasing the number of pixels associated with the one or more target object(s). By doing so, several training images with varying backgrounds can be obtained based on one generated backgrounds. This can be energy-efficient, especially when generating thousands of training images.

[0047] In an embodiment, any one of the methods may further comprise providing a target object mask indicative of a contour of the one or more target object(s). The one or more target object(s) are extracted from the one or more synthetic object image(s) according to the target object mask. By doing so, only the desired object will be inserted into the background image and the interaction between the object and the object background can be included into the training data while enabling to adapt the background to the needs of the user.

[0048] In an embodiment, the target object mask may be generated and / or provided by a mask generator. The mask generator may be parametrized and / or trained based on one or more mask training image(s) associated with the one or more target object(s) and an indication on the contour of the one or more target object(s). By doing so, the contour of the object is learned precisely by the mask generator. Therefore, only the desired object will be inserted into the background image and the interaction between the object and the object background can be included into the training data while enabling to adapt the background to the needs of the user.

[0049] In an embodiment, the object image generator may be further parametrized and / or trained to generate a plurality of synthetic object image(s), wherein at least a part of the one or more synthetic training image(s) may be associated with an object different from and / or of a different type than the one or more target object(s). Any one of the methods may further comprise generating one or more synthetic object image(s) associated with the one or more target object(s) and one or more object(s) different from the one or more target object(s). Placing the one or more target object(s) on one or more background(s) may comprise extracting the one or more target object(s) and the one or more object(s) different from the one or more target object(s) from the one or more synthetic object image(s) and placing the one or more target object(s) and the one or more object(s) different from the one or more target object(s) on the one or more background(s). Two or more synthetic object image(s) may be generated by the object image generator. The two or more synthetic object image(s) may be associated with the one or more target object(s) and the one or more object(s) different from the one or more target object(s).

[0050] In an embodiment, the one or more training images may be associated with two or more different target objects, in particular at least partially overlapping target objects. The object image generator may generate and / or provide two or more different synthetic object images. Additionally or alternatively, the one or more target object(s) associated with the one or more synthetic object image(s) may be placed on the one or more background(s) and the one or more target object may be changed, in particular by the one or more image augmentation techniques, and placed on the one or more background(s). The two or more different target object(s) are obtained by placing the one or more target object(s) associated with the one or more synthetic object image(s) on the one or more background(s), augmenting the one or more target object(s) and placing the one or more augmented target object(s) on the one or more background(s). Augmenting the one or more target object(s) may comprise applying one or more image augmentation techniques to the one or more target object(s). The one or more image augmentation technique(s) may change an orientation of the one or more target object(s), change a scaling associated with the one or more target object(s) and / or change one or more pixel value(s) associated with the one or more target object(s). Changing the one or more pixel value(s) associated with the one or more target object(s) may result in changing one or more hue(s) associated with the one or more target object(s). Examples for changing the orientation of the one or more target object(s) may include rotating or mirroring the one or more target object(s), in particular as compared to the one or more synthetic object image(s). Examples for changing the scaling associated with the one or more target object(s) may include increasing and / or decreasing the number of pixels associated with the one or more target object(s). By doing so, several training images with varying backgrounds can be obtained based on one generated backgrounds. This can be energy-efficient, especially when generating thousands of training images. By doing so, training images with a plurality and / or overlapping objects may be generated. This enables to carry out accurate object detection in scenarios with a plurality of objects to be detected and / or analyzed such as field trials. In an embodiment, any one of the methods as described herein may further comprise providing a selection of the one or more target object(s), in particular via a user interface. Providing the selection of the one or more target object(s) may initiate generating the one or more synthetic object image(s). The selection of the one or more target object(s) may be provided to the object image generator. The object image generator may be further parametrized and / or trained to generate the one or more synthetic object image(s) according to the provided selection. This allows the user to interact and tailor the training data to be generated to examples available. Hence, this feature enables a precise adaption of the training data to current use cases resulting in an improved accuracy of the object detection model.

[0051] In an embodiment, any one of the methods as described herein may further comprise providing the one or more training image(s), in particular via a user interface.

[0052] In an embodiment, between 50 % and 90% of an image area associated with one or more synthetic object images may comprise the one or more target object(s). Preferably, between 60 % and 85 %, more preferably between 65 % and 78 % of the image area associated with one or more synthetic object images may comprise the one or more target object(s). By setting the focus on the one or more target object(s), the object image generator is enabled to solidly generate realistic target object images. Less image area associated with the target objects would lead to a decreased quality of the generated target objects. More image area associated with the target objects would lead to a decreased quality of the contours of the generated target objects. Followingly, choosing the image area associated with the target objects accordingly, the quality of the inlay and the contours of the target objects is balanced leading to overall realistic synthetic images.

[0053] In an embodiment, providing the one or more synthetic object image(s) to the mask generator may comprise providing a numerical representation of the one or more synthetic object image(s) to the mask generator. The numerical representation of the one or more synthetic object image(s) may be provided by the object image generator, in particular by one or more hidden and / or intermediate layer(s) of the object image generator. The numerical representation of the one or more synthetic object image(s) may be associated with a different data structure than the random tensor and / or the one or more synthetic image(s). The numerical representation of the one or more synthetic object image(s) may be obtained by processing the random tensor by at least a part of the object image generator. The one or more synthetic object image(s) may be obtained by processing the numerical representation of the one or more synthetic object image(s) by at least a part of the object image generator.

[0054] In the following, terminology as used herein and / or the technical field of the present disclosure will be outlined by ways of definitions and / or examples. Where examples are given, it is to be understood that the present disclosure is not limited to said examples. These and other objects, which become apparent upon reading the following description, are solved by the subject matters of the independent claims. The dependent claims refer to embodiments of the invention.

[0055] BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS

[0056] In the following, the present disclosure is further described with reference to the enclosed figures. The same reference numbers in the drawings and this disclosure are intended to refer to the same or like elements, components, and / or parts.

[0057] FIG. 1 illustrates an embodiment of an operating system 102 of a chemical production facility.

[0058] FIG. 2 illustrates an embodiment of monitoring and / or controlling a chemical production facility 206 based on an indication on the presence of a target object within an image scene.

[0059] FIG. 3 illustrates an embodiment of a method for generating one or more training image(s) for parametrizing an object detection model.

[0060] FIG. 4 illustrates an embodiment of generating one or more object images.

[0061] FIG. 5 illustrates an embodiment of generating one or more one or more background(s)s.

[0062] FIG. 6 illustrates an embodiment of generating one or more object masks.

[0063] FIG. 7 illustrates an embodiment of generating one or more object masks.

[0064] FIG. 8A illustrates an embodiment for generating a training image by extracting an object from an object image.

[0065] FIG. 8B illustrates an embodiment for generating a training image by placing the target object on the one or more background(s).

[0066] FIG. 9 illustrates an embodiment of an object detection model. DETAILED DESCRIPTION

[0067] The following embodiments are mere examples for implementing the method, the system or application device disclosed herein and shall not be considered limiting.

[0068] FIG. 1 illustrates an embodiment of an operating system 102 of a chemical production facility.

[0069] The chemical production facility may be controlled by the operating system 102. The operating system 102 may comprise an object image providing engine 104. The object image providing engine 104 may be configured for providing one or more synthetic object image(s) associated with one or more target object(s). The object image providing engine 104 may include an object image generator providing engine 106 and / or an object image providing engine 108. The object image generator providing engine 106 may be configured for providing an object image generator. The object image generator may be configured for generating the one or more synthetic object image(s), in particular in response to receiving a trigger for generating the one or more synthetic object image(s). The trigger for generating the one or more synthetic object image(s) may be a random tensor. For example, the object image generator may be a generative adversarial network (GAN) such as a conditional GAN, a normalizing flow, a variational autoencoder, a diffusion model or the like. These machine learning models are explained in more detail in the context of FIG. 4. The one or more synthetic object image(s) may be provided by the object image providing engine 108 to an object extraction engine 130. The object extraction engine 130 may be configured for extracting the one or more target object(s) from the one or more synthetic object image(s). The one or more target object(s) may be extracted from the one or more synthetic object image(s) by using a target object mask. The target object mask may be indicative of the contour of the one or more target object(s) associated with the one or more synthetic object image(s). The object extraction engine 130 may receive the target object mask from an object mask providing engine 124. The object mask providing engine 124 may provide the target object mask generated by a mask generator providing engine 122. The mask generator providing engine 122 may include a mask generator configured for generating the target object mask. The mask generator may be configured for generating and / or providing the target object mask in response to receiving the one or more synthetic object image(s). The mask generator may be described in further detail in the context of FIG. 6 and Fig. 7. The mask generator providing engine 122 and object mask providing engine 124 may be included in a mask providing engine 118. The mask providing engine 1 18 may be configured for providing the target object mask to the training image providing engine 126, in particular the object extraction engine 130. The extracted one or more target object(s) may be provided to a training image generating engine 132. The training image generating engine 132 and the object extraction engine 130 may be included in a training image providing engine 126. The training image providing engine 126 may be configured for providing the one or more training image(s) to an object detection engine 134, in particular for parametrizing and / or training an object detection model. The training image generating engine 132 may be configured for generating the one or more training image(s) by receiving the one or more synthetic object image(s) from the object extraction engine 130 and placing the one or more target object(s), in particular as extracted by the object extraction engine 130, on the one or more background(s). The one or more background(s) may be provided by a background providing engine 1 10. The background providing engine 110 may include a background generator providing engine 112 and a background generator engine 114. The background generator providing engine 112 may be configured for providing a background generator. The background generator may be configured for generating the one or more background(s). In particular, the background generator providing engine 112 may be configured for parametrizing the background generator based on one or more historical background(s), in particular a plurality of different historical background(s). The background generator may be received by the background generator engine 114. The background generator engine 114 may be configured for generating one or more background(s) by providing one or more random tensors to the background generator. The background generator may be configured to generate the one or more background(s) from the one or more random tensors. Hence, the one or more background(s) may be one or more synthetic background(s). The background generator may be configured to receive a random tensor, e.g. a tensor indicative of a random distribution of numerical values. The random tensor may be associated with a noise image, e.g. an image showing a distribution of pixel values according to the random distribution of numerical values. The background generator may be a GAN. The background generator may be further described in FIG. 5. The background generator engine 114 may provide the generated one or more background(s) to the training image generating engine 132. The generated one or more training image(s) may be provided to an object detection engine 134. The object detection providing engine 138 may be configured for providing an object detection model. The object detection model may be parametrized and / or trained for detecting a target object in an image scene. The object detection model may be a classification model. The classification model may be configured for classifying the image scene according to the presence of the one or more target object(s). The object detection model may generate the indication on the presence of the one or more target object(s) in response to receiving the image scene. The classification model may comprise one or more input layer(s) for receiving the image scene. The one or more input layer(s) may be configured.

[0070] Further, the object detection providing engine 138 may be configured for providing the object detection model to an indication providing engine 142. The indication providing engine 142 may be configured for providing an indication on a presence of the one or more target object(s) within the image scene in response to receiving the image scene. An input and / or output interface 140 may provide the image scence to the indication providing engine 142. Further, the indication providing engine 142 may provide the indication to the input and / or output interface 140 for controlling and / or monitoring the chemical production facility. The indication on the presence of the target object within the image scene may be generated by the object detection model in response to receiving the image scene.

[0071] FIG. 2 illustrates an embodiment of monitoring and / or controlling a chemical production facility 206 based on an indication on the presence of a target object within an image scene. The chemical production facility 206 may be configured for producing the chemical product 212 from the input material A 204. The input material A 204 may be the target object to be detected. Efficient production of the chemical product 212 may require reliable detection of input material A 204. The input material A 204 may be associated with a manufacturing defect. The manufacturing defect may result in changing the chemical properties associated with the input material A 204. Hence, the manufacturing defect may hinder the production of the chemical product 212. The potential manufacturing defect may visible in an image scene 202 associated with the input material A 204. Hence, the manufacturing defect may be detected by the object detection model 208 upon receiving the image scene 202 associated with the input material A 204. The object detection model 208 may generate an indication on the presence of the input material A within the image scene 210 if the input material A 204 may be independent of manufacturing defects. Otherwise, the object detection model may provide an indication on the absence of the input material A within the image scene 202. Providing indication on the presence of the input material A within the image scene 210 may be referred to as detecting the input material A 204 in the image scene. The indication on the presence of the input material A within the image scene 210 may be provided to the chemical production facility 206, in particular a control engine for controlling the chemical production facility 206 and / or a monitoring engine for monitoring the chemical production facility 206. Upon receiving indication on the presence of the input material A within the image scene 210, the chemical production facility 206 may produce the chemical product 212 from the input material A 204 associated with the image scene from which indication on the presence of the input material A within the image scene 210 was generated.

[0072] FIG. 3 illustrates an embodiment of a method for generating one or more training image(s) for parametrizing an object detection model.

[0073] One or more background(s) may be provided, e.g. by a background generator, 316 as described within the context of FIG. 1.

[0074] One or more synthetic object image(s) associated with the one or more target object(s) may be provided 318, e.g. by an object image generator, as described within the context of FIG. 1 .

[0075] A target object mask associated with a contour of the one or more target object(s) may be provided 320, e.g. by a mask generator, as described within the context of FIG. 1.

[0076] The one or more target object(s) may be extracted from the one or more synthetic object image(s) by applying the target object mask to the one or more synthetic object image(s) 322. The target object mask may indicate the location of the one or more target object(s) within the one or more synthetic object image(s). Hence, the location of the one or more target object(s) within the one or more synthetic object image(s) may be identified based on the target object mask. This may be described in further detail in FIG. 8A. The one or more extracted target object(s) may be augmented 324. Augmenting the one or more extracted target object(s) may include for example rotating and / or scaling. Augmenting the one or more target object(s) may generate two or more different target object(s) from the one or more target object(s). Hence, the number of synthetically generated target objects with different orientations can be increased. This enables tailoring the synthetic data according to the use case and generating training image(s) of rare events in a sufficient quantity and quality.

[0077] The one or more target object(s) may be placed on the one or more background(s) 326. The one or more target object(s) may be placed in an image area larger than the image area associated with the one or more target object(s) and filling the image area associated with background by one or more hue(s), e.g. according to a selection received via a user interface and / or as instructed by a request for generating one or more training image(s). In another embodiment, the one or more background(s) may be generated by randomly selecting one or more background(s) from a plurality of backgrounds.

[0078] The one or more training image(s) may be provided 330 e.g. via a user interface and / or to an object detection model for parametrizing and / or training the object detection model as described in the context of FIG. 1.

[0079] FIG. 4 illustrates an embodiment of generating one or more object images 410.

[0080] Hence, the object image 410 may be generated by an object image generator 402. For this purpose, the object image generator 402 may receive a trigger for generating the one or more object image 806. The trigger can be a noise image. The noise image may show a random distribution of pixel values and / or the noise image may comprise gaussian noise. The object image generator 402 may generate the one or more object images 806 based on being provided with the trigger for generating the one or more object images 806. In this context, generating the one or more object images 410may refer to transforming the received trigger into the one or more object images 410 by passing the trigger through the layers of the object image generator 402. Hence, the object image generator 402 may comprise two or more layers. The trigger may be received at an input layer of the object image generator 402. The trigger may be passed to a hidden layer of the object image generator 402. Passing the trigger to another layer and / or passing the trigger through the object image generator 402 may comprise performing one or more mathematical operations to the received trigger. Transforming the trigger and / or performing one or more mathematical operations may result in changing one or more parts of the trigger, preferably changing a plurality of pixel values. In particular, at least two parts of the trigger may be changed, wherein the degree of change is different between the at least two parts. By passing the trigger through the object image generator 402 the object being comprised in the one or more to be generated object images 806 may evolve. Examples for object image generators 402 may include a generator trained based on a detector trained for distinguishing real data from synthetic data according to the concept of generative adversarial networks (GAN), a variational autoencoder, a diffusion model or the like. Where the object image generator 402 may be and / or may comprise a generator of a GAN, the object image generator 402 may be trained based on the feedback from a detector of a GAN. The feedback from the detector may be indicative of whether the generated data may be classified as real data or synthetic data. Hence, the object image generator 402 may be trained according to a classification of previously generated data e.g. example object image 410. The object image generator 402 may be trained and / or parametrized based on one or more historical images 406. GANs are known in the art such as XXX. An example of a GAN implementation can be seen in FIG. 4. The object image generator 402 may be a conditional GAN. The conditional GAN may be parametrized and / or trained to receive a the random tensor 412 and the indication on the synthetic data to be generated 414, in particular concatenating the random vector the indication on the synthetic data to be generated 414 into one vector. The indication on the synthetic data to be generated 414 may be a one-hot vector. The indication on the synthetic data to be generated 414 may be a label indicating the synthetic data to be generated. Further, the conditional GAN (cGAN) may be parametrized to receive the concatenated vector, i.e. the random tensor 412 and the indication on the synthetic data to be generated 414. The cGAN may comprise a plurality of expanding blocks. The expanding blocks may comprise one or more ReLU (rectified linear unit) layers 418 and one or more transposed convolutional layers 420. The one or more expanding block(s) may be configured to change a data format associated with the concatenated input data to a data format associated with the one or more synthetic object image(s).

[0081] The object image generator 402 may be parametrized and / or trained by providing random tensors 412 and optionally indications on the synthetic data to be generated 414 and generating one or more synthetic object image(s) from the random tensors 412, optionally indications on the synthetic data to be generated 414. During parametrizing and / or training, the generated one or more synthetic object image(s) may be provided to a discriminator model. The discriminator model may be a classification model. The discriminator model may be configured, in particular initialized for classifying whether the received input data to the discriminator model may be synthetically generated, i.e. generated by the object image generator or not. A deviation, i.e. a loss, between the provided classification by the discriminator model and a predefined and / or correct classification may be determined. Parameters associated with the discriminator model and the object image generator may be updated based on the determined deviation, e.g. by deploying a backpropagation algorithm. This may include determining a discriminator loss and a generator loss from the deviation between the provided classification and the predefined classification.

[0082] In an embodiment, the indication on the synthetic data to be generated 414 may comprise text indicative of the synthetic data to be generated. The object image generator may comprise a text encoder trained and / or parametrized to transform the text comprised in the trigger into a machine-processable representation of the indication on the synthetic data to be generated 414. The machine-processable representation may be a tensor. The one or more synthetic object image(s) 410 may be generated based on the machine-processable representation and the random tensor 412. In other embodiments, the indication on the synthetic data to be generated 414 may include an image, a sketch or the like. Here, the object image generator may include an image encoder.

[0083] In an embodiment, a numerical representation of the one or more synthetic object image(s) may be obtained by processing the random tensor and optionally the indication on the synthetic data to be generated by the object image generator. Preferably, the numerical representation may be output data of the one or more transposed convolutional layer(s). The numerical representation of the one or more synthetic object image(s) may be associated with a different data structure than the synthetic object image(s). The one or more object image(s) may be obtained by processing the numerical representation of the one or more synthetic object image(s).

[0084] FIG. 5 illustrates an embodiment of generating one or more one or more background(s)s 506. One or more background(s) 506 may be a synthetic one or more background(s). Hence, the one or more background(s) 506 may be generated by a background generator 504. For this purpose, the background generator 504 may receive a trigger for generating the one or more one or more background(s)s 818. The trigger can be a noise image. The noise image may show a random distribution of pixel values and / or the noise image may comprise gaussian noise. The background generator may generate the one or more one or more background(s)s 506 based on being provided with the trigger for generating the one or more one or more background(s)s 506. In this context, generating the one or more one or more background(s)s 506 may refer to transforming the received trigger into the one or more one or more background(s)s 506 by passing the trigger through the layers of the background generator 504. Hence, the background generator 504 may comprise two or more layers. The trigger may be received at an input layer of the background generator 504. The trigger may be passed to a hidden layer of the background generator 504. Passing the trigger to another layer and / or passing the trigger through the background generator 504 may comprise performing one or more mathematical operations to the received trigger. Transforming the trigger and / or performing one or more mathematical operations may result in changing one or more parts of the trigger, preferably changing a plurality of pixel values. In particular, at least two parts of the trigger may be changed, wherein the degree of change is different between the at least two parts. By passing the trigger through the background generator 504 the background being comprised in the one or more to be generated one or more background(s)s 506 may evolve.

[0085] Examples for background generators 504 may include a generator trained based on a detector trained for distinguishing real data from synthetic data according to the concept of generative adversarial networks (GAN), a variational autoencoder, a diffusion model or the like. Where the background generator 504 may be and / or may comprise a generator of a GAN, the background generator 504 may be trained based on the feedback from a detector of a GAN. The feedback from the detector may be indicative of whether the generated data may be classified as real data or synthetic data. Hence, the background generator 504 may be trained according to a classification of previously generated data e.g. example one or more background(s)s 506. The background generator 504 may be trained and / or parametrized based on one or more historical one or more background(s)s 502.

[0086] Preferably, the background generator 504 may be a conditional GAN. The conditional GAN may be parametrized and / or trained to receive a trigger comprising text, wherein the text may be indicative of the to be generated one or more background(s)s 506. For this purpose the conditional GAN may comprise a text encoder trained and / or parametrized to transform the text comprised in the trigger into a machine-processable representation of the trigger to generate the one or more background(s)s 506. The one or more one or more background(s)s 506 may be generated based on the machine-processable representation. Conditional GANs maybe parametrized and / or trained to received further input such as a noise image, a sketch or the like. Hence, the conditional GAN may receive two or more modalities of trigger for generating the one or more one or more background(s)s 506. Similarly, variational autoencoders may receive two or more modalities of trigger. In the case of a diffusion model, the trigger may comprise noise, in particular a noise image and / or the trigger may be monomodal.

[0087] FIG. 6 illustrates an embodiment of generating one or more target object masks 604.

[0088] For generating one or more target object masks 604 one or more object images 606 may be received. The object image 606 may show the one or more target object(s) and the at least one object background. The object image 606 may be associated with and / or may compose of one or more part(s). The target object mask 604 may indicate whether the one or more part(s) may be associated with the one or more target object(s) and / or the one or more object background(s). An example of a target object mask 604 may be shown in FIG. 6. The target object mask 604 may comprise one or more labels indicative of whether the one or more part(s) associated with the one or more label(s) may be associated with the one or more target object(s) and / or one or more object background(s). For example, a 1 may indicate the presence of the one or more target object(s) within the one or more part(s) of the object image associated with the label 1. Analogously, 0 may indicate that the one or more part(s) of the object image associated with the label 0 may be associated with the object background and / or may be independent of the target object. In particular, the pixels of the object image may be assigned with a label indicative of the presence of the object within the part of the object image and / or an indication on whether the one or more parts associated with the indication may comprise the object separately. In particular, the target object mask 604 may comprise and / or assign an indication on the presence of the one or more target object(s) to the one or more part(s) of the object image 606. In an embodiment, the target object mask 614 may comprise and / or assign an indication on the presence of a plurality of target object(s) to the one or more part(s) of the object image 606. This may be known as semantic segmentation.

[0089] Followingly, the mask generator 602 may be a classification model. The classification model may be trained and / or parametrized to classify the one or more parts, in particular the one or more pixels, of the received object image. Hence, the mask generator 602 may be trained and / or parametrized to provide one or more indications on whether the one or more parts associated with the indication may comprise the object. The mask generator 602 may be trained on object images 606 to classify parts of the object images. Based on the indication on whether the one or more parts associated with the indication may comprise the 608, the 608 may be extracted, preferably cutted, from the object image comprising the 608 and background. For example, the indication on whether the one or more parts associated with the indication may comprise the object may be of the same format as the object image. Further, the parts of the object image may be assigned with a value indicating whether the one or more parts may comprise the 608.

[0090] In an example, the mask generator 602 may be comprise one or more blocks for obtaining a contracted representation of the object image 606 and one or more blocks for obtaining the target object mask 614 from the contracted representation of the object image 606. The contracted representation of the object image 606 may be a numerical representation of the object image 606, e.g. in a latent space. The object image 606 may be associated and or may be a 2-dimensional tensor. The contracted representation may be associated and / or may be a higherdimensional tensor than the object image 606.

[0091] In an example, the mask generator 602 may be a U-Net (doi 1505.04597v1 ). Hence, the mask generator 602 may comprise one or more contracting block(s) for generating the contracted representation of the object image 606 from the object image 606. The contracting block(s) may comprise a 3x3 convolutional layer 616 followed by a rectified linear unit (ReLU) layer 618, a 3x3 convolutional layer 616, a ReLU layer 618 and a pooling layer 620. Preferably, the mask generator 602 may comprise a plurality of contracting blocks, e.g. 4 as in the example of U- Net. The one or more contracting block(s) may be configured to increase the dimensionality of received input data. The contracted representation of the object image 606 may be provided to one or more expanding block(s). The expanding block(s) may comprise a 3x3 convolutional layer 616 followed by a rectified linear unit (ReLU) layer 618, a 3x3 convolutional layer 616, a ReLU layer 618 and an upsampling layer 626. The upsampling layer 626 may be a 2x2 convolutional layer. The expanding blocks, in particular the upsampling layer 626 may reduce the dimensionality of received input data. The output of the upsampling layer 626 may be concatenated with output from the one or more contracting block(s). This may include cropping the output of the one or more contracting block(s) to match a format of the output of the upsampling layer 626. Preferably, the output of the one or more contracting block(s) of a first contracting block may be concatenated with the output of the upsampling layer 626 of the last expanding block. The number of contracting blocks may be equal to the number of expanding blocks. Output of the contracting blocks of a number M may be concatenated with output of the expanding blocks of a number N-M. N may be the total number of contracting blocks.

[0092] Further convolutional layer(s) 628 may format the data received from the one or more expanding block(s) to a target image format associated with the object image 606. FIG. 7 illustrates an embodiment of generating one or more target object masks 704.

[0093] The mask generator may be provided with the numerical representation of the one or more synthetic object image(s) as described in the context of Fig.4. The numerical representation may be a hidden state of the object image generator. The mask generator may be configured for receiving the numerical representation of the one or more synthetic object image(s). The mask generator may comprise an upsampling layer for generating a feature map associated with the one or more synthetic object image(s). The feature map may be a multidimensional representation of the one or more synthetic object image(s). The feature map may comprise a number of parts. A part of the feature map may correspond to a pixel associated with the one or more synthetic object image(s). Further, the mask generator may comprise one or more classification block(s) for classifying the pixel associated with the one or more synthetic object image(s), in particular pixel-wise. Hence, one or the parts may be provided to the one or more classification block(s) and / or an indication on the presence of the one or more target object(s) may be provided and / or generated by the one or more classification block(s). For example, the classification block(s) may comprise of a fully-connected layer, a ReLU layer and / or a batch normalization layer. Providing the parts of the feature map part-by-part to the mask generator allows to classify the feature map according to the presence of the one or more target object(s) by providing indications on the presence of the one or more target object(s) associated with the parts of the feature map. By collecting the indications on the presence of the one or more target object(s) associated with the parts of the feature map the target object mask indicative on the presence of the one or more target object image(s) within the one or more synthetic object image(s) may be obtained.

[0094] For generating one or more target object masks 704 one or more object images 706 may be received. The object image 706 may show the one or more target object(s) and the at least one object background. The object image 706 may be associated with and / or may compose of one or more part(s). The target object mask 704 may indicate whether the one or more part(s) may be associated with the one or more target object(s) and / or the one or more object background(s). An example of a target object mask 704 may be shown in FIG. 7. The target object mask 704 may comprise one or more labels indicative of whether the one or more part(s) associated with the one or more label(s) may be associated with the one or more target object(s) and / or one or more object background(s). For example, a 1 may indicate the presence of the one or more target object(s) within the one or more part(s) of the object image associated with the label 1. Analogously, 0 may indicate that the one or more part(s) of the object image associated with the label 0 may be associated with the object background and / or may be independent of the target object. In particular, the pixels of the object image may be assigned with a label indicative of the presence of the object within the part of the object image and / or an indication on whether the one or more parts associated with the indication may comprise the object separately. In particular, the target object mask 704 may comprise and / or assign an indication on the presence of the one or more target object(s) to the one or more part(s) of the object image 706. In an embodiment, the target object mask 714 may comprise and / or assign an indication on the presence of a plurality of target object(s) to the one or more part(s) of the object image 706. This may be known as semantic segmentation.

[0095] Followingly, the mask generator 702 may be a classification model. The classification model may be trained and / or parametrized to classify the one or more parts, in particular the one or more pixels, of the received object image. Hence, the mask generator 702 may be trained and / or parametrized to provide one or more indications on whether the one or more parts associated with the indication may comprise the object. The mask generator 702 may be trained on object images 706 to classify parts of the object images. Based on the indication on whether the one or more parts associated with the indication may comprise the 708, the 708 may be extracted, preferably cutted, from the object image comprising the 708 and background. For example, the indication on whether the one or more parts associated with the indication may comprise the object may be of the same format as the object image. Further, the parts of the object image may be assigned with a value indicating whether the one or more parts may comprise the 708.

[0096] In an example, the mask generator 702 may be comprise one or more blocks for obtaining a contracted representation of the object image 706 and one or more blocks for obtaining the target object mask 714 from the contracted representation of the object image 706. The contracted representation of the object image 706 may be a numerical represenation of the object image 706, e.g. in a latent space. The object image 706 may be associated and or may be a 2-dimensional tensor. The contracted representation may be associated and / or may be a higherdimensional tensor than the object image 706.

[0097] In an example, the mask generator 702 may be a U-Net (doi 1505.04597v1 ). Hence, the mask generator 702 may comprise one or more contracting block(s) for generating the contracted representation of the object image 706 from the object image 706. The contracting block(s) may comprise a 3x3 fully-connected layer 716 followed by a rectified linear unit (ReLU) layer 718, a 3x3 fully-connected layer 716, a ReLU layer 718 and a Batch normalization layer 720. Preferably, the mask generator 702 may comprise a plurality of contracting blocks, e.g. 4 as in the example of U-Net. The one or more contracting block(s) may be configured to increase the dimensionality of received input data. The contracted representation of the object image 706 may be provided to one or more expanding block(s). The expanding block(s) may comprise a 3x3 fully-connected layer 716 followed by a rectified linear unit (ReLU) layer 718, a 3x3 fully-connected layer 716, a ReLU layer 718 and an upsampling layer 726. The upsampling layer 726 may be a 2x2 convolutional layer. The expanding blocks, in particular the upsampling layer 726 may reduce the dimensionality of received input data. The output of the upsampling layer 726 may be concatenated with output from the one or more contracting block(s). This may include cropping the output of the one or more contracting block(s) to match a format of the output of the upsampling layer 726. Preferably, the output of the one or more contracting block(s) of a first contracting block may be concatenated with the output of the upsampling layer 726 of the last expanding block. The number of contracting blocks may be equal to the number of expanding blocks. Output of the contracting blocks of a number M may be concatenated with output of the expanding blocks of a number N- M. N may be the total number of contracting blocks.

[0098] Further convolutional layer(s) 728 may format the data received from the one or more expanding block(s) to a target image format associated with the object image 706.

[0099] FIG. 7 illustrates an embodiment of generating one or more object masks.

[0100] FIG. 8A illustrates an embodiment for generating a training image 810 by extracting a target object 802 from an object image 806. The target object 802 may be extracted from the object image 806 by cutting the target object 802 based on the target object 802 mask from the object image 806. Further image augmentation techniques may be applied to the object image 806 prior to extracting and / or prior to inserting the target object 802 into the one or more background(s) 818 as described within the context of FIG. 8B. The size of the object mask 804 may be equal to the size of the target object 802 within the object image 806. Further, the object mask 804 may have the shape of the target object 802 within the object image 806.

[0101] By doing so, only the desired target object 802 will be inserted into the one or more background(s) 818 and the interaction between the target object 802 and the object background can be included into the training data while enabling to adapt the one or more background(s) to the needs of the user.

[0102] FIG. 8B illustrates an embodiment for generating a training image by placing the target object on the one or more background(s).

[0103] Hence, the target object 816 as generated based on the synthetic object image and the target object mask as described within the context of FIG. 8A may replace and / or fill an area of the one or more background(s) 818 equal to the size of the target object 816. The area of the one or more background(s) 818 equal to the size of the target object 808 may be merged, e.g. by adding the pixel values or substituting the pixel values of the area associated with at least a part of the one or more background(s) by the pixel values associated with the one or more target object(s). The area of the one or more background(s) 818 equal to the size of the target object 808 may have the same shape as the target object 808.

[0104] FIG. 9 illustrates an embodiment of an object detection model 902.

[0105] The object detection model 902 may be configured to classify if the one or more target object(s) 920 may be present in the image scene 916. In the example, the image scene 916 may show the target object 920. The object detection model 902 may provide an indication on the presence of the one or more target object(s) within the image scene. The indication on the presence of the one or more target object(s) within the image scene may be a vector. The vector may comprise a plurality of numerical values, e.g. at least two. The numerical values may indicate a certainty and / or a confidence score associated with the indication on the presence of the one or more target object(s) within the image scene. In the example, the object detection model may provide an indication that the target object present in the image scene 916.

[0106] The object detection model 902 may comprise of one or more convolutional block(s). The convolutional block(s) may comprise one or more convolutional layer(s) 904 and one or more pooling layer(s) 906. The convolutional block(s) 904 may be configured for reducing the dimensionality associated with the image scene. The one or more convolutional block(s) may provide a machine-processable representation of the image scene 916 such as a tensor. In an example, two convolutional blocks with one fully-connected layer 716 and one Batch normalization layer 720 per convolutional block may be implemented as LeNet or two convolutional layers with one fully-connected layer 716 and one Batch normalization layer 720 per convolutional block and one convolutional block with three fully- connected layers 716 and one Batch normalization layer 720 may be implemented as AlexNet.

[0107] The tensor may be provided to one or more fully connected layer(s) 908. The one or more fully connected layer(s) 908 may process the machine-processable representation of the image scene 916 to a vector with a number of entries equal to a number of classes associated with the object detection model 902. A softmax function 910 may determine numbers in a range between 0 and 1 associated with a distribution equal to the output of the one or more fully connected layer(s) 908. The output of the softmax function 910 may be the certainties and / or confidence scores associated with the classes of the object detection model 902.

[0108] In an embodiment, the object detection model 902 may be associated with a plurality of classes indicative of the presence of the one or more target object(s) within the image scene and indicative of the location of the one or more target object(s) within the image scene. The object detection model 902 may be parametrized and / or trained to provide an indication on the presence and the location of the one or more target object(s) within the image scene. To do so, the one or more fully connected layer(s) 908 may be configured to change a format associated with the machine-processable representation of the image scene to a format of the indication on the presence and the location of the one or more target object(s). For example, The indication on the presence and the location of the one or more target object(s) may be a vector with a plurality of entries, wherein a part of the vector may indicate one or more classes related to the presence of the one or more target object(s) within the image scene and one or more classes related to the location of the one or more target object(s) within at least a part of the image scene. The number of classes related to the location may be equal to a number of parts of the image scene, in particular may be equal to the number of pixels associated with the image scene. The object detection model 902 may be parametrized and / or trained based on historical image scenes and corresponding indications on the presence and the location of the one or more target object(s). In an embodiment, the indication on the location of the one or more target object(s) may be provided by the object localization model. The object localization model may be parametrized based on historical image scenes and corresponding indications on the location of the one or more target object(s). The object localization model may be parametrized and / or trained analogous to the object detection model 902. The object localization model may be a classification model analogous to the object detection model 902.

[0109] The present disclosure has been described in conjunction with preferred embodiments and examples as well. However, other variations can be understood and effected by those persons skilled in the art and practicing the claimed invention, from the studies of the drawings, this disclosure and the claims. Notably, in particular, the any steps presented can be performed in any order, i.e. the present invention is not limited to a specific order of these steps. Moreover, it is also not required that the different steps are performed at a certain place or at one node of a distributed system, i.e. each of the steps may be performed at different nodes using different equipment / data processing.

[0110] As used herein ..determining" also includes ..initiating or causing to determine", “generating" also includes ..initiating and / or causing to generate" and “providing” also includes “initiating or causing to determine, generate, select, send and / or receive”. “Initiating or causing to perform an action” includes any processing signal that triggers a computing node or device to perform the respective action.

[0111] In the claims as well as in the description the word “comprising” or “including” or similar wording does not exclude other elements or steps and shall not be construed limiting to the elements or steps lined out. The indefinite article “a” or “an” does not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation or further elements may be included.

[0112] Providing in the scope of this disclosure may include any interface configured to provide data. This may include an application programming interface, a human-machine interface such as a display and / or a software module interface. Providing may include communication of data or submission of data to the interface, in particular display to a user or use of the data by the receiving entity.

[0113] Any disclosure and embodiments described herein relate to the methods, the systems, devices, the computer program element lined out above and vice versa. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples and vice versa.

Claims

CLAIMSWhat is claimed is:

1. A method, in particular a computer-implemented method, for detecting one or more target object(s), the method comprising: providing an image scene including one or more target object(s), providing the image scene to an object detection model for determining an indication on the presence of the one or more target object(s) within the image scene, wherein the object detection model is parametrized based on one or more training image(s) and one or more corresponding indication(s) on the presence of the one or more target object(s) within the one or more training image(s), and wherein at least a part of the one or more training image(s) are synthetic images obtained by providing one or more synthetic object image(s) comprising one or more target object(s) and object background, and placing the target object extracted from the one or more synthetic object image(s) on one or more background(s), providing the indication on the presence of the target object within the image scene.

2. The method according to claim 1 , wherein the object detection model is further parametrized based on indications on the location of the one or more target object(s) within the one or more training image(s).

3. The method according to claim 2, wherein the indication on the location of the one or more target object(s) is provided and / or generated by an object localization model, wherein the object localization model is parametrized based on localization training images including the one or more target object(s) and corresponding indications on the location of the one or more target object(s).

4. The method according to claim 2, wherein the object detection model is an object detection and localization model, wherein the object detection and localization model is further configured for providing the indication on the location of the one or more target object(s) within the image scene.

5. A computer-implemented method for generating one or more training image(s) for training an object detection model for determining an indication on the presence of one or more target object(s) within the image scene, the method comprising: providing one or more synthetic object image(s) comprising one or more target object(s) and object background, andplacing the target object extracted from the one or more synthetic object image(s) on one or more background(s), providing the one or more training image(s) for parametrizing the object detection model for determining an indication on the presence of the one or more target object(s) within the image scene.

6. The method according to any one of claims 1 to 5, wherein between 50 % and 90% of an image area associated with one or more synthetic object images comprises the one or more target object(s).

7. The method according to any one of claims 1 to 6, further comprising applying one or more image augmentation technique(s) to the extracted one or more target object(s) prior to placing the one or more target object(s) on the one or more background(s).

8. The method according to any one of claims 1 to 7, further comprising providing a target object mask indicative of a contour of the one or more target object(s), and wherein the one or more target object(s) are extracted from the one or more synthetic object image(s) according to the target object mask.

9. The method according to claim 8, wherein the target object mask is generated and / or provided by a mask generator, wherein the mask generator is parametrized and / or trained based on one or more mask training image(s) associated with the one or more target object(s) and an indication on the contour of the one or more target object(s).

10. The method according to claim 9, wherein the target object mask is generated by providing a numerical representation of the one or more synthetic object image(s) to a mask generator, wherein the mask generator is parametrized and / or trained based on one or more mask training numerical representations of synthetic images of the one or more target object(s) and corresponding indications on the contour of the one or more target object(s).1 1. A computer-implemented method for training an object detection model for detecting one or more target object(s) within an image scene, the method comprising: providing one or more training image(s) associated with the one or more target object(s) as generated by any one of any one of claims 5 to 10, providing one or more indications on the presence of the one or more target object(s) within the one or more training image(s), training the object detection model based on the one or more training image(s) and the one or more indications on the presence of the one or more target object(s) within the one or more training image(s), optionally providing the object detection model.

12. Use of an object detection model as trained by claim 11 for detecting and / or localizing a target object within an image scene.

13. Use of an indication on the presence of a target object within an image scene as obtained by any one of claims 1 to 4 and claims 6 to 11 where dependent on claims 1 to 4 for controlling and / or monitoring a production of a chemical product and / or an application of a chemical product.

14. An apparatus comprising: a processor; and a memory storing instructions that, when executed by the processor, configure the apparatus to perform the any one of the methods according to any one of any one of claims 1 -10.

15. A training data set comprising one or more training image(s) generated according to any one of claims 5 to 9.

Citation Information

Patent Citations

  • System and method for generating training data for computer vision systems based on image segmentation

    US11475246B2

  • Teaching GAN (generative adversarial networks) to generate per-pixel annotation

    US20210089845A1

  • Generation of synthetic image data for computer vision models

    US10860836B1

  • Image composites using a generative adversarial neural network

    US20190251401A1

  • Data augmentation including background modification for robust prediction using neural networks

    US20220101047A1