synthetic data for the chemical industry

By generating synthetic object images to train an object detection model, the problem of detecting objects with different appearances in multi-object scenes is solved, improving detection accuracy and production efficiency, and optimizing the production process of chemical products.

CN122641872APending Publication Date: 2026-08-25BASF SE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202580011955.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-29
Filing Date
2025-01-21
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

In multi-object scenarios, it is difficult to reliably detect individual objects with different appearances, and existing technologies struggle to provide efficient and customized object detection solutions.

Method used

By generating synthetic object images that include both the target object and the background, an object detection model is trained. Realistic synthetic images are generated using models such as generative adversarial networks and variational autoencoders. Combined with image enhancement techniques, training data suitable for different backgrounds is generated, thereby improving the performance of the object detection model.

Benefits of technology

It achieves time-efficient and barrier-free object detection, improves the accuracy and applicability of object detection models, saves energy and time, and optimizes the production process of chemical products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

A method, in particular a computer-implemented method, for detecting one or more target objects, the method comprising: providing an image scene comprising one or more target objects; providing the image scene to an object detection model for determining an indication about a presence of the one or more target objects within the image scene, wherein the object detection model is parameterized based on one or more training images and one or more corresponding indications about a presence of the one or more target objects within the one or more training images, and wherein at least a part of the one or more training images is a synthetic image obtained by providing one or more synthetic object images comprising one or more target objects and an object background, and placing target objects extracted from the one or more synthetic object images on one or more backgrounds; providing an indication about a presence of the target objects within the image scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a computer-implemented method for detecting objects, a method for training an object detection model, a method for generating training data, the use of an object detection model, the use of training images, the use of indicating the presence of objects in an image, an apparatus for detecting objects, an apparatus for training an object detection model, and an apparatus for generating training data. Background Technology

[0002] Image recognition based on convolutional neural networks has been a continuously evolving field for many years. It is used in various technological applications, such as agriculture and industrial production. However, reliably detecting and recognizing individual objects with diverse appearances remains challenging in multi-object scenarios.

[0003] US11475246B2 describes a system and method for training a model using a training dataset. The training dataset can consist only of real data, only of synthetic data, or any combination of synthetic and real data. Images are segmented to define objects with known labels. The objects are pasted onto a background to generate a synthetic dataset. Various aspects of the invention include generating data to supplement or enhance real data. Labels or attributes can be automatically added to the data during data generation. This data can be generated using seed data. This data can be generated using synthetic data. This data can be generated from any source, including user ideas or memories. Using the training dataset, adaptive models for various domains can be trained.

[0004] US20210089845A1 describes a method and apparatus for jointly synthesizing an image with pixel-by-pixel annotations using a generative adversarial network (GAN). The method includes: obtaining a first image from a GAN by inputting data into the GAN; inputting first feature values ​​obtained from at least one intermediate layer of the GAN based on the input of data into the GAN into a decoder; and obtaining a first semantic segmentation mask from the decoder based on the input of the first feature values ​​into the decoder.

[0005] WO 2023242236A1 relates to image processing. The computer-implemented method described herein can be used to improve computer vision techniques for application in agricultural technology and production environments. Summary of the Invention

[0006] On the other hand, this disclosure relates to a method for detecting computer implementations of one or more target objects, particularly a method for detecting computer implementations, the method comprising:

[0007] Provide an image scene that includes one or more target objects.

[0008] The image scene is provided to an object detection model for determining indications of one or more target objects present in the image scene, wherein the object detection model is parameterized based on one or more training images and one or more corresponding indications of one or more target objects present in the one or more training images, and wherein at least a portion of the one or more training images are synthetic images obtained in the following manner:

[0009] Provide one or more composite object images, including one or more target objects and object backgrounds, and

[0010] The target object extracted from the one or more composite object images is placed on one or more backgrounds.

[0011] Provides indication of the target object within the image scene.

[0012] On the other hand, this disclosure relates to a method, particularly a computer-implemented method, for generating one or more synthetic object images including one or more target objects and object backgrounds, wherein the synthetic object images are used to generate one or more training images for training an object detection model, the method comprising:

[0013] A request for generating one or more composite object images is provided to an object image generator, wherein the object image generator is parameterized to generate composite object images in response to receiving the request for generating one or more composite object images.

[0014] The object image generator generates one or more composite object images.

[0015] Provide one or more images of the composite object.

[0016] In another aspect, this disclosure relates to a computer-implemented method for generating training data, the training data including training images suitable for training an object detection model, the method comprising: receiving a request to generate a training image including a target object and a first background; generating a first object image including the target object and a target object background based on an object image generator, wherein the object image generator is parameterized and / or trained based on a first historical image including the target object and a first historical object background; receiving one or more backgrounds including the first background; combining the one or more backgrounds and the first object image to obtain the training image; and providing the training image.

[0017] On the other hand, this disclosure relates to the use of training images obtained by any of the methods described herein for training an object detection model.

[0018] On the other hand, this disclosure relates to the use of an indication of an object present in an image, obtained by any of the methods described herein, for the control and / or monitoring of the production and / or application of a chemical product.

[0019] On the other hand, this disclosure relates to a computer-implemented method for training an object detection model suitable for detecting objects within an image, the method comprising: receiving a training image generated as by any of the methods described herein; receiving an indication of the presence of an object within the training image; and training the object detection model based on the training image and the indication of the presence of an object within the training image.

[0020] On the other hand, this disclosure relates to the use of training images generated by any of the methods described herein for training an object detection model.

[0021] In another aspect, this disclosure relates to an apparatus comprising: a processor; and a memory storing instructions that, when executed by the processor, configure the apparatus to perform any of the methods described herein.

[0022] In another aspect, this disclosure relates to a computer-implemented method for detecting at least a portion of a crop in an image, the method comprising: receiving an image showing at least a portion of the crop; detecting the at least a portion of the crop within the image by providing the image to an image detection model; and receiving an indication regarding the presence of the at least a portion of the crop within the image, wherein the object detection model has been trained based on a plurality of training images, and wherein the plurality of training images are generated by: receiving one or more historical images including at least one or more portions of the crop; generating an object image based on an object image generator, wherein the object image includes at least one of the crop and a background; wherein the object image generator is parameterized and / or trained based on the one or more historical images to generate a plurality of object images; receiving one or more backgrounds; combining the one or more backgrounds and the object image to obtain the training image; and providing an indication regarding the presence of the at least a portion of the crop within the image. Example

[0023] Any disclosures, embodiments, and examples described herein relate to the methods, systems, apparatuses, uses, models, and training datasets listed above and below. Advantageously, the benefits provided by any embodiment and example also apply to all other embodiments and examples.

[0024] From raw material supply and growth monitoring to finished product quality control, object inspection plays a crucial role in producing reliable and high-quality products. Therefore, implementing use case-specific and reliable object inspection is essential.

[0025] The invention presented herein enables the customization of object detection to suit the needs of a use case by training an object detection model based on a first object image including the target object and its background. The resulting synthetic image provides realistic image features, and therefore, the object detection model trained on this synthetic image can provide improved performance. Furthermore, complete control over the generation of synthetic data for training the object detection model is provided. Realistically appearing synthetic target objects can be pasted in a customized manner, allowing for realistic overlapping objects on different backgrounds. Furthermore, existing workflows for evaluating the production and performance of chemical products can be improved, saving energy and time required for new development. This enables resource-efficient application and production of chemical products while shortening time-to-market during chemical product development.

[0026] Therefore, the present invention provides a time-efficient and barrier-free solution to the above-mentioned problems.

[0027] In this embodiment, the background may include scenery and / or one or more objects unrelated to the target object. The background may be defined based on one or more target objects. The background may include at least a portion of the environment surrounding one or more target objects.

[0028] In embodiments, a background generator can refer to a generative model. The background generator can be configured to generate one or more backgrounds, preferably multiple different backgrounds. Examples of generative models can include and / or can be based on generative adversarial networks, variational autoencoders, diffusion models, etc. The background generator can be parameterized and / or trained based on one or more historical backgrounds, particularly multiple different historical backgrounds. The historical backgrounds can differ from the one or more backgrounds. Preferably, the one or more backgrounds generated by the background generator can be synthetic backgrounds. The background generator can receive triggers for generating one or more backgrounds. Triggers for generating one or more backgrounds can initiate the generation of one or more backgrounds. Triggers for generating one or more backgrounds can be included in a request for generating training images and / or can be generated based on a request for generating training images. Triggers for generating one or more backgrounds can include text, audio, noise (particularly noisy images), and / or combinations thereof. Noisy images can refer to data representing noise, particularly data associated with the same format as the one or more backgrounds. Triggers for generating one or more backgrounds can be included in a request for generating one or more training images and / or can be generated in response to a request for generating one or more training images. The background generator can include at least one machine learning architecture and model parameters. Parameterizing the background generator can include initializing and / or defining model parameters, for example, by assigning random numbers to the model parameters. Training the background generator can include updating the model parameters. For example, the machine learning architecture can be or can include one or more of the following: linear regression, logistic regression, random forest, piecewise linear, nonlinear classifier, support vector machine, Naive Bayes classification, nearest neighbor, neural network, convolutional neural network, generative adversarial network, diffusion model, support vector machine or gradient boosting algorithm, etc. In the case of neural networks, the model can be a multi-scale neural network or a recurrent neural network (RNN), such as, but not limited to, gated recurrent unit (GRU) recurrent neural network, visual transformer or long short-term memory (LSTM) recurrent neural network.

[0029] In an embodiment, a target object mask may indicate the outline of one or more target objects. Further, the target object mask may indicate one or more portions of an image scene and / or training image associated with one or more target objects. For example, the target object mask may indicate portions of an image scene and / or training image associated with one or more target objects using numerical and / or Boolean values. Therefore, the target object mask may be associated with the same image format and / or an equal number of pixel values ​​as one or more target objects and / or one or more synthetic object images. The target object mask may further indicate the outline of one or more backgrounds. For example, the target object mask may indicate the outline of one or more target objects, particularly within one or more synthetic object images, using positive Boolean and / or numerical values ​​(e.g., 1 or values ​​> or equal to 0.5). The target object mask may indicate the background, particularly the object background, using negative Boolean and / or numerical values ​​(e.g., 0 or values ​​< 0.5).

[0030] In embodiments, image enhancement techniques can refer to techniques suitable for changing the orientation of one or more target objects, changing the scaling ratio associated with one or more target objects, and / or changing one or more pixel values ​​associated with one or more target objects. For example, image enhancement techniques may include at least one of scaling, cropping, rotating, blurring, distorting, shearing, resizing, folding, changing contrast, changing brightness, adding noise, multiplying by at least a portion of pixel values, filtering, adjusting color, applying convolution, imprinting, sharpening, flipping, averaging pixel values, or combinations thereof. For a non-exhaustive list of image enhancement techniques, see TOM McREYNOLDS and DAVID BLYTHE, “Advanced Graphics Programming Using OpenGL - A volume in The Morgan Kaufmann Series in Computer Graphics” (2005), ISBN 9781558606593, https: / / doi.org / 10.1016 / B978-1-55860-659-3.50030-5.

[0031] In embodiments, the mask generator may be parameterized and / or trained based on one or more mask training images associated with one or more target objects and indications of the contours of the one or more target objects. The mask generator may be a generative model. The mask generator may be configured to generate target object masks, preferably for generating multiple target object masks, and more preferably multiple different target object masks. Examples of generative models may include and / or may be based on generative adversarial networks, variational autoencoders, diffusion models, etc. The mask generator may receive triggers for generating target object masks. Triggers for generating target object masks may initiate the generation of target object masks. Triggers for generating target object masks may be included in a request for generating one or more training images and / or may be generated based on a request for generating training images. Triggers for generating target object masks may include text, audio, noise (especially noisy images), and / or combinations thereof. Noisy images may refer to data representing noise, particularly data associated with the same format as the target object masks.

[0032] In an embodiment, the mask generator may be further parameterized on a representation associated with one or more composite object images and / or a request for generating one or more composite object images. The representation of the one or more composite object images may be a numerical representation of the one or more composite object images, particularly a vector representation of the one or more composite object images. The representation of the one or more composite object images may be obtained by an object image generator. The mask generator may be parameterized to generate a target object mask based on the representation associated with one or more composite object images and / or a request for generating one or more composite object images.

[0033] In embodiments, the memory can be physical system memory, which can be volatile, non-volatile, or a combination thereof. The memory can include non-volatile mass storage devices, such as physical storage media. The memory can be a computer-readable storage medium (such as RAM, ROM, EEPROM, CD-ROM) or other optical disc storage, disk storage, or other magnetic storage devices, non-disk storage (such as solid-state drives), or any other physical tangible storage medium that can be used to store desired program code means in the form of computer-executable instructions or data structures and that can be accessed by the computing system. Furthermore, the memory can be a computer-readable medium carrying computer-executable instructions (also referred to as a transmission medium). Further, upon arrival at various computing system components, program code means in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to the storage medium (and vice versa). For example, computer-executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (e.g., a "NIC") and then ultimately transferred to the computing system RAM and / or a less volatile storage medium at the computing system. Therefore, it should be understood that storage media may be included in computing components that also (or even primarily) utilize transmission media.

[0034] In embodiments, an object can refer to a thing, a part of a thing, a part of an organism (such as a body part), and / or an organism. An object can be characterized by a shape associated with the object and / or a part of an object. An object can be a crop and / or may include attributes of a crop. A target object can be an object to be detected, particularly an object to be detected by an object detection model. A target object can be associated with one or more types of objects, particularly with one or more different types of objects. The type of object can be associated with a set of objects that share a common designation.

[0035] In embodiments, the object background can refer to the background associated with the object. Preferably, the object background can be the background in an image showing the object. The object background can be a background that at least partially surrounds the object. The object background can be the target object background and / or a second object background. In particular, the object background can be a background within a predefined distance from one or more target objects in one or more composite object images.

[0036] In an embodiment, the object detection model can be parameterized and / or trained to receive an image scene associated with one or more target objects and provide an indication of the presence of one or more target objects within the image scene. The object detection model can be parameterized and / or trained based on one or more training images. The object detection model can be a classification model trained and / or parameterized to classify images based on the presence of one or more target objects. The classification model may include at least one machine learning architecture and model parameters.

[0037] In an embodiment, the object image may include an object and an object background. The object image may be a synthetic object image. Synthetic object images can be generated by an object image generator, particularly an object image generator model. The object image may be associated with an image region that shows the object and the object background. For example, an image region may be associated with multiple pixel values. At least 1% of the image region may show the object background and / or at least 1% of the image's pixel values ​​may be associated with the object background, preferably at least 2%, more preferably at least 3%, more preferably at least 4%, more preferably at least 5%, more preferably at least 7%, and most preferably at least 10%. Considering some background when generating object images for training the object image generator increases the likelihood that the object image generated by the object image generator will be classified as a real image. Adding some background to the object image helps the object image generator learn to correctly generate the object's boundaries and mimic the relationship between the object and the background. Therefore, this improves the quality of the object image, thereby improving the accuracy of the object detection model trained on the object image. Ultimately, this can increase yields gained through crop treatment based on detected diseases, improve the performance of product lines that would otherwise be incorrectly supplied to product consumers, or reduce the amount of resources used to produce products by correctly identifying poor supply. The object image can include one or more target objects and an object background located at the edges. The object image can be a cutout of one or more target objects from an image scene, particularly a historical image scene. Further, the object image can include the remaining background resulting from cutting out one or more target objects.

[0038] In an embodiment, the object image generator may be an object image generator model.

[0039] In embodiments, the object localization model can be parameterized to provide and / or generate indications of the location of one or more objects (preferably two or more distinct objects including one or more target objects) within an image scene. The two or more distinct objects may include two or more target objects of two or more different types, or include one or more target objects and one or more objects different from the one or more target objects. The object localization model can be parameterized based on a localization training image including one or more target objects (optionally including two or more objects containing the one or more target objects) and corresponding indications of the location of the one or more target objects (optionally including two or more objects containing the one or more target objects) within the localization training image. The object localization model and / or object image generator may include at least one machine learning architecture and model parameters. Parameterizing the object localization model may include initializing and / or defining model parameters, for example, by assigning random numbers to the model parameters. Training the object localization model may include updating the model parameters. For example, a machine learning architecture can be or may include one or more of the following: linear regression, logistic regression, random forest, piecewise linear, nonlinear classifier, support vector machine, Naive Bayes classification, nearest neighbor, neural network, convolutional neural network, generative adversarial network, diffusion model, support vector machine, or gradient boosting algorithm, etc. In the case of neural networks, the model can be a multi-scale neural network or a recurrent neural network (RNN), such as, but not limited to, gated recurrent unit (GRU) recurrent neural network, visual transformer, or long short-term memory (LSTM) recurrent neural network.

[0040] In embodiments, a processor can refer to any logical circuit system configured to perform basic operations of a computer or system, and / or generally refers to a device configured to perform computations or logical operations. Specifically, a processor or computer processor can be configured to process the basic instructions that drive a computer or system. A processor can be a semiconductor-based processor, a quantum processor, or any other type of processor configured to process instructions. As an example, a processor can be or can include a Central Processing Unit (“CPU”). A processor can be a Graphics Processing Unit (“GPU”), a Tensor Processing Unit (“TPU”), a Complex Instruction Set Computing Microprocessor (“CISC”), a Reduced Instruction Set Computing (“RISC”) microprocessor, a Very Long Instruction Word (“VLIW”) microprocessor, or a processor implementing other instruction sets or multiple processors implementing combinations of instruction sets. A processing device can also be one or more special-purpose processing devices, such as an Application-Specific Integrated Circuit (“ASIC”), a Field-Programmable Gate Array (“FPGA”), a Complex Programmable Logic Device (“CPLD”), a Digital Signal Processor (“DSP”), a Network Processor, etc. The methods, systems, and devices described herein can be implemented as software in a DSP, microcontroller, or any other auxiliary processor, or as hardware circuitry within an ASIC, CPLD, or FPGA. It should be understood that the term "processor" can also refer to one or more processing devices, such as a distributed processing device system located across multiple computer systems (e.g., cloud computing), and is not limited to a single device unless otherwise stated. A processor can also be an interface to a remote computer system, such as a cloud service. A processor can include or be a Secure Isolation Zone (SEP). An SEP can be secure circuitry configured to process the spectrum. "Secure circuitry" is circuitry that protects isolated internal resources from direct access by external circuitry. A processor can be an image signal processor (ISP) and can include circuitry suitable for processing images, particularly images containing personal and / or confidential information.

[0041] In embodiments, providing one or more synthetic object images may include generating one or more synthetic object images by an object image generator. The object image generator may be parameterized to generate one or more synthetic object images in response to receiving a request for generating one or more synthetic images. The object image generator may be parameterized and / or trained based on multiple historical object images. The request may be received via a user interface. The request may be triggered and / or may include providing one or more synthetic object images. Any of these methods may further include providing a request for generating one or more training images associated with a target object and one or more backgrounds. The request may include a random tensor (in particular a random vector). The object image generator may be parameterized and / or trained to receive a random tensor (in particular a random vector) and generate one or more synthetic object images based on the random tensor (in particular a random vector) (in particular by removing noise from the random tensor and / or random vector). Removing noise from the random tensor may include passing the random tensor through one or more layers of a neural network. The random tensor may be associated with a value distribution that follows a random value distribution. The random tensor may be associated with a random image scene that illustrates a random pixel value distribution, in particular representing that random image scene. One or more layers of a neural network can be configured to increase the difference from a value distribution that follows a random value distribution. An object image generator can be parameterized and / or trained to generate representations (particularly multidimensional representations) associated with one or more synthetic images and / or requests for generating one or more synthetic images. Further, the object image generator can be parameterized and / or trained to generate one or more synthetic images based on the representations (particularly multidimensional representations).

[0042] In an embodiment, placing one or more composite object images on one or more backgrounds may include: providing one or more backgrounds; and placing one or more composite object images on the provided one or more backgrounds. Further, placing one or more composite object images on one or more backgrounds may include: determining whether an object background matches one or more backgrounds, particularly by comparing the object background with one or more backgrounds; and placing one or more composite object images in response to determining that the object background matches one or more backgrounds. Comparing the object background with one or more backgrounds may include: determining whether at least a portion of pixel values ​​associated with the object background corresponds to at least a portion of pixel values ​​associated with one or more backgrounds. At least a portion of pixel values ​​associated with one or more backgrounds corresponds to at least a portion of pixel values ​​associated with the object background if at least a portion of pixel values ​​associated with one or more backgrounds is within a 15% tolerance range associated with at least one pixel value among the pixel values ​​associated with the object background. Placing one or more composite object images may include: placing one or more composite object images on one or more backgrounds; and filling one or more backgrounds before or after placing one or more composite object images on one or more backgrounds.

[0043] In an embodiment, providing one or more backgrounds may include randomly selecting one or more backgrounds from a plurality of backgrounds or selecting one or more backgrounds based on one or more target objects. The plurality of backgrounds may be associated with a variety of different colors, textures, patterns, etc. This increases the diversity of the training data.

[0044] In an embodiment, the object detection model may be parameterized based on one or more training images and one or more corresponding indications about the presence of one or more target objects in the one or more training images.

[0045] In embodiments, any of the methods described herein may further include providing an indication of the location of one or more target objects within an image scene to an object detection model. The object detection model may further be parameterized based on the indication of the location of the one or more target objects. Any of these methods may further include providing one or more training images and / or image scenes to an object localization model to generate an indication of the location of one or more target objects within the image scene and / or the one or more training images. The indication of the location of one or more target objects may be provided and / or generated by the object localization model. The indication of the location of one or more objects, including one or more target objects, within an image scene may be location data indicating a portion of the image scene associated with the one or more objects including the one or more target objects. For example, the indication of the location of one or more objects, including one or more target objects, within an image scene may be bounding boxes. Thus, the indication of location (particularly the location of a portion of the image scene associated with the one or more objects including one or more target objects) may include one or more objects containing the one or more target objects, and at least a portion of the background associated with the image scene. This is advantageous because the object localization model may be readily available and may not require further training. Such a model can be trained to localize multiple objects. Therefore, this allows for efficient utilization of existing capabilities. Therefore, resources that would otherwise be used for further training and time that would otherwise be used to build the training dataset can be saved.

[0046] In an embodiment, the object detection model can be further parameterized and / or trained to generate indications about the location of objects within an image. The object detection model can be further trained and / or parameterized based on one or more training images associated with (particularly including) the object and its background, and the indications about the location of the object within the one or more training images, optionally providing indications about the location of the object within the image. Thus, the object detection model can be trained and / or parameterized to identify the location of objects within an image. The object detection model can be an object detection and localization model. The object detection and localization model can be further configured to provide indications about the location of one or more target objects within an image scene. This is advantageous because the performance of a model configured to localize and detect objects is generally superior to using a first model to localize objects and providing the identified locations to a second model for object detection. Therefore, by training a model for localizing and detecting objects within an image, improved performance for correctly detecting objects can be achieved.

[0047] In the example, an indication of the presence of an object within an image can be represented as a label used for training and / or parameterizing an object detection model. An indication of the presence of an object within an image scene and / or one or more training images can be associated with Boolean and / or numerical values ​​indicating true and / or false values. In the example, this indication can be a Boolean value, "yes or no," and / or a numerical value (e.g., a value between 0 and 1). Specifically, the numerical value can indicate a confidence score indicating whether an object can exist within the image scene and / or training images. In the example, an indication of the location of an object within an image scene and / or one or more training images can be represented as one or more labels associated with at least a portion of the image. Preferably, the indication of the object's location can be assigned to a portion of the image scene and / or one or more training images. The indication of location can indicate whether a portion of the image can be associated with an object. Multiple indications of the location of one or more target objects and / or one or more objects different from the one or more target objects can be associated with one or more training images and / or the image scene, particularly assigned to the one or more training images and / or the image scene. This can be referred to as semantic segmentation. For example, an indication of the location of an object may include one or more values ​​associated with one or more portions of an image. The values ​​associated with one or more portions of the image may be, for example, numeric values, one or more Boolean values, and / or one or more "yes or no" values. Numerical values ​​may be any value between, for example, 0 and 1, where the numerical value may indicate a confidence score indicating whether a portion of the image scene and / or one or more training images associated with the numerical value can include at least a portion of the object and / or show at least a portion of the object.

[0048] In embodiments, any of the methods described herein may further include: preferably providing an image scene showing the object via a user interface and / or preferably providing an indication of the presence of the object within the image via a user interface. Requests to generate one or more synthetic images may be received via the user interface, and / or one or more training images may be provided via the user interface. This allows the user to interact and customize the training data to be generated according to current needs. Therefore, this feature enables the training data to be precisely adapted to the use case, thereby improving the accuracy of the object detection model.

[0049] In one embodiment, the request to generate a first object image may include a first historical image. This allows the user to interact and customize the training data to be generated based on available examples. Therefore, this feature enables the training data to be precisely adapted to the current use case, thereby improving the accuracy of the object detection model.

[0050] In an embodiment, any of the methods described herein may further include: training an object image generator based on a first historical image.

[0051] In an embodiment, generating a first object image based on an object image generator may include: providing a trigger for generating the first object image to the object image generator; and receiving the first object image from the object image generator. The trigger allows for precise definition of the required training data. For example, a conditional GAN ​​may be instructed to generate certain synthetic data. This is particularly necessary when it is necessary to detect specific objects and / or when the generated real-world data does not represent certain configurations of the objects in sufficient quantities to train a robust object detection model. For example, synthetic data will reflect situations where specific outputs are typically produced with high accuracy and precision. Therefore, influencing precisely which type of data is generated has the benefit of generating the desired data for each specific use case. Consequently, the accuracy of detection trained on the training data will be improved.

[0052] In embodiments, requests for generating training images and / or images may be provided specifically via a user interface, and / or instructions regarding the presence of objects within the images may be received specifically via a user interface. This allows the user to interact and customize the training data to be generated based on available examples. Therefore, this feature enables the training data to be precisely adapted to the current use case, thereby improving the accuracy of the object detection model.

[0053] In embodiments, one or more backgrounds may be provided and / or generated by a background generator. The background generator may be parameterized and / or trained to generate backgrounds. The background generator may be parameterized to generate backgrounds based on one or more historical backgrounds. Furthermore, one or more backgrounds may be selected from a plurality of backgrounds based on a synthetic object image and / or a selection of one or more backgrounds. Any of these methods may further include: preferably providing the selection of one or more backgrounds via a user interface. This is advantageous because it allows the background to be customized according to a specific use case. Therefore, this feature enables training data to be precisely adapted to the current use case, thereby improving the accuracy of the object detection model. Alternatively or additionally, any of these methods may further include: applying one or more image enhancement techniques to the background image, particularly before combining the first object image and / or the second object image with the background image. The one or more backgrounds may be one or more background images and / or background specification data associated with instructions for obtaining one or more backgrounds. The instructions for obtaining one or more backgrounds may indicate one or more hues and one or more corresponding positions within one or more backgrounds.

[0054] In embodiments, placing one or more synthetic object images onto one or more backgrounds may include: extracting target objects from the one or more synthetic object images; and placing the one or more target objects onto the one or more backgrounds. This feature allows the target object to be introduced into the background, thereby defining the relationship between the target object and the background. Any of these methods may further include: applying one or more image enhancement techniques to the extracted one or more target objects, particularly before placing the one or more target objects onto the one or more backgrounds. The one or more image enhancement techniques may change the orientation of the one or more target objects, change the scaling associated with the one or more target objects, and / or change one or more pixel values ​​associated with the one or more target objects. Changing the one or more pixel values ​​associated with the one or more target objects may cause a change in one or more hues associated with the one or more target objects. Examples of changing the orientation of the one or more target objects may include rotating or mirroring the one or more target objects, particularly relative to one or more synthetic object images. Examples of changing the scaling associated with the one or more target objects may include increasing and / or decreasing the number of pixels associated with the one or more target objects. By doing so, several training images with different backgrounds can be obtained based on a generated background. This can be energy-efficient, especially when generating thousands of training images.

[0055] In embodiments, any of these methods may further include: providing a target object mask that indicates the contours of one or more target objects; extracting one or more target objects from one or more synthetic object images based on the target object mask. By doing so, only the desired objects are inserted into the background image, and the interaction between the object and the object background can be included in the training data, while enabling the background to be adapted to the user's needs.

[0056] In this embodiment, the target object mask may be generated and / or provided by a mask generator. The mask generator may be parameterized and / or trained based on one or more mask training images associated with one or more target objects and indications of the contours of those target objects. By doing so, the mask generator can accurately learn the contours of the objects. Therefore, only the desired objects are inserted into the background image, and the interaction between the object and its background can be included in the training data, while allowing the background to be adapted to the user's needs.

[0057] In embodiments, the object image generator may be further parameterized and / or trained to generate multiple synthetic object images, wherein at least a portion of one or more synthetic training images may be associated with objects different from and / or objects of different types than one or more target objects. Any of these methods may further include generating one or more synthetic object images associated with one or more target objects and one or more objects different from the one or more target objects. Placing one or more target objects on one or more backgrounds may include: extracting one or more target objects and one or more objects different from the one or more target objects from one or more synthetic object images; and placing one or more target objects and one or more objects different from the one or more target objects on one or more backgrounds. Two or more synthetic object images may be generated by the object image generator. Two or more synthetic object images may be associated with one or more target objects and one or more objects different from the one or more target objects.

[0058] In embodiments, one or more training images may be associated with two or more distinct target objects, particularly at least partially overlapping target objects. An object image generator may generate and / or provide two or more distinct synthetic object images. Alternatively or additionally, one or more target objects associated with one or more synthetic object images may be placed on one or more backgrounds, and the placement of one or more target objects on one or more backgrounds may be modified, particularly by one or more image enhancement techniques. Two or more distinct target objects are obtained by: placing one or more target objects associated with one or more synthetic object images on one or more backgrounds; enhancing one or more target objects; and placing the enhanced one or more target objects on one or more backgrounds. Enhancing one or more target objects may include applying one or more image enhancement techniques to one or more target objects. These one or more image enhancement techniques may change the orientation of one or more target objects, change the scaling associated with one or more target objects, and / or change one or more pixel values ​​associated with one or more target objects. Changing one or more pixel values ​​associated with one or more target objects may result in changing one or more hues associated with the one or more target objects. Examples of changing the orientation of one or more target objects may include rotating or mirroring one or more target objects, particularly relative to one or more synthetic object images. Examples of changing the scaling associated with one or more target objects can include increasing and / or decreasing the number of pixels associated with one or more target objects. By doing so, several training images with different backgrounds can be obtained based on a generated background. This can be energy-efficient, especially when generating thousands of training images. By doing so, training images with multiple and / or overlapping objects can be generated. This enables accurate object detection in scenarios where multiple objects need to be detected and / or analyzed, such as field trials.

[0059] In embodiments, any of the methods described herein may further include, in particular, providing a selection of one or more target objects via a user interface. Providing a selection of one or more target objects may initiate the generation of one or more synthetic object images. The selection of one or more target objects may be provided to an object image generator. The object image generator may be further parameterized and / or trained to generate one or more synthetic object images based on the provided selection. This allows the user to interact and customize the training data to be generated based on available examples. Therefore, this feature enables the training data to be precisely adapted to the current use case, thereby improving the accuracy of the object detection model.

[0060] In embodiments, any of the methods described herein may further include providing one or more training images specifically via a user interface.

[0061] In an embodiment, 50% to 90% of the image region associated with one or more composite object images may include one or more target objects. Preferably, 60% to 85%, more preferably 65% ​​to 78%, of the image region associated with one or more composite object images may include one or more target objects. By focusing on one or more target objects, the object image generator can robustly generate realistic target object images. The smaller the image region associated with the target object, the lower the quality of the generated target object. The larger the image region associated with the target object, the lower the quality of the generated target object's contour. Therefore, appropriately selecting the image region associated with the target object balances the quality of the target object's inlay and contour, thereby producing an overall realistic composite image.

[0062] In an embodiment, providing one or more synthetic object images to a mask generator may include providing numerical representations of the one or more synthetic object images to the mask generator. The numerical representations of the one or more synthetic object images may be provided by the object image generator, particularly by one or more hidden layers and / or intermediate layers of the object image generator. The numerical representations of the one or more synthetic object images may be associated with data structures different from random tensors and / or the one or more synthetic images. The numerical representations of the one or more synthetic object images may be obtained by at least a portion of the object image generator by processing random tensors. The one or more synthetic object images may be obtained by at least a portion of the object image generator by processing the numerical representations of the one or more synthetic object images.

[0063] In the following sections, the terminology and / or technical fields used herein and / or the scope of this disclosure will be outlined by way of definition and / or examples. Where examples are given, it should be understood that this disclosure is not limited to those examples.

[0064] These and other objectives are addressed by the subject matter of the independent claims, and will become apparent upon reading the following description. The dependent claims relate to embodiments of the invention. Attached Figure Description

[0065] The disclosure will be further described below with reference to the accompanying drawings. In the drawings and the disclosure, the same reference numerals are intended to refer to the same or similar elements, components and / or portions.

[0066] Figure 1 An embodiment of the operating system 102 for a chemical production facility is shown.

[0067] Figure 2 An embodiment of monitoring and / or controlling a chemical production facility 206 based on indications of the presence of a target object within an image scene is demonstrated.

[0068] Figure 3 An embodiment of a method for generating one or more training images for parameterizing an object detection model is shown.

[0069] Figure 4 An example of generating one or more object images is shown.

[0070] Figure 5 Examples of generating one or more backgrounds are shown.

[0071] Figure 6 An example of generating one or more object masks is shown.

[0072] Figure 7 An example of generating one or more object masks is shown.

[0073] Figure 8A An example is shown for generating training images by extracting objects from an object image.

[0074] Figure 8B An example is shown for generating training images by placing a target object on one or more backgrounds.

[0075] Figure 9 An example of an object detection model is shown. Detailed Implementation

[0076] The following embodiments are merely examples for implementing the methods, systems, or application devices disclosed herein and should not be considered limiting.

[0077] Figure 1 An embodiment of the operating system 102 for a chemical production facility is shown.

[0078] The chemical production facility can be controlled by an operating system 102. The operating system 102 may include an object image providing engine 104. The object image providing engine 104 can be configured to provide one or more synthetic object images associated with one or more target objects. The object image providing engine 104 may include an object image generator providing engine 106 and / or an object image providing engine 108. The object image generator providing engine 106 can be configured to provide an object image generator. The object image generator can be configured to generate one or more synthetic object images specifically in response to receiving a trigger for generating one or more synthetic object images. The trigger for generating one or more synthetic object images can be a random tensor. For example, the object image generator can be a generative adversarial network (GAN), such as a conditional GAN, normalized flow, variational autoencoder, diffusion model, etc. These machine learning models in Figure 4 This will be explained in more detail in the context of [the previous sentence]. The object image providing engine 108 can provide one or more composite object images to the object extraction engine 130. The object extraction engine 130 can be configured to extract one or more target objects from the one or more composite object images. One or more target objects can be extracted from the one or more composite object images using a target object mask. The target object mask can indicate the outline of one or more target objects associated with the one or more composite object images. The object extraction engine 130 can receive a target object mask from the object mask providing engine 124. The object mask providing engine 124 can provide a target object mask generated by the mask generator providing engine 122. The mask generator providing engine 122 can include a mask generator configured to generate target object masks. The mask generator can be configured to generate and / or provide target object masks in response to receiving one or more composite object images. The mask generator can [follow the previous sentence]. Figure 6 and Figure 7The following is described in further detail within the context of the above. Mask generator providing engine 122 and object mask providing engine 124 may be included in mask providing engine 118. Mask providing engine 118 may be configured to provide target object masks to training image providing engine 126, particularly object extraction engine 130. One or more extracted target objects may be provided to training image generation engine 132. Training image generation engine 132 and object extraction engine 130 may be included in training image providing engine 126. Training image providing engine 126 may be configured to provide one or more training images to object detection engine 134, particularly for parameterizing and / or training an object detection model. Training image generation engine 132 may be configured to generate one or more training images by receiving one or more synthetic object images from object extraction engine 130; and placing one or more target objects, particularly those extracted by object extraction engine 130, onto one or more backgrounds. One or more backgrounds may be provided by background providing engine 110. Background providing engine 110 may include background generator providing engine 112 and background generator engine 114. Background generator providing engine 112 can be configured to provide a background generator. The background generator can be configured to generate one or more backgrounds. Specifically, background generator providing engine 112 can be configured to parameterize the background generator based on one or more historical backgrounds, particularly multiple different historical backgrounds. The background generator can be received by background generator engine 114. Background generator engine 114 can be configured to generate one or more backgrounds by providing one or more random tensors to the background generator. The background generator can be configured to generate one or more backgrounds based on one or more random tensors. Therefore, the one or more backgrounds can be one or more synthetic backgrounds. The background generator can be configured to receive random tensors (e.g., tensors indicating random numerical distributions). Random tensors can be associated with noisy images (e.g., images showing pixel value distributions according to random numerical distributions). The background generator can be a GAN. The background generator can... Figure 5Further description follows. Background generator engine 114 can provide one or more generated backgrounds to training image generation engine 132. The generated one or more training images can be provided to object detection engine 134. Object detection engine 138 can be configured to provide an object detection model. The object detection model can be parameterized and / or trained to detect target objects in an image scene. The object detection model can be a classification model. The classification model can be configured to classify the image scene based on the presence of one or more target objects. The object detection model can generate an indication of the presence of one or more target objects in response to receiving an image scene. The classification model can include one or more input layers for receiving the image scene. The one or more input layers can be configured.

[0079] Furthermore, the object detection providing engine 138 can be configured to provide an object detection model to the indication providing engine 142. The indication providing engine 142 can be configured to provide an indication of the presence of one or more target objects within an image scene in response to a received image scene. The input and / or output interface 140 can provide the image scene to the indication providing engine 142. Furthermore, the indication providing engine 142 can provide indications to the input and / or output interface 140 for controlling and / or monitoring a chemical production facility. The indication of the presence of target objects within the image scene can be generated by the object detection model in response to the received image scene.

[0080] Figure 2 An embodiment of monitoring and / or controlling a chemical production facility 206 based on indications of the presence of a target object within an image scene is demonstrated.

[0081] Chemical production facility 206 can be configured to produce chemical product 212 from input material A 204. Input material A 204 can be a target object to be detected. Efficient production of chemical product 212 may require reliable detection of input material A 204. Input material A 204 may be associated with manufacturing defects. Manufacturing defects may cause changes in the chemical properties associated with input material A 204. Therefore, manufacturing defects may hinder the production of chemical product 212. Potential manufacturing defects may appear in image scene 202 associated with input material A 204. Therefore, upon receiving image scene 202 associated with input material A 204, object detection model 208 can detect manufacturing defects. If input material A 204 may not have manufacturing defects, object detection model 208 can generate indication 210 regarding the presence of input material A in the image scene. Otherwise, object detection model can provide indication regarding the absence of input material A in image scene 202. Providing indication 210 regarding the presence of input material A in the image scene can be referred to as detecting input material A 204 in the image scene. Instruction 210 regarding the presence of input material A within an image scene can be provided to chemical production facility 206, particularly to a control engine for controlling chemical production facility 206 and / or a monitoring engine for monitoring chemical production facility 206. Upon receiving instruction 210 regarding the presence of input material A within an image scene, chemical production facility 206 can produce chemical product 212 from input material A 204 associated with the image scene that generated instruction 210 regarding the presence of input material A within an image scene.

[0082] Figure 3 An embodiment of a method for generating one or more training images for parameterizing an object detection model is shown.

[0083] One or more backgrounds 316 can be provided, for example, by a background generator, as in Figure 1 Described in the context of.

[0084] One or more composite object images 318 associated with one or more target objects can be provided, for example, by an object image generator, as in Figure 1 Described in the context of.

[0085] A target object mask 320, for example, associated with the outline of one or more target objects, can be provided by a mask generator, as in... Figure 1 Described in the context of.

[0086] One or more target objects can be extracted from one or more composite object images by applying a target object mask. The target object mask indicates the location of one or more target objects within one or more composite object images. Therefore, the location of one or more target objects within one or more composite object images can be identified based on the target object mask. This can be done in... Figure 8A The text describes this in further detail.

[0087] One or more extracted target objects can be augmented.324 Augmentation of the extracted target objects can include, for example, rotation and / or scaling. Augmentation of one or more target objects can generate two or more distinct target objects from one or more target objects. Therefore, the number of synthesized target objects with different orientations can be increased. This makes it possible to customize the synthesized data according to the use case and generate training images of rare events with sufficient quantity and quality.

[0088] One or more target objects can be placed on one or more backgrounds 326. One or more target objects can be placed in an image region larger than the image region associated with the one or more target objects, and the image region associated with the background can be filled with one or more hues, for example, according to a selection received via a user interface and / or as indicated by a request for generating one or more training images. In another embodiment, one or more backgrounds can be generated by randomly selecting one or more backgrounds from a plurality of backgrounds.

[0089] One or more training images can be provided (e.g., provided via a user interface and / or provided to an object detection model for parameterization and / or training of the object detection model)330, as in Figure 1 Described in the context of.

[0090] Figure 4 An embodiment of generating one or more object images 410 is shown.

[0091] Therefore, object image 410 can be generated by object image generator 402. For this purpose, object image generator 402 can receive triggers for generating one or more object images 806. The trigger can be a noisy image. The noisy image can show a random pixel value distribution and / or the noisy image can include Gaussian noise. Object image generator 402 can generate one or more object images 806 based on the triggers provided for generating one or more object images 806. In this context, generating one or more object images 410 can mean transforming the received trigger into one or more object images 410 by passing the received trigger through the layers of object image generator 402. Therefore, object image generator 402 can include two or more layers. The trigger can be received at the input layer of object image generator 402. The trigger can be passed to the hidden layer of object image generator 402. Passing the trigger to another layer and / or passing the trigger through object image generator 402 can include performing one or more mathematical operations on the received trigger. Transforming the trigger and / or performing one or more mathematical operations can change one or more parts of the trigger, preferably changing multiple pixel values. Specifically, at least two parts of the trigger can be changed, wherein the degree of change between the at least two parts is different. By causing the trigger to pass through the object image generator 402, objects included in one or more object images 806 to be generated can be evolved.

[0092] Examples of object image generator 402 may include generators trained based on detectors trained to distinguish between real and synthetic data according to the concepts of generative adversarial networks (GANs), variational autoencoders, diffusion models, etc. Where object image generator 402 is and / or may include a GAN generator, object image generator 402 may be trained based on feedback from the GAN's detector. Feedback from the detector can indicate whether the generated data can be classified as real data or synthetic data. Therefore, object image generator 402 can be trained based on the classification of previously generated data (e.g., example object image 410). Object image generator 402 may be trained and / or parameterized based on one or more historical images 406. GANs are known in the art, such as XXX. Examples of GAN implementations may be found in... Figure 4As seen in the diagram, the object image generator 402 can be a conditional GAN. The conditional GAN ​​can be parameterized and / or trained to receive a random tensor 412 and an indication 414 regarding the synthetic data to be generated, specifically concatenating the random vector with the indication 414. The indication 414 regarding the synthetic data to be generated can be a one-hot vector. The indication 414 can also be a label indicating the synthetic data to be generated. Further, the conditional GAN ​​(cGAN) can be parameterized to receive the concatenated vector, i.e., the random tensor 412 and the indication 414 regarding the synthetic data to be generated. The cGAN can include multiple expansion blocks. Each expansion block can include one or more ReLU (Rectified Linear Unit) layers 418 and one or more transposed convolutional layers 420. One or more expansion blocks can be configured to change the data format associated with the concatenated input data to the data format associated with one or more synthetic object images.

[0093] The object image generator 402 can be parameterized and / or trained by providing a random tensor 412 and optionally an indication 414 regarding the synthetic data to be generated; and by generating one or more synthetic object images based on the random tensor 412 (optionally the indication 414 regarding the synthetic data to be generated). During parameterization and / or training, the generated one or more synthetic object images can be provided to a discriminator model. The discriminator model can be a classification model. The discriminator model can be configured, and in particular initialized, to classify whether the input data received by the discriminator model can be synthetically generated (i.e., generated by the object image generator). The deviation between the classification provided by the discriminator model and a predefined and / or correct classification, i.e., the loss, can be determined. The parameters associated with the discriminator model and the object image generator can be updated based on the determined deviation, for example, by deploying a backpropagation algorithm. This can include determining the discriminator loss and the generator loss based on the deviation between the provided classification and a predefined classification.

[0094] In one embodiment, the indication 414 regarding the synthetic data to be generated may include text indicating the synthetic data to be generated. The object image generator may include a text encoder trained and / or parameterized to transform the text included in the trigger into a machine-processable representation of the indication 414 regarding the synthetic data to be generated. The machine-processable representation may be a tensor. One or more synthetic object images 410 may be generated based on the machine-processable representation and the random tensor 412. In other embodiments, the indication 414 regarding the synthetic data to be generated may include an image, a sketch, etc. Here, the object image generator may include an image encoder.

[0095] In an embodiment, an object image generator can obtain numerical representations of one or more synthetic object images by processing random tensors and optionally instructions regarding the synthetic data to be generated. Preferably, the numerical representation can be the output data of one or more transposed convolutional layers. The numerical representations of one or more synthetic object images can be associated with data structures different from those of the synthetic object images. One or more object images can be obtained by processing the numerical representations of one or more synthetic object images.

[0096] Figure 5 An embodiment for generating one or more backgrounds 506 is illustrated. The one or more backgrounds 506 can be synthesized backgrounds. Therefore, one or more backgrounds 506 can be generated by a background generator 504. For this purpose, the background generator 504 can receive a trigger for generating one or more backgrounds 818. The trigger can be a noisy image. The noisy image can show a random pixel value distribution and / or the noisy image can include Gaussian noise. The background generator can generate one or more backgrounds 506 based on the trigger provided for generating one or more backgrounds 506. In this context, generating one or more backgrounds 506 can mean transforming the received trigger into one or more backgrounds 506 by passing the received trigger through the layers of the background generator 504. Therefore, the background generator 504 can include two or more layers. The trigger can be received at the input layer of the background generator 504. The trigger can be passed to the hidden layers of the background generator 504. Passing the trigger to another layer and / or passing the trigger through the background generator 504 can include performing one or more mathematical operations on the received trigger. Transforming the trigger and / or performing one or more mathematical operations can alter one or more parts of the trigger, preferably altering multiple pixel values. Specifically, at least two parts of the trigger can be altered, wherein the degree of change between the at least two parts differs. By passing the trigger through the background generator 504, the background included in one or more backgrounds 506 to be generated can be evolved.

[0097] Examples of background generator 504 may include generators trained based on detectors trained to distinguish real data from synthetic data according to the concepts of generative adversarial networks (GANs), variational autoencoders, diffusion models, etc. Where background generator 504 is and / or may include a GAN generator, it can be trained based on feedback from the GAN's detector. Feedback from the detector can indicate whether the generated data can be classified as real data or synthetic data. Therefore, background generator 504 can be trained based on the classification of previously generated data (e.g., example, one or more backgrounds 506). Background generator 504 may be trained and / or parameterized based on one or more historical backgrounds 502.

[0098] Preferably, the background generator 504 can be a conditional GAN. The conditional GAN ​​can be parameterized and / or trained to receive triggers including text, where the text may indicate one or more backgrounds 506 to be generated. For this purpose, the conditional GAN ​​may include a text encoder that is trained and / or parameterized to transform the text included in the trigger into a machine-processable representation of the trigger, thereby generating one or more backgrounds 506. The one or more backgrounds 506 may be generated based on the machine-processable representation. The conditional GAN ​​can be parameterized and / or trained to receive further inputs, such as noisy images, sketches, etc. Thus, the conditional GAN ​​can receive two or more trigger modalities for generating one or more backgrounds 506. Similarly, a variational autoencoder can receive two or more trigger modalities. In the case of a diffusion model, the trigger may include noise (particularly noisy images) and / or the trigger may be unimodal.

[0099] Figure 6 An example of generating one or more target object masks 604 is shown.

[0100] To generate one or more target object masks 604, one or more object images 606 may be received. Object image 606 may show one or more target objects and at least one object background. Object image 606 may be associated with one or more portions and / or may be composed of said one or more portions. Target object mask 604 may indicate whether said one or more portions can be associated with one or more target objects and / or one or more object backgrounds. Figure 6An example of a target object mask 604 may be shown. Target object mask 604 may include one or more labels indicating whether one or more portions associated with one or more tags can be associated with one or more target objects and / or one or more object backgrounds. For example, 1 may indicate the presence of one or more target objects within one or more portions of an object image associated with label 1. Similarly, 0 may indicate that one or more portions of an object image associated with label 0 can be associated with an object background and / or can be unrelated to a target object. Specifically, pixels of the object image may be assigned to indicate the presence of an object's label within a portion of the object image and / or to indicate whether one or more portions associated with the label can individually include that object. Specifically, target object mask 604 may include indications regarding the presence of one or more target objects and / or to assign indications regarding the presence of one or more target objects to one or more portions of object image 606. In an embodiment, target object mask 614 may include indications regarding the presence of multiple target objects and / or to assign indications regarding the presence of multiple target objects to one or more portions of object image 606. This may be referred to as semantic segmentation.

[0101] Therefore, the mask generator 602 can be a classification model. The classification model can be trained and / or parameterized to classify one or more portions (specifically one or more pixels) of a received object image. Thus, the mask generator 602 can be trained and / or parameterized to provide one or more indications regarding whether one or more portions associated with an indication can include the object. The mask generator 602 can be trained on the object image 606 to classify portions of the object image. Based on the indications regarding whether one or more portions associated with an indication can include 608, 608 can be extracted, preferably cut, from the object image including 608 and the background. For example, the indications regarding whether one or more portions associated with an indication can include the object can have the same format as the object image. Further, portions of the object image can be assigned values ​​indicating whether one or more portions can include 608.

[0102] In the example, mask generator 602 may include one or more blocks for obtaining a contracted representation of object image 606 and one or more blocks for obtaining a target object mask 614 based on the contracted representation of object image 606. The contracted representation of object image 606 may be a numerical representation of object image 606, such as a representation in latent space. Object image 606 may be an associated and / or a 2D tensor. The contracted representation may be an associated and / or a tensor of a higher dimension than object image 606.

[0103] In the example, mask generator 602 may be U-Net (doi 1505.04597v1). Therefore, mask generator 602 may include one or more shrink blocks for generating a shrunken representation of object image 606 based on object image 606. A shrink block may include a 3×3 convolutional layer 616, followed by a rectified linear unit (ReLU) layer 618, a 3×3 convolutional layer 616, a ReLU layer 618, and a pooling layer 620. Preferably, mask generator 602 may include multiple shrink blocks, for example, four shrink blocks as in the U-Net example. One or more shrink blocks may be configured to increase the dimension of the received input data. The shrunken representation of object image 606 may be provided to one or more expansion blocks. An expansion block may include a 3×3 convolutional layer 616, followed by a rectified linear unit (ReLU) layer 618, a 3×3 convolutional layer 616, a ReLU layer 618, and an upsampling layer 626. The upsampling layer 626 may be a 2×2 convolutional layer. The expansion blocks, particularly the upsampling layer 626, can reduce the dimensionality of the received input data. The output of the upsampling layer 626 can be concatenated with the outputs from one or more contraction blocks. This can include cropping the outputs of one or more contraction blocks to match the output format of the upsampling layer 626. Preferably, the outputs of one or more contraction blocks of the first contraction block can be concatenated with the output of the upsampling layer 626 of the last expansion block. The number of contraction blocks can be equal to the number of expansion blocks. The outputs of M contraction blocks can be concatenated with the outputs of NM expansion blocks, where N can be the total number of contraction blocks.

[0104] The additional convolutional layer 628 can format data received from one or more extended blocks into a target image format associated with the object image 606.

[0105] Figure 7 An example of generating one or more target object masks 704 is shown.

[0106] Mask generators can be provided as follows: Figure 4The mask generator is a numerical representation of one or more synthetic object images described in the context of the object image generator. The numerical representation can be a hidden state of the object image generator. The mask generator can be configured to receive the numerical representation of one or more synthetic object images. The mask generator can include an upsampling layer for generating a feature map associated with one or more synthetic object images. The feature map can be a multidimensional representation of one or more synthetic object images. The feature map can include multiple parts. A portion of the feature map can correspond to a pixel associated with one or more synthetic object images. Further, the mask generator can include one or more classification blocks for classifying (particularly pixel-wise) the pixels associated with one or more synthetic object images. Thus, one or more parts can be provided to one or more classification blocks, and / or indications about the presence of one or more target objects can be provided and / or generated by one or more classification blocks. For example, classification blocks can include fully connected layers, ReLU layers, and / or batch normalization layers. Providing the parts of the feature map to the mask generator piece by piece allows the feature map to be classified based on the presence of one or more target objects by providing indications about the presence of one or more target objects associated with each part of the feature map. By collecting indications about the existence of one or more target objects associated with each part of the feature map, a target object mask indicating the presence of one or more target object images within one or more synthetic object images can be obtained.

[0107] To generate one or more target object masks 704, one or more object images 706 may be received. Object image 706 may show one or more target objects and at least one object background. Object image 706 may be associated with one or more portions and / or may be composed of said one or more portions. Target object mask 704 may indicate whether said one or more portions can be associated with one or more target objects and / or one or more object backgrounds. Figure 7An example of a target object mask 704 may be shown. Target object mask 704 may include one or more labels indicating whether one or more portions associated with one or more tags can be associated with one or more target objects and / or one or more object backgrounds. For example, 1 may indicate the presence of one or more target objects within one or more portions of an object image associated with label 1. Similarly, 0 may indicate that one or more portions of an object image associated with label 0 can be associated with an object background and / or can be unrelated to a target object. Specifically, pixels of the object image may be assigned to indicate the presence of an object's label within a portion of the object image and / or to indicate whether one or more portions associated with the label can individually include that object. Specifically, target object mask 704 may include indications regarding the presence of one or more target objects and / or to assign indications regarding the presence of one or more target objects to one or more portions of object image 706. In an embodiment, target object mask 714 may include indications regarding the presence of multiple target objects and / or to assign indications regarding the presence of multiple target objects to one or more portions of object image 706. This may be referred to as semantic segmentation.

[0108] Therefore, the mask generator 702 can be a classification model. The classification model can be trained and / or parameterized to classify one or more portions (specifically one or more pixels) of a received object image. Thus, the mask generator 702 can be trained and / or parameterized to provide one or more indications regarding whether one or more portions associated with an indication can include the object. The mask generator 702 can be trained on the object image 706 to classify portions of the object image. Based on the indications regarding whether one or more portions associated with an indication can include 708, 708 can be extracted, preferably cut, from the object image including 708 and the background. For example, the indications regarding whether one or more portions associated with an indication can include the object can have the same format as the object image. Further, portions of the object image can be assigned values ​​indicating whether one or more portions can include 708.

[0109] In the example, mask generator 702 may include one or more blocks for obtaining a contracted representation of object image 706 and one or more blocks for obtaining a target object mask 714 based on the contracted representation of object image 706. The contracted representation of object image 706 may be a numerical representation of object image 706, such as a representation in latent space. Object image 706 may be an associated and / or a 2D tensor. The contracted representation may be an associated and / or a tensor of a higher dimension than object image 706.

[0110] In the example, mask generator 702 can be U-Net (doi 1505.04597v1). Therefore, mask generator 702 can include one or more shrink blocks for generating a shrunken representation of object image 706 based on object image 706. A shrink block can include a 3×3 fully connected layer 716, followed by a rectified linear unit (ReLU) layer 718, a 3×3 fully connected layer 716, a ReLU layer 718, and a batch normalization layer 720. Preferably, mask generator 702 can include multiple shrink blocks, for example, four shrink blocks as in the U-Net example. One or more shrink blocks can be configured to increase the dimension of the received input data. The shrunken representation of object image 706 can be provided to one or more expansion blocks. An expansion block can include a 3×3 fully connected layer 716, followed by a rectified linear unit (ReLU) layer 718, a 3×3 fully connected layer 716, a ReLU layer 718, and an upsampling layer 726. The upsampling layer 726 can be a 2×2 convolutional layer. The expansion blocks, particularly the upsampling layer 726, can reduce the dimensionality of the received input data. The output of the upsampling layer 726 can be concatenated with the outputs from one or more contraction blocks. This can include cropping the outputs of one or more contraction blocks to match the output format of the upsampling layer 726. Preferably, the outputs of one or more contraction blocks of the first contraction block can be concatenated with the output of the upsampling layer 726 of the last expansion block. The number of contraction blocks can be equal to the number of expansion blocks. The outputs of M contraction blocks can be concatenated with the outputs of NM expansion blocks, where N can be the total number of contraction blocks.

[0111] The additional convolutional layer 728 can format the data received from one or more extended blocks into a target image format associated with the object image 706.

[0112] Figure 7 An example of generating one or more object masks is shown.

[0113] Figure 8A An embodiment for generating a training image 810 by extracting a target object 802 from an object image 806 is illustrated. The target object 802 can be extracted from the object image 806 by cutting the target object 802 from the object image 806 based on a mask of the target object 802. Further image enhancement techniques, such as […], can be applied to the object image 806 before extracting the target object 802 and / or before inserting it into one or more backgrounds 818. Figure 8B The object mask 804 can be equal in size to the target object 802 within the object image 806. Furthermore, the object mask 804 can have the shape of the target object 802 within the object image 806.

[0114] By doing so, only the desired target object 802 will be inserted into one or more backgrounds 818, and the interaction between the target object 802 and the object background can be included in the training data, while enabling one or more backgrounds to be adapted to the user's needs.

[0115] Figure 8B An example is shown for generating training images by placing a target object on one or more backgrounds.

[0116] Therefore, as based on as in Figure 8A The target object 816 generated from the composite object image and target object mask described in the context can replace and / or fill one or more regions in the background 818 equal to the size of the target object 816. Regions in the background 818 equal to the size of the target object 808 can be merged, for example, by adding pixel values ​​or replacing pixel values ​​of regions associated with at least a portion of the background with pixel values ​​associated with one or more target objects. Regions in the background 818 equal to the size of the target object 808 can have the same shape as the target object 808.

[0117] Figure 9 An embodiment of object detection model 902 is shown.

[0118] Object detection model 902 can be configured to classify whether one or more target objects 920 exist in image scene 916. In the example, image scene 916 may show target objects 920. Object detection model 902 can provide an indication of the presence of one or more target objects in the image scene. The indication of the presence of one or more target objects in the image scene can be a vector. The vector can include multiple numerical values, for example, at least two numerical values. The numerical values ​​can indicate the certainty and / or confidence scores associated with the indication of the presence of one or more target objects in the image scene. In the example, object detection model can provide an indication of the presence of target objects in image scene 916.

[0119] The object detection model 902 may include one or more convolutional blocks. Each convolutional block may include one or more convolutional layers 904 and one or more pooling layers 906. The convolutional block 904 may be configured to reduce the dimensionality associated with the image scene. One or more convolutional blocks may provide a machine-processable representation of the image scene 916, such as a tensor. In the example, two convolutional blocks, each with one fully connected layer 716 and one batch normalization layer 720, can be implemented as LeNet, or two convolutional layers, each with one fully connected layer 716 and one batch normalization layer 720, and one convolutional block with three fully connected layers 716 and one batch normalization layer 720, can be implemented as AlexNet.

[0120] A tensor can be provided to one or more fully connected layers 908. The one or more fully connected layers 908 can process a machine-processable representation of the image scene 916 into a vector with an entry number equal to the number of classes associated with the object detection model 902. A softmax function 910 can determine a number in the range of 0 to 1 associated with a distribution equal to the outputs of the one or more fully connected layers 908. The output of the softmax function 910 can be a deterministic and / or confidence score associated with the classes of the object detection model 902.

[0121] In an embodiment, object detection model 902 may be associated with multiple categories indicating the presence of one or more target objects within an image scene and indicating the location of the one or more target objects within the image scene. Object detection model 902 may be parameterized and / or trained to provide indications about the presence of one or more target objects within an image scene and the location of said one or more target objects within the image scene. To this end, one or more fully connected layers 908 may be configured to change the format associated with a machine-processable representation of the image scene to a format indicating the presence and location of the one or more target objects. For example, the indications about the presence and location of the one or more target objects may be a vector with multiple entries, wherein a portion of the vector may indicate one or more categories associated with the presence of one or more target objects within the image scene and one or more categories associated with the location of the one or more target objects within at least a portion of the image scene. The number of categories associated with location may be equal to the number of portions of the image scene, and in particular, may be equal to the number of pixels associated with the image scene. Object detection model 902 may be parameterized and / or trained based on historical image scenes and corresponding indications about the presence and location of one or more target objects.

[0122] In this embodiment, the indication of the location of one or more target objects can be provided by an object localization model. The object localization model can be parameterized based on a historical image scene and the corresponding indication of the location of one or more target objects. The object localization model can be parameterized and / or trained similarly to object detection model 902. The object localization model can be a classification model, similar to object detection model 902.

[0123] This disclosure has also been described in conjunction with various preferred embodiments and examples. However, by studying the accompanying drawings, this disclosure, and the claims, those skilled in the art, as well as those practicing the claimed invention, will understand and implement other variations. It is particularly noteworthy that any of the proposed steps can be performed in any order; that is, the invention is not limited to a specific order of these steps. Furthermore, it is not required that different steps be performed at a specific location or node in a distributed system; that is, each step can be performed on different nodes using different devices / data processing.

[0124] As used herein, "determine" also includes "initiating or causing determination," "generate" also includes "initiating and / or causing generation," and "provide" also includes "initiating or causing determination, generation, selection, sending, and / or receiving." "Initiating or causing an action" includes any processing signal that triggers a computing node or device to perform a corresponding action.

[0125] In the claims and specification, the word "comprising" or "including" or similar wording does not exclude other elements or steps and should not be construed as limiting oneself to the listed elements or steps. The indefinite article "a" or "an" does not exclude multiple. A single element or other unit may perform the function of several entities or items recited in the claims. The fact that certain measures are recited only in mutually different dependent claims does not indicate that a combination of these measures cannot be used in advantageous implementations or that additional elements may be included.

[0126] Within the scope of this disclosure, provision may include any interface configured to provide data. This may include application programming interfaces, human-machine interfaces (such as displays), and / or software module interfaces. Provision may include transmitting or submitting data to the interface, particularly displaying data to a user or having data used by a receiving entity.

[0127] Any disclosures and embodiments described herein relate to the methods, systems, devices, and computer program elements listed above, and vice versa. Advantageously, the benefits provided by any embodiment and example also apply to all other embodiments and examples, and vice versa.

Claims

1. A method for detecting one or more target objects, particularly a computer-implemented method, the method comprising: Provide an image scene that includes one or more target objects. The image scene is provided to an object detection model for determining indications of one or more target objects present in the image scene, wherein the object detection model is parameterized based on one or more training images and one or more corresponding indications of one or more target objects present in the one or more training images, and wherein at least a portion of the one or more training images are synthetic images obtained in the following manner: Provide one or more composite object images, including one or more target objects and object backgrounds, and The target object extracted from the one or more composite object images is placed on one or more backgrounds. Provides indication of the target object within the image scene.

2. The method according to claim 1, wherein, The object detection model is further parameterized based on indications of the location of the one or more target objects within the one or more training images.

3. The method according to claim 2, wherein, The indication of the location of the one or more target objects is provided and / or generated by an object localization model, wherein the object localization model is parameterized based on localization training images of the one or more target objects and corresponding indications of the location of the one or more target objects.

4. The method according to claim 2, wherein, The object detection model is an object detection and localization model, wherein the object detection and localization model is further configured to provide an indication of the location of the one or more target objects within the image scene.

5. A computer-implemented method for generating one or more training images for training an object detection model to determine an indication of the presence of one or more target objects within an image scene, the method comprising: Provide one or more composite object images, including one or more target objects and object backgrounds, and The target object extracted from the one or more composite object images is placed on one or more backgrounds. Provide one or more training images for parameterizing an object detection model used to determine indications of the one or more target objects present in the scene of the image.

6. The method according to any one of claims 1 to 5, wherein, 50% to 90% of the image region associated with one or more composite object images includes the one or more target objects.

7. The method according to any one of claims 1 to 6, further comprising applying one or more image enhancement techniques to the one or more target objects before placing the extracted target objects on the one or more backgrounds.

8. The method according to any one of claims 1 to 7, further comprising providing a target object mask indicating the outline of the one or more target objects, wherein, The one or more target objects are extracted from the one or more composite object images based on the target object mask.

9. The method according to claim 8, wherein, The target object mask is generated and / or provided by a mask generator, which is parameterized and / or trained based on one or more mask training images associated with the one or more target objects and indications of the contours of the one or more target objects.

10. The method according to claim 9, wherein, The target object mask is generated by providing a numerical representation of the one or more synthetic object images to a mask generator, which is parameterized and / or trained based on one or more mask training numerical representations of the synthetic images of the one or more target objects and corresponding indications of the contours of the one or more target objects.

11. A computer-implemented method for training an object detection model for detecting one or more target objects within an image scene, the method comprising: Provide one or more training images associated with the one or more target objects, as generated by any one of claims 5 to 10. Provide one or more indications about the one or more target objects that exist within the one or more training images. The object detection model is trained based on the one or more training images and one or more indications about the one or more target objects existing in the one or more training images, and optionally the object detection model is provided.

12. The use of the object detection model trained by claim 11 for detecting and / or locating target objects within an image scene.

13. The indication of the presence of a target object in an image scene, obtained by any one of claims 1 to 4 and any one of claims 6 to 11 when subordinate to claims 1 to 4, for the control and / or monitoring of the production and / or application of a chemical product.

14. An apparatus comprising: processor; as well as A memory for storing instructions that, when executed by the processor, configure the device to perform any one of the methods according to any one of claims 1 to 10.

15. A training dataset comprising one or more training images generated according to any one of claims 5 to 9.

Citation Information

Patent Citations

  • System and method for generating training data for computer vision systems based on image segmentation

    US11475246B2

  • Teaching GAN (generative adversarial networks) to generate per-pixel annotation

    US20210089845A1

  • Synthetic generation of training data

    WO2023242236A1