Using a text-to-image model to generate synthetic images for training and / or validating anomaly detection models.
By using text-to-image and image-to-image models to generate and modify images with anomalies, the challenge of obtaining positive ground truth images for anomaly detection in industrial facilities is addressed, enhancing model accuracy and robustness.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-10-02
- Publication Date
- 2026-04-01
AI Technical Summary
Existing anomaly detection ML models for industrial facilities face challenges in obtaining a large and diverse set of positive ground truth images due to the infrequent occurrence and potential dangers of creating anomalies, making fine-tuning and validation difficult.
Utilize text-to-image models to generate synthetic images of industrial facilities with anomalies, and image-to-image models to modify real-world images with anomalies, enabling the creation of a diverse set of positive ground truth images for training and validation of anomaly detection models.
This approach allows for the generation of a large number of diverse synthetic images reflecting anomalies, improving the accuracy and robustness of anomaly detection models by enabling effective fine-tuning and validation without the need for dangerous real-world anomaly creation.
Smart Images

Figure 0007838613000001 
Figure 0007838613000002 
Figure 0007838613000003
Abstract
Description
[Technical Field]
[0001] This invention relates to a technique for processing images using machine learning (ML) models. [Background technology]
[0002] Complex industrial facilities such as petrochemical refineries and chemical plants may contain numerous components used in handling liquids and / or other substances involved in the industrial processes of the facility. To ensure that the components involved in the industrial processes are functioning as intended, and / or that the substances involved in the industrial processes are in their intended state, it is important to monitor for anomalies within the industrial facility (e.g., the presence of petroleum or other liquids on the ground, corroded piping, etc.) so that one or more corrective actions (e.g., sounding an alarm, stopping an automated process, etc.) can be taken in a timely manner when an anomaly is detected. [Overview of the project] [Means for solving the problem]
[0003] Various techniques have been proposed for processing images using machine learning (ML) models to generate outputs indicating whether conditions exist within an image. For example, techniques have been proposed for processing images capturing one or more components of an industrial facility using an anomaly detection ML model (sometimes referred to herein as the “anomaly detection model”) to generate outputs indicating whether anomalies exist in one or more components captured by the image. However, for an anomaly detection ML model to be effective and have sufficient predictive accuracy, it needs to be trained on a large and diverse collection of “positive” ground truth images, each reflecting the presence of a corresponding anomaly.
[0004] Separately, text-to-image models have been released that provide text prompts and enable the generation of detailed and realistic composite images based on those prompts. For example, a text prompt such as "an orange cat riding a donkey" can be processed to generate a realistic composite image that includes an orange cat riding a donkey. Some of these models also enable image-to-image conversions based on text prompts. For example, a base image containing a black cat can be provided with the prompt "make the cat orange" to generate a converted composite image that includes an orange cat instead of the black cat, but otherwise largely matches the base image. Non-restrictive examples of text-to-image models and image-to-image models (many models can do both) include "Stable Diffusion" and "DALL-E".
[0005] As mentioned above, for anomaly detection ML models to be effective, they need to be trained on a large number of diverse "positive" ground truth images, each reflecting the presence of a corresponding anomaly. In the surrounding environment (setting) of industrial facilities, obtaining such a large number of diverse real-world images can be difficult because corresponding anomalies occur infrequently and / or may be dangerous.
[0006] Furthermore, before utilizing such anomaly detection ML models in a specific industrial facility, there is currently no way to (a) fine-tune the model to suit that particular industrial facility, and / or (b) validate the model's performance in detecting anomalies within that industrial facility. For example, with regard to (a), fine-tuning an anomaly detection ML model for a specific industrial facility (i.e., limited additional training) may not be possible, as capturing anomalies and obtaining an actual image of "positive" ground truth specific to that industrial facility can be difficult. For example, creating such anomalies in an industrial facility may not be safe or desirable. Also, for example, an industrial facility may be under construction but not yet actually in use. As another example, with regard to (b), it may also not be possible to validate that an anomaly detection ML model is effective for a specific industrial facility before utilizing it in that facility. This may be because such validation would require determining that the ML model can detect the occurrence of anomalies in that particular facility. Again, actually creating anomalies in a particular facility may not be safe or desirable.
[0007] The implementations disclosed herein enable the generation of a large number of synthesized and diverse "positive" ground truth images, each reflecting the presence of corresponding anomalies within an industrial facility. These implementations generate the synthesized images using a text-to-image model. More specifically, the text-to-image model is prompted with prompts describing anomalies and the surrounding environment of the industrial facility, resulting in images of the surrounding environment of the industrial facility, including the anomalies. For example, prompts such as "exterior of chemical plant with oil puddles" and / or "interior of a chemical plant with oil on floor." Note that these implementations can generate multiple different synthesized images by processing the same prompt multiple times. The same prompt is used in each iteration, but different "seeds" are randomly used by the model in each iteration, resulting in the generation of various synthesized images.
[0008] The implementations disclosed herein additionally or alternatively enable the generation of a diverse array of synthesized “positive” ground truth images, each reflecting the presence of corresponding anomalies within an industrial facility, and each being specific to a particular industrial facility. This makes it possible to use such specific images to (a) fine-tune an ML model before use in a particular industrial facility (which can improve accuracy and / or robustness), and / or (b) validate an ML model before use in a particular industrial facility (to ensure it is effective for that particular industrial facility).
[0009] These implementations can capture actual images of a specific industrial facility and then process those images using an image-to-image transformation model, conditional on text prompts describing anomalies. For example, processing an actual image with the text prompt "add an oil puddle" can yield a composite image that closely matches the actual image but includes an oil puddle on the ground. As a specific example, hundreds or thousands of actual images of a particular industrial facility can be captured (e.g., by a robot) while the facility is free of anomalies. These actual images can then be processed using an image-to-image model, conditional on text prompts describing anomalies, to generate corresponding composite images that include the anomalies. These corresponding composite images can then be used for fine-tuning and / or verification, as described herein.
[0010] The implementation uses a text-to-image model to generate a large number of composite, diverse "positive" ground truth images, each reflecting the presence of corresponding anomalies within the industrial facility. More specifically, the text-to-image model is prompted with prompts describing the anomalies and the surrounding environment of the industrial facility, resulting in images of the surrounding environment of the industrial facility that include the anomalies. For example, a prompt such as "chemical plant room with pipes that include corrosion." Some of these implementations can also use actual images without anomalies from the industrial facility and an image-to-image transformation model to generate additional composite "positive" ground truth images. For example, an actual image without anomalies can be processed with a prompt such as "add some corrosion" to generate a composite image that matches the actual image but includes some corrosion on objects in the image.
[0011] The generated synthetic ground truth images can then be optionally labeled (for example, by a human reviewer). For example, each label may indicate whether an anomaly is present in the image, and optionally, the region where the anomaly occurred and / or the type of anomaly. The synthetic ground truth images and their labels can then be used to train an anomaly detection ML model. The anomaly detection ML model can then be used for anomaly detection in industrial facilities.
[0012] Several implementations additionally or alternatively generate a variety of transformed, synthesized "positive" ground truth images, each reflecting the presence of corresponding anomalies within the industrial facility, and each being specific to a particular industrial facility. This allows such identified images to be used to (a) fine-tune an anomaly detection ML model before use in a particular industrial facility (which can improve accuracy and / or robustness), and / or (b) validate the anomaly detection ML model before use in a particular industrial facility (to ensure it is effective for that particular industrial facility).
[0013] These implementations can capture actual images of a specific industrial facility and then process those images using an image-to-image transformation model, conditional on text prompts describing anomalies. For example, processing an actual image with the text prompt "add corrosion to one component" can yield a transformed composite image that closely matches the actual image but includes corrosion on the component in the image. As a specific example, hundreds or thousands of actual images of a specific industrial facility can be captured (e.g., by a robot) while the industrial facility is free of anomalies. These actual images can then be processed using an image-to-image model, conditional on text prompts describing anomalies, to generate corresponding transformed composite images that include the anomalies. These corresponding transformed composite images can then be used for fine-tuning and / or verification.
[0014] For example, when validating an anomaly detection ML model for a specific industrial facility, the anomaly detection ML model can be used to process each transformed composite image to generate a corresponding output indicating whether an anomaly is present. These corresponding outputs can then be compared to a ground truth output (indicating whether an anomaly actually exists in the corresponding transformed composite image) to generate an accuracy metric for the output (e.g., 90% if 900 out of 1000 predictions were correct). Optionally, the anomaly detection ML model is deployed for use within a specific industrial facility only if the accuracy metric and / or other validation metrics meet the required thresholds.
[0015] Various implementations provide a method implemented by the processor, which includes the step of performing multiple iterations of processing a text string for each of a plurality of text strings, each describing an anomaly and the surrounding environment of the industrial facility, in each iteration using a text-to-image model to generate a corresponding composite image (for example, conditional on the text string).
[0016] The surrounding environment of an industrial facility may include industrial automation facilities that implement any number of at least partially automated processes. For example, industrial automation facilities can take the form of chemical processing plants, oil or natural gas refineries, catalyst plants, manufacturing facilities, offshore oil platforms, or any other applicable industrial environment.
[0017] Multiple text strings can each describe an anomaly present in one or more components within the surrounding environment of an industrial facility. For example, different text strings can describe different anomalies in the same component within or around the surrounding environment of an industrial facility. As another example, different text strings can describe different anomalies in different components within or around the surrounding environment of an industrial facility. As yet another example, different text strings can describe different anomalies in different components within different surrounding environments of industrial facilities. A composite image can be a detailed, realistic image that captures the presence of each anomaly in one or more components within or around the surrounding environment of an industrial facility.
[0018] In various implementations, the method may further include a step of training an anomaly detection machine learning (ML) model using the generated synthetic image and the corresponding supervised label of the generated synthetic image. For example, in the first iteration, the generated synthetic image M is created using a first text string (which describes or indicates a first anomaly associated with a component in or around the surrounding environment of a corresponding industrial facility). 1-1 For label L 1-1 It can generate label L. 1-1 These can be automatically generated based on a first text string, or they can be manually generated (or inspected) by a human technician or expert trained to identify labels for objects in the surrounding environment of an industrial facility. In this example, for example, the corresponding anomaly A 1-1To generate a corresponding output indicating the type, name, and / or location (which can be highlighted using a bounding box) of, the synthetic image M 1-1 can be applied as a training instance input and processed using an anomaly detection ML model. Then, the corresponding output can be compared with the label L 1-1 (ground truth label that describes the ground truth type, name, and / or location of anomaly A 1-1 ) to determine the differences, and based on those differences, one or more parameters / weights of the anomaly detection ML model can be adjusted. Similarly, for the synthetic image M 1-2 generated in a second iteration using a first text string, a label L 1-2 can be generated. For example, to generate a corresponding output indicating the type, name, and / or location (which can be highlighted using a bounding box) of the corresponding anomaly A 1-2 , the synthetic image M 1-2 can be applied as a training instance input and processed using an anomaly detection ML model. Then, additional differences can be determined by comparing the corresponding output with the label L 1-2 , and based on the additional differences, one or more of the parameters of the anomaly detection ML model can be adjusted.
[0019] Alternatively, or additionally, for the synthetic image M 2-1 generated in a first iteration using a second text string different from the first text string, a label L 2-1 can be generated. For example, to generate a corresponding output indicating the type, name, and / or location (which can be highlighted using a bounding box) of the corresponding anomaly A 2-1 , the synthetic image M 2-1 can be applied as a training instance input and processed using an anomaly detection ML model. Then, the corresponding output can be compared with the label L 2-1Compared with [something], further differences can be determined, and one or more of the parameters of the anomaly detection ML model can be adjusted based on the further differences.
[0020] In some implementations, the second text string and the first text string can describe different types of anomalies. In some implementations, the second text string and the first text string are of the same type of anomaly but can describe different components within the same or various industrial facilities. The anomaly detection ML model can be trained using various synthetic images and corresponding labeled labels until the difference between the model output and the corresponding label is minimized or meets a difference threshold.
[0021] In various implementations, the method can further include the step of providing a trained anomaly detection ML model for use in anomaly detection within a specific industrial facility.
[0022] In various implementations, before providing a trained anomaly detection ML model for use in anomaly detection within a specific industrial facility, for each of a plurality of actual images of the specific industrial facility, to generate a corresponding transformed synthetic image, using an image-to-image transformation model to process the actual image and a prompt describing the anomaly, and further fine-tuning the trained anomaly detection ML model by further training the anomaly detection ML model based on the transformed synthetic image.
[0023] In some implementations, one or more vision sensors mounted on a mobile robot can be used to capture multiple real-world images. The mobile robot can be a quadruped walking robot (e.g., a robot dog), a wheeled robot, an unmanned aerial vehicle, or any other applicable robot that can move within an industrial automation facility. The one or more vision sensors can be a monographic RGB camera, a stereographic camera, a thermal camera, or any other applicable vision sensor. Accordingly, the real-world images can be RGB images, RGB-D images, thermal images, or any other applicable images.
[0024] In some implementations, the method can further include using a trained anomaly detection ML model when processing the real-world images, and using a trained anomaly detection ML model when processing the real-world images includes using the trained anomaly detection ML model to process one of the real-world images to generate a model output indicating whether there is an anomaly in one or more of the components, and performing one or more repair actions in response to the model output indicating that there is an anomaly in one or more of the components.
[0025] In various implementations, the method can further include, for each of a plurality of real-world images of a particular industrial facility, using an image-to-image conversion model to process the real-world image and a prompt describing an anomaly to generate a corresponding transformed synthetic image, and validating a trained anomaly detection ML model based on the transformed synthetic image, before providing the trained anomaly detection ML model for use in anomaly detection within the particular industrial facility. The step of providing the trained anomaly detection ML model for use in anomaly detection within the particular industrial facility is performed in response to a determination that the validation meets one or more criteria.
[0026] In various implementations, the step of validating a trained anomaly detection ML model based on a transformed synthetic image is implemented by determining an accuracy metric for anomaly prediction based on the output from the anomaly detection ML model, based on the processing of the transformed synthetic image. In those implementations, the step of determining that the validation satisfies one or more conditions may include a step of determining that the accuracy metric satisfies a threshold accuracy metric.
[0027] Furthermore, some implementations include one or more processors in one or more computing devices, the one or more processors being operable to perform instructions stored in associated memory, and the instructions being configured to perform one of the methods described above. Some implementations also include one or more non-temporary computer-readable storage media that store computer instructions that can be performed by one or more processors to perform one of the methods described above.
[0028] In some implementations, the system includes one or more processors and memory for storing instructions, and the instructions cause one or more processors to perform multiple iterations of processing text strings, each of which describes an anomaly and the surrounding environment of an industrial facility, using a text-to-image model to generate a corresponding composite image in each iteration; to train an anomaly detection machine learning (ML) model using the generated composite image and the corresponding supervised label of the generated composite image; and to provide a trained anomaly detection ML model for use in anomaly detection within a specific industrial facility.
[0029] In various implementations, the system may further include instructions for processing actual images and prompts describing anomalies using an image-to-image transformation model to generate a corresponding transformed composite image for each of several actual images of a particular industrial facility, and for fine-tuning the trained anomaly detection ML model by further training the anomaly detection ML model based on the transformed composite image, before providing a trained anomaly detection ML model for use in anomaly detection within a particular industrial facility.
[0030] In various implementations, the system may further include instructions to process actual images and prompts describing anomalies using an image-to-image transformation model to generate a corresponding transformed composite image for each of several actual images of a particular industrial facility, before providing a trained anomaly detection ML model for use in anomaly detection within that particular industrial facility, and to validate the trained anomaly detection ML model based on the transformed composite image. In those implementations, providing a trained anomaly detection ML model for use in anomaly detection within a particular industrial facility is performed upon determination that the validation satisfies one or more conditions. In some implementations, validating the trained anomaly detection ML model based on the transformed composite image includes determining an accuracy metric for anomaly predictions made based on the output from the anomaly detection ML model, based on the processing of the transformed composite image. One or more conditions may include, for example, a threshold accuracy metric.
[0031] In some implementations, providing a trained anomaly detection ML model for use in anomaly detection within a specific industrial facility involves using the trained anomaly detection ML model when processing real-world images captured via a visual sensor located inside or outside the specific industrial facility. In some implementations, the visual sensor can be mounted on a mobile robot moving around the specific industrial facility. In some of these implementations, using the trained anomaly detection ML model when processing real-world images may involve downloading the trained anomaly detection ML model locally to the mobile robot moving around the specific industrial facility for monitoring and anomaly detection. By having the trained anomaly detection ML model locally on the mobile robot, images captured using the mobile robot's visual sensor can be processed using the trained anomaly detection ML model for rapid anomaly detection and / or notification.
[0032] In some implementations, using a trained anomaly detection ML model when processing real-world images involves downloading the trained anomaly detection ML model to one or more on-site computers (e.g., servers) located in the industrial facility. In some of these implementations, images captured by visual sensors are provided to one or more on-site computers within a specific industrial facility and processed for anomaly detection using the trained anomaly detection ML model downloaded to those servers.
[0033] It should be understood that all combinations of the aforementioned concepts and any additional concepts described in more detail herein are intended to be part of the subject matter disclosed herein. For example, all combinations of claimed subject matter described at the end of this disclosure are intended to be part of the subject matter disclosed herein. [Brief explanation of the drawing]
[0034] [Figure 1A] This is a schematic diagram of an exemplary environment in which a selected aspect of the present disclosure can be implemented by various embodiments. [Figure 1B] This is an exemplary flowchart for carrying out a selected aspect of the present disclosure by various embodiments. [Figure 1C] This figure shows examples of machine learning (ML) models for performing selected aspects of the present disclosure, according to various embodiments. [Figure 2A] This is a schematic diagram of examples of images generated using the techniques described herein, according to various embodiments. [Figure 2B] This is a schematic diagram of an example of another image generated using the techniques described herein, according to various embodiments. [Figure 2C] This is a schematic diagram of examples of further images generated using the techniques described herein, according to various embodiments. [Figure 2D] This is a schematic diagram of examples of additional images generated using the techniques described herein, according to various embodiments. [Figure 3] This figure shows an exemplary method for carrying out a selected aspect of the present disclosure. [Figure 4] This is a schematic diagram of an exemplary computer architecture that can implement selected aspects of the present disclosure. [Modes for carrying out the invention]
[0035] An implementation of this disclosure utilizes one or more image generation models to generate a large number of images for training, fine-tuning, and / or validation of anomaly detection ML models used to detect one or more anomalies in components and / or materials within an industrial environment (e.g., an industrial automation facility). One or more image generation models may include, for example, text-to-image models that can be used to process text prompts and generate realistic synthetic images conditioned on those text prompts.
[0036] For example, a text-to-image model can be used to process text prompts (sometimes referred to herein as “prompts” or “text strings”) that indicate or describe anomalies, and to generate detailed, realistic composite images conditioned on those text prompts. In some implementations, the text-to-image model can be used to process the same text prompt multiple times / iterations to generate multiple different composite images, each reflecting an anomaly.
[0037] In some implementations, various text prompts can describe various anomalies within or around a particular industrial automation facility or a particular type of industrial automation facility (sometimes referred to herein as “industrial facility”), and accordingly, various composite images generated using a text-to-image model can capture various anomalies within or around a particular (or particular type) industrial automation facility. In some implementations, various text prompts can describe various anomalies within or around various industrial facilities, and accordingly, various composite images generated using a text-to-image model can capture various anomalies within or around various industrial facilities. Various composite images generated based on the same text prompt over multiple iterations, and / or various composite images generated based on various text prompts, can be applied to train, fine-tune, and / or validate anomaly detection ML models.
[0038] In various implementations, one or more image generation models may additionally or alternatively include image-to-image models that can be used to process a base image and a text prompt to generate a transformed / modified composite image that largely matches the base image but includes modifications that match the text prompt. For example, an image-to-image model could be provided with an image capturing one or more components of an industrial facility (e.g., a real-world image or a realistic composite image that does not reflect any anomalies) and a text prompt describing / indicating an anomaly present in one or more of the components. In this example, the image-to-image model is used to process the image and the text prompt to generate a modified image that largely matches the image but reflects the anomaly described or indicated in the text prompt. For example, if the image captures a floor without oil, two tanks, and piping, and the text prompt is "Add oil to the ground," the modified image might similarly include the floor, two tanks, and piping, but with oil on the floor. Therefore, after processing the base image and text prompts, a modified image can be generated based on the model output of the image-to-image model, and the modified image is a corrected version of the real-world image that reflects the anomaly indicated in the text prompt.
[0039] In some implementations, an image-to-image model can be used to generate various modified images using the same image capturing the same component of an industrial facility and various text prompts describing different anomalies in that component. Each of the various modified images can reflect the corresponding anomaly described or indicated by one of the various text prompts, and the various modified images (or parts thereof) can be applied as diverse "positive" ground truth images for training, fine-tuning, and / or validation of anomaly detection ML models when detecting anomalies in a particular component within an industrial facility.
[0040] In some implementations, as a practical example, one or more real-world vision sensors can be used to acquire various images capturing different components of an industrial facility (or multiple industrial facilities) in a normal state (i.e., without abnormalities). For example, one or more vision sensors can be mounted on a robot placed within the industrial facility or positioned at various locations within the industrial facility. One or more vision sensors may be monographic RGB cameras, stereographic cameras, thermal cameras, and / or any other applicable vision sensors. The robot may be a mobile robot, such as a robotic dog that moves around or through the industrial facility. Note that images captured using real-world vision sensors or other real-world sensor devices are often called “real images,” “real-world images,” or “images of the real world,” while images generated using the aforementioned image generation models (e.g., text-to-image models or image-to-image models) are often called “synthetic images.”
[0041] Continuing with the practical example above, we can receive various text prompts describing various anomalies, for example, from user input or from a file of collected anomaly text descriptions. Each of the various images and the corresponding text prompts can be processed using the image-to-image model so that the image-to-image model outputs a modified image that reflects the anomaly described in the corresponding text prompt. In this way, we can generate various modified images describing various anomalies in various components and / or various industrial facilities.
[0042] Various modified images (or parts thereof) can be applied as diverse “positive” ground truth images for training, fine-tuning, and / or validation of anomaly detection ML models when detecting anomalies in various components within the same or different industrial facilities. Additionally, or alternatively, one or more real-world images can be selected from various images acquired using one or more of the visual sensors, each to be assigned a “normal” supervised label (by capturing the component in a normal state). These one or more real-world images can be applied as “negative” ground truth images for training, fine-tuning, and / or validation of anomaly detection ML models. Note that the term “positive” is used herein to indicate the presence of anomalies, and the term “negative” is used herein to indicate the absence of anomalies. Furthermore, if various modified images (or parts thereof) are used to train, fine-tune, and / or validate the same ML model (e.g., an anomaly detection ML model), it should be noted that the various modified images can optionally be divided into various sets (e.g., a first set for training, a second set for fine-tuning, and / or a third set for validation), and that no modified image is included in more than two of these sets.
[0043] In some implementations, instead of, or in addition to, real-world images captured using sensors, realistic synthetic images capturing one or more components (within an industrial facility) that do not display anomalies can also be processed as input using an image-to-image model, along with text prompts describing foreseeable anomalies in the industrial facility, to generate modified images for training, fine-tuning, and / or validation purposes. The realistic synthetic images can be generated, for example, using an ML model (e.g., the aforementioned text-to-image model) that performs text-to-image conversion based on a text string simply describing an industrial facility (e.g., "chemical plant").
[0044] By utilizing one or more image generation models to generate synthetic images that reflect anomalies within one or more industrial facilities (or to modify real, anomaly-free images captured from one or more industrial facilities to reflect anomalies, or to modify realistic, anomaly-free synthetic images), it is possible to generate a large number of diverse synthetic images, each reflecting a corresponding anomaly within one or more industrial facilities, without humans creating (and later cleaning up and removing) dangerous anomalies. This also saves time and effort in capturing or waiting for image captures when certain types of anomalies (e.g., anomalies that are not easily detected) occur. Furthermore, training anomaly detection ML models with diverse and realistic synthetic images can improve the accuracy of the anomaly detection ML models in detecting anomalies.
[0045] Additionally, or alternatively, a large number of synthetic images (or a portion thereof) can be applied to fine-tune the anomaly detection ML model, allowing the fine-tuned model to be applied to detect various anomalies in specific components or specific industrial facilities. Additionally, or alternatively, a portion of a large number of synthetic images can be applied to validate the anomaly detection ML model before deployment, ensuring the quality of the anomaly detection ML model when detecting anomalies.
[0046] In some implementations, given a total of N (where N is greater than 1) diverse composite images generated using text-to-image models and / or image-to-image models, a first set of diverse composite images can be applied to initially train a first anomaly detection ML model for use in anomaly detection within or around an industrial facility. A second set of diverse composite images can then be applied additionally or alternatively to fine-tune the trained anomaly detection ML model, for example, to enhance anomaly detection for specific types of anomalies, specific industrial environments, and / or specific types of industrial environments. A third set of diverse composite images can then be applied additionally or alternatively to validate the trained and optionally fine-tuned anomaly detection ML model before it is deployed to detect anomalies within a specific industrial environment.
[0047] As described herein, in various implementations, the second and / or third sets of composite images may include (for example, limited to) modified composite images generated using an image-to-image transformation model and based on corresponding actual images of a particular industrial environment. For example, the second set of composite images may include images generated based on processing using an image-to-image transformation model, corresponding actual images of a particular industrial environment, and corresponding text prompts. In these and other ways, fine-tuning and / or validation may be specific to a particular industrial environment. Fine-tuning an anomaly detection ML model based on such images can improve the accuracy and / or robustness of the anomaly detection ML model for a particular industrial environment. Validating the anomaly detection ML model based on such images ensures that it is effective for a particular industrial environment.
[0048] Accordingly, the implementations described herein relate to training, fine-tuning, and / or validating an anomaly detection machine learning (ML) model to monitor and detect anomalies in one or more components (e.g., liquid tanks) within an industrial facility (e.g., an industrial automation facility) based on text prompts and / or images (realistic synthetic images and / or real-world images captured using one or more visual sensors such as cameras) that capture one or more (or some of) components.
[0049] In various implementations, anomaly detection ML models can be trained, refined, and / or validated using multiple composite images, each reflecting anomalies present in one or more of their components. For example, multiple composite images could include a first set of composite images generated using a text-to-image ML model (sometimes called a “text-to-image model”) based on one or more text strings describing anomalies and the surrounding environment of the industrial facility where the anomalies reside. Alternatively, or additionally, multiple composite images could include a second set of composite images generated using an image-to-image ML model (sometimes called an “image-to-image model”) based on real-world images, each conditioned on a text string (for example, introducing / adding anomalies to each image of the real-world image).
[0050] It should be noted that in some implementations, image generation models can function as both text-to-image and image-to-image models. In other words, image generation models can perform both text-to-image and image-to-image conversions. For example, one or more visual sensors mounted on a mobile robot (e.g., a robotic dog) can be used to capture images of the real world. By training, fine-tuning, or validating an anomaly detection ML model using a wide range of (or selected) composite images (e.g., realistic composite images) generated by text-to-image and / or image-to-image models, each reflecting corresponding anomalies in the surrounding environment of an industrial facility, it becomes unnecessary to wait for time-consuming anomalies (e.g., corrosion of pipes or other components) to occur and be discovered. Nor is it necessary to force the occurrence of hazardous anomalies, such as oil or other liquids spilling onto the ground in the surrounding environment of an industrial facility. Furthermore, the costs and effort associated with capturing appropriate images when anomalies occur can be saved or reduced.
[0051] Referring to the drawings, Figure 1A schematically illustrates exemplary environments in which selected aspects of the disclosure can be implemented in various embodiments. Figure 1B illustrates exemplary flowcharts for performing selected aspects of the disclosure in various embodiments. Figure 1C illustrates examples of machine learning (ML) models for performing selected aspects of the disclosure in various embodiments. Referring again to Figure 1A, exemplary environment 100 is schematically shown in which various aspects of the disclosure can be implemented. Exemplary environment 100 may be an industrial automation facility, or may include an industrial automation facility, and can take various forms. For example, exemplary environment 100 may be designed to implement any number of at least partially automated processes. An industrial automation facility may take the form of a chemical processing plant, an oil or natural gas refinery, a catalyst plant, a manufacturing facility, an offshore oil platform, etc.
[0052] An exemplary environment 100 may include one or more client devices (e.g., local client devices 103-A and 103-B) operably coupled to a process automation network 106 within an industrial automation facility. Client devices 103-A and 103-B may be implemented as computers (e.g., laptops, desktops, notebooks), tablets, robots, smart appliances (e.g., smartphones), messaging devices, wearable devices (e.g., watches), or any other applicable devices. The process automation network 106 may be implemented using a variety of wired and / or wireless communication technologies, including but not limited to cellular networks such as the Institute of Electrical and Electronics Engineers (IEEE) 802.3 standard (Ethernet), IEEE 802.11 (Wi-Fi), 3GPP Long-Term Evolution ("LTE"), or other wireless protocols designated as 3G, 4G, 5G and later, as well as other types of communication networks in various types of topologies (e.g., mesh).
[0053] An exemplary environment 100 may further include a mobile robot 101 having or carrying a vision component 1011. The mobile robot 101 may be a quadruped robot (e.g., a robotic dog), a wheeled robot, an unmanned aerial vehicle, or any other applicable robot capable of moving within an industrial facility. The vision component 1011 may be a monographic camera, a stereographic camera, a thermal camera, or any other applicable vision sensor for capturing images of one or more images of one or more specific components of an industrial automation facility (e.g., a liquid tank T or tube 102 for storing or transporting liquid substances). The vision component 1011 may be detachably coupled to or integrated into the mobile robot 101. In some implementations, the vision component 1011 may change its location and / or orientation relative to the mobile robot 101, for example, by rotation or other movement. In addition to the vision component 1011, the mobile robot 101 may include one or more additional vision components for sensing static or dynamic objects and / or capturing images as it moves through the industrial facility.
[0054] The exemplary environment 100 may further include a server computing device 105 (sometimes simply referred to as the “server device”). The server computing device 105 may include a machine learning (ML) engine 1051, an anomaly detection engine 1052, and an image generation engine 1057. The server computing device 105 may further include, or otherwise access, one or more machine learning (ML) models 1053. One or more ML models 1053 may include anomaly detection ML models (for example, M1, M2, and / or M3 in Figure 1B). Alternatively, or additionally, one or more ML models 1053 may include image generation models that perform text-to-image and / or image-to-image conversions (for example, text-to-image model 111A or image-to-image model 111B in Figure 1B). In some implementations, one or more ML models 1053 may include a first model for anomaly detection at a first site and a second model for anomaly detection at a second site, where the first site is different from the second site.
[0055] The server computing device 105 can connect to multiple client devices. The server computing device 105 can communicate with one or more local client devices (e.g., 103-A and 103-B) and / or with one or more remote client devices (not shown). Local client devices 103-A or 103-B can connect to the server computing device 105 via one or more local area networks (e.g., process automation network 106), and remote client devices can connect to the server computing device 105 via one or more wide area networks (e.g., the Internet). The local and remote client devices may be operated by personnel such as a system integrator to constitute and / or interact with various aspects of the exemplary environment 100.
[0056] In some implementations, the server computing device 105 may include, in addition to the ML engine 1051 and the anomaly detection engine 1052, a database (not shown) that stores information used by the ML engine 1051 and / or the anomaly detection engine 1052 and / or the image generation engine 1057 to implement selected embodiments of the present disclosure. In some implementations, the server computing device 105 may include, in addition to the ML engine 1051 and the anomaly detection engine 1052, an image preprocessing engine 1055 that processes different images (e.g., images captured using different visual sensors) so that they have the same image dimensions. Various embodiments of the server computing device 105, including the ML engine 1051, the anomaly detection engine 1052, the image generation engine 1057, and / or the image preprocessing engine 1055, can be implemented using any combination of hardware and software.
[0057] In some implementations, the ML engine 1051, the anomaly detection engine 1052, the image preprocessing engine 1055, or the trained ML model 1053 can be implemented across multiple computer systems as part of what is often called a “cloud infrastructure” or simply “cloud.” However, this is not mandatory, and as shown in Figure 1A, for example, the ML engine 1051 is implemented within an industrial facility, for example, within a single building, or across a single premises of a building or other industrial infrastructure. In such implementations, the ML engine 1051 can be implemented on one or more local computing systems, such as one or more server computers.
[0058] In some implementations, the mobile robot 101 can move through an industrial facility and capture images at one or more spots (e.g., designated spots). For example, as shown in Figure 1A, the visual component 1011 of the mobile robot 101 can be configured (but not necessarily) in a given pose to capture an image of the liquid tube 102 in that pose. In some implementations, the images captured by the visual component 1011 can be processed as input to an anomaly detection model (after being trained and / or validated) to generate a model output indicating whether an anomaly exists in the captured component in the image.
[0059] For example, the model output may include classification results that predict the probability of each of a predetermined number of anomalies being present in an image. In a non-limiting example, the classification result may be, for example, (oil spill, 0.7), (no oil spill, 0.3). In this non-limiting example, it can be determined that an anomaly corresponding to "oil spill" has been detected in an image with a confidence score of approximately "0.7" (satisfying a predefined confidence score threshold, e.g., 0.6). In response to the detection of the anomaly "oil spill," one or more mitigation / remediation actions may be performed, including, but not limited to, generating a warning message (e.g., 107 in Figure 1A, rendered via client device 103-A) or emitting a warning sound, to notify the person responsible for the oil spill.
[0060] As another non-limiting example, the classification result could be, for example, (oil spill, 0.1), (corrosion, 0.8), (fire, 0.1). In this non-limiting example, it can be determined that an anomaly corresponding to "corrosion" was detected from an image with a confidence score of approximately "0.8" (meeting a predefined confidence score threshold, e.g., 0.6). In response to the detection of the anomaly "oil spill," one or more mitigation / remediation actions can be taken, including, but not limited to, generating a warning message or sounding an alarm to notify the person responsible for the corrosion.
[0061] In some implementations, the image captured by the visual component 1011 can be processed as input, along with a text prompt (or text string) that introduces an anomaly into the image, using an image-to-image model to generate a modified image that reflects an anomaly described or indicated in the text prompt. As a non-limiting example, the image captured by the visual component 1011 may be a real-world image capturing one or more oil tanks in a storage room of an industrial facility. Such a real-world image can be provided to the aforementioned image generation model, along with a text prompt (e.g., "add an oil spill on the ground"), to perform an image-to-image transformation conditioned by the text prompt, where the output of the image generation model (e.g., the image-to-image model) corresponds to a modified image (a realistic composite image) that modifies the real-world image to show that there is an oil spill on the ground in the storage room where one or more oil tanks are located.
[0062] In the non-limiting example above, supervised labels can be generated based on the modified image and / or based on a text prompt (e.g., "add oil spill to the ground"). A human operator trained to recognize anomalies in the surrounding environment of an industrial facility can scan the modified image to determine a supervised label for the anomaly (e.g., "oil spill"). Alternatively, the supervised label (e.g., "oil spill") can be automatically extracted from a text prompt (e.g., "add oil spill to the ground") and / or validated by a human operator by inspecting the modified image. The modified image and supervised labels can then be applied to generate training instances (here, "positive" training instances because an anomaly is present), or instances for fine-tuning or validation.
[0063] For example, a modified image (e.g., a realistic composite image showing oil spilling onto the ground of a storage room where one or more oil tanks are stored) can be applied and saved as the training instance input ("positive" ground truth image) of a training instance, and supervised labels can be applied and saved as the ground truth labels of the training instance. The training instance can be applied to train an anomaly detection ML model, in which case the modified image can be processed using the anomaly detection ML model to generate the model output (e.g., whether the anomaly of "oil spill" was detected in the modified image and / or a confidence score indicating that the anomaly of "oil spill" was detected in the modified image). The model output can be compared to the ground truth labels of "oil spill," and based on this, one or more weights / parameters of the anomaly detection ML model can be fitted (or fine-tuned, if the modified image and ground truth labels are saved for the purpose of fine-tuning the model).
[0064] Additionally, or alternatively, the aforementioned real-world images capturing one or more oil tanks in an industrial facility storage room can be assigned a supervised label of "no anomaly." To train an anomaly detection ML model, the real-world images capturing one or more oil tanks in an industrial facility storage room and the supervised label of "no anomaly" can be saved as a "negative" training instance. Additionally, or alternatively, the real-world images capturing one or more oil tanks in an industrial facility storage room and the supervised label of "no anomaly" can be applied to generate instances for fine-tuning and / or validating additional anomaly detection ML models different from the anomaly detection ML model.
[0065] Optionally, two or more real-world images can be captured using the visual component 1011 (and / or additional sensors). As a non-limiting example, the visual component 1011 can be used to capture a total of 1,000 real-world images. To generate 100,000 composite images reflecting a first anomaly in a first type of facility, each of the 1,000 real-world images (or a portion thereof) can be conditioned based on a first text prompt describing the first anomaly in the first type of facility and processed 100 iterations using an image generation model (image-to-image model 111B). To generate 50,000 additional composite images reflecting a second anomaly within a first (or second) type facility, a total of 1,000 real-world images (or a portion thereof) can be processed 50 iteratively using an image generation model (image-to-image model 111B), each conditioned on a second text prompt describing the second anomaly within a first type facility (or a second type facility distinct from the first type). Further composite images can be generated in a similar manner. Duplicate descriptions are omitted herein.
[0066] Note that the number of iterations performed for each text prompt may vary depending on the type of anomaly described or indicated in the corresponding text prompt, and / or on the functionality / usage of the anomaly detection ML model being trained (or fine-tuned or validated). In some implementations, the number of iterations performed for a particular text prompt can be increased by identifying that the anomalies described by that particular text prompt are particularly important for training, fine-tuning, or validating the anomaly detection ML model.
[0067] Continuing with the non-limiting examples above, in some implementations, 100,000 synthetic images reflecting a first anomaly within a first type of facility can be assigned a first supervised label (e.g., type, name, location, etc.) to identify the first anomaly, and an additional 50,000 synthetic images reflecting a second anomaly within a first (or second) type of facility can be assigned a second supervised label (e.g., type, name, location, etc.) to identify the second anomaly. The 100,000 synthetic images and the 50,000 additional synthetic images can be divided into three groups, each for training, fine-tuning, and validation purposes.
[0068] For example, 100,000 composite images and 50,000 additional composite images can be divided into a first training group of composite images containing 80,000 of the 100,000 composite images and 40,000 of the 50,000 additional composite images; a second fine-tuning group of composite images containing 15,000 of the 100,000 composite images and 8,000 of the 50,000 additional composite images; and a third validation group of composite images containing 5,000 of the 100,000 composite images and 2,000 of the 50,000 additional composite images. For example, the first training group of composite images can be applied to generate "positive" training instances to train an anomaly detection model, the second fine-tuning group of composite images can be applied to generate "positive" instances to fine-tune the same anomaly detection model, and the third validation group of composite images can be applied to generate "positive" instances to validate the same anomaly detection model.
[0069] Note that when training and fine-tuning (or validation) various anomaly detection models, 100,000 composite images (or a portion thereof) and 50,000 additional composite images (or a portion thereof) can be applied to train the first anomaly detection model, and the same 100,000 composite images (or a portion thereof) and the same 50,000 additional composite images can be applied to fine-tune a second anomaly detection model that is different from or distinct from the first anomaly detection model.
[0070] Continuing with the non-restrictive example above, each of the aforementioned 1,000 real-world images (and / or additional real-world images) can be assigned a supervised label of "no anomaly." For example, the 1,000 real-world images (and / or additional real-world images) can be divided into a first training group of real-world images containing 7,500 real-world images, a second fine-tuning group of real-world images containing 2,000 real-world images, and a third validation group of real-world images containing 500 real-world images. The first training group can be applied to generate "negative" training instances to train the anomaly detection model, the second fine-tuning group of real-world images can be applied to generate "negative" fine-tuning instances to fine-tune the same anomaly detection model, and the third validation group of real-world images can be applied to generate "negative" validation instances to validate the same anomaly detection model. When training and fine-tuning (or validation) various anomaly detection models, note that one image from 1,000 real-world images (and / or additional real-world images) can be used to train (fine-tune or validate) the first model, and simultaneously to train (fine-tune or validate) a second model different from the first model.
[0071] Referring here to Figure 1C, the ML model 1053 stored in the ML model database may include a text-to-image model 111A that performs text-to-image conversion, an image-to-image model 111B that performs image-to-image conversion, an anomaly detection ML model M1 for detecting a first specific type of anomaly, an anomaly detection ML model M2 for detecting a second specific type of anomaly different from the first specific type, and / or an anomaly detection ML model M3 capable of detecting a predefined number of anomalies (e.g., including the first and second specific types of anomalies), and / or other ML models. In some implementations, instead of storing a text-to-image model 111A and an image-to-image model 111B that is separate from or independent of the text-to-image model 111A, the ML model database may include an image generation model that performs both text-to-image and image-to-image conversion, or may have access to an image generation model. Note that the image-to-image conversion may be conditional on a text prompt describing a corresponding anomaly specific to the industrial facility. The anomalies described in text prompts may vary depending on the type and location of the facility (or other factors). Furthermore, different text prompts can describe different anomalies.
[0072] Referring here to Figure 1B, a text string 108 (sometimes called a “prompt” or “text prompt”) can be provided to the text-to-image model 111A. The text string 108 might be, for example, “inside a chemical plant with oil on the floor”. The text string 108 can be processed multiple times / iteratively using the text-to-image model 111A. For example, in a first iteration, the text string 108 can be processed as a single input using the text-to-image model 111A to generate an image 120_1 showing a first chemical plant with oil on the floor. In a second iteration, the text string 108 can be processed as a single input using the text-to-image model 111A to generate an additional image 120_2 showing a second chemical plant with oil on the floor. The second chemical plant may be the same as or different from the first chemical plant. Alternatively, the oil on the floor of the first chemical plant may be of a different size, location, or color (or other features) than that of the second chemical plant. Similarly, in the Nth iteration, a text string 108 can be processed as a single input using the text-to-image model 111A to generate an image 120_N showing an Nth chemical plant with oil on the floor. Optionally, images 120_1, 120_2, ..., 120_N can be stored in an image database 140. Images 120_1, 120_2, ..., 120_N can be examined (e.g., by a technician, human operator, etc.) to determine the corresponding supervised labels (e.g., "oil on the ground", "oil spill", "oil spill on the left", etc.). Alternatively, the supervised labels (e.g., "oil on the ground") can be determined or extracted from the text string 108 (e.g., "inside a chemical plant with oil on the floor") and / or validated by a human operator.
[0073] Once supervised labels are determined for each of images 120_1 to 120_N, multiple training instances can be generated. For example, a first training instance can be generated that includes image 120_1 as the input to the first training instance and includes the first supervised label (e.g., "oil spill") as a ground truth label for comparison with the model output of an anomaly detection ML model (e.g., M1 in Figure 1A) processing the first training instance as input. Similarly, a second training instance can be generated that includes image 120_2 as the input to the second training instance and includes the second supervised label (e.g., "oil spill") as a ground truth label for comparison with the model output of an anomaly detection ML model (e.g., M1 in Figure 1A) processing the second training instance as input. The Nth training instance can be generated by including image 120_N as the Nth training instance input and including the Nth supervised label (e.g., "oil spill") as a ground truth label for comparison with the model output of an anomaly detection ML model (e.g., M1 in Figure 1A) that processes the Nth training instance as input. Based on the comparison, one or more weights / parameters of the anomaly detection ML model M1 can be fitted / modified / updated.
[0074] In some implementations, images 120_1~120_N (or a portion thereof) and their corresponding supervised labels can be applied to generate one or more tune-up instances, instead of applying them to generate training instances, or in addition to that. The generated one or more tune-up instances (if they do not contain images used to train model M1) can be applied to tune model M1, i.e., to fine-tune or adjust one or more weights / parameters of one or more layers of the anomaly detection ML model M1. The generated one or more tune-up instances (if they contain images used to train model M1) can be applied to tune a differently trained anomaly detection ML model (e.g., M2 in Figure 1B) than the anomaly detection ML model M1. For example, image 120_1 can be selected as the first tune-up instance input to the first tune-up instance, and the supervised label corresponding to image 120_1 can be selected as the first ground truth label to the first tune-up instance. Image 120_1 can be processed using a trained anomaly detection ML model to generate model outputs that indicate, for example, the name and location of the anomaly. The model outputs can be compared to the supervised labels corresponding to Image 120_1 to determine the differences, and based on this, one or more weights of the trained anomaly detection ML model can be fine-tuned.
[0075] In some implementations, instead of applying them to generate training instances, or in addition to that, images 120_1 to 120_N (or a portion thereof) and their corresponding supervised labels can be applied to generate one or more validation instances (sometimes called "test instances"). For example, image 120_1 can be selected as the input to the first validation instance for the first validation instance, and the supervised labels corresponding to image 120_1 can be selected as the first ground truth labels for the first validation instance. The first validation instance can be applied to validate an anomaly detection ML model (e.g., M3 in Figure 1B) by determining whether the difference between the first ground truth labels and the model output of the anomaly detection ML model corresponding to image 120_1 satisfies a difference threshold. If images 120_1 to 120_N (or a portion thereof) and their corresponding supervised labels can be applied to generate test instances, the performance of the anomaly detection ML model (e.g., M3 in Figure 1B) can be evaluated, for example, based on whether the output of model M3 satisfies one or more conditions (e.g., accuracy metrics described in other aspects of this disclosure).
[0076] In some implementations, still referring to Figure 1B, a visual sensor can be used to capture an image 110 of a component of an industrial facility in a normal state (e.g., no anomalies) (thus, image 110 is a real image). Image 110 can be provided to an image-to-image model 111B along with a text string 109 that introduces the anomaly of “oil puddle” to the industrial facility (e.g., “add oil puddle”). Image 110 and the text string 109 (e.g., “add oil puddle”) can be processed as input using the image-to-image model 111B to generate a modified image (e.g., real and composite, 130_1) that modifies the real image 110 to indicate an oil puddle with respect to a component of the industrial facility.
[0077] In some implementations, image 110 and text string 109 (for example, "add an oil reservoir" or "add an oil reservoir near an oil tank") can be processed multiple times / iteratively as input using the image-to-image model 111B, where in the first iteration, image 130_1 is generated by the image-to-image model 111B based on image 110 conditional on text string 109; in the second iteration, image 130_2 is generated by the image-to-image model 111B based on image 110 conditional on text string 109; and in the Nth iteration, image 130_N is generated by the image-to-image model 111B based on image 110 conditional on text string 109. The corrected images 130_1, 130_2, ..., 130_N (or a portion thereof), along with supervised labels (e.g., those extracted or determined from text string 109), can each be applied to train, fine-tune, and / or validate one or more anomaly detection ML models.
[0078] Additionally, or alternatively, a different text string (for example, "Add corrosion to the oil tank") can be provided to the image-to-image model 111B along with the image 110. Such text strings and images 110 can be iteratively processed multiple times using the image-to-image model 111B to generate multiple modified images. Similarly, supervised labels can be created and validated for each of the multiple modified images, and such supervised labels can be determined, for example, from the text string (for example, "Add corrosion to the oil tank" or "Add corrosion to the tank on the left").
[0079] It should be noted that in some implementations, image 110 does not need to be an actual image. For example, image 110 could be a realistic composite image generated based on a text string describing the surrounding environment of an industrial facility, such as a text-to-image model 111A. In these implementations, the text string processed by the text-to-image model 111A does not describe anomalies, and as a result, the generated realistic composite image does not show or reflect anomalies other than those in the surrounding environment of the industrial facility.
[0080] In some implementations, additionally or alternatively, the image database 140 may store one or more real images 150 captured using physical sensors, each reflecting a corresponding anomaly. In some implementations, additionally or alternatively, the image database 140 may store one or more real images 160 captured using physical sensors, each not reflecting an anomaly. One or more real images 150 may be stored in the image database 140 associated with a corresponding supervised label that identifies or indicates the corresponding anomaly. One or more real images 160 may each be stored in the image database 140 associated with a supervised label, such as “No Anomaly”. One or more real images 150 (or selected portions thereof) and their corresponding supervised labels may be applied as a “positive” ML instance to train, fine-tune, and / or validate an anomaly detection model, such as model M1 (or other anomaly detection ML models such as M2 or M3). One or more real images (or selected portions thereof) and their corresponding supervised labels can be applied as a “negative” training instance to train, fine-tune, and / or validate an anomaly detection model such as Model M1 (or other anomaly detection ML models such as M2 or M3).
[0081] In some implementations, one or more images 113A can be selected from the image database 140 to train an anomaly detection ML model (e.g., M1), and one or more images 113A may include images from composite images (e.g., images 120_1~120_N) generated using the text-to-image model 111A, images from composite images (e.g., images 130_1~130_N) generated using the image-to-image model 111B, images from actual images 150 that have captured anomalies, images from actual images 160 that have not captured anomalies, and / or other applicable images. For example, the image database 140 may also include composite images that exhibit anomalies, and such non-anomalous composite images may be generated using the image-to-image model 111B based on an initial composite image (non-anomalous) generated using the text-to-image model 111A (which receives text input that simply describes a scenario or industrial facility). The images included in the image database 140 are not limited to those described herein.
[0082] In some implementations, one or more images 113B can be selected from the image database 140 to fine-tune the anomaly detection ML model (e.g., M2), and one or more images 113B may include images from composite images generated using the text-to-image model 111A (e.g., images 120_1~120_N), images from composite images generated using the image-to-image model 111B (e.g., images 130_1~130_N), images from actual images 150 that captured anomalies, images from actual images that did not capture anomalies, and / or other applicable images (e.g., the images mentioned above).
[0083] In some implementations, one or more images 113C can be selected from the image database 140 to validate an anomaly detection ML model (e.g., M3), and one or more images 113C may include images from composite images generated using the text-to-image model 111A (e.g., images 120_1~120_N), images from composite images generated using the image-to-image model 111B (e.g., images 130_1~130_N), images from actual images 150 that captured anomalies, images from actual images that did not capture anomalies, and / or other applicable images (e.g., the images mentioned above).
[0084] Figure 2A schematically shows examples of images generated using the techniques described herein in various embodiments. Figure 2B schematically shows another example of images generated using the techniques described herein in various embodiments. Figure 2C schematically shows further examples of images generated using the techniques described herein in various embodiments. Referring here to Figure 2A, a mobile robot 101 carrying a visual sensor is used to capture images of the industrial facility 200 over a period of time (e.g., several weeks, several months, or longer). From the images captured during the period, a set of images that each captured anomalies within the industrial facility is selected and labeled, and the labeled set of images is applied to train an anomaly detection ML model 221. Additional sets of images that did not capture anomalies can be selected from the images captured during the period. For example, the additional set of images could include image 201 (a real-world image) that captures multiple oil tanks (e.g., parts of tanks A, B, C, and D visible within the area of industrial facility 200 enclosed by dashed lines). Image 201, along with text string A (for example, "Add an oil puddle to the ground"), can be processed as input once or multiple times / iteratively using the image-to-image model 211.
[0085] For example, as shown in Figure 2A, in the first iteration, a realistic composite image 213_A1 can be derived from the model output of the image-to-image model 211 corresponding to image 201 and text string A. The realistic composite image 213_A1 may be a modified version of the real-world image 201 by adding an oil puddle "O1" to the ground next to tank A. A supervised label L_A1 (e.g., "oil spill" or "oil puddle") can be determined for the realistic composite image 213_A1, and a training instance A1 can be generated that includes the realistic composite image 213_A1 as the training instance input and the supervised label L_A1 as the ground truth label. The anomaly detection ML model 221 can be trained (or validated or fine-tuned, if trained) using training instance A1. For example, a realistic synthetic image 213_A1 can be used to generate a model output 223_A1 for comparison with a supervised label L_A1, and then processed as input to an anomaly detection ML model 221 to determine the difference between the model output 223_A1 and the supervised label L_A1. Based on this difference, one or more weights / parameters of the anomaly detection ML model 221 can be fitted (or fine-tuned).
[0086] In the second iteration, referring to Figure 2B, the same image 201 and the same text string A can be used as input to produce a model output that can generate / derive a realistic composite image 213_A2, using the image-to-image model 211. The realistic composite image 213_A2 could be a modified version of the real-world image 201 by adding an oil puddle "O2" to the ground next to tank A. The oil puddle "O2" in the realistic composite image 213_A2 may differ in size, shape, and / or location from the oil puddle "O1" in the realistic composite image 213_A1. A supervised label L_A2 (e.g., "oil spill" or "oil puddle") can be determined for the realistic composite image 213_A2, and a training instance A2 can be generated that includes the realistic composite image 213_A2 as the training instance input and the supervised label L_A2 as the ground truth label. The anomaly detection ML model 221 can be trained (or validated or fine-tuned) using training instance A2. For example, a real synthetic image 213_A2 can be used as input to the anomaly detection ML model 221 to generate a model output 223_A2 for comparison with supervised label L_A2, and to determine the difference between the model output 223_A2 and the supervised label L_A2. Based on this difference, one or more weights / parameters of the anomaly detection ML model 221 can be modified / fitted (or fine-tuned).
[0087] In the third iteration, referring to Figure 2C, the same image 201 and the same text string A can be used as input to generate a model output that can produce / derive a realistic composite image 213_A3. The realistic composite image 213_A3 may be a modified version of the real-world image 201 by adding an oil puddle "O3" to the ground next to tank D. The oil puddle "O3" in the realistic composite image 213_A3 may differ in size, shape, and / or location from the oil puddle "O1" in the realistic composite image 213_A1 (and / or the oil puddle "O2" in the realistic composite image 213_A2). Supervised labels L_A3 (e.g., "oil spill" or "oil pool") can be determined for a realistic synthetic image 213_A3, and a training instance A3 can be generated that includes the realistic synthetic image 213_A3 as the training instance input and the supervised labels L_A3 as the ground truth labels. The anomaly detection ML model 221 can be trained (or validated or fine-tuned) using training instance A3. For example, the realistic synthetic image 213_A3 can be used to generate a model output 223_A3 for comparison with the supervised labels L_A3, and processed as input to the anomaly detection ML model 221 to determine the difference between the model output 223_A3 and the supervised labels L_A3. Based on this difference, one or more weights / parameters of the anomaly detection ML model 221 can be fitted (or fine-tuned, if fine-tuning has been performed). Note that with additional iterations, more realistic synthetic images can be generated. Redundant explanations are omitted herein.
[0088] In some implementations, the anomaly detection ML model 221 may be an ML model trained (or validated, or fine-tuned) on a realistic composite image (e.g., image 201) generated / derived from processing real-world images (if any anomalies are captured) and / or real-world images (if no anomalies are captured) conditional on a text string describing an anomaly within an industrial facility. For example, referring to Figure 2D, the image-to-image model 211 can be used to process image 201 (e.g., a real-world image) along with text string B (e.g., "oil spill") as input to generate a model output that can generate / derive a realistic composite image 213_B1. The realistic composite image 213_B1 may be a modified version of real-world image 201 by adding an oil spill from tank B (or another tank) (indicated by "L1" in Figure 2D).
[0089] A supervised label L_B1 (e.g., "oil spill") can be determined for a realistic synthetic image 213_B1, and a training instance B1 can be generated that includes the realistic synthetic image 213_B1 as the training instance input and the supervised label L_B1 as the ground truth label. The anomaly detection ML model 221 can be trained using the training instance B1. For example, the realistic synthetic image 213_B1 can be processed as input to the anomaly detection ML model 221 to generate a model output 223_B1 for comparison with the supervised label L_B1, thereby determining the difference between the model output 223_B1 and the supervised label L_B1. Based on this difference, one or more weights / parameters of the anomaly detection ML model 221 can be modified / fitted. Note that if additional iterations are performed, more realistic synthetic images can be generated. Redundant explanations are omitted herein.
[0090] Figure 3 is a flowchart illustrating an exemplary method 300 for implementing a selected aspect of the disclosure in an implementation available herein. For convenience, the operations in the flowchart are described with reference to a system performing the operations. This system may include various components of various computer systems, such as one or more components of a server computing device 105 (and / or additional computing devices such as client devices 103-A), including an ML engine 1051, an anomaly detection engine 1052, and / or a database. Furthermore, while the operations of method 300 are shown in a particular order, this is not intended to limit them. One or more operations may be rearranged, omitted, or added.
[0091] In block 302, the system may perform multiple iterations of processing text strings for each of a plurality of text strings, each describing (or indicating) an anomaly and the surrounding environment of an industrial facility, via a server such as server computing device 105, in each iteration using a text-to-image model to generate a corresponding composite image. The surrounding environment of an industrial facility may include, but is not limited to, one or more of the following: a chemical processing plant, an oil or natural gas refinery, a catalyst plant, a manufacturing facility, an offshore oil platform, or any other applicable facility (capable of implementing one or more partially automated processes).
[0092] In some implementations, the system can determine multiple text strings, each describing an anomaly and the surrounding environment of the corresponding industrial facility, before performing multiple iterations of text string processing for each of the multiple text strings. As a non-limiting example, the multiple text strings could include a first text string, “oil spill,” stored in association with a first chemical plant (or “oil spill at a chemical plant”), a second text string, “pipe corrosion,” stored in association with a second chemical plant (which may be the same as or different from the first chemical plant), and / or a third text string, “broken sensor wire,” stored in association with a manufacturing plant.
[0093] The first text string, second text string, third text string, and / or additional text strings can be received from user input or retrieved from a database or file that stores / lists text descriptions collected about various anomalies. For example, the database (or file) may include a first entry to store the first text string, a second entry to store the second text string, and a third entry to store the third text string. In some implementations, database entries can be grouped based on various anomaly types, the severity levels of various anomalies, the industrial facilities where each anomaly is foreseeable, and / or past locations or areas where various anomalies have been witnessed or recorded.
[0094] Continuing with the non-restrictive example above, given a first text string, "oil spill at a chemical plant," a text-to-image model can be used in a first iteration to process the first text string as input to generate an image N_11 (a realistic composite image) showing a part of the chemical plant (e.g., the ground inside the chemical plant) that reflects the oil spill. In this way, the text-to-image model can be used repeatedly (e.g., m times / iteration) to generate a set of images (e.g., a total of m images) (e.g., images N_11, N_12, ..., N_1m) each reflecting the corresponding oil spill within the chemical plant. In some implementations, the generated set of images can be stored (e.g., locally, on a server device, or in the cloud) associated with the first text string "oil spill at a chemical plant" for subsequent use.
[0095] For example, a first training instance can be generated and stored with image N_11 as the training instant input and a first text string "oil spill at a chemical plant" (or the keyword "oil spill" extracted from the first text string) as the supervised label; a second training instance can be generated and stored with image N_12 as the training instant input and a first text string "oil spill at a chemical plant" (or the keyword "oil spill" extracted from the first text string) as the supervised label; and the mth training instance can be generated and stored with image N_1m (where m is a positive integer greater than or equal to 1) as the training instant input and a first text string "oil spill at a chemical plant" (or the keyword "oil spill" extracted from the first text string) as the supervised label.
[0096] Optionally, the first, second, ..., and mth training instances can be examined (for example, by an expert or technician trained to identify oil spills) to correct / modify supervised labels or to remove one or more of the generated training instances (if the image in the corresponding training instance does not adequately reflect "oil spill") before the first, second, ..., and mth training instances are stored, for example, in a database. In other words, a first text string can be processed to automatically extract supervised labels for images generated by a text-to-image model corresponding to the first text string, and one or more supervised labels (if necessary) can be validated or modified by an expert before the corresponding training instances containing the supervised labels are stored.
[0097] Continuing with the non-restrictive example above, given a second text string "pipe corrosion" (or "pipe corrosion in a chemical plant" or "pipe corrosion at a manufacturing site"), in a second iteration, the text-to-image model could be used to process the second text string as input to generate an image N_21 (a realistic composite image) showing "pipe corrosion" (e.g., on a single pipe). Optionally, the text-to-image model could be used to iterate (e.g., p times / iteration) with the second text string "pipe corrosion" as input to generate a set of images (e.g., images N_21, N_22, ..., N_2p) each reflecting pipe corrosion on one or more pipes (e.g., a total of "p" pipes, where p is a positive integer greater than or equal to "1", and "p" may be the same as or different from "m"). Similarly, a set of generated images (e.g., image N_21, image N_22, ..., image N_2p) and a second text string, "pipe corrosion", can be used to generate and / or store a total of p training instances in the aforementioned database. Duplicate descriptions are omitted herein.
[0098] In block 304, the system can train an anomaly detection machine learning (ML) model using the generated composite image and the corresponding supervised labels of the generated composite image, for example, via a server such as server computing device 105. The anomaly detection ML model may differ from the text-to-image model described above. While the text-to-image model can perform text-to-image conversion, the anomaly detection ML model may be a classifier used to detect one or more anomalies (for example, to predict the probability that each of the predefined anomalies exists in a real-world image of the surrounding environment of an industrial facility), and thus, in response to the detection of one of the predefined anomalies, an appropriate and rapid remediation action can be taken accordingly. In some implementations, the anomaly detection ML model can be used to detect anomalies that are not included in the anomalies reflected in the composite image (generated by the text-to-image model based on multiple text strings), and / or to detect anomalies that are not described by multiple text strings.
[0099] In some implementations, the system can train an anomaly detection ML model by processing each generated composite image as input, thereby producing an output (e.g., a classification output) indicating multiple possibilities where each of a predefined number of anomalies exists in the surrounding environment of the industrial facility captured in each composite image. The output can be compared to the supervised label corresponding to each composite image to determine the difference, and based on the determined difference, one or more weights / parameters of the anomaly detection ML model can be adjusted or modified.
[0100] While the generated composite images described herein are intended for use in training an anomaly detection ML model, it should be noted that the generated composite images (or parts thereof) can be additionally or alternatively applied to fine-tune or validate additional anomaly detection ML models. Alternatively, the generated composite images can be divided into two or three groups, one used for training the anomaly detection ML model, one used for fine-tuning the anomaly detection ML model, and / or one used for validating the anomaly detection ML model, and the group of composite images used for training the anomaly detection ML model may have the largest number of composite images among the two or three groups from which the generated composite images were divided.
[0101] In block 306, the system can provide a trained anomaly detection ML model for use in anomaly detection within a specific industrial facility, for example, via a server such as server computing device 105. For example, one or more visual sensors within a specific industrial facility can be used to monitor the facility by capturing images or videos. Each image (or image frame from a video) or selected image (or image frame) captured for a specific industrial facility can be processed as input using a trained anomaly detection ML model to generate an output indicating the name of the anomaly, the likelihood of the anomaly being present within the specific industrial facility, and / or the location of the anomaly in the captured image (e.g., highlighted by a bounding box). Depending on the output of the trained anomaly detection ML model indicating the presence of the anomaly (e.g., based on the likelihood of meeting a probability threshold), one or more remedial actions can be performed. One or more remedial actions may include, for example, a first remedial action that generates and delivers a warning message, a second remedial action that pauses or stops one or more industrial processes, and a third remedial action that cuts off power. It should be noted that one or more remediation actions may be performed based on the type (or name), severity level, and / or location (or size or quantity estimated based on the output described herein), and are not limited to those described herein.
[0102] In some implementations, one or more vision sensors may be included in the mobile robot, which may capture images of one or more specific components in the surrounding environment of the industrial facility. These one or more specific components may, in non-limiting examples, include liquid tanks and / or the liquids they carry. The mobile robot may be a quadruped robot (e.g., a robotic dog), a wheeled robot, an unmanned aerial vehicle, or any other applicable robot capable of moving within the surrounding environment of the industrial facility.
[0103] One or more vision sensors may include, for example, a monographic RGB camera, a stereographic camera, a thermal camera, and / or any other applicable vision sensor. Similarly, images captured by the vision sensor may be RGB images, RGB-D images, thermal images, or any other applicable images. In some implementations, the images may be high-resolution images. In some implementations, the vision sensor may also be integrated with a mobile robot and may be detachably coupled to the mobile robot.
[0104] In some implementations, the step of providing a trained anomaly detection ML model for use in anomaly detection within a specific industrial facility includes the step of using the trained anomaly detection ML model when processing real-world images captured via a visual sensor located inside or outside the specific industrial facility. In some implementations, the visual sensor can be mounted on a mobile robot (e.g., a robotic dog) that moves around within the specific industrial facility. In these implementations, the step of using the trained anomaly detection ML model when processing real-world images may include the step of locally downloading the trained anomaly detection ML model to the mobile robot moving around the specific industrial facility for continuous or periodic monitoring and anomaly detection. By having the trained anomaly detection ML model locally on the mobile robot, images captured using the mobile robot's visual sensor can be processed using the trained anomaly detection ML model for rapid anomaly detection and / or notification.
[0105] Additionally, or alternatively, the step of using a trained anomaly detection ML model when processing real-world images includes the step of downloading the trained anomaly detection ML model to one or more on-site computers (e.g., servers). In these implementations, images captured by visual sensors are provided to one or more on-site computers located within a specific industrial facility and processed for anomaly detection using the trained anomaly detection ML model downloaded to those servers.
[0106] In some implementations, for example, two or more anomaly detection ML models can be trained using one or more steps described throughout this disclosure. The two or more anomaly detection ML models can be stored on a remote server (or distributed across one or more remote servers). As a non-limiting example, the two or more anomaly detection ML models may include a first anomaly detection ML model fine-tuned or validated for a first site (e.g., a first industrial facility) and a second anomaly detection ML model fine-tuned or validated for a second site different from the first site (e.g., a second industrial facility). Actual images captured by sensors at the first site can be transferred / transmitted to the remote server, for example, along with an identifier for the first site. Additional actual images captured by sensors at the second site can be transferred / transmitted to the remote server, for example, along with an identifier for the second site.
[0107] In the non-limiting example above, upon receiving an actual image bearing the identifier of the first site, a first anomaly detection ML model, fine-tuned or validated for the first site, can be accessed and used when processing the actual image, thus determining whether anomalies exist in the actual image. Similarly, upon receiving an additional actual image bearing the identifier of the second site, a second anomaly detection ML model, fine-tuned or validated for the second site, can be accessed and used when processing the additional actual image, thus determining whether anomalies exist in the additional actual image. Additionally, or alternatively, trained anomaly detection ML models (and / or one or more additional trained anomaly detection ML models) can be accessed on the cloud. For example, they can be accessed for use when processing real-world images, including the step of having the trained anomaly detection ML models be used.
[0108] In various implementations, before providing a trained anomaly detection ML model for use in anomaly detection within a specific industrial facility, for each of several actual images of a particular industrial facility, the system processes the actual image and a prompt describing the anomaly using an image-to-image transformation model in block 3051A to generate a corresponding transformed composite image (for example, one containing an image representation of the anomaly described in the prompt). In block 3051B, the trained anomaly detection ML model is fine-tuned by further training the anomaly detection ML model based on the transformed composite image.
[0109] By fine-tuning a trained anomaly detection ML model using transformed synthetic images generated from actual images of specific industrial facilities, one or more weights / parameters of the trained anomaly detection ML model can be adjusted, thereby improving the performance of the trained anomaly detection ML model in detecting anomalies in specific industrial facilities (and / or in detecting specific anomalies).
[0110] In some implementations, as described above, the trained anomaly detection ML model can be trained on a large number of diverse synthetic images (e.g., synthetic images generated by the text-to-image model in multiple iterations in block 302) each reflecting a corresponding anomaly in the surrounding environment of the industrial facility. In these implementations, the anomalies artificially created / introduced in the transformed synthetic images (derived from actual images capturing one or more components or substances within a particular industrial facility) and described in the prompts (e.g., in block 3051A) may have little to no overlap with the anomalies described in the multiple text strings (e.g., in block 302). As a non-limiting example, the multiple text strings could describe anomalies such as "oil spill," "pipe corrosion," and "wiring failure," while the prompts used to refine the actual images could describe anomalies specific to a particular industrial facility (e.g., "water leak").
[0111] In various implementations, before providing a trained anomaly detection ML model for use in anomaly detection within a specific industrial facility, the system may use an image-to-image transformation model to process the actual image and a prompt describing the anomaly for each of several actual images of the specific industrial facility in order to generate a corresponding transformed composite image, and then validate (or test) the trained anomaly detection ML model based on the transformed composite image. In these implementations, providing a trained anomaly detection ML model for use in anomaly detection within a specific industrial facility is done upon determining that the validation satisfies one or more conditions.
[0112] In some implementations, the step of validating a trained anomaly detection ML model based on the transformed synthetic image includes determining an accuracy metric for anomaly predictions made based on the output from the anomaly detection ML model, based on the processing of the transformed synthetic image. One or more conditions may include, for example, a threshold accuracy metric.
[0113] As a non-restrictive example, suppose an image-to-image transformation model is used to generate 50 transformed composite images, and these 50 transformed composite images are used to validate a trained anomaly detection ML model. In this non-restrictive example, if 42 of the 50 outputs of the trained anomaly detection ML model (corresponding to the 50 transformed composite images) match the supervised labels of the corresponding 50 transformed composite images, and the accuracy metric is approximately 0.84 (satisfying a threshold accuracy metric, e.g., 0.8), then the trained anomaly detection ML model can be considered validated and ready to be deployed for anomaly monitoring and detection in a specific industrial facility.
[0114] The methods described herein, or as described in other aspects of this disclosure, may be performed via one or more processors mounted on a mobile robot, or they may be included in one or more computing devices separate from and not attached to the mobile robot. In this case, real-world images captured by the mobile robot to monitor anomalies may be transmitted by the mobile robot and, after transmission by the mobile robot, identified by one or more computing devices for subsequent use and processing.
[0115] Figure 4 is a block diagram of an exemplary computing device 410 that can be optionally used to perform one or more embodiments of the techniques described herein. The computing device 410 typically includes at least one processor 414 that communicates with a number of peripheral devices via a bus subsystem 412. These peripheral devices may include, for example, a storage subsystem 424 including a memory subsystem 425 and a file storage subsystem 426, a user interface output device 420, a user interface input device 422, and a network interface subsystem 416. The input and output devices allow a user to interact with the computing device 410. The network interface subsystem 416 provides an interface to an external network and is coupled to a corresponding interface device in another computing device.
[0116] The user interface input device 422 may include pointing devices such as keyboards, mice, trackballs, touchpads, or graphics tablets, audio input devices such as scanners, touchscreens integrated into displays, speech recognition systems, microphones, and / or other types of input devices. Generally, the use of the term “input device” is intended to include all possible types of devices and methods for inputting information into the computing device 410 or a communication network.
[0117] The user interface output device 420 may include a non-visual display such as a display subsystem, printer, fax machine, or audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), liquid crystal display (LCD), projection device, or any other mechanism for creating a visible image. The display subsystem may also provide a non-visual display via an audio output device, etc. In general, the use of the term “output device” is intended to include all possible types of devices and methods for outputting information from the computing device 410 to the user or to another machine or computing device.
[0118] The storage subsystem 424 stores programming and data structures that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 424 may include logic to perform a selected embodiment of the method shown in Figure 3, as well as logic to implement the various components shown in Figures 1 and 2.
[0119] These software modules are typically executed by processor 414 alone or in combination with other processors. The memory 425 used in the storage subsystem 424 may include multiple memories, including main random access memory (RAM) 430 for storing instructions and data during program execution, and read-only memory (ROM) 432 for storing fixed instructions. The file storage subsystem 426 can provide persistent storage for program and data files and may include hard disk drives, floppy disk drives with associated removable media, CD-ROM drives, optical drives, or removable media cartridges. Modules implementing a particular implementation of a function may be stored by the file storage subsystem 426 within the storage subsystem 424, or on other machines accessible by processor 414.
[0120] The bus subsystem 412 provides a mechanism for various components and subsystems of the computing device 410 to communicate with each other as intended. Although the bus subsystem 412 is schematically shown as a single bus, multiple buses can be used in alternative implementations of the bus subsystem.
[0121] The computing device 410 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Because computers and networks are constantly changing, the description of the computing device 410 shown in Figure 4 is intended only as a specific example to illustrate several implementation forms. Many other configurations of the computing device 410 are possible, having more or fewer components than the computing device shown in Figure 4. [Explanation of symbols]
[0122] 100 exemplary environments 101 Mobile Robots 102 Liquid tubes, tubes 103-A Local client device 103-B Local client device 105 Server Computing Devices 106 Process Automation Network 108 Text strings 109 Text string 110 images 111A Text-to-Image Model 111B Image-to-Image Model 113A Image 113B Image 113C Image 120_1 Image 120_2 Image 120_N Image 130_1 Image 130_2 Image 130_N Image 140 Image Database 150 Actual Images 160 Actual Images 200 industrial facilities 201 images 211 Image-to-Image Models 213_A1 Realistic composite image 213_A2 Realistic composite image 213_A3 Realistic composite image 213_B1 Realistic composite image 221 Anomaly Detection ML Model 223_A1 Model Output 223_A2 Model Output 223_A3 Model Output 223_B1 Model Output 300 ways 410 Computing Devices 412 Bus subsystem 414 processors 416 Network Interface Subsystem 420 User Interface Output Devices 422 User Interface Input Devices 424 Storage subsystems 425 Memory subsystem 426 File Storage Subsystem 430 Main Random Access Memory (RAM) 432 Read-only memory (ROM) 1011 Visual Components 1051 Machine Learning (ML) Engine 1052 Anomaly detection engine 1053 Machine Learning (ML) Models 1055 Image preprocessing engine 1057 Image generation engine
Claims
1. A method implemented by one or more processors, For each of the multiple text strings that describe the anomaly and the surrounding environment of the corresponding industrial facility, A step of performing multiple iterations of processing the text string, in each of the iterations, using a text-to-image model to generate a corresponding composite image, The steps to be performed include: the plurality of text strings including a first text string describing a first anomaly in the surrounding environment of the first industrial facility and a second text string describing a second anomaly in the surrounding environment of the first industrial facility; The steps include training an anomaly detection machine learning (ML) model using the generated composite image and the corresponding supervised label of the generated composite image, The steps include providing the trained anomaly detection ML model for use in anomaly detection within a specific industrial facility, and Methods that include...
2. Before providing the trained anomaly detection ML model for use in anomaly detection within the aforementioned specific industrial facility, For each of the multiple actual images of the aforementioned specific industrial facility, The steps include: processing the actual image and a prompt describing the anomaly using an image-to-image transformation model to generate a corresponding transformed composite image; The steps include: fine-tuning the trained anomaly detection ML model by further training the anomaly detection ML model based on the converted composite image; and The method according to claim 1, further comprising:
3. The step of fine-tuning the aforementioned trained anomaly detection ML model is: The method according to claim 2, further comprising the step of fine-tuning the weights of one or more layers of the trained anomaly detection ML model.
4. Before providing the trained anomaly detection ML model for use in anomaly detection within the aforementioned specific industrial facility, For each of the multiple actual images of the aforementioned specific industrial facility, The steps include: processing the actual image and a prompt describing the anomaly using an image-to-image transformation model to generate a corresponding transformed composite image; The steps include: verifying the trained anomaly detection ML model based on the converted composite image; It further includes, The method according to claim 1, wherein the step of providing the trained anomaly detection ML model for use in anomaly detection within the specific industrial facility is performed upon determination that the verification satisfies one or more conditions.
5. The step of validating the trained anomaly detection ML model based on the transformed composite image includes determining an accuracy metric for anomaly prediction based on the output from the anomaly detection ML model, based on the processing of the transformed composite image. The method according to claim 4, wherein the step of determining that the verification satisfies one or more conditions includes the step of determining that the accuracy index satisfies a threshold accuracy index.
6. The method according to claim 1, wherein the plurality of text strings further comprises one text string describing one anomaly in the surrounding environment of one industrial facility, and additional text strings describing the one anomaly in the surrounding environment of an additional industrial facility.
7. The method according to claim 1, wherein the corresponding supervised label of the generated composite image is automatically determined based on the plurality of text strings that describe at least the anomaly.
8. The step of providing the trained anomaly detection ML model for use in anomaly detection within the aforementioned specific industrial facility is: The method according to claim 1, further comprising the step of using the trained anomaly detection ML model when processing actual images captured via a visual sensor within the specified industrial facility.
9. The process further comprises using the trained anomaly detection ML model when processing the actual image, and the process of using the trained anomaly detection ML model when processing the actual image is The steps include processing one of the actual images using the trained anomaly detection ML model to generate a model output indicating whether an anomaly exists in one or more of the components, Depending on whether the model output indicates that there is an abnormality in one or more of the components, the step of performing one or more repair actions: The method according to claim 8, including the method described in claim 8.
10. The method according to claim 9, wherein the vision sensor is mounted on a mobile robot, and the actual image is captured by the mobile robot at a designated location within the specific industrial facility.
11. The method according to claim 10, wherein the one or more repair actions include a warning message that alerts to the detection of the anomaly in one or more of the components located at the designated location.
12. A method implemented by one or more processors, For each of the multiple actual images of a particular industrial facility, A step of processing the actual image and a corresponding prompt describing a corresponding anomaly associated with a particular industrial facility, using an image-to-image transformation model to generate a corresponding transformed composite image, The corresponding prompt includes a first prompt containing a first text string describing a first anomaly of the particular industrial facility, and a second prompt containing a second text string describing a second anomaly associated with the particular industrial facility, and the processing step of: The steps include: fine-tuning the trained anomaly detection ML model by further training the trained anomaly detection ML model based on the converted composite image; and Methods that include...
13. The method according to claim 12, wherein the trained anomaly detection ML model is pre-trained on one or more real images, each capturing a corresponding anomaly.
14. The method according to claim 13, wherein the corresponding anomaly is captured using a visual sensor within the specific industrial facility.
15. The method according to claim 12, wherein the trained anomaly detection ML model is pre-trained on a plurality of composite images, each representing a realistic scene of a corresponding anomaly in the surrounding environment of an industrial facility, and the plurality of composite images are generated using a text-to-image model, each based on one or more text strings, each describing an anomaly.
16. A method implemented by one or more processors, For each of the multiple actual images of a particular industrial facility, A step of processing the actual image and a corresponding prompt describing a corresponding anomaly associated with a particular industrial facility, using an image-to-image transformation model to generate a corresponding transformed composite image, The corresponding prompt includes a first prompt containing a first text string describing a first anomaly associated with the particular industrial facility, and a second prompt containing a second text string describing a second anomaly associated with the particular industrial facility, and the processing step of: The steps include: verifying the trained anomaly detection ML model based on the converted composite image; and Methods that include...
17. The steps include determining whether the verification satisfies one or more conditions, Depending on the determination that the verification satisfies one or more of the conditions, the step is to provide the trained anomaly detection ML model for use in anomaly detection within the specific industrial facility. The method according to claim 16, further comprising:
18. The method according to claim 17, wherein the step of validating the trained anomaly detection ML model based on the transformed synthetic image includes determining an accuracy metric for anomaly prediction made based on the output from the anomaly detection ML model based on processing the transformed synthetic image, wherein one or more conditions include a threshold accuracy metric.
19. The method according to claim 16, wherein the trained anomaly detection ML model is pre-trained on a plurality of composite images, each representing a realistic scene of a corresponding anomaly in the surrounding environment of an industrial facility, and the plurality of composite images are generated using a text-to-image model, each based on one or more text strings, each describing an anomaly.
20. The method according to claim 16, wherein the trained anomaly detection ML model is pre-trained on one or more real images, each capturing a corresponding anomaly in the surrounding environment of an industrial facility.
21. A method implemented by one or more processors, For each of the multiple text strings that describe the anomaly and the surrounding environment of the corresponding industrial facility, A step of performing multiple iterations of processing the text string, in each of the iterations, using a text-to-image model to generate a corresponding composite image, The steps to be performed include: the plurality of text strings include one text string describing one anomaly in the surrounding environment of one industrial facility, and an additional text string describing the same anomaly in the surrounding environment of an additional industrial facility; The steps include training an anomaly detection machine learning (ML) model using the generated composite image and the corresponding supervised label of the generated composite image, The steps include providing the trained anomaly detection ML model for use in anomaly detection within a specific industrial facility, and Methods that include...
Citation Information
Patent Citations
Visual inspection device
JP1991252883A
Fire monitoring system
JP2018088630A
Learning data collection device, learning device, learning data collection method, and program
JP2022038941A
Abnormality determination model generation method, abnormality determination model generation device and inspection device
JP2022040531A
Method of changing character part with image, computer equipment, and computer program
JP2022107580A