Generative sampling for imbalanced classes in image and video datasets

By employing synthetic data and transfer learning with expert advice, the method addresses the scarcity of industrial defect data, improving model performance and reducing costs for training machine learning models in industrial applications.

JP2026509137APending Publication Date: 2026-03-17HITACHI VANTARA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-02-13
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

The scarcity of industrial defects and high costs associated with collecting and managing data for training machine learning models for fault detection in industrial processes pose challenges, particularly for smaller companies, leading to inaccurate models and limited data sharing.

Method used

Utilizing synthetic data and transfer learning techniques to generate additional training data, along with expert advice, to balance the representation of classifications and improve model performance.

Benefits of technology

This approach reduces the need for raw data collection and costs associated with data management and effectively improves the performance and accuracy of machine learning models by generating additional training data, thereby enhancing the effectiveness of the models and enhancing the accuracy of the models and reducing errors in the machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509137000001_ABST
    Figure 2026509137000001_ABST
Patent Text Reader

Abstract

In exemplary implementations described herein, there are systems and methods for training a machine learning model for industrial failure detection based on an initial set of training data comprising multiple input-output datasets, each containing input data and output data for at least one classification of the input data. The method may include identifying at least one underrepresented classification of the input data among the multiple classifications. The method may also include automatically generating additional input-output datasets for the identified at least one classification to balance the representation of the classifications in the modified set of training data for inclusion in the modified set of training data. Finally, the method may include training the machine learning model based on the modified set of training data, which includes at least a subset of the initial set of training data and the additional input-output datasets.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates, in general terms, to the generation of datasets for fault detection associated with industrial processes. [Background technology]

[0002] This disclosure describes solutions to problems related to the monitoring of industrial defects. While industrial defect monitoring may be data-driven in some aspects, industrial defects may be rare events. The scarcity of data may, in some aspects, make it difficult to develop industrial models (e.g., artificial intelligence and / or machine learning (AI / ML) models) and / or limit the number of models that can be used for defect monitoring in relation to one or more industrial processes. In some aspects, the scarcity of industrial defects may also make it more difficult to improve the accuracy of the models.

[0003] The rarity of industrial defects and / or failure events can, in some aspects, be due to the high quality standards of industrial components. Therefore, capturing failure events can be challenging in relation to industrial processes, which present challenges in capturing a sufficient number of data points (e.g., images, audio, video, or other data) to train accurate models (e.g., machine learning models such as AI / ML models) for defect and / or failure event detection. In addition, the costs of capturing, managing, and annotating datasets (e.g., images, audio, video, or other data) can be excessively high for smaller companies, and in many cases, datasets cannot be shared between companies to reduce and / or share data management and associated costs, or to improve the quality of data available to each company. Due to the difficulties in developing machine learning models, it is not uncommon for subject matter experts, such as engineers, to be involved in an advisory role in identifying industrial defects and / or failure events within the data (e.g., images or other data). [Overview of the Initiative] [Means for solving the problem]

[0004] The exemplary implementations described herein include innovative methods for using synthetic data to increase the amount of training data for machine learning models relevant to industrial applications, thereby improving the performance of the machine learning models and / or reducing errors in the machine learning models. In some embodiments, the disclosed methods may help companies / users struggling with a lack of data for model training. In some embodiments, the methods provide solutions to problems related to the scarcity of industrial defects and / or failure events, and / or the high costs associated with collecting and maintaining a sufficient amount of defect and / or failure event data for training one or more accurate machine learning models.

[0005] For example, this method can significantly reduce the need for raw data and the costs associated with collecting it by using synthetic data to increase the amount of training data for industrial applications. In addition, the high costs associated with annotating images (or other data) and managing data for AI / ML projects (e.g., for training machine learning models) can be significantly reduced by automatically annotating images (or other data) using transfer learning techniques or expert advice from subject matter specialists. Thus, increasing the amount of training data for industrial applications using synthetic data can improve the statistical confidence and generalization ability of models (e.g., improve model performance). In some embodiments, new data can be used to teach the model new patterns, thereby further improving the performance of the machine learning model and reducing errors in the machine learning model.

[0006] Aspects of this disclosure include methods, non-temporary computer-readable media, systems, or devices for training a machine learning model for industrial fault detection based on an initial set of training data. In some aspects, the initial set of training data may include multiple input-output datasets, each containing input data relating to industrial objects and output data relating to at least one classification of the input data. The method may include identifying, within the initial set of training data, at least one classification of the input data that is underrepresented among several classifications of the input data. The method may further include automatically generating additional input-output datasets for the identified at least one classification to be included in a modified set of training data, along with the initial set of training data, thereby balancing the representation of the multiple classifications in the modified set of training data, and training a machine learning model based on the modified set of training data, which includes at least a subset of the initial set of training data and the additional input-output datasets.

[0007] Aspects of this disclosure include a non-temporary computer-readable medium for storing instructions for execution by a processor, which may include instructions for identifying at least one classification of the input data that is underrepresented among a plurality of classifications of the input data within an initial set of training data. The instructions may further include instructions for training a machine learning model based on the modified set of training data, which includes at least a subset of the initial set of training data and the additional input-output dataset, by automatically generating additional input-output datasets for the identified at least one classification to be included in a modified set of training data along with the initial set of training data.

[0008] Aspects of the present disclosure include a system that may include means for identifying at least one classification of the input data that is underrepresented among a plurality of classifications of the input data within an initial set of training data. The means may further include means for automatically generating additional input-output datasets for the identified at least one classification to be included in a modified set of training data along with the initial set of training data, thereby balancing the classification representations of the plurality of classifications in the modified set of training data, and for training a machine learning model based on the modified set of training data, which includes at least a subset of the initial set of training data and the additional input-output datasets.

[0009] Aspects of the present disclosure include a device that may include a processor configured to identify at least one underrepresented classification of the input data among a plurality of classifications of the input data within an initial set of training data. The processor may further be configured to automatically generate additional input-output datasets for the identified at least one classification to be included in a modified set of training data, along with the initial set of training data, to balance the classification representations of the plurality of classifications in the modified set of training data, and to train a machine learning model based on the modified set of training data, which includes at least a subset of the initial set of training data and the additional input-output datasets. [Brief explanation of the drawing]

[0010] [Figure 1] Figure 1 shows a set of functional components of a system according to several aspects of this disclosure.

[0011] [Figure 2] Figure 2 shows several embodiments of the data management component and the automated labeling / annotation component according to several aspects of the present disclosure.

[0012] [Figure 3]Figure 3 is a diagram showing a set of related operations according to some aspects of the present disclosure.

[0013] [Figure 4] Figure 4 is a diagram showing a set of operations related to a data analysis component according to some aspects of the present disclosure.

[0014] [Figure 5] Figure 5 is a diagram showing a set of operations related to an image generation component according to some aspects of the present disclosure.

[0015] [Figure 6] Figure 6 is a diagram showing a set of operations related to a generative model learning component and a tensor framework image generation component according to some aspects of the present disclosure.

[0016] [Figure 7] Figure 7 is a diagram showing elements of a morphism according to some aspects of the present disclosure.

[0017] [Figure 8] Figure 8 is a diagram showing an adversarial generative network (GAN) according to some aspects of the present disclosure.

[0018] [Figure 9] Figure 9 is a diagram showing a variational autoencoder according to some aspects of the present disclosure.

[0019] [Figure 10] Figure 10 is a diagram showing a diffusion model according to some aspects of the present disclosure.

[0020] [Figure 11] Figure 11 is a flowchart showing a method according to some aspects of the present disclosure.

[0021] [Figure 12]Figure 12 shows a method according to several embodiments of the present disclosure.

[0022] [Figure 13] Figure 13 illustrates how to automatically or programmatically generate additional input-output datasets according to several aspects of this disclosure.

[0023] [Figure 14] Figure 14 shows an exemplary computing environment with exemplary computer equipment suitable for use in several exemplary implementations. [Modes for carrying out the invention]

[0024] The following detailed description provides further details of the drawings and implementation examples of this application. For clarity, redundant reference numbers and descriptions of elements between the drawings have been omitted. Terms used throughout this description are provided as examples and are not intended to be limiting. For example, the use of the term “automatic” may include a fully automatic implementation or a semi-automatic implementation that includes user or administrator control over certain aspects of the implementation, depending on the desired implementation for those skilled in the art practicing the implementation of this application. Selection may be made by the user via a user interface or other input means, or may be implemented by a desired algorithm. The implementation examples described herein may be used individually or in combination, and the functionality of the implementation examples may be implemented by any means depending on the desired implementation.

[0025] Figure 1 is a figure 100 showing a set of functional components of the system in several embodiments of this disclosure. In some embodiments, the system may include a data management component 102, a statistical inference / data analysis component 103, a distributed model learning component 104, a three-dimensional (3D) image generation component 105, a generative model learning component 106, a tensor framework image generation component 107, and an auto-labeling / annotation component 108. Some components (e.g., the statistical inference / data analysis component 103, the generative model learning component 106, and the auto-labeling / annotation component 108) may interact with a user 109. Different functions and subcomponents of the components shown in Figure 100 will be described below in relation to Figures 2 to 6.

[0026] In some embodiments, the data management component 102 may manage multiple types of datasets and metadata associated with different types of datasets. Figure 2 is Figure 200, which shows several embodiments of the data management component 102 and the auto-labeling / annotation component 108 of Figure 1, according to several embodiments of the present disclosure. In some embodiments, the data management component 102 may be associated with multiple types of data, such as ultraviolet (UV), infrared (IR), or thermal images (e.g., UV / IR / thermal image data 205), X-ray, magnetic resonance imaging (MRI), or positron emission tomography (PET) scan images (e.g., X-ray / MRI / PET scan image data 210), or general images (visible light image data 215). Different images may be monochrome (e.g., grayscale) or multi-channel (e.g., blue / red or red / green / blue (RGB)). Image data may further include 2D or 3D images (e.g., RGB and depth). The data management component 102 may generate an aggregated dataset (e.g., an imbalanced dataset 220) containing one or more data types (or other types of data to which the methods described in this disclosure may be applied).

[0027] The imbalanced dataset 220 may, in some embodiments, be provided as input for the image caption operation 225. The image caption operation 225 may, in some embodiments, provide additional information about where the image was acquired or other context for the image. The capted image data generated by the image caption operation 225 may, in some embodiments, be provided as input to element 505 in Figure 5, as described later. In some embodiments, the imbalanced dataset 220 may also be provided to the image metadata operation 230 to add metadata such as location (e.g., GPS), camera type, focal length, equipment manufacturer, equipment model, f-number, lens dimensions, or other metadata (e.g., used to calculate the distance between the camera and the object). The image data with metadata generated by the image metadata operation 230 may, in some embodiments, be provided as input to element 520 in Figure 5, as described later. The imbalanced dataset 220 may be further provided to one or more of the auto-labeling / annotating components 255 and / or merging components 245 to generate an annotated balanced dataset 250 which may then be provided to element 320 in Figure 3. In some embodiments, the imbalanced dataset 220 may be further provided to elements 310 and 410 in Figures 3 and 4. In some embodiments, the imbalanced dataset 220 may be provided to other elements in Figures 2 to 6 in relation to different stages (or functions) of the method.

[0028] In some embodiments, after generating an additional set of images to balance the dataset in element 635 of Figure 6 (described later), the additional set of images or data (e.g., a new complementary dataset 235) may be provided to the auto-labeling / annotation component 255 along with the imbalanced dataset 220. In some embodiments, the auto-labeling / annotation component 255 may annotate the unannotated (or unlabeled) data from the imbalanced dataset 220 and the new complementary dataset 235 to generate a new annotated complementary dataset 240. In some embodiments, the auto-labeling / annotation component 255 may use a machine learning model received from element 325 of Figure 3. The (annotated) imbalanced dataset 220 and the new annotated complementary dataset 240 may then be provided to the merge component 245 to generate an annotated balanced dataset 250. In some embodiments, the auto-labeling may be validated by user 202.

[0029] Biased datasets, such as imbalanced dataset 220, are a very common problem in data science. For example, imbalanced dataset 220 may, in some embodiments, have more examples of properly functioning industrial components than examples of industrial components that are defective or experiencing failure events. In some cases, imbalanced dataset 220 is provided to element 310 during the initial training of a model (e.g., a machine learning model), as described in relation to Figure 3. Figure 3 is a figure 300 illustrating a set of operations associated with a distributed model training component 104 according to some embodiments of the present disclosure. In some embodiments, initial model training may begin by receiving imbalanced dataset 220 and using preparatory operations 310 to identify within a plurality of classifications at least one classification (e.g., a label or class) associated with imbalanced dataset 220 (e.g., an initial training data set) that is underrepresented (and conversely, overrepresented by another classification) with respect to imbalanced dataset 220. Next, the dataset analysis and preparation operation 310 may perform one or more downsampling of data associated with overrepresented classifications or upsampling of data associated with underrepresented classifications (e.g., on the imbalanced dataset 220) to generate a more balanced set of data for training. Upsampling may, in some embodiments, involve duplicating a subset of data associated with underrepresented classes, while downsampling may involve ignoring a subset of data associated with overrepresented classes (e.g., images associated with or labeled as normal). For example, in the context of industrial applications where defects and / or failures may be rare, underrepresented classifications may be associated with images classified or labeled as defect or failure events, while images classified or labeled as normal may be overrepresented.

[0030] For initial model training, the imbalanced dataset 220 may be provided for the model training operation 320 after the dataset analysis and preparation operation 310. In some embodiments, for subsequent training, the annotated balanced dataset 250 may be provided for model training and / or updating (refinement). The input dataset (e.g., the imbalanced dataset 220 or the annotated balanced dataset 250) may, in some embodiments, be split to generate a training dataset and a validation and / or test dataset. The model training operation 320 may generate a model that accepts input image data (whatever type it is trained on and / or configured for) and may provide classifications into two or more classes. For example, different classifications in an industrial context may include normal / OK, a first type of defect (indicating a potential or likely failure in the near future), or a second type of defect (a complete failure or failure event). The model training operation 320 may, in some embodiments, use the training dataset to perform iterative model training operations to train the model (e.g., to adjust the weights of the neural network or other parameters for other machine learning models) until some set of criteria is met, such as a threshold accuracy of the trained model (e.g., more than 95% of the samples are correctly labeled). The trained model may also be provided for a model unfolding operation 325, which unfolds the model by providing the trained model to the annotation component 255 in Figure 2 as an object detection model used to automatically label input images (or other trained data types). In some embodiments, a trained model based on an imbalanced dataset may be a starting point for model labeling. For example, by using the trained model, the system may label the images generated in the current iteration. The generated and labeled images in the current iteration may, in some embodiments, be used for subsequent iterations of the model training operation 320.

[0031] In some embodiments, the model learning associated with the model learning operation 320 may be based on a genetic algorithm. The genetic algorithm (as an example of a machine learning algorithm for generating a machine learning model) may, in some embodiments, be set up using hyperparameter choices. In the context of a machine learning algorithm, hyperparameters may, in some embodiments, generally mean parameters used to train the model, such as determining the likelihood of identifying a local minimum of a cost function associated with learning, or the magnitude of changes between steps or other embodiments that affect speed, rather than the machine learning model generated by the machine learning algorithm. The hyperparameter choices may, in some embodiments, include a set of (predefined) random hyperparameters that can be used to generate a matrix of possible solutions. In some embodiments, the genetic algorithm may apply a Natural Evolutionary Strategy (NES) algorithm. The NES algorithm may, in some embodiments, be an optimization algorithm that finds an optimized set of parameters for a machine learning model. In some embodiments, the NES algorithm may compute gradients associated with the changes (or mutations) to the model between each iteration to improve the model outcome by using information derived from both successful (e.g., better or more accurate) and unsuccessful (e.g., worse or less accurate) mutations to the model in previous iterations. In some embodiments, the computed gradients may be natural gradients. In some embodiments, the computed gradients may be used in conjunction with Monte Carlo approximations to determine the set of mutations (or the set of parameters or likelihoods for a particular mutation or the type of mutation) for subsequent generation / iterations of the genetic algorithm (e.g., a natural evolutionary strategy algorithm). The NES algorithm may be executed iteratively until a stopping criterion (e.g., an accuracy threshold) is met, as described above. In some embodiments, the use of natural gradients can prevent premature convergence to local optima while ensuring large update steps.

[0032] In some embodiments, the model training operation 320 may distribute the data (e.g., an imbalanced dataset 220 or an annotated balanced dataset 250) to different agents, each agent performing an independent training sub-operation. Each sub-operation may produce different configurations of the machine learning model and sets of different configurations of the machine learning model, or the outputs of different sets of machine learning models may be aggregated or combined in some way to produce an output for a set of different machine learning models. In some embodiments, the best machine learning model from the sets of different machine learning models may be selected as the model to be deployed. A cost function may be used for each iteration or generation of the NES algorithm to determine the fit of each "child" compared to the parent function or other "children" of the model. In some embodiments, the cost function may be based on accuracy, recall, and intersection overunion (cross-rate).

[0033] In some embodiments, the learning model may be provided for a model accuracy calculation operation 315. The model accuracy calculation operation 315 may calculate accuracy and / or correctness for each identified class (or classification), such as normal / OK, first type defect, second type defect, or other labels / classes identified in the training data (e.g., an imbalanced dataset 220 or an annotated balanced dataset 250). In some embodiments, the model accuracy calculation operation 315 may further include calculating a threshold number 305 (or percentage) for instances (e.g., data points or images) of each class. In some embodiments, the threshold number for each class may be the minimum (or optimal) number or percentage of instances for a balanced dataset (e.g., an annotated balanced dataset 250). The threshold number or percentage of instances may be used to identify classes that are overrepresented and / or underrepresented based on the total number of instances contained in the complete dataset (e.g., an imbalanced dataset 220 for initial training or an annotated balanced dataset 250 for subsequent training). In some embodiments, the total number of instances associated with an over-represented class may be used to determine the minimum (or optimal) number of instances for an under-represented class (for example, based on the determined minimum percentage of instances). The threshold number may then be provided as a parameter of the data generation process in element 405 of Figure 4.

[0034] Figure 4 is a figure 400 illustrating a set of operations associated with a data analysis component 103 according to several embodiments of the present disclosure. In some embodiments, the set of operations may include a threshold number 305 for determining whether a sufficient number of instances (images or data points) exist in a dataset (e.g., an imbalanced dataset 220 or an annotated balanced dataset 250) for a generation process (e.g., a generative adversarial network (GAN) and / or variational autoencoder, conditional latent diffusion model) and a determination 405 based on the number of instances associated with each (e.g., one or more) underrepresented class. If the determination 405 identifies that a sufficient number of instances associated with a particular underrepresented class does not exist, the method may, in some embodiments, proceed to a process based on a human-computer process using a 3D application, an insufficient number of data points (instances) and a subject-specific expert (SME) to generate additional images for use in model training (e.g., a subsequent generation process described in relation to Figure 6, for generating further additional images for model training). Based on the identification that there are no sufficient instances associated with a particular underrepresented class, the method may proceed to a class distribution operation 410. In some embodiments, the class distribution operation 410 may be based on an imbalanced dataset 220, and for each identified class, the method may proceed to element 505 of Figure 5 to generate additional images, which are described in relation to Figure 5 below.

[0035] If determination 405 identifies that there are sufficient instances associated with a particular underrepresented class, the method may, in some embodiments, proceed to selection of underrepresented classes 415 for a generation process, as described in relation to Figure 6, for example, to generate further additional images for model training. Selection of underrepresented classes 415 may be further represented by a 2D image generated by a class distribution calculation 410 for identifying underrepresented classes and element 510 in Figure 5. Based on selection of underrepresented classes 415 (and the number of instances generated by model training in Figure 5), the method may proceed to a generation process for generating additional images (represented, for example, by element 605) in Figure 6. The generation process may, in some embodiments, be a human-computer process using deep learning, data points (instances) associated with underrepresented classes, and SME. For each selected class, the method may proceed to a balancing operation 420 which computes several additional instances to balance the representation of the selected class to provide to the new instance creation operation 635 in Figure 6 (for example, to determine when the new instance creation operation 635 stops creating additional instances and / or data points).

[0036] Figure 5 is a figure 500 illustrating a set of operations associated with an image generation component 105 according to several embodiments of the present disclosure. The set of operations shown in Figure 500 may, in some embodiments, be based on several 2D images of the same object which can be used to render a 3D virtual object representing the object. In some embodiments, the 3D virtual object may be generated on a generative model of high-quality 3D images (or virtual objects). The generative model may, in some embodiments, include a combination of a convolutional neural network (CNN) and a generative neural network (e.g., a GAN) which can be used to generate different situational scenarios for the image. Different scenarios may, in some embodiments, be associated with different locations and / or orientations of a virtual camera for compositing 2D images of the 3D virtual object. In some embodiments, different scenarios may be further associated with different states of the 3D virtual object, such as time or material, as described later in relation to the 2D image generation operation 515, in relation to different states and irradiance algorithms, such as different lighting or light sources. The first CNN may, in some embodiments, be used to generate a shape differential surface representation which imparts shape sides of hidden faces of the object to the image. In some embodiments, the second CNN may be used to generate textures for an image. For example, the second CNN may be used to apply color and a 2D silhouette to an image using differentiable rendering. In some embodiments, both the first and second CNNs may use pre-trained models that have representations of the object's shape and texture.

[0037] For example, Figure 500 shows that additional image generation may be initiated by a set of 2D raw images 505 (in this case, the images may be of any of the types described above). In some embodiments, the set of 2D raw images 505 may include a subset of images associated with corresponding objects (e.g., industrial components) within a set of one or more different objects. In some embodiments, the subset of images associated with a particular object within the set of 2D raw images 505 may include multiple images taken from different angles and / or camera positions, which may be determined based on captions provided by the image caption operation 225. In some embodiments, the different objects may all be objects associated with a particular label, class, or classification, and the process described later may be performed for each label, class, or classification that is determined not to have enough instances to perform the generation process for generating additional images.

[0038] For each object associated with a subset of images from the set of 2D raw images 505, the method may provide the subset of images from the set of 2D raw images 505 to the additional feature generation operation 510 in order to create (generate or identify) additional features (e.g., parameters and / or modification algorithms) of the object in order to simulate aging (e.g., exposure to one or more environmental conditions over a period of time of one or more years) or to modify the image to represent a similar object made from a different material, for example, by using a ray autoencoder, as described later in relation to Figure 7. The subset of images from the set of 2D raw images 505 may also be provided to the 2D image generation operation 515 (along with the features or modification algorithms provided by the additional feature generation operation 510). In some embodiments, the 2D image generation operation 515 may generate additional 2D images for different sets of features (e.g., different combinations of environmental conditions, duration, and material type) based on a particular object and the subset of images associated with the additional features.

[0039] In some embodiments, a subset of the set of 2D raw images 505, including objects and associated images (and each additional set of 2D images for different sets of features), may be provided to a 3D object generation operation 520 to generate a 3D model of the object (or a 3D model for each object of different sets of features). In some embodiments, the 3D object generation operation 520 may receive image metadata from an image metadata operation 230 for use in generating a 3D model of the object. The generated 3D model (or multiple 3D models) may then be provided to a hyperparameter configuration operation 525 to generate one or more sets of parameters used with the 3D model to generate additional 2D images (instances and / or data points). In some embodiments, the set of (hyper)parameters generated by the hyperparameter configuration operation 525 may include luminance, brightness, or other (hyper)parameters that may affect the image generated from the 3D model. The 3D model (or multiple models) may then be labeled and / or classified by a subject matter expert (SME) as part of a manual labeling operation 530 (for example, input from an SME may be received).

[0040] In some embodiments, the set of (hyper)parameters may include one or more parameters relating to the visual quality of an image or an object in an image, such as (1) the luminance of an object, relating to the level of visual perception at which an object is observed to emit or reflect light; (2) the contrast, relating to the luminance or color difference that makes an object distinguishable from other objects in the image; (3) the saturation, relating to the intensity of the colors in the image; (4) the color scheme, relating to whether the image is a color or grayscale image; (5) the hue, relating to the (changing) color channels of the input image; and (6) blurring, relating to the mixing of neighboring pixel values ​​to mimic a lack of focus. The set of (hyper)parameters may, in some embodiments, include one or more parameters related to image orientation or the at least one processor, such as (1) image rotation, (2) image flipping, (3) image pruning, (4) image resizing (e.g., adjusting the aspect ratio), (5) cutout (e.g., random covering of an area of ​​an image with a random set of pixels or pixels having the average pixel value of a training set), (6) mosaic associated with tilting different images to generate a new image, (7) cut mix associated with random cutting of a portion of a first image and its placement on another image, or (8) mixup associated with generating weighted combinations of random image pairs. These sets of (hyper)parameters may provide more samples for the dataset and expose the model to new sets of data. These more samples improve the generalization of the model by acting as a normalization of the model to avoid overfitting in some embodiments. Therefore, for each original image (raw 2D image or composite 2D image generated from a 3D model), in some embodiments, several images may be generated based on a selection of (hyper)parameters chosen by the user.

[0041] Based on the (hyper)parameters generated in the hyperparameter configuration operation 525 (based on user selection) and the labels and / or classifications applied in the manual labeling operation 530, an additional set of 2D images may be generated in the 2D image generation operation 535. In some embodiments, the additional set of 2D images may include a subset of images from the set of 2D raw images 505 and, for each associated object, a plurality of additional composite images (e.g., from a plurality of virtual camera positions and / or orientations relative to a 3D model). In some embodiments, the plurality of composite images may be further multiplied by the application of a set of operations based on morphisms (a set of features based on an autoencoder or other modifications / adjustments) and one or more (hyper)parameters. In some embodiments, the additional set of images may simulate several physical scenarios associated with the images for use in training a machine learning-trained autolabeling model and a more accurate and diverse (e.g., non-overfitting) machine learning model.

[0042] In some embodiments (indicated by dotted connectors), the manual labeling operation 530 may be performed after the 2D image generation operation 535 to ensure that the labels are accurate not only for the object as a whole, but also for each specific image (or set of images associated with the same viewpoint), to ensure that defects associated with the object are detectable / identifiable based on the observation angle, set of features and / or (hyper)parameters used to generate the image. In some embodiments, the SME may label after the generation of a set of core images (e.g., 2D images generated from a 3D model) used for data generation (and augmentation) via the manual labeling operation 530. For example, the SME may generate a seed annotation for each core image. In some embodiments, this seed annotation may be associated with an image generated in the data generation process described above (or later in relation to the generation process described in relation to Figure 6). For example, if image A is generated by bounding box (x1, y1, x2, y2) and class c, then images generated using image A based on the above-described morphism or (hyper)parameters may have the same annotations, such as bounding box (x1, y1, x2, y2) and class c. As a result, the seed annotation may, in some embodiments, be a reference for all images generated from the original image, and the images may be included in a set of data (such as an unbalanced dataset 220 or an annotated balanced dataset 250) along with the instructions of the seed annotation. Additional images (or information about the number of additional images generated associated with their respective labels and / or classifications) may be provided to the underrepresented class selection 415, which is used to select underrepresented classes, and a generation process for generating additional images is performed.

[0043] Figure 600 shows a set of operations associated with the generative model learning component 106 and the tensor framework image generation component 107 according to some aspects of the present disclosure. Figure 600 shows that the set of operations may be initiated by filtering an imbalanced dataset 220 (e.g., original image data or generated original image data and additional images as described in relation to Figure 5) to receive a selection of classes from a selection of underrepresented classes 415 for a filtering operation 605 and to isolate a dataset for at least one (or each) underrepresented label / class classification. In some aspects, the filtering operation is pseudocoded as follows: g=max(x k ) and r k =gx k , may include processes described by k∈[1,n], where n is the number of classes and x k is the number of records in class (k), g is the number of samples required for each class, and r k is the number of elements that need to be upsampled for the k-th class. An exemplary filtering process results in a balanced dataset containing the same number of instances (data points such as images and videos) for each class.

[0044] The filtered data may be provided to one or more of the following: GAN training operation 610, variational autoencoder training operation 615, and diffusion model training operation 620, or other training operations for the model / network to generate images. In some embodiments, after iterations of the training operation (one of the GAN training operation 610, variational autoencoder training operation 615, or diffusion model training operation 620), one or more images generated by the training model or network may be provided to the similarity determination component 625. The similarity determination component 625 may perform a similarity determination (e.g., FID or other measure of similarity) to determine whether the images generated by the machine learning model or network are sufficiently similar to a set of test images generated to terminate training (e.g., whether the Fréchet starting distance (FID) is less than a threshold). In some embodiments, the similarity determination component 625 may also be used to determine whether to use each of the machine learning models and / or networks, in addition or instead.

[0045] If the similarity determination component 625 determines that the generated images are not sufficiently similar, the similarity determination component 625 may indicate that the associated learning operation should continue learning (e.g., perform additional iterations of the learning process). This process may be repeated until the similarity score satisfies a configured threshold (e.g., FID < threshold). If the similarity determination component 625 determines that the generated images are sufficiently similar to one or more machine learning models or networks, the associated one or more machine learning models or networks may be expanded by the model expansion component 630. Once expanded, the one or more machine learning models or networks may be used for the new instance generation operation 635 corresponding to the tensor framework image generation component 107. In some embodiments, the new instance generation operation 635 may generate new instances (e.g., images or data points) for underrepresented labels / classes / classifications. The new instance generation operation 635 may generate additional instances based on the number of additional instances in order to balance the representation of selected classes provided by the balancing operation 420. In some embodiments, the new instance generation operation 635 may generate a new complementary dataset 235 (e.g., a new dataset to complement or balance the imbalanced dataset 220) as shown in Figure 2. As shown in Figure 2, the new complementary dataset 235 may, in some embodiments, not be labeled / classified / annotated. As described above in relation to Figures 2 and 3, the data generated by the new instance generation operation 635 may be added to the imbalanced dataset 220 to generate a balanced dataset, which may then be used to retrain and / or update (e.g., refine) the initial training model trained on the imbalanced dataset 220.

[0046] FIG. 7 is FIG. 700 showing shooting elements according to some aspects of the present disclosure. Shooting is applied using deep learning techniques (e.g., Sham network or twin network) that, in some aspects, involve learning two different neural networks (e.g., first neural network 715 and second neural network 725) having both an encoder and a decoder architecture simultaneously. The set of weights (e.g., network configuration) associated with the encoder layers (e.g., encoder (E A ) 712 and encoder (E B ) 722) may, in some aspects, be a set of shared weights. For example, the set of shared weights may be learned to recognize elements common to a set of images (associated with latent image 714 and latent image 724).

[0047] In some aspects, a first group of images of a first object may be used to learn a first neural network (as an example of a machine learning network), while a second group of images of a second different object may be used to learn a second neural network. The first and second neural networks (e.g., each of first neural network 715 and second neural network 725) may have independently learned decoders (e.g., neural network or other machine learning network or model), while being two autoencoder neural networks that share weights for the encoder (e.g., neural network or other machine learning network or model). The first and second neural networks 715 and 725 each receive input data (e.g., an image) from a corresponding set of training data, generate a representation of the input data (e.g., a latent image) using an encoder network, and may be learned to reproduce the input data from the representation of the input data using a corresponding decoder network. For example, the encoder (e.g., encoder (E A ) 712 or encoder (E B ) 722) may, in some aspects, be associated with a decoder (e.g., decoder (DA )716 or decoder (D B The input image (original image 710 or original image 720) may be processed to determine features (e.g., latent images 714 and 724) that can be used to generate an approximation of the original image (e.g., reconstructed image 718 or reconstructed image 728) using )726). The learning may define a metric for the accuracy of the reconstruction (e.g., FID or other similarity metrics) and a threshold similarity for the trained network (autoencoder).

[0048] In some embodiments, the morphism may generate additional data for underrepresented labels / classes / classifications by using the neural network 735 to process a first dataset (first set of images) used to train a first neural network 715 to generate a third set of reconstructed images (including reconstructed images 738) that shares features of a second set of data used to train a neural network 735 (e.g., decoder (DB) 726). In some embodiments, the neural network 735 may effectively be a second neural network 725 trained on a second dataset including the original images 720, because the encoder (E A )712 is an encoder (E B ) is the same as 722, and the decoder is the decoder of the second neural network 725 (D B)726). Similarly, the first neural network 715 may also be used to process a second set of data to generate a fourth set of reconstructed data (additional images) that share features of the first set of data. For example, in some embodiments, the first and second autoencoder networks (e.g., the first and second neural networks 715 and 725) may be trained using a first set of images of a porcelain insulator having a partially broken disk (e.g., including the original image 710) and a second set of images of a glass insulator (e.g., including the original image 720), respectively. The second neural network (associated with the glass insulator) may then be used to generate an image resembling a broken glass disk insulator from the image of the broken porcelain insulator. Such reconstructed images may allow or tolerate rare events, such as partial breakage of the glass disk insulator, to be reflected in the training of a machine learning model for labeling the images. For example, magnetic insulators are used over longer periods than glass disk insulators, and glass disk insulators are more likely to experience complete failure; therefore, the amount of data associated with partial failure of glass disk insulators may be too small to reliably train a machine learning model to recognize such defects and / or failure events. The selection of a particular set of input images and the trained autoencoders may, in some embodiments, be based on the type of images to be generated, such as the type of image identified as underrepresented within the imbalanced dataset 220 in Figure 2. For example, in some embodiments, both trained autoencoders may be used to generate additional data if both of the different input datasets have features that are less likely (or present in fewer numbers) to be present in the other dataset, but which the user would want to reflect or consider when training an auto-labeling machine learning model.

[0049] Figure 8 is a diagram illustrating a GAN according to several embodiments of the present disclosure. In some embodiments, the GAN may be a neural network (NN) architecture that uses two NNs, namely a discriminator NN806 and a generator NN804, to generate a synthetic instance of the data (e.g., sample 802) based on real data 801. When training the GAN, data from an initial dataset (e.g., the imbalanced dataset 220 in Figure 2) may be prepared by splitting the input data (e.g., images, videos, audio, etc., contained within the real data 801) from the relevant outputs (e.g., labels / classes / classifications). The split / separated data may then be used to train the GAN. In some embodiments, the GAN may be trained in a conventional mode that uses annotation data to distinguish between labels / classes / classifications. In some embodiments, the GAN may be trained in an unconditional mode that ignores labels / classes / classifications. Both conditional and unconditional modes of training may be used in some embodiments, and then the trained model having better results (e.g., for a set of test results) is selected.

[0050] In some embodiments, GAN training may include one or more iterations in which the generator NN804 and the discriminator NN806 are trained alternately. For example, while the discriminator NN806 is kept constant for the training of the generator NN804, the generator NN804 (e.g., the neurons and associated weights of the generator NN804) is updated based on the results of the classification / analysis performed by the discriminator NN806. During the training of the generator NN804 (based on a random seed 803), the results of the classification / analysis performed by the discriminator NN806 on the samples 805 generated by the generator NN804 may, in some embodiments, be associated with a generator loss 808 used as feedback for updating the generator NN804. The feedback may, in some embodiments, be in the form of gradient descent learning (e.g., mini-batch stochastic gradient descent learning) or other training methods. Similarly, the generator loss 808 may be calculated using any appropriate loss function. After the training period for the generator NN804, the training period for the discriminator NN806 may begin.

[0051] During the training period for the classifier NN806, the generator NN804 may be kept constant, while the classifier NN806 (e.g., the neurons and associated weights of the classifier NN806) is updated based on the results of the classification / analysis performed by the classifier NN806. The results of the classification / analysis performed by the classifier NN806 on sample 805 generated by the generator NN804 and sample 802 from real data 801 during the training of the classifier NN806 may, in some embodiments, be associated with a generator loss 808 used as feedback for updating the classifier NN806. The feedback may, in some embodiments, be used in the form of gradient descent learning (e.g., mini-batch stochastic gradient descent learning) or other training methods. Similarly, the classifier loss 807 may be calculated using any appropriate loss function. Multiple iterations of alternating training of the generator NN804 and the classifier NN806 may be performed until the images generated by the generator NN804 are indistinguishable from the real data by the classifier NN806. Once the generator NN804 is trained, it may be used to generate additional data (e.g., an input-output dataset). For example, a generator NN804 trained for a first class (based on the first class and its associated data) may be used to generate an input-output dataset, thereby generating additional input data associated with the first class's labels / classes / classifications.

[0052] Figure 9 is a figure 900 illustrating variational autoencoders (VAEs) according to several embodiments of the present disclosure. In some embodiments, the VAE consists of an autoencoder NN that generates a latent space (e.g., latent states 905) from a decoder 902 and an encoder 901, and adds noise to a neural network (after encoder 901 and before decoder 902) to generate data. In some embodiments, the VAE is a type of generative model that uses unsupervised techniques to model the distribution of input data and then uses random data generation to generate a new image. In some embodiments, the VAE is a principled framework for training deep latent variable models and corresponding inference models. Thus, in some embodiments, the VAE uses a decoder layer (e.g., decoder 902) to learn to reconstruct a noisy input image.

[0053] In some embodiments, the VAE model may include a probabilistic encoder layer (encoder 901), a latent variable layer (z) (layer variable 910), and a probabilistic decoder layer (decoder 902). In some embodiments, the VAE learns a combined probability distribution of observable data and latent space that assigns model parameters using backpropagation. In some embodiments, Kullback-Leibler (KL) divergence may be used to calculate the distance between the approximate posterior probability and the true posterior probability. The KL divergence may also be used to compute the gap between the evidence lower bound (ELBO) and the boundary tightness. In some embodiments, the objective function of the VAE is the maximization of the ELBO, and the ELBO approach, in some embodiments, maximizes the difference in the log-likelihoods of the observed dataset and (e.g., θ) * ,φ * =argmaxL θ,φThe goal may be to find one or more optimal parameters for the approximate and exact posterior probabilities by minimizing the divergence between the approximate and exact posterior probabilities (using a differentiable loss function such as (x)). In some embodiments, gradient descent optimization techniques are used to search the space of parameter choices and to maximize the ELBO function. In some embodiments, other loss functions and / or methods may be used to search the space of parameter choices.

[0054] Figure 10 is a diagram illustrating diffusion models according to several embodiments of the present disclosure. In some embodiments, the diffusion model may include a cross-attention architecture of encoders and decoders that learn how to reconstruct the original image as a noisy image. In some embodiments, the diffusion model may be a conditional latent diffusion model (CLDM). In some embodiments, the diffusion model may include a probabilistic model designed to learn a data distribution p(x) by gradually removing noise from a normally distributed variable, which corresponds to learning the inverse process of a fixed Markov chain.

[0055] In some embodiments, based on multiple types of machine learning models for generating additional images, the system and / or method may generate N samples from each method (e.g., GAN, VAE, and CLDM) and compare the input data with the generated data from the three methods. The metric used for image comparison may, in some embodiments, be a FID that compares the distribution of the generated images with the distribution of the observed images. For example, the FID is a distance formula that measures the sum of the elements on the diagonal of the result of twice the square root of the sum of the covariance matrices X and Y - the sum of the elements on the diagonal of the product of the covariance matrices X and Y.

number

[0056] In some embodiments, FID serves two purposes within the system: the first being during model training, as described in relation to Figure 6, and the second relating to generative model selection after model training. For example, in some embodiments, FID may function as a stopping criterion for model training. To compare multiple types of machine learning models, the method or system may compute a set of FIDs for each of the sets of data generated by different machine learning models. For example, in the case of a set of generated images (or other data types),

number

number

number

[0057] In some embodiments, the conditional model selection step (i.e., selection of the best machine learning model) is performed for each label / class / classification. Therefore, the selection of the best machine learning model is, in some embodiments,

number

[0058] Figure 11 is a flowchart 1100 illustrating a method according to several embodiments of the present disclosure. In some embodiments, the method is performed by an inference engine or analytical device (e.g., the system or computer device 1405 in Figure 100) that performs various analytical, machine learning operations, data augmentation operations, and inference (e.g., classification) operations based on collected data relating to industrial processes and / or components. The method may be for industrial failure detection based on an initial set of training data. The initial set of training data may include a plurality of input-output datasets, each containing input data relating to industrial objects and output data relating to at least one classification of the input data. In 1110, the device may identify, within the initial set of training data, at least one classification of the input data that is underrepresented among a plurality of classifications of the input data. Referring, for example, to Figures 1 and 3, 1110 may be performed by a distributed model learning component 104 or a dataset analysis and preparation operation 310, as described in relation to Figures 1 to 6. In some embodiments, the input data may include at least one of the following: image data, video data, audio data, X-ray data, MRI data, PET scan data, infrared image data, UV image data, thermal data, or any other type of industrial data that is sensitive to analysis by a machine learning network for classification.

[0059] In addition to identifying at least one underrepresented classification of the input data among multiple classifications of the input data, the device may also identify at least one additional overrepresented classification of the input data among multiple classifications of the input data within the initial training data set. In some embodiments, the identification of underrepresented and underrepresented classifications is based on a target distribution of data points (or instances such as images or other data) among multiple classifications for generating an accurate machine learning model. In some embodiments, the target distribution may represent a range (proportion or percentage) of acceptable distributions for each of multiple classifications, which can be translated into several data points based on the total number of data points in the initial (e.g., imbalanced) dataset for machine learning-based model training. For example, referring to Figures 1 and 3, the identification of at least one additional overrepresented classification of the input data among multiple classifications may be performed by the distributed model learning component 104 or the dataset analysis and preparation operation 310, as described in relation to Figures 1 to 6.

[0060] Based on the identification of underrepresented (and overrepresented) classifications in 1110, the apparatus may, in some embodiments, update the initial dataset by at least one of downsampling associated with overrepresented classes or upsampling associated with underrepresented classes. Downsampling may, in some embodiments, include removing (or ignoring) a first number of input-output data from the initial training data set for training a machine learning model. Upsampling may, in some embodiments, include duplicating data points associated with underrepresented classifications to ensure that there are enough examples of underrepresented classifications to affect the configuration of the machine learning model (e.g., the set of weights associated with the machine learning model). For example, referring to Figures 1 and 3, upsampling and / or downsampling may be performed by the distributed model learning component 104 or the data analysis and preparation operation 310, as described in relation to Figures 1 to 6.

[0061] In some embodiments, the device may train an initial machine learning model based on an initial set of updated data. In some embodiments, the initial training of the machine learning model may be based on a genetic algorithm as described in relation to Figure 2. The initial training of the machine learning model may be validated using a subset of the initial training data set that is not used to train the machine learning model (e.g., a related set of validation data derived from a larger common dataset which is subdivided into a training data set and a related set of validation and / or test data). Validation may, in some embodiments, be used to determine when to terminate the training operation and deploy the machine learning model. For example, referring to Figures 1 and 3, the initial training may be performed by a distributed model training component 104 or a model training operation 320 as described in relation to Figures 1 to 6.

[0062] In some embodiments, the device may determine the minimum number of input-output datasets for each classification to generate a balanced set of training data. In some embodiments, the determination may also be based on analysis performed in 1110. The minimum number of input-output datasets for a particular classification may be based on the desired (or known) ratios, proportions, and / or percentages associated with different classifications for generating (or training) an accurate model, as well as the total number of data points (e.g., input-output datasets or instances). For example, referring to Figures 1 and 3, the determination may be performed by the distributed model training component 104 or the model accuracy calculation operation 315 to generate a threshold number 305, as described in relation to Figures 1 to 6.

[0063] In 1160, the device may automatically (or programmatically) generate additional input-output datasets for at least one identified classification to be included in the modified set of training data, along with the initial set of training data, to balance the classification representations of multiple classifications in the modified set of training data. Figure 13 shows how to automatically or programmatically generate additional input-output datasets according to some aspects of the present disclosure. Referring to Figures 1 and 4-6, the methods of 1160 and Figure 13 may, in some aspects, be performed by one or more of the data analysis component 103, the image generation component 105, the generative model training component 106, the tensor framework image generation component 107, elements 505-535 in Figure 5, elements 605-635 in Figure 6, or in relation to the decision 405. The generation of additional input-output datasets may, in some aspects, include one or more different types of image generation algorithms. For example, a first set of non-machine learning algorithms, such as algorithms based on morphisms or algorithms based on 3D model generation, as described in relation to elements 505-535 of Figure 5, or a second set of machine learning algorithms, such as those described in relation to elements 605-635 of Figure 6.

[0064] As part of the generation of additional input-output sets in 1160, the device may, in 1361, determine whether a sufficient number of input-output datasets associated with the underrepresented classifications (if multiple underrepresented classifications are identified, the currently selected underrepresented classifications) exist to train one or more of the second sets of machine learning algorithms. In some embodiments, the first set of non-machine learning algorithms may be used (as shown in Figure 13) if the number of data points / instances for the underrepresented classifications is not sufficient (e.g., does not meet the threshold) to train one or more of the second sets of machine learning algorithms. If the number of data points / instances (input-output datasets) is sufficient, the device may bypass the first set of non-machine learning algorithms (as shown in Figure 13) and generate additional input-output datasets using one or more of the second sets of machine learning algorithms.

[0065] For example, if the device determines that there are not enough input-output datasets associated with underrepresented classifications to train one or more of a second set of machine learning algorithms, the device may identify a first set of two-dimensional images or videos for each industrial object within a set of at least one industrial object. In some embodiments, the set of at least one industrial object may include industrial objects for which there are enough images to generate a 3D model. For example, referring to Figure 5, the device may identify a set of 505 2D raw images.

[0066] After identifying the first set of two-dimensional images or videos, the apparatus may, in 1363, generate a three-dimensional representation of each of at least one industrial object based on the first set of two-dimensional images or videos for each of at least one industrial object. For example, referring to Figure 5, the apparatus may use a set of two-dimensional raw images 505 to perform the 3D object generation operation 520. In some embodiments, the first set of two-dimensional images or videos may be expanded based on morphology or other modifications / adjustments to generate a virtual object from which a model can be generated.

[0067] After generating a 3D representation, the device may generate a set of 2D images or videos based on the 3D representation of at least one industrial object, which will be included as input datasets for additional input-output datasets. For example, referring to Figure 5, the device may use a set of 3D models generated by the 3D object generation operation 520 to generate an additional set of 2D images by the 2D image generation operation 535. In some embodiments, the first set of 2D images or videos may be expanded based on morphology or other modifications / adjustments to generate virtual objects from which the models can be generated. After generating additional input-output datasets using the first set of non-machine learning algorithms, the device may return to determining whether there are enough input-output datasets associated with underrepresented classifications.

[0068] If the device determines that there are enough input-output datasets associated with underrepresented classifications to train one or more machine learning algorithms, the device may proceed to train at least one of a second set of machine learning algorithms (e.g., at least one of GANs, variational autoencoders, or diffusion models). Training may, in some embodiments, depend on the type of machine learning algorithm being trained and may be performed through multiple iterations of processing the input data of the input-output datasets to determine the accuracy of the output of the algorithm or model in the current training step and to update the algorithm or model to improve accuracy. Training may continue until a stopping criterion is met (e.g., a threshold number of iterations without convergence or a threshold accuracy criterion is met). Specific machine learning algorithms are described in relation to Figures 8–10, but other machine learning algorithms or models may be used according to the needs of the specific application of the method in some embodiments of this disclosure. For example, referring to Figure 6, the device may perform a GAN learning operation 610, a variational autoencoder learning operation 615, and one or more other learning operations for a model / network for generating images in relation to a diffusion model learning operation 620 or a similarity determination component 625.

[0069] Once trained, the device may generate additional 2D images or videos (associated with underrepresented classifications) to be included as input datasets for additional input-output datasets by using a (machine-learned) generation process. For example, referring to Figure 6, one or more GANs, variational autoencoders, or diffusion models trained by the GAN training operation 610, variational autoencoder training operation 615, and diffusion model training operation 620, respectively, may be expanded by the model expansion component 630 to generate additional instances using the new instance generation operation 635. The additional 2D images or videos generated in 1160 may, in some embodiments, already be associated with classifications (or labels) based on the “parent” input images or videos (or set of images or videos) used to generate the additional 2D images or videos. In some embodiments, the labels may be generated and / or validated by an SME that examines or validates one or more of the initial machine learning models or at least a subset of the generated additional 2D images or videos as described in relation to the auto-labeling / annotation component 255 in Figure 2.

[0070] The device may then determine whether a sufficient number of additional input-output datasets have been generated. In some embodiments, the determination may be based on a determination of the minimum number of input-output datasets for each classification to generate a balanced set of training data and the current number of input-output datasets associated with the underrepresented classification (based on a second set of machine learning algorithms or, respectively, the first and second sets of non-machine learning and machine learning algorithms). For example, referring to Figures 3, 4, and 6, the new instance generation operation 635 may continue until the number of instances associated with a classification satisfies or exceeds the minimum number of instances indicated by the threshold number 305 or the balancing operation 420. If the device determines that a sufficient number of additional input-output datasets have not been generated, the device may return to generating additional 2D images or videos to be included as input datasets for the additional input-output datasets. The generation of additional input-output datasets may be performed for each of several classifications identified as underrepresented, based at least on steps or operations 1110 and 1160.

[0071] If the device determines that a sufficient number of additional input-output datasets have been generated, the device may, in 1160, automatically (or programmatically) generate additional input-output datasets for at least one identified classification to be included in the modified set of training data along with the initial set of training data, thereby balancing the classification representations of multiple classifications in the modified set of training data, and in 1170, train a machine learning model based on the modified set of training data having at least a subset of the initial set of training data and the additional input-output datasets. As described in relation to the training performed by the model training operation 320, model training may use a genetic algorithm (e.g., the NES algorithm) or other training algorithm to train (or update) a machine learning model based on a balanced dataset that includes at least a subset of the initial set of training data (e.g., the initial training dataset with or without associated data that is overrepresented, ignored, or removed) and the additional input-output datasets generated in 1160 (for one or more underrepresented classifications).

[0072] After training the machine learning model in 1170, the device may use the machine learning model to generate at least one corresponding classification for at least one input dataset relating to at least one industrial object to identify whether at least one industrial object is associated with a failure event or a non-failure event. In some embodiments, the at least one industrial object may not have been labeled / classified in association with industrial defect and / or failure detection operations (e.g., it may be newly collected data) or may be from a set of test data.

[0073] Figure 12 is a flowchart 1200 illustrating a method according to several embodiments of the present disclosure. In some embodiments, the method is performed by an inference engine or analytical device (e.g., the system or computer device 1405 in Figure 100) that performs various analytical, machine learning operations, data augmentation operations, and inference (e.g., classification) operations based on collected data relating to industrial processes and / or components. The method may be for industrial failure detection based on an initial set of training data. The initial set of training data may include a plurality of input-output datasets, each containing input data relating to industrial objects and output data relating to at least one classification of the input data. In 1210, the device may identify, within the initial set of training data, at least one classification of the input data that is underrepresented among a plurality of classifications of the input data. Referring, for example, to Figures 1 and 3, 1210 may be performed by a distributed model learning component 104 or a dataset analysis and preparation operation 310, as described in relation to Figures 1 to 6. In some embodiments, the input data may include at least one of the following: image data, video data, audio data, X-ray data, MRI data, PET scan data, infrared image data, UV image data, thermal data, or any other type of industrial data that is sensitive to analysis by a machine learning network for classification.

[0074] In 1220, the device may, within the initial set of training data, identify at least one additional classification of the input data that is overrepresented among multiple classifications of the input data. In some embodiments, the identification of underrepresented and underrepresented classifications is based on a target distribution of data points (or instances such as images or other data) among multiple classifications for generating an accurate machine learning model. In some embodiments, the target distribution may represent an acceptable range (proportion or percentage) of the distribution of each classification among multiple classifications, which may be converted into a number of data points based on the total number of data points in the initial (e.g., imbalanced) dataset for machine learning-based model training. For example, referring to Figures 1 and 3, 1210 may be performed by the distributed model training component 104 or the database analysis and preparation operation 310, as described in relation to Figures 1 to 6.

[0075] Based on the identification of underrepresented classifications in 1210 and overrepresented classifications in 1220, the device may update the initial dataset in 1230 with at least one of downsampled data associated with the overrepresented classes or upsampled data associated with the underrepresented classes. Downsampling may, in some embodiments, involve removing (or ignoring) a first number of input-output datasets from the initial training data set for training a machine learning model. Upsampling may, in some embodiments, involve duplicating data points associated with underrepresented classifications to ensure that there are enough examples of the underrepresented classifications to affect the configuration of the machine learning model (e.g., the set of weights associated with the machine learning model). For example, referring to Figures 1 and 3, 1230 may be performed by the distributed model learning component 104 or the dataset analysis and preparation operation 310, as described in relation to Figures 1 to 6.

[0076] The device, in 1240, trains an initial machine learning model based on an initial set of updated data generated in 1230. In some embodiments, the initial training of the machine learning model may be based on a genetic algorithm, as described in relation to Figure 2. The initial training of the machine learning model may be validated using a subset of the initial training data set that is not used to train the machine learning model (e.g., a related set of validation data derived from a larger common dataset that is subdivided into a training data set and a related set of validation and / or test data). Validation may, in some embodiments, be used to determine when to terminate the training operation and deploy the machine learning model. For example, referring to Figures 1 and 3, 1240 may be performed by a distributed model training component 104 or a model training operation 320, as described in relation to Figures 1 to 6.

[0077] In step 1250, the device may determine the minimum number of input-output datasets for each classification to generate a balanced set of training data. In some embodiments, the determination in step 1250 may also be based on analysis performed in one or more of steps 1210 and / or 1220. The minimum number of input-output datasets for a particular classification may be based on the desired (or known) ratios, proportions and / or percentages associated with different classifications for generating (or training) an accurate model, as well as the total number of data points (e.g., input-output datasets or instances). For example, referring to Figures 1 and 3, step 1250 may be performed by a distributed model learning component 104 or a model accuracy calculation operation 315 to generate a threshold number 305, as described in relation to Figures 1 to 6.

[0078] In 1260, the device may automatically (or programmatically) generate additional input-output datasets for at least one identified classification to be included in the modified set of training data, along with the initial set of training data, to balance the classification representations of multiple classifications in the modified set of training data. Figure 13 shows how additional input-output datasets are generated automatically or programmatically according to some aspects of the present disclosure. Referring to Figures 1 and 4-6, the methods of 1260 and Figure 13 may, in some aspects, be performed by one or more of the data analysis component 103, the image generation component 105, the generative model training component 106, the tensor framework image generation component 107, elements 505-535 in Figure 5, elements 605-635 in Figure 6, or in relation to the decision 405. For example, the generation of additional input-output datasets may, in some aspects, include one or more different types of image generation algorithms. For example, a first set of non-machine learning algorithms, such as algorithms based on morphisms or algorithms based on 3D model generation, as described in relation to elements 505-535 of Figure 5, or a second set of machine learning algorithms, as described in relation to elements 605-635 of Figure 6.

[0079] As part of the generation of additional input-output sets in 1260, the device may, in 1361, determine whether a sufficient number of input-output datasets associated with the underrepresented classifications (if multiple underrepresented classifications are identified, the currently selected underrepresented classifications) exist to train one or more of the second sets of machine learning algorithms. In some embodiments, the first set of non-machine learning algorithms may be used (as shown in Figure 13) if the number of data points / instances for the underrepresented classifications is not sufficient (e.g., does not meet the threshold) to train one or more of the second sets of machine learning algorithms. If the number of data points / instances (input-output datasets) is sufficient, the device may bypass the first set of non-machine learning algorithms (as shown in Figure 13) and generate additional input-output datasets using one or more of the second sets of machine learning algorithms.

[0080] For example, if the device determines in 1361 that there are not enough input-output datasets associated with underrepresented classifications to train one or more second sets of machine learning algorithms, the device may, in 1362, identify a first set of two-dimensional images or videos for each industrial object within a set of at least one industrial object. In some embodiments, the set of at least one industrial object may include industrial objects for which there are enough images to generate a 3D model. For example, referring to Figure 5, the device may identify a set of 2D raw images 505.

[0081] After identifying a first set of two-dimensional images or videos in 1362, the device may, in 1363, generate a three-dimensional representation of each of at least one industrial object based on the first set of two-dimensional images or videos for each of at least one industrial object. For example, referring to Figure 5, the device may use a set of two-dimensional raw images 505 to perform the 3D object generation operation 520 corresponding to 1363. In some embodiments, the first set of two-dimensional images or videos may, in 1363, be expanded to generate virtual objects for which a model can be generated based on ray or other modification / adjustment.

[0082] After generating a three-dimensional representation in 1363, the device may, in 1364, generate a set of two-dimensional images or videos based on the three-dimensional representation of at least one industrial object, which will be included as input datasets for additional input-output datasets. For example, referring to Figure 5, the device may use a set of 3D models generated by the 3D object generation operation 520 to generate an additional set of two-dimensional images by the 2D image generation operation 535 corresponding to 1364. In some embodiments, the first set of two-dimensional images or videos may be extended in 1363 to generate virtual objects from which models can be generated, based on morphology or other modifications / adjustments. After generating additional input-output datasets using the first set of non-machine learning algorithms, the device may, in 1361, return to determining whether there are a sufficient number of input-output datasets associated with the underrepresented classification.

[0083] In 1361, if the device determines that there are enough input-output datasets associated with underrepresented classifications to train one or more of a second set of machine learning algorithms, the device may proceed in 1365 to train at least one of the second set of machine learning algorithms (e.g., at least one of GANs, variational autoencoders, or diffusion models). The training in 1365 may, in some embodiments, depend on the type of machine learning algorithm being trained and may be performed through multiple iterations of processing the input data of the input-output datasets to determine the accuracy of the output of the algorithm or model in the current training step and to update the algorithm or model to improve the accuracy. Training may continue until a stopping criterion is met (e.g., a threshold number of iterations without convergence or a threshold accuracy criterion is met). While specific machine learning algorithms are described above in relation to Figures 8 to 10, other machine learning algorithms or models may be used according to the needs of a particular application of the method in accordance with some embodiments of this disclosure. For example, referring to Figure 6, the device may perform, in 1365, one or more other learning operations for a model / network for generating images in relation to a similarity determination component 625, which corresponds to the learning of at least one of a second set of machine learning algorithms, such as a GAN learning operation 610, a variational autoencoder learning operation 615, and a diffusion model learning operation 620.

[0084] Once trained in 1365, the device may, in 1366, generate additional 2D images or videos (associated with underrepresented classifications) to be included as input datasets for additional input-output datasets by using a (machine-learned) generation process. For example, referring to Figure 6, to generate additional instances using a new instance generation operation 635 corresponding to 1366, one or more GANs, variational autoencoders, or diffusion models trained by the GAN training operation 610, variational autoencoder training operation 615, and diffusion model training operation 620 may be unfolded by the model unfolding component 630. The additional 2D images or videos generated in 1260 and / or 1366 may, in some embodiments, already be associated with classifications (or labels) based on the “parent” input images or videos (or sets of images or videos) used to generate the additional 2D images or videos. In some embodiments, the labels may be generated and / or validated by an SME that examines or validates one or more of the initial machine learning models trained in 1240 or at least a subset of the generated additional 2D images or videos described in relation to the automated labeling / annotation component 255 in Figure 2.

[0085] Next, the device may determine at 1367 whether a sufficient number of additional input-output datasets have been generated. In some embodiments, the determination may be based on the determination of the minimum number of input-output datasets for each classification to generate a balanced set of training data at 1250 and the current number of input-output datasets associated with the underrepresented classifications (based on a second set of machine learning algorithms or, respectively, a first and second set of non-machine learning or machine learning algorithms). For example, referring to Figures 3, 4, and 6, the new instance generation operation 635 may continue until the number of instances associated with a classification satisfies or exceeds the minimum number of instances indicated by the threshold number 305 or the balancing operation 420. If the device determines at 1367 that a sufficient number of additional input-output datasets have not been generated, the device may return to 1366 to generate additional 2D images or videos to be included as input datasets for the additional input-output datasets. The generation of additional input-output datasets may be performed for each of the multiple classifications identified as underrepresented, based on at least steps or operations 1210, 1250, and 1260.

[0086] In 1367, if the device determines that a sufficient number of additional input-output datasets have been generated, the device may terminate the method in Figure 13 and proceed from 1260, where it automatically (or programmatically) generates additional input-output datasets for at least one identified classification to be included in the modified set of training data along with the initial set of training data, thereby balancing the classification representations of multiple classifications in the modified set of training data, to 1270, where it trains a machine learning model based on the modified set of training data having at least a subset of the initial set of training data and the additional input-output datasets. As described in relation to the training performed by the model training operation 320, model training may use a genetic algorithm (e.g., the NES algorithm) or other training algorithm to train (or update) a machine learning model based on a balanced dataset including at least a subset of the initial set of training data (e.g., the initial training dataset having or not having data associated with overrepresented or removed classifications) and the additional input-output datasets generated in 1260 (for one or more underrepresented classifications).

[0087] After training the machine learning model in 1270, the device may use the machine learning model in 1280 to generate at least one corresponding classification for at least one input data relating to at least one industrial object in order to identify whether at least one industrial object is associated with a failure event or a non-failure event. In some embodiments, the at least one industrial object may not have been labeled / classified in association with industrial defect and / or failure detection operations (for example, it may be newly collected data) or may be from a set of test data.

[0088] The disclosed methods, apparatus, and systems can, in some aspects, improve the training of machine learning networks in the presence of imbalanced datasets where underrepresented classes (e.g., labels or classifications) can be effectively ignored due to overrepresented classes. For example, a dataset containing inputs of a predetermined percentage associated with a first classification that exceeds a precision threshold (e.g., a 95% precision threshold) used to terminate the training operation (e.g., 99% of images may be associated with a normal classification) can be trained to recognize only the first classification, even if all other classifications are mislabeled as 100% of the time. This is achieved by a machine learning model or algorithm trained to accurately identify (label or classify) only normal states that meet the precision threshold (e.g., have 95.2% precision), or if the machine learning model or algorithm labels everything as normal, achieving 99% precision. Thus, generating additional data to balance imbalanced datasets provides an improvement to machine learning networks for identifying rare events / classifications.

[0089] In addition, the methods, apparatus, and systems may, in some embodiments, offer the benefit of being able to learn accurate models from limited collected data. Automated labeling may also, in some embodiments, save resources by reducing the need for human SME involvement, which is costly and much slower than algorithmic labeling. Improved models may also offer the benefit of being able to accurately identify defects or failure events associated with industrial equipment that can lead to significant losses for businesses in the form of downtime, or, more importantly, catastrophic failure events such as downed power lines that can lead to forest fires or other damage to nearby people or property. In addition, the methods, apparatus, and systems may, in some embodiments, be applied to any type of data using appropriate machine learning networks for data generation and / or analysis / inference.

[0090] Figure 11 shows an exemplary computing environment having exemplary computer equipment suitable for use in several exemplary implementations. The computer equipment 1405 within the computing environment 1400 may include one or more processing units, cores, or processors 1410, memory 1415 (e.g., RAM, ROM, and / or similar), internal storage 1420 (e.g., magnetic, optical, solid-state storage, and / or organic) and / or I / O interfaces 1425, any of which may be connected on a communication mechanism or bus 1430 for transmitting information, or may be incorporated within the computer equipment 1405. The I / O interface 1425 may also be configured to receive images from a camera or provide images to a projector or display, depending on the desired implementation.

[0091] Computer device 1405 may be communicatively connected to input / user interface 1435 and output device / interface 1440. One or both of input / user interface 1435 and output device / interface 1440 may be wired or wireless interfaces and may be detachable. Input / user interface 1435 may include any physical or virtual device, component, sensor, or interface that can be used to provide input (e.g., buttons, touchscreen interfaces, keyboards, pointing / cursor controls, microphones, cameras, Brailles, motion sensors, accelerometers, optical readers, and / or similar). Output device / interface 1440 may include displays, televisions, monitors, printers, speakers, Brailles, or similar. In some exemplary implementations, input / user interface 1435 and output device / interface 1440 may be integrated with or physically connected to computer device 1405. In other exemplary implementations, other computer devices may function as or provide input / user interfaces 1435 and output devices / interfaces 1440 for computer device 1405.

[0092] Examples of computer devices 1405 may include, but are not limited to, highly mobile devices (e.g., smartphones, devices in vehicles or other machines, devices carried by humans and animals, and similar devices), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, and similar devices), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions, radios, and similar devices having one or more processors built in and / or connected thereto).

[0093] Computer device 1405 may be communicably connected (for example, via I / O interface 1425) to external storage 1445 and network 1450 for communication with any number of network-connected components, devices, and systems, including one or more computer devices of the same or different configurations. Computer device 1405 or any connected computer device may function as, provide, or be referred to as a server, client, thin server, general-purpose machine, special-purpose machine, or other label.

[0094] The IO interface 1425 may include, but is not limited to, wired and / or wireless interfaces that use any communication or IO protocol or standard (e.g., Ethernet, 1402.11x, Universal Serial Bus, WiMAX, modem, cellular network protocol, and similar) to communicate information to and from at least all connected components, devices, and networks within the computing environment 1400. The network 1450 may be any network or combination of networks (e.g., the Internet, local area network, wide area network, telephone network, cellular network, satellite network, and similar).

[0095] The computer device 1405 may use and / or communicate using computer-usable or computer-readable media, including temporary and non-temporary media. Temporary media include transmission media (e.g., metal cables, optical fibers), signals, carrier waves, and similar entities. Non-temporary media include magnetic media (e.g., disks and tapes), optical media (e.g., CD-ROMs, digital video discs, Blu-ray discs), solid-state media (e.g., RAM, ROMs, flash memory, solid-state storage), and other non-volatile storage or memory.

[0096] Computer device 1405 may be used to implement techniques, methods, applications, processes, or computer executable instructions in several exemplary computing environments. Computer executable instructions may be obtained from temporary media and stored on and retrieved from non-temporary media. Executable instructions may originate from one or more arbitrary programming, scripting, and machine languages ​​(e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, and others).

[0097] The processor 1410 may run under any operating system (OS) (not shown) in a native or virtual environment. One or more applications, including a logical unit 1460, an application programming interface (API) unit 1465, an input unit 1470, an output unit 1475, and an inter-unit communication mechanism 1495 for different units to communicate with each other, may be deployed with the OS and other applications (not shown). The units and elements described may be modified in design, function, configuration, or implementation, and are not limited to the description provided. The processor 1410 may be in the form of a hardware processor, such as a central processing unit (CPU), or a combination of hardware and software units.

[0098] In some exemplary implementations, information or execution instructions, upon being received by the API unit 1465, may be transmitted to one or more other units (e.g., a logical unit 1460, an input unit 1470, and an output unit 1475). In some cases, the logical unit 1460 may be configured to control the flow of information between units and to control the services provided by the API unit 1465, the input unit 1470, and the output unit 1475 in some exemplary implementations described above. For example, the flow of one or more processes or implementations may be controlled by the logical unit 1460 alone or in conjunction with the API unit 1465. The input unit 1470 may be configured to take input for a computation described in an exemplary implementation, and the output unit 1475 may be configured to provide output based on a computation described in an exemplary implementation.

[0099] The processor 1410 may be configured to identify at least one underrepresented classification among multiple classifications of the input data in the initial training dataset. The processor 1410 may be configured to automatically generate additional input / output datasets for the identified at least one classification, along with the initial training dataset of the corrected training dataset, to balance the classification representations of the multiple classifications in the corrected training dataset. The processor 1410 may be configured to train a machine learning model on the corrected training dataset, which includes at least a subset of the initial training dataset and the additional input / output datasets. The processor 1410 may be configured to use the machine learning model to generate at least one corresponding classification for at least one input dataset related to at least one industrial object in the test dataset, in order to identify whether at least one industrial object is associated with a failure event or a non-failure event. The processor 1410 may be configured to identify a first set of two-dimensional images or videos for each of the at least one industrial object. The processor 1410 may be configured to create a two-dimensional representation of each of the at least one industrial object based on the first set of two-dimensional images or videos for each of the at least one industrial object. The processor 1410 may be configured to generate a set of two-dimensional images or videos contained in the input data of an additional input / output dataset, based on the three-dimensional representation of at least one industrial object. The processor 1410 may be configured to determine the minimum number of input / output datasets for each classification in order to generate a balanced training dataset. The processor 1410 may be configured to identify at least one additional classification that is overrepresented among multiple classifications of the input data in the initial training dataset. The processor 1410 may be configured to remove a first number of input / output datasets from the initial training dataset, and a subset of the initial training dataset consists of the initial training dataset after the removal of the first number of input / output datasets.

[0100] Some parts of the detailed explanation have been presented concerning algorithms and symbolic representations of operations within a computer. These algorithmic descriptions and symbolic representations are means used by those skilled in data processing technology to convey the essence of the innovation to others skilled in the art. An algorithm is a set of predefined steps that produce a desired end state or result. In one implementation example, the steps performed require the physical manipulation of tangible quantities to achieve a specific result.

[0101] Unless otherwise specified, explanations that use terms such as “processing,” “computing,” “calculating,” “determining,” and “displaying” throughout the explanation, as is evident from the explanation, should be understood to include actions and processes of a computer system or other information processing device that manipulate data represented as physical (electronic) quantities in the registers and memory of a computer system and convert it into other data similarly represented as physical quantities in the memory or registers of a computer system, or in other information storage, transmission, or display devices.

[0102] Implementation examples may also relate to apparatus for performing the operations described herein. Such apparatus may be specifically constructed for a required purpose and may include one or more general-purpose computers that are selectively activated or reconfigured by one or more computer programs. Such computer programs may be stored in computer-readable media such as computer-readable storage media or computer-readable signal media. Computer-readable storage media may include, but are not limited to, tangible media such as optical disks, magnetic disks, read-only memory, random-access memory, solid-state devices and drives, or any other type of tangible or non-temporary media suitable for storing electronic information. Computer-readable signal media may include media such as carrier waves. The algorithms and representations presented herein are not specific to any particular computer or other apparatus. Computer programs may include pure software implementations containing instructions for performing the operations in a desired implementation form.

[0103] Various general-purpose systems may be used with the programs and modules illustrated herein, or it may be convenient to construct more specialized devices for performing desired method steps. Furthermore, the implementation examples do not describe any particular programming language. It will be understood that various programming languages ​​may be used to implement the implementations described herein. Instructions in a programming language may be executed by one or more processing units, such as a central processing unit (CPU), processor, or controller.

[0104] As is known in the art, the operations described above can be performed by hardware, software, or any combination of software and hardware. Various aspects of the implementation examples may be implemented using circuits and logic devices (hardware), while other aspects may be implemented using instructions (software) stored on a machine-readable medium, which, when executed by a processor, cause the processor to execute a method for performing the implementation of the present application. Furthermore, some implementation examples of the present application may be performed by hardware alone, while others may be performed by software alone. Moreover, the various functions described may be performed within a single unit or distributed across several components in any number of ways. When performed by software, the method may be executed by a processor such as a general-purpose computer based on instructions stored on a computer-readable medium. If desired, the instructions may be stored on the medium in compressed and / or encrypted form.

[0105] Furthermore, other implementations of the Application will become apparent to those skilled in the art by examining this Specification and practicing the Techniques of the Application. Various aspects and / or components of the implementations described herein may be used individually or in any combination. This Specification and the implementations are intended to be considered merely as examples, and the true scope and spirit of the Application are indicated by the appended claims.

Claims

1. A method for training a machine learning model for industrial fault detection based on an initial training data set that includes a plurality of input-output datasets, each containing input data relating to an industrial object and output data relating to at least one classification of the input data, Within the initial set of training data, identify at least one classification of the input data that is underrepresented among multiple classifications of the input data, To include in the modified set of training data, along with the initial set of training data, an additional input-output dataset for the identified at least one classification is automatically generated to balance the representation of at least one of the multiple classifications in the modified set of training data. Training the machine learning model based on the modified set of training data, which includes at least a subset of the initial set of training data and the additional input-output datasets, A method that includes this.

2. The method according to claim 1, further comprising using the machine learning model to generate at least one corresponding classification for at least one input dataset relating to at least one industrial object in a set of test data, in order to identify whether the at least one industrial object is associated with a failure event or a non-failure event.

3. The input data of the aforementioned additional input-output dataset includes a set of two-dimensional images or videos associated with at least one industrial object, and the automatic generation of the aforementioned additional input-output dataset is: For each of the at least one industrial object, the first plurality of two-dimensional images or videos are identified, Based on the first plurality of two-dimensional images or videos of each of the at least one industrial object, create at least one three-dimensional representation corresponding to the at least one industrial object. Based on the three-dimensional representation of each of the at least one industrial object, a set of two-dimensional images or videos included in the input data of the additional input-output dataset is generated. The method according to claim 1, including the method described in claim 1.

4. The method according to claim 3, wherein the set of two-dimensional images included in the input data of the additional input-output dataset includes a second plurality of two-dimensional images associated with the at least one three-dimensional representation, and generating the second plurality of two-dimensional images includes changing a set of parameters used to generate each of the second plurality of two-dimensional images, the set of parameters including at least one of the position relative to the at least one three-dimensional representation associated with the two-dimensional image, the brightness associated with the two-dimensional image, the saturation associated with the two-dimensional image, or the focus associated with the two-dimensional image.

5. The method according to claim 1, wherein the automatic generation is based on at least one input data associated with the at least one classification and a corresponding morphism.

6. The method according to claim 5, wherein the injection is associated with at least one first autoencoder associated with a first industry object type associated with the at least one input data and a second autoencoder associated with a second industry object type, and at least one additional input-output dataset of the additional input-output datasets for the at least one identified classification includes an input dataset associated with the second industry object type based on the at least one input data, the first industry object type includes a specific industry object manufactured from a first material, and the second industry object type includes the specific industry object manufactured from a second material.

7. The method according to claim 1, wherein the automatically generated step is performed using at least one of a generative adversarial network, a variational encoder, or a diffusion model.

8. The method according to claim 1, wherein the input data includes at least one of image data, video data, audio data, X-ray data, magnetic resonance imaging (MRI) data, positron emission tomography (PET) scan data, infrared image data, or thermal data.

9. The method according to claim 1, further comprising determining the minimum number of input-output datasets for each classification to generate a balanced set of training data, identifying the at least one classification of the input data that is underrepresented by determining that the at least one classification is associated with a first number of input-output datasets in the initial set of training data that is less than the minimum number of input-output datasets, and automatically generating the additional input-output datasets for the identified at least one classification by automatically generating a second number of additional input-output datasets that, when added to the first number, exceed the minimum number.

10. Within the initial set of training data, identify at least one additional classification of the input data that is overrepresented among the multiple classifications of the input data, The method according to claim 1, further comprising removing a first number of input-output datasets from the initial set of training data, wherein the subset of the initial set of training data includes the initial set of training data after removing the first number of input-output datasets.

11. An apparatus for training a machine learning model for industrial fault detection based on an initial training data set that includes a plurality of input-output datasets, each containing input data relating to an industrial object and output data relating to at least one classification of the input data, Memory and Connected to the memory and based at least partially on the information stored in the memory, at least one processor and The at least one processor includes, Within the initial set of training data, identify at least one classification of the input data that is underrepresented among multiple classifications of the input data. To include in the modified set of training data, along with the initial set of training data, an additional input-output dataset is automatically generated for at least one of the identified classifications in order to balance the representation of the classification among the multiple classifications in the modified set of training data. A device configured to train the machine learning model based on the modified set of training data, which includes at least a subset of the initial set of training data and the additional input-output dataset.

12. The apparatus according to claim 11, wherein the at least one processor is further configured to use the machine learning model to generate at least one corresponding classification for at least one input dataset relating to at least one industrial object in a set of test data, in order to identify whether the at least one industrial object is associated with a failure event or a non-failure event.

13. The input data of the additional input-output dataset includes a set of two-dimensional images or videos associated with at least one industrial object, and the at least one processor configured to automatically generate the additional input-output dataset is: For each of the at least one industrial object, a first plurality of two-dimensional images or videos are identified. Based on the first plurality of two-dimensional images or videos of each of the at least one industrial object, at least one three-dimensional representation corresponding to the at least one industrial object is created. The apparatus according to claim 11, configured to generate a set of two-dimensional images or videos included in the input data of the additional input-output dataset, based on the three-dimensional representation of each of the at least one industrial object.

14. The apparatus according to claim 13, wherein the set of two-dimensional images included in the input data of the additional input-output dataset includes a second plurality of two-dimensional images associated with the at least one three-dimensional representation, and the at least one processor configured to generate the second plurality of two-dimensional images is configured to modify a set of parameters used to generate each of the second plurality of two-dimensional images, the set of parameters including at least one of the position relative to the at least one three-dimensional representation associated with the two-dimensional image, the brightness associated with the two-dimensional image, the saturation associated with the two-dimensional image, or the focus associated with the two-dimensional image.

15. The apparatus according to claim 11, wherein the at least one processor is configured to automatically generate additional input-output datasets using at least one input data associated with the at least one classification and associated morphisms.

16. The apparatus according to claim 15, wherein the injection is associated with at least one first autoencoder associated with a first industrial object type associated with the at least one input data and a second autoencoder associated with a second industrial object type, and at least one additional input-output dataset of the additional input-output datasets for the at least one identified classification includes an input dataset associated with the second industrial object type based on the at least one input data, the first industrial object type includes a specific industrial object manufactured from a first material, and the second industrial object type includes the specific industrial object manufactured from a second material.

17. The apparatus according to claim 11, wherein the at least one processor is configured to automatically generate additional input-output datasets using at least one of a generative adversarial network, a variational encoder, or a diffusion model.

18. The apparatus according to claim 11, wherein the input data includes at least one of image data, video data, audio data, X-ray data, magnetic resonance imaging (MRI) data, positron emission tomography (PET) scan data, infrared image data, or thermal data.

19. Apparatus according to claim 11, wherein the at least one processor is further configured to determine a minimum number of input-output datasets for each classification to generate a balanced set of training data, the at least one processor is configured to identify the at least one classification of the input data that is underrepresented, and is configured to determine that the at least one classification is associated with a first number of input-output datasets in the initial set of training data that is less than the minimum number of input-output datasets, and the at least one processor is configured to automatically generate the additional input-output datasets for the identified at least one classification, and is configured to automatically generate a second number of additional input-output datasets that, when added to the first number, exceeds the minimum number.

20. The aforementioned at least one processor is Within the initial set of training data, identify at least one additional classification of the input data that is overrepresented among the multiple classifications of the input data. The apparatus according to claim 11, further configured to remove a first number of input-output datasets from the initial set of training data, wherein the subset of the initial set of training data includes the initial set of training data after the removal of the first number of input-output datasets.