Systems, methods, and computer program products that facilitate generation of synthetic training data

By generating synthetic training data, using element enhancement, modal enhancement and geometric enhancement technologies, the problem of insufficient training data in the existing technology is solved, and the generalization ability and performance of machine learning models is significantly improved.

CN114155402BActive Publication Date: 2025-05-09GE PRECISION HEALTHCARE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110965176.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-21
Filing Date
2021-08-20
Publication Date
2025-05-09
Estimated Expiration
2041-08-20

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively improve the generalization ability of machine learning models, mainly due to the insufficient authenticity, capacity, type and speed of training data.

Method used

By generating synthetic training data, using element enhancement, modal enhancement and geometric enhancement techniques, we create diverse and rich training images, thereby improving the generalization capabilities of machine learning models.

Benefits of technology

It significantly improves the generalization capability and performance of machine learning models, and can more accurately analyze input data in the real world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155402B_ABST
    Figure CN114155402B_ABST
Patent Text Reader

Abstract

The present invention is entitled "Systems, methods and computer program products for facilitating the generation of synthetic training data". An element enhancement component can generate a set of annotated preliminary training images based on annotated source images. The annotated preliminary training images can be formed by inserting at least one element of interest or at least one background element into the annotated source images. A modality enhancement component can generate a set of annotated intermediate training images based on the set of annotated preliminary training images. The annotated intermediate training images can be formed by changing at least one modality-based characteristic of the annotated preliminary training images. A geometry enhancement component can generate a set of annotated deployable training images based on the set of annotated intermediate training images. The annotated deployable training images can be formed by changing at least one geometric characteristic of the annotated intermediate training images. A training component can train a machine learning model on the set of annotated deployable training images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The subject disclosure generally relates to the training of machine learning models, and more specifically to the generation of synthetic training data for improving the generalization ability of machine learning models. Background Art

[0002] The efficacy and / or generalization ability of a machine learning model depends on the authenticity, volume, variety, and / or velocity of the data used to train the machine learning model. In other words, the implementation of high-quality, larger volume, more varied / diverse, and / or more readily available training data can result in the creation of a machine learning model that is not affected by the various challenges faced in real-world operating scenarios. Conversely, the implementation of low-quality, smaller volume, less varied / poorly diverse, and / or less readily available training data can result in the creation of a machine learning model that is easily hindered by the various challenges faced in real-world operating scenarios. Therefore, systems and / or techniques that can increase the authenticity, volume, variety, and / or velocity of available training data may be desirable. Summary of the invention

[0003] The following presents a summary of the invention to provide a basic understanding of one or more embodiments of the invention. The summary of the invention is not intended to identify key or important elements, nor is it intended to delineate any scope of a specific embodiment or any scope of the claims. Its sole purpose is to present concepts in a simplified form as a prelude to a more detailed description presented later. In one or more embodiments described herein, devices, systems, computer-implemented methods, apparatus, and / or computer program products are provided that facilitate the generation of synthetic training data to achieve improved generalization capabilities of machine learning models.

[0004] According to one or more embodiments, a system is provided. The system may include a memory that can store computer executable components. The system may also include a processor that can be operably coupled to the memory and can execute the computer executable components stored in the memory. In various embodiments, the computer executable components may include an element enhancement component that can generate a set of annotated preliminary training images based on an annotated source image. In various aspects, the annotated preliminary training image can be formed by inserting at least one element of interest or at least one background element into the annotated source image. In various cases, the computer executable component may include a modality enhancement component that can generate a set of annotated intermediate training images based on the set of annotated preliminary training images. In various cases, the annotated intermediate training images can be formed by changing at least one modality-based characteristic of the annotated preliminary training images. In various aspects, the computer executable component may include a geometry enhancement component that can generate a set of annotated deployable training images based on the set of annotated intermediate training images. In various cases, the annotated deployable training images can be formed by changing at least one geometric characteristic of the annotated intermediate training images. In various embodiments, the computer-executable components may include a training component that may train a machine learning model on the set of annotated deployable training images.

[0005] According to one or more embodiments, the above-described system may be implemented as a computer-implemented method and / or a computer program product.

[0006] According to one or more embodiments, a computer program product may be provided. In various cases, the computer program product may include a computer readable memory having program instructions embodied therein. In various cases, the program instructions may be executed by a processor to enable the processor to perform various operations. In some cases, such operations may include parameterizing the simulation space of the data segment by defining a set of enhanced subspaces, wherein each enhanced subspace includes a corresponding set of enhanceable parameters. In various cases, each enhanceable parameter may have a corresponding parameter range of possible values ​​or states. In various aspects, the operation may also include receiving a source data segment. In various embodiments, the operation may also include sampling the parameter range of possible values ​​or states corresponding to the enhanceable parameter for each enhanceable parameter. In some cases, this may generate a set of sampled ranges representing the value or state of the simulation space. In various aspects, the operation may also include generating a set of training data segments by applying the set of sampled ranges of value or state to a copy of the source data segment. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] This patent or patent application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0008] Figure 1 A block diagram of an exemplary, non-limiting system that facilitates synthetic training data generation to achieve improved machine learning model generalization capabilities according to one or more embodiments described herein is shown.

[0009] Figure 2 A block diagram of an exemplary, non-limiting system including an element catalog that facilitates synthetic training data generation to achieve improved machine learning model generalization capabilities according to one or more embodiments described herein is shown.

[0010] Figure 3 to Figure 4 A block diagram of an exemplary, non-limiting preliminary training image formed from annotated source images according to one or more embodiments described herein is shown.

[0011] Figure 5 A block diagram of an exemplary, non-limiting system including modality-based features that facilitates synthetic training data generation to achieve improved machine learning model generalization capabilities in accordance with one or more embodiments described herein is shown.

[0012] Figure 6 A block diagram illustrating an exemplary, non-limiting intermediate training image formed from a preliminary training image according to one or more embodiments described herein.

[0013] Figure 7 A block diagram of an exemplary, non-limiting system including geometric transformations that facilitate synthetic training data generation to achieve improved machine learning model generalization capabilities is shown according to one or more embodiments described herein.

[0014] Figure 8 A block diagram is shown of an exemplary, non-limiting deployable training image formed from intermediate training images according to one or more embodiments described herein.

[0015] Fig. 9 Block diagrams illustrating exemplary, non-limiting variations of modality-based and geometric properties according to one or more embodiments described herein.

[0016] Fig.10 Exemplary, non-limiting experimental results according to one or more embodiments described herein are shown.

[0017] Figures 11 to 20 A block diagram of an exemplary, non-limiting image enhancement is shown according to one or more embodiments described herein.

[0018] Fig.21 A flowchart is shown of an exemplary, non-limiting computer-implemented method for facilitating synthetic training data generation to achieve improved machine learning model generalization capabilities according to one or more embodiments described herein.

[0019] Fig. 22 A flowchart is shown of an exemplary, non-limiting computer-implemented method for facilitating synthetic training data generation to achieve improved machine learning model generalization capabilities according to one or more embodiments described herein.

[0020] Fig.23 A block diagram of an exemplary, non-limiting enhanced spatial hierarchy that facilitates synthetic training data generation to achieve improved machine learning model generalization capabilities according to one or more embodiments described herein is shown.

[0021] Fig.24 A block diagram of an exemplary, non-limiting operating environment is shown in which one or more embodiments described herein may be facilitated.

[0022] Fig.25 An exemplary networking environment is shown that is operable to perform various implementations described herein. DETAILED DESCRIPTION

[0023] The following specific embodiments are merely exemplary and are not intended to limit the application or use of the embodiments and / or the embodiments. In addition, it is not intended to be bound by any express or implied information set forth in the aforementioned "background technology" or "content of the invention" section or "specific embodiments" section.

[0024] One or more embodiments are now described with reference to the accompanying drawings, wherein the same reference numerals are used to represent the same elements throughout. In the following description, for the purpose of explanation, many specific details are set forth in order to provide a more thorough understanding of one or more embodiments. However, it is apparent that in various cases, one or more embodiments may be practiced without these specific details.

[0025] A machine learning model may be any suitable artificial intelligence model and / or algorithm capable of mapping a set (e.g., one or more) of input variables to a set (e.g., one or more) of output variables. In various aspects, each output variable may be referred to as a class, classification, label, category, segmentation, detection, etc. In other words, a machine learning model may receive input data and may determine to which category the input data belongs (e.g., the input data may be classified). In other cases, a machine learning model may produce any suitable segmentation, determination, decision, prediction, inference, regression, etc. as output. In various aspects, a machine learning model may be designed and / or configured to receive any suitable type of input data (e.g., scalars, vectors, matrices, and / or tensors) of any suitable dimension, and generate any suitable type of output data (e.g., scalars, vectors, matrices, and / or tensors) of any suitable dimension. As some non-limiting examples, the machine learning model may be configured to perform image recognition, classification, and / or segmentation (e.g., identifying strings of characters, digital objects, and / or alphanumeric objects depicted in an image; identifying flora and / or fauna depicted in an image; identifying anatomical structures depicted in an image; identifying inanimate objects depicted in an image), may be configured to perform sound recognition, classification, and / or segmentation (e.g., identifying spoken letters, words, and / or speech present in audio data; identifying voices present in audio data; identifying faunal sounds present in audio data; identifying the sounds of inanimate objects present in audio data), and / or any other suitable type of data recognition, classification, segmentation, prediction, determination, and / or detection (e.g., distinguishing spam emails from non-spam emails; distinguishing customers who are likely to transact from customers who are unlikely to transact; distinguishing transactions that are likely to be fraudulent from transactions that are unlikely to be fraudulent; etc.). In various aspects, any suitable output dimensionality may be achieved (e.g., binary classification, ternary classification, quaternary classification, and / or any suitable high-order classification).

[0026] In various aspects, a machine learning model may be trained (e.g., via supervised training, unsupervised training, and / or reinforcement learning) to classify, label, and / or make any other determination, prediction, and / or inference about received input data. When supervised training is implemented, each piece of training data may have a corresponding annotation. In various aspects, the corresponding annotation may indicate which true classification the piece of training data is known to belong to (e.g., may represent a true value). During supervised training, a piece of training data may be fed to the machine learning model, and the machine learning model may generate a resulting classification accordingly. In various cases, the difference between the resulting classification and the known annotation (e.g., in backpropagation) may be used to update the parameters of the machine learning model. Updating the parameters of the machine learning model in this manner may help the machine learning model more accurately analyze future input data similar to the training data.

[0027] In various cases, the efficacy of a machine learning model may depend on the quality of the training that the machine learning model undergoes. In other words, when a machine learning model is trained on better and / or higher quality training data, the machine learning model may perform better (e.g., may analyze input data more accurately). In various aspects, the quality of the training data may be described in terms of authenticity, volume, variety, and / or velocity. In various cases, the authenticity of the training data may be related to the accuracy of the known annotations corresponding to the training data (e.g., the parameters of the machine learning model may be accurately updated / adjusted only when accurate annotations of the training data are involved; therefore, if the known annotations of the training data are inaccurate, the training may be invalid, and when the machine learning model is deployed in real life, the machine learning model may not accurately analyze the input data). In various cases, the volume of the training data may be related to the amount of training data (e.g., when more training data is available, the parameters of the machine learning model may be updated / adjusted more comprehensively / appropriately; therefore, if only very little training data is available to feed the machine learning model, the training may be invalid, and when the machine learning model is deployed in real life, the machine learning model may not accurately analyze the input data). In various aspects, the variety of training data may relate to the diversity of features present within the training data (e.g., a machine learning model may be trained to detect and / or ignore only those features present within the training data; thus, if there is not much real-world variety in the features depicted in the training data, the training may be ineffective and the machine learning model may not accurately analyze the input data when it is deployed in real life). In some cases, the velocity of training data may relate to how quickly training data and associated annotations can be collected from the source of the training data (e.g., a machine learning model may be trained only when annotated training data is available; thus, if it takes days, weeks, or months to generate the annotated training data, then it may be necessary to wait days, weeks, or months before training and / or deploying the machine learning model).

[0028] In short, insufficient training data can lead to inadequate machine learning models (e.g., a 5% to 40% drop in performance accuracy when the model operates on a data set not represented by the training data set). Therefore, improving the authenticity, volume, variety, and / or speed of training data can help improve the generalization ability of machine learning models. For example, a machine learning model trained on high-authenticity, high-volume, rich variety, and / or high-speed training data can accurately analyze the input data regardless of the real-world variability of the input data. Conversely, a machine learning model trained on low-authenticity, low-volume, less diverse, and / or low-speed training data can easily run into problems due to the real-world variability of the input data, and therefore may not be able to accurately analyze the input data. Therefore, in various aspects, systems and / or technologies that can improve the authenticity, volume, variety, and / or speed of training data can be desirable.

[0029] Various embodiments of the subject innovation may address one or more of these issues / problems. One or more embodiments described herein include systems, computer-implemented methods, apparatus, and / or computer program products that may facilitate the generation of synthetic training data to achieve improved generalization capabilities of machine learning models. In various cases, embodiments of the subject innovation may be viewed as computerized tools for rapidly generating real, large-volume, and / or diverse training data for any suitable machine learning model. In various aspects, embodiments of the subject innovation may then train a machine learning model on the rapidly generated, real, large-volume, and / or diverse training data, thereby improving the efficacy and / or generalization capabilities of the machine learning model.

[0030] For ease of explanation, the teachings herein about rapidly generating real, large-volume, and / or diverse training data are discussed with respect to a machine learning model configured to classify / label two-dimensional medical images in a clinical setting. However, it should be understood that this is exemplary and non-limiting. In various aspects, the teachings herein can be used to rapidly generate real, large-volume, and / or diverse training data for any suitable machine learning model (e.g., a machine learning model configured to receive two-dimensional and / or three-dimensional image data as input, a machine learning model configured to receive one-dimensional and / or multi-dimensional sound data as input, and / or a machine learning model configured to receive any other suitable data having any suitable dimensionality as input) configured to generate any suitable type of results (e.g., classification, segmentation, determination, inference, prediction, etc.) in any suitable operating context.

[0031] In various cases, embodiments of the subject innovation may electronically receive annotated source images. In various aspects, the annotated source image may be a medical image of a patient (e.g., an X-ray image of a patient, a computed tomography (CT) image of a patient, a magnetic resonance imaging (MRI) image of a patient, a positron emission tomography (PET) image of a patient, a visible spectrum photograph of a patient, etc.). In various aspects, the annotated source image may be generated and / or captured by any suitable imaging device and / or apparatus, and the annotated source image may be received directly from the imaging device and / or apparatus. In various other aspects, the annotated source image may be stored in any suitable database and / or data structure, and the annotated source image may be retrieved from the database and / or data structure. In various cases, as the name implies, the annotated source image may be associated with the annotation. In various aspects, the annotations may be any suitable indication of a class, classification, category, and / or label known to be applied to the annotated source image (e.g., the annotation may indicate that the annotated source image depicts a patient with a brain lesion, the annotation may indicate that the annotated source image depicts a patient with dental caries, the annotation may indicate that the annotated source image depicts a patient with a blocked blood vessel, the annotation may indicate that the annotated source image depicts a patient with a particular skin condition, the annotation may indicate that the annotated source image depicts a patient with lung cancer, etc.). In various aspects, the annotations may be at any suitable granularity and / or level of specificity (e.g., the annotation may indicate only the condition affecting the patient, and / or the annotation may more specifically indicate any other information characterizing the condition affecting the patient, such as the localization / laterality of the condition, the severity of the condition, the time frame of the condition, the prognosis associated with the condition, etc.). In various cases, the annotations may be generated and / or created by any suitable technique, such as manually generated and / or created by a clinician and / or medical professional.

[0032] As described herein, various embodiments of the subject innovation can electronically receive an annotated source image and can electronically generate a plurality of realistic, high-volume and / or diverse training images based on the annotated source image. In various aspects, this can be accomplished by copying the annotated source image and by performing three different types of augmentations on the copy of the annotated source image.

[0033] Specifically, in various cases, embodiments of the subject innovation may generate a set of annotated preliminary training images (e.g., also referred to as preliminary training images) based on annotated source images via an element enhancement component. In various aspects, the annotated preliminary training images may be formed by inserting at least one element / feature of interest and / or at least one background element / feature into the annotated source image. In various aspects, the element / feature of interest (at least with respect to the image) may be any suitable visual object and / or visual characteristic of an element / feature that can be added to the annotated source image (e.g., it can be added to a copy of the annotated source image) and that the machine learning model to be trained should learn, predict, detect and / or classify. For example, if the machine learning model to be trained should learn, predict, detect and / or classify different types of brain lesions, the element / feature of interest may be an independent image of a specific brain lesion that can be inserted into the annotated source image. For another example, if the machine learning model to be trained should learn, predict, detect and / or classify skin growths, the element / feature of interest may be an independent image of a specific skin growth that can be inserted into the annotated source image. In some cases, the element / feature of interest may be referred to as a positive element / feature. In various aspects, a background element / feature (at least with respect to the image) can be any suitable visual object and / or visual characteristic that can be added to an annotated source image (e.g., it can be added to a copy of the annotated source image) and is an element / feature that the machine learning model to be trained should not learn, predict, detect and / or classify. Conversely, in various cases, the background element / feature can adversely affect the classification generated by the machine learning model (e.g., can interfere with and / or hinder the machine learning model). For example, if the machine learning model to be trained should learn, predict, detect and / or classify a specific type of morbidity, the background element / feature can be an independent image of an unrelated comorbidity that can be inserted into the annotated source image. As another example, in some cases, the background element / feature can be an independent image of a specific medical device that can be inserted into the annotated source image. In various cases, when the element / feature is inserted into a copy of the annotated source image, the copy can now be referred to as a preliminary training image. In various aspects, any suitable number of preliminary training images can be generated based on the annotated source image.

[0034] In various cases, different elements / features may be positioned and / or arranged differently in different ways within an annotated source image (e.g., within a copy of an annotated source image), thereby producing different preliminary training images. In some cases, the inserted elements / features may be randomly positioned and / or arranged within any suitable range of biologically feasible positions / positions in the annotated source image. For example, assume that the annotated source image is an X-ray image of a patient's chest and abdomen. Thus, the annotated source image may depict the patient's chest cavity, the patient's intestines / abdominal cavity, etc. In addition, assume that the machine learning model to be trained should learn, predict, detect and / or classify lung cancer (and / or otherwise perform lung segmentation). In various aspects, elements / features may be added to and / or inserted into any biologically feasible position / position in the annotated source image. For example, assume that the element / feature is a cancerous lung growth (e.g., an element / feature of interest). In various aspects, a first copy of the X-ray may be made, and the cancerous lung growth may be inserted into the first copy at any suitable location within the depicted chest cavity and may not be inserted into the depicted abdominal cavity (e.g., a cancerous lung growth may form in the patient's lungs and therefore in the chest cavity; however, a cancerous lung growth may not form in the patient's abdominal cavity). In various cases, the first copy may now be considered a first preliminary training image. As another example, assume that the element / feature is gastric gas (e.g., a background element / feature). In various aspects, a second copy of the X-ray may be made, and the gastric gas may be inserted into the second copy at any suitable location within the depicted abdominal cavity and may not be inserted into the depicted chest cavity (e.g., gastric gas may form in the patient's abdominal cavity; however, gastric gas may not form in the patient's chest cavity). In this way, the insertable element / feature may be positioned in any suitable biologically feasible location / positioning in the annotated source image. In various cases, the second copy may now be considered a second preliminary training image. Thus, different preliminary training images may be formed by inserting different elements / features into the annotated source image (eg, into a copy of the annotated source image).

[0035] In various aspects, the same element / feature may be positioned differently within the annotated source image (e.g., within a copy of the annotated source image), thereby producing different preliminary training images. For example, consider again the example above, where the first preliminary training image includes an inserted cancerous lung growth. Assume that the cancerous lung growth is inserted into the top portion of the right lung in the chest cavity of the depicted patient. In some cases, a third copy of the X-ray may be made, and the cancerous lung growth may be inserted into the bottom portion of the left lung in the chest cavity depicted. In various aspects, the third copy may now be considered a third preliminary training image. Therefore, both the first preliminary training image and the third preliminary training image may be formed by inserting an image of the cancerous lung growth into the annotated source image, but they may be different preliminary training images because the cancerous lung growth may be positioned differently therein. For another example, consider again the example above, where the second preliminary training image includes an inserted stomach gas. Assume that the stomach gas is inserted into the upper left portion of the abdominal cavity of the depicted patient. In some cases, a fourth copy of the X-ray may be made, and the gastric gas may be inserted into the lower right portion of the depicted abdominal cavity. In various aspects, the fourth copy may now be considered a fourth preliminary training image. Thus, both the second preliminary training image and the fourth preliminary training image may be formed by inserting the image of the gastric gas into the annotated source image, but they may be different preliminary training images because the gastric gas may be positioned differently therein. In this way, the same element / feature may be inserted into different places / positions / positions in different copies of the annotated source image, thereby generating different preliminary training images.

[0036] It should be understood that when inserting an insertable element / feature into an annotated source image (e.g., into a copy of an annotated source image), any suitable characteristics of the insertable element / feature may be changed. For example, the same element / feature may be inserted into two different copies of an annotated source image such that the same element / feature has different spatial dimensions (e.g., length, width, height, thickness), different spatial orientations (e.g., inverted orientation, sideways orientation, backward orientation), and / or different intensities in different copies of the annotated source image.

[0037] Note that when background elements / features are inserted, the preliminary training images generated based on the annotated source images may in some cases share the annotations of the annotated source images (e.g., the classification and / or label of the preliminary training images may be the same as the classification and / or label of the annotated source images when the background elements / features are inserted). For example, suppose the annotated source image depicts a chest X-ray of a patient, and suppose the annotations indicate that the patient has pneumonia. In this case, inserting background elements / features (e.g., gastric gas, medical cables / tubes / wires, pacemakers, etc.) does not change the fact that the patient has pneumonia.

[0038] Note that when inserting an element / feature of interest, the preliminary training image generated based on the annotated source image may in some cases have annotations based on the inserted element / feature of interest. For example, suppose the annotated source image depicts a CT scan of a patient's head, and suppose the annotation indicates that the patient has a blocked blood vessel on the left side. In this case, inserting the element / feature of interest may require a corresponding change / update of the annotation. For example, if another blocked blood vessel is inserted to the right side of the depicted cranial cavity, the annotation may be updated to indicate the presence of a left blocked blood vessel and a right blocked blood vessel in the updated image.

[0039] Thus, in various embodiments, annotations for any preliminary training images may be learned / created based on annotations for annotated source images and / or based on elements / features inserted into the annotated source images.

[0040] In various aspects, a catalog of pre-made, pre-drawn, pre-illustrated, and / or pre-generated elements / features may be maintained, and any suitable number of pre-made, pre-drawn, pre-illustrated, and / or pre-generated elements / features from the catalog may be inserted into the annotated source image (e.g., into a copy of the annotated source image) at any suitable location and / or in any suitable orientation to generate the set of preliminary training images. In various aspects, the catalog may be any suitable database and / or data structure (e.g., a relational database, a graphical database, a hybrid database). In various aspects, the elements / features stored in the catalog may be created via any suitable technique (e.g., the elements / features stored in the catalog may be electronic copies of hand-drawn elements / features, may be electronic images of a two-dimensional computer-aided design model, may be the two-dimensional computer-aided design model itself, may be a two-dimensional projection of a three-dimensional computer-aided design model, may be the three-dimensional computer-aided design model itself, etc.).

[0041] In a medical context, such permuted insertion of elements / features may help to more completely simulate and / or approximate real-world biological variability (e.g., a single x-ray scan may not adequately represent the entire space of biological variability experienced by a real-world patient; thus, various biological structures and / or medical device structures may be added to and / or superimposed on a single x-ray scan and / or copies of a single x-ray scan in order to help simulate and / or span the entire space of biological variability experienced by a real-world patient).

[0042] In various cases, embodiments of the subject innovation may generate a set of annotated intermediate training images (e.g., also referred to as intermediate training images) based on the set of preliminary training images via a modality enhancement component. In various aspects, the intermediate training images may be formed by changing at least one modality-based characteristic of the preliminary training images. In various aspects, the modality-based characteristic (at least with respect to the image) may be any suitable image attribute that depends on the device modality (e.g., image capture device) that generates and / or captures the annotated source image. For example, different image capture device modalities may exhibit different gamma / radiation levels, different brightness / contrast levels, different motion / blur levels, different noise levels, different resolutions, different fields of view, different magnification levels, different visual textures, different imaging artifacts (e.g., glare; scratches, dust, and / or any other occluding material on the camera lens), and the like. It should be understood that in various aspects, some modality-based characteristics may vary continuously, while other modality-based characteristics may vary discretely. In various cases, intermediate training images may be formed from preliminary training images by changing the gamma / radiation level, brightness / contrast level, motion / blur level, noise level, resolution, field of view, magnification level, visual texture, and / or imaging artifacts of the preliminary training images. In various cases, any suitable number of intermediate training images may be formed from each preliminary training image by changing one or more modality-based characteristics of the preliminary training images. For example, consider a preliminary training image (e.g., one of many images generated from annotated source images) that exhibits an existing gamma / radiation level. In various aspects, a first copy of the preliminary training image may be made, and the existing gamma / radiation level of the first copy may be changed to a first gamma / radiation level. In various cases, the first copy of the preliminary training image may now be considered a first intermediate training image. In various aspects, a second copy of the preliminary training image may be made, and the existing gamma / radiation level of the second copy may be changed to a second gamma / radiation level. In various cases, the second copy of the preliminary training image may now be considered a second intermediate training image. As another example, assume that the preliminary training image exhibits an existing brightness / contrast level. In various aspects, a third copy of the preliminary training image may be made, and the existing brightness / contrast level of the third copy may be changed to the first brightness / contrast level. In various cases, the third copy of the preliminary training image may now be considered a third intermediate training image. In various aspects, a fourth copy of the preliminary training image may be made, and the existing brightness / contrast level of the fourth copy may be changed to the second brightness / contrast level. In various cases, the fourth copy of the preliminary training image may now be considered a fourth intermediate training image. As another example, assume that the preliminary training image exhibits existing glare. In various aspects, a fifth copy of the preliminary training image may be made, and the existing glare of the fifth copy may be removed, supplemented, and / or changed to the first glare.In various cases, the fifth copy of the preliminary training image may now be considered a fifth intermediate training image. In various aspects, a sixth copy of the preliminary training image may be made, and the existing glare of the sixth copy may be removed, supplemented, and / or changed to a second glare. In various cases, the sixth copy of the preliminary training image may now be considered a sixth intermediate training image. In this way, any suitable number of intermediate training images may be generated by changing at least one modality-based characteristic of each of the preliminary training images in a permutation manner. In various cases, any suitable strategy and / or scheme for changing the modality-based characteristics of the preliminary training images may be implemented.

[0043] In a medical context, such permuted variations of modality-based characteristics may help to more completely simulate and / or approximate real-world device modality variability (e.g., a single x-ray scan may be generated by a single type / model of x-ray machine and, therefore, may not adequately represent the entire space of x-ray machine variability that exists in real-world medical / clinical environments; therefore, various modality-based characteristics of a single x-ray scan and / or copies of a single x-ray scan may be adjusted / changed to help simulate and / or span the entire space of x-ray machine variability that exists in real-world medical / clinical environments).

[0044] In various aspects, embodiments of the subject innovation may generate a set of annotated deployable training images (e.g., also referred to as deployable training images) based on the set of intermediate training images via a geometric enhancement component. In various aspects, deployable training images may be formed by applying at least one geometric transformation to the intermediate training images. In various aspects, the geometric transformation (at least with respect to the image) may be any suitable mathematical transformation and / or operation that may spatially change and / or transform the pixel grid of the image. For example, the geometric transformation may include reflecting the image around any suitable axis, rotating the image around any suitable axis, cropping any suitable portion of the image, translating the image, tilting the image, enlarging the image and / or reducing the image, applying affine and / or elastic transformations to the image, distorting the image away from a straight line projection, and the like. In various cases, deployable training images may be formed from the intermediate training images by flipping, rotating, cropping, translating, tilting, scaling, and / or distorting the intermediate training images. In various cases, any suitable number of deployable training images may be formed from each intermediate training image by applying one or more geometric transformations to the intermediate training images. For example, consider an intermediate training image that exhibits an existing orientation (e.g., one of many intermediate training images generated from an intermediate training image). In various aspects, a first copy of the intermediate training image may be made, and the existing orientation of the first copy may be reflected, rotated, translated, tilted, and / or scaled to a first orientation. In various cases, the first copy of the intermediate training image may now be considered a first deployable training image. In various aspects, a second copy of the intermediate training image may be made, and the existing orientation of the second copy may be reflected, rotated, translated, tilted, and / or scaled to a second orientation. In various cases, the second copy of the intermediate training image may now be considered a second deployable training image. As another example, assume that the intermediate training image exhibits an existing appearance. In various aspects, a third copy of the intermediate training image may be made, and the existing appearance of the third copy may be distorted via a first affine and / or elastic transformation. In various cases, the third copy of the intermediate training image may now be considered a third deployable training image. In various aspects, a fourth copy of the intermediate training image may be made, and the existing appearance of the fourth copy may be distorted via a second affine and / or elastic transformation. In various cases, the fourth copy of the intermediate training image may now be considered a fourth deployable training image. In this way, any suitable number of intermediate training images may be generated by changing at least one modality-based characteristic of each of the preliminary training images. In various cases, any suitable strategy and / or scheme for applying geometric transformations to intermediate training images may be implemented.

[0045] In a medical context, such permutative application of geometric transforms can help to more completely simulate and / or approximate real-world image variability (e.g., a single x-ray scan may have certain geometric properties and, therefore, may not adequately represent the entire space of x-ray properties that exist in a real-world medical / clinical environment; therefore, to help simulate and / or span the entire space of x-ray properties that exist in a real-world medical / clinical environment, various geometric transforms may be applied to a single x-ray scan and / or copies of a single x-ray scan).

[0046] In various cases, embodiments of the subject innovation may train a machine learning model on the set of deployable training images via a training component. Note that, as described herein, a single annotated source image may be used to automatically and quickly generate multiple deployable training images. Specifically, the multiple deployable training images may be formed by making different copies of the annotated source image, inserting different elements / features into different positions / orientations in different copies, changing different modality-based properties of different copies, changing the same modality-based properties of different copies differently, and / or applying different combinations of geometric transformations to different copies. In other words, a single annotated source image may be used to create multiple synthetically generated training images that help account for real-world variability through element / feature insertion, through modality-based modulation, and / or through geometric transformations (e.g., there may be multiple arrangements of different insertable elements / structures, different insertion positions and / or orientations and / or sizes, different modality-based properties, and / or different geometric transformations). Thus, training a machine learning model on multiple deployable training images may improve the performance and / or efficacy of the machine learning model compared to training on only a single annotated source image.

[0047] To help illustrate some of the above discussion, consider the following non-limiting example. Assume that it is desired to train a machine learning model on an initial training data set. Further, assume that the initial training data set includes annotated chest X-ray images (e.g., source images) received from an X-ray machine, and assume that the machine learning model is assumed to learn, predict, detect and / or classify lung cancer in chest X-ray images. In various aspects, a set of preliminary training X-ray images can be formed based on inserting various elements / features into the annotated chest X-ray images. For example, in some cases, a first preliminary training X-ray image may be formed by inserting an image of a pacemaker into the heart position of the annotated chest X-ray image, a second preliminary training X-ray image may be formed by placing an image of a medical tube in the upper left portion of the annotated chest X-ray image, a third preliminary training X-ray image may be formed by inserting an image of the medical tube into the upper right portion of the annotated chest X-ray image (e.g., same element orientation, different position), a fourth preliminary training X-ray image may be formed by inserting an image of a different orientation / size of the medical tube into the upper left portion of the annotated chest X-ray image (e.g., different element orientation / size, same position), a fifth preliminary training X-ray image may be formed by inserting an image of gastric gas into the lower portion of the annotated chest X-ray image, and a sixth preliminary training X-ray image may be formed by not inserting elements / features into the annotated chest X-ray. That is, in various cases, element / feature insertion may be implemented to generate six preliminary training X-ray images based on a single annotated X-ray image.

[0048] In various aspects, a set of intermediate training X-ray images may be formed based on adjusting various modality-based characteristics of each of the six preliminary training X-ray images. For example, in some cases, it is assumed that three possible gamma / radiation levels may be represented in the X-ray images (e.g., high gamma / radiation, medium gamma / radiation, low gamma / radiation), it is assumed that three possible blur levels may be represented in the X-ray images (e.g., high blur, medium blur, low blur), and it is assumed that two possible artifacts may be represented in the X-ray images (e.g., lens glare and no lens glare). In this case, eighteen intermediate training X-ray images may be formed from each of the preliminary training X-ray images (e.g., three gamma / radiation levels multiplied by three blur levels multiplied by two artifact levels), thus forming a total of one hundred and eight intermediate training X-ray images (e.g., eighteen intermediate training X-ray images per preliminary training X-ray image multiplied by six preliminary training X-ray images).

[0049] In various aspects, a set of deployable training x-ray images can be formed based on applying various geometric transformations to each intermediate training x-ray image. For example, in some cases, assume that the available geometric transformations include three potential reflections (e.g., reflections around a horizontal axis, reflections around a vertical axis, and / or no reflection at all), four potential rotations (e.g., 15 degrees clockwise, 45 degrees clockwise, 75 degrees clockwise, and / or no rotation at all), two possible crops (e.g., applying a center crop and not applying a center crop), and two possible distortions (e.g., applying a barrel distortion and not applying a barrel distortion). In this case, forty-eight different deployable training x-ray images can be formed from each intermediate training x-ray image (e.g., three possible reflections times four possible rotations times two possible crops times two possible distortions), thus forming a total of 5,184 deployable training x-ray images (e.g., forty-eight deployable training x-ray images per intermediate training x-ray image times one hundred and eight intermediate training x-ray images). That is, by applying the teachings disclosed herein, a single annotated X-ray image in an initial training dataset can be used to synthetically generate a very large number (e.g., 5,184) of deployable training X-ray images that simulate real-world categories and on which a machine learning model can be trained. If the initial training dataset includes one hundred annotated X-ray images instead of just one, various embodiments of the subject innovation can therefore generate 518,400 deployable training X-ray images (e.g., 5,184 deployable training X-ray images per annotated X-ray image in the initial training dataset multiplied by the 100 annotated X-ray images in the initial training dataset). In various aspects, training a machine learning model on the set of deployable training X-ray images can produce significantly improved efficacy and / or performance compared to training the machine learning model on only the set of initial training data. In fact, by inserting various elements / features, by varying different modal-based properties, and / or by applying different geometric transformations, embodiments of the subject innovation can electronically create a set of training data that can cause a machine learning model to become invariant and / or robust to such various elements / features, different modal-based properties, and / or different geometric transformations.

[0050] It should be understood that the numbers and / or details in the above examples are exemplary, non-limiting, and for purposes of illustration.

[0051] Various embodiments of the subject innovations may be used to solve problems that are highly technical in nature (e.g., to facilitate generation of synthetic training data for improving the generalization capabilities of machine learning models) using hardware and / or software that are not abstract and cannot be performed as a set of mental behaviors of a human. In addition, some of the processes performed may be performed by a dedicated computer (e.g., a trained machine learning model) for performing defined tasks related to generation of synthetic training data for improving the generalization capabilities of a machine learning model (e.g., generating a set of annotated preliminary training images based on annotated source images, wherein the annotated preliminary training images are formed by inserting at least one element of interest or at least one background element into the annotated source images; generating a set of annotated intermediate training images based on the set of annotated preliminary training images, wherein the annotated intermediate training images are formed by changing at least one modality-based characteristic of the annotated preliminary training images; generating a set of annotated deployable training images based on the set of annotated intermediate training images, wherein the annotated deployable training images are formed by changing at least one geometric characteristic of the annotated intermediate training images; and training a machine learning model on the set of annotated deployable training images). Such defined tasks are not conventionally performed manually by humans. Furthermore, neither the human mind nor a person using paper and pen can electronically insert elements / features into an image, nor can they electronically change the modality-based properties of an image, nor can they electronically adjust the geometric properties of an image. Alternatively, various embodiments of the subject innovation are inherently and inextricably related to computer technology and cannot be implemented outside of a computing environment (e.g., embodiments of the subject innovation constitute a computerized tool that synthetically generates many different training images based on a given annotated source image; such a computerized device may exist only in a computing environment).

[0052] In various cases, embodiments of the present invention may integrate the disclosed teachings on synthetic training data generation to improve the generalization ability of machine learning models into practical applications. In fact, in various embodiments, the disclosed teachings may provide a computerized system that receives one or more annotated source images (e.g., real-world medical / clinical images of patients with associated annotations created by real-world medical / clinical professionals) as input, and generates a plurality of training images based on the one or more annotated source images as output, wherein the plurality of training images are formed by copying the one or more annotated source images, by electronically inserting elements / features of interest and / or background elements / features into the copies, by electronically changing modality-based properties of the copies, and / or by electronically applying geometric transformations to the copies. The resulting plurality of training images is a very diverse set of images that approximate and / or simulate real-world variability (e.g., element / feature insertion may help approximate real-world biological variability; modality-based property changes may help approximate real-world device modality variability; and geometric property changes may help further approximate real-world variability). Training a machine learning model on such multiple training images may result in significantly improved performance and / or efficacy compared to training the machine learning model on only one or more annotated source images. Thus, such a computerized system is clearly a useful and practical application of computers.

[0053] In addition, various embodiments of the present invention can provide technical improvements to problems that arise in the field of machine learning model training and solve these problems. As described above, the efficacy and / or performance of machine learning models can be limited by the authenticity, capacity, variety and / or speed of training data (for example, insufficient training data that simulates real-world variability can lead to insufficient machine learning models). The embodiments of the subject innovation solve this technical problem by providing a computerized system that can quickly and synthetically generate real, large-volume and / or multiple different training data (for example, element / feature insertion, modality-based characteristic changes and geometric transformations can all help simulate and / or approximate real-world variability). Training machine learning models on such real, large-volume and / or multiple different training data can result in significantly improved model performance. Precisely because the embodiments of the subject innovation can improve the computational performance of machine learning models, the embodiments of the subject innovation constitute technical improvements.

[0054] In addition, various embodiments of the subject innovation can control real-world devices based on the disclosed teachings. For example, embodiments of the subject innovation can electronically receive annotated source images of the real world (e.g., X-ray scans, CT scans, MRI scans, PET scans, ultrasound scans, visible spectrum photographs). Embodiments of the subject innovation can electronically insert real-world images of elements / features of interest and / or real-world images of background elements / features into annotated source images of the real world. Embodiments of the subject innovation can electronically change the real-world modality-based characteristics of the real-world annotated source images. In addition, embodiments of the subject innovation can electronically change the real-world geometric characteristics of the real-world annotated source images. Such electronic insertions and / or electronic changes can produce multiple real-world training images that more completely and / or more completely simulate the variability of real-world images. Training a real-world machine learning model on such multiple real-world training images can improve the efficacy / performance of the machine learning model, which is a concrete and tangible technical improvement.

[0055] It should be understood that the drawings herein are illustrative and non-limiting.

[0056] Figure 1 A block diagram of an exemplary, non-limiting system 100 is shown that can facilitate synthetic training data generation to achieve improved machine learning model generalization capabilities in accordance with one or more embodiments described herein. As shown, it may be desirable to train a machine learning model 106 on annotated source images 104. However, in various aspects, the annotated source images 104 may not fully and / or adequately represent the entire space of real-world image variability. In various cases, a synthetic training data generation system 102 can address this problem by electronically generating a set of training images based on the annotated source images 104, and a machine learning model 106 can be trained on the set of training images.

[0057] In various aspects, the machine learning model 106 may be any suitable computationally implemented artificial intelligence model and / or algorithm (e.g., support vector machine, neural network, expert system, Bayesian belief network, fuzzy logic, data fusion engine, etc.) designed to receive one or more images as input and produce one or more classifications, labels, and / or predictions as output based on the input one or more images. In various aspects, any suitable machine learning model and / or algorithm may be implemented, such as a model and / or algorithm for performing classification, for performing segmentation, for performing detection, for performing regression, for performing reconstruction, for performing image-to-image (and / or data-to-data) transformation, and / or for performing any other suitable machine learning function.

[0058] In various aspects, the annotated source image 104 can be any suitable image that the machine learning model 106 is designed to analyze. For example, if the machine learning model 106 is designed to classify medical images, the annotated source image 104 can be any suitable medical image (e.g., an X-ray scan of a patient, a CT scan of a patient, an MRI scan of a patient, a PET scan of a patient, an ultrasound scan of a patient, a visible spectrum photograph of a patient). As described above, the annotated source image 104 can have corresponding and / or associated annotations (e.g., classifications and / or labels that are considered to be ground truth values ​​for the annotated source image 104).

[0059] In various embodiments, the synthetic training data generation system 102 may electronically receive / retrieve (e.g., via any suitable wired and / or wireless electronic connection) the annotated source image 104. In various aspects, the synthetic training data generation system 102 may electronically receive / retrieve the annotated source image 104 from any suitable database and / or data structure accessible to the synthetic training data generation system 102. In various aspects, the synthetic training data generation system 102 may electronically receive / retrieve the annotated source image 104 directly from an image capture device that generated, captured, and / or created the annotated source image 104 (e.g., directly from an X-ray scanner, from a CT scanner, from a PET scanner, from an MRI scanner).

[0060] In various embodiments, the synthetic training data generation system 102 may include a processor 108 (e.g., a computer processing unit, a microprocessor) and a computer-readable memory 110 operatively and / or operatively and / or communicatively connected / coupled to the processor 108. The memory 110 may store computer-executable instructions that, when executed by the processor 108, may cause the processor 108 and / or other components of the synthetic training data generation system 102 (e.g., the element enhancement component 112, the modal enhancement component 114, the geometric enhancement component 116, the training component 118) to perform one or more actions. In various embodiments, the memory 110 may store computer-executable components (e.g., the element enhancement component 112, the modal enhancement component 114, the geometric enhancement component 116, the training component 118), and the processor 108 may execute the computer-executable components.

[0061] In various embodiments, the synthetic training data generation system 102 may include an element enhancement component 112. In various aspects, the element enhancement component 112 may generate a set of preliminary training images based on the annotated source image 104. Specifically, the element enhancement component 112 may include an element catalog that electronically stores images of elements (e.g., elements of interest and / or background elements) that can be inserted into the annotated source image 104 (e.g., can be inserted into a copy of the annotated source image 104). In various aspects, the element enhancement component 112 may form / generate preliminary training images by making an electronic copy of the annotated source image 104 and by inserting at least one element from the element catalog into the electronic copy of the annotated source image 104.

[0062] It should be understood that when the present disclosure discusses inserting an element into the annotated source image 104 , this may include inserting the element into one or more copies of the annotated source image 104 .

[0063] In various cases, the elements stored in the element catalog may depend on the nature of the machine learning model 106. For example, the element catalog may include images of elements of interest and may include images of background elements. In various aspects, the elements of interest may be any suitable visual object that the machine learning model 106 should learn, predict, detect, and / or classify. In various cases, the background elements may be any suitable visual object that the machine learning model 106 does not need to learn, predict, detect, and / or classify, but may hinder and / or interfere with the machine learning model 106. For example, if the machine learning model 106 is configured to learn, predict, detect, and / or classify lung growths, the elements of interest may include various malignant lung growths and / or various benign lung growths, and the background elements may include various medical devices (e.g., pacemakers, intravenous tubes, stents, implants, electrocardiogram leads), various comorbidities (e.g., heart defects, blood vessel blockages), gastric gas, etc. In other words, in a medical context, the element of interest may be any suitable anatomical structure and / or biological symptom manifestation that the machine learning model 106 should learn, predict and / or detect, and the background element may be any other suitable anatomical structure and / or biological symptom manifestation that may interfere with and / or hinder the machine learning model 106, and / or may be any suitable medical device that may interfere with and / or hinder the machine learning model 106.

[0064] In various aspects, the element augmentation component 112 may insert any suitable combination of elements from the element catalog into the annotated source images 104 to create preliminary training images (e.g., each preliminary training image may have one inserted element, each preliminary training image may have multiple inserted elements, different preliminary training images may have different numbers of inserted elements, and / or at least one preliminary training image may have no inserted elements).

[0065] In various aspects, the element augmentation component 112 may position the inserted elements in the annotated source image 104 in any suitable biologically feasible location / position. For example, if the element augmentation component 112 inserts an image of a lung lesion, the image of the lung lesion may be placed in the depicted thorax of the annotated source image 104, and may avoid being placed in the depicted abdominal cavity of the annotated source image 104 (e.g., a lung lesion may form in the thorax, but not in the abdominal cavity). In this way, the same element may be input into different locations / positions of the annotated source image 104, thereby generating different preliminary training images.

[0066] In various cases, the element augmentation component 112 can control the orientation of the inserted elements in the annotated source image 104. For example, if the element augmentation component 112 is inserting an image of a lung lesion, the image can be oriented as depicted in the element catalog, can be oriented upside down, can be oriented backwards, can be oriented sideways, can be reflected / rotated in any suitable manner, etc. In this way, the same element can be oriented differently at the same location in the annotated source image 104, thereby producing different preliminary training images.

[0067] In some cases, the element enhancement component 112 can control the size and / or intensity of the inserted elements in the annotated source image 104. For example, if the element enhancement component 112 inserts an image of a lung lesion, the image of the lung lesion can be expanded, shrunk, lengthened, widened, thickened, manipulated in any other suitable manner, etc. In this way, the same element can have different sizes at the same location and / or the same orientation in the annotated source image 104, thereby generating different preliminary training images.

[0068] In various cases, when the preliminary training images are formed by inserting only background elements into the annotated source images 104, the annotations of the preliminary training images may be the same as the annotations of the annotated source images 104 (e.g., if an X-ray image is annotated as depicting one type of lung cancer, then adding gas to the X-ray image may not affect the accuracy / completeness of the annotations). In various aspects, when the preliminary training images are formed by inserting elements of interest into the annotated source images 104, the annotations of the preliminary training images may be initialized to the annotations of the annotated source images 104 and then may be updated based on the inserted elements of interest (e.g., if an X-ray image is annotated as depicting one type of lung cancer, then adding a second type of lung cancer to the X-ray image may affect the accuracy / completeness of the annotations; therefore, the annotations may be updated to indicate that the X-ray image now depicts two types of lung cancer).

[0069] In various embodiments, the synthetic training data generation system 102 may include a modality enhancement component 114. In various aspects, the modality enhancement component 114 may generate a set of intermediate training images based on the set of preliminary training images generated by the element enhancement component 112. Specifically, the modality enhancement component 114 may include a list of various modality-based characteristics. In various aspects, the modality-based characteristics may be any suitable image attribute related to and / or dependent on the device modality that captured and / or generated the annotated source image 104. For example, the modality-based characteristics may include gamma / radiation levels (e.g., because gamma / radiation is used to generate X-rays and / or CT scans), brightness levels, contrast levels, blur levels, noise levels, image textures, image fields of view, image resolution, and / or image artifacts (e.g., glare on lenses, scratches on lenses, dust on lenses). In other words, the modality-based characteristics may represent parameters of the image capture device that may affect the quality / attributes of the captured image. In various cases, the modality enhancement component 114 may form / generate the intermediate training image by making an electronic copy of the preliminary training image and by changing / adjusting at least one modality-based characteristic of the preliminary training image.

[0070] It should be understood that when the present disclosure discusses different modality-based characteristics of preliminary training images, this may include different modality-based characteristics of electronic copies of the preliminary training images.

[0071] In various aspects, the modality enhancement component 114 may change / adjust / manipulate / modify any suitable combination of modality-based characteristics of the preliminary training images to create the intermediate training images (e.g., each intermediate training image may be formed by changing one modality-based characteristic, each intermediate training image may be formed by changing multiple modality-based characteristics, different intermediate training images may be formed by changing different numbers of modality-based characteristics, and / or at least one intermediate training image may not involve changes to any modality-based characteristics).

[0072] In various cases, changing and / or modifying the modality-based characteristics may have no impact on the accuracy and / or completeness of the annotations.Thus, an intermediate training image formed from a preliminary training image may have the same annotations as the preliminary training image.

[0073] In various embodiments, the synthetic training data generation system 102 may include a geometry enhancement component 116. In various aspects, the geometry enhancement component 116 may generate a set of deployable training images based on the set of intermediate training images generated by the modality enhancement component 114. Specifically, the geometry enhancement component 116 may include a list of various geometric transformations that may be applied to the images. In various aspects, the geometric transformation may be any suitable mathematical operation that may transform the spatial properties of the image. For example, the geometric transformation may include reflection of the image around any suitable axis, rotation of the image around any suitable axis, translation and / or tilting of the image to change the two-dimensional projection and / or perspective of the image, cropping any suitable portion of the image, enlarging and / or reducing the image, optically distorting the image (e.g., barrel distortion, pincushion distortion, mustache distortion, and / or any other suitable distortion away from a straight line projection), and the like. In various cases, the geometry enhancement component 116 may form / generate the deployable training images by making an electronic copy of the intermediate training images and by applying at least one geometric transformation to the electronic copy of the intermediate training images.

[0074] It should be understood that when this disclosure discusses applying a geometric transform to an intermediate training image, this may include applying the geometric transform to an electronic copy of the intermediate training image.

[0075] In various aspects, the geometry enhancement component 116 may apply any suitable combination of geometric transformations to the intermediate training images to create the deployable training images (e.g., each deployable training image may be formed by applying one geometric transformation, each deployable training image may be formed by applying multiple geometric transformations, different deployable training images may be formed by applying different numbers of geometric transformations, and / or at least one deployable training image may not undergo a geometric transformation).

[0076] In various cases, applying the geometric transformation may have no effect on the accuracy and / or completeness of the annotations.Thus, a deployable training image formed from an intermediate training image may have the same annotations as the intermediate training image.

[0077] In various embodiments, the synthetic training data generation system 102 can include a training component 118. In various aspects, the training component 118 can actually train (e.g., via back-propagation and / or any other suitable technique) the machine learning model 106 on the set of deployable training images generated by the synthetic training data generation system 102.

[0078] Figure 2 A block diagram of an exemplary, non-limiting system 200 including an element catalog is shown that can facilitate synthetic training data generation to achieve improved machine learning model generalization capabilities according to one or more embodiments described herein. As shown, in some cases, system 200 can include the same components as system 100, and can also include element catalog 202 and preliminary training images 204.

[0079] In various aspects, the element enhancement component 112 may include an element catalog 202. In various cases, the element enhancement component 112 may utilize the element catalog 202 to generate a preliminary training image 204 based on the annotated source image 104. As described above, the element catalog 202 may electronically store and / or maintain images of elements / features that can be inserted into the annotated source image 104. In particular, the element catalog 202 may include elements of interest and / or background elements. In various cases, the elements of interest may be any suitable visual object that the machine learning model 106 should learn, predict, and / or detect (e.g., if the machine learning model 106 is configured to detect blocked blood vessels in a patient's brain, the elements of interest may be various images of blocked blood vessels). In various aspects, the background elements may be any suitable visual object that may obstruct and / or interfere with the machine learning model 106 (e.g., if the machine learning model 106 is configured to detect blocked blood vessels in a patient's brain, the background elements may be various images of brain lesions and / or various images of intracranial and / or brain implants).

[0080] In various cases, the elements stored in the element catalog 202 may be generated via any suitable technique and / or may be stored in any suitable computerized format. In some cases, the elements stored in the element catalog 202 may be scanned images of hand-drawn drawings (e.g., a medical professional may hand-draw a sketch of a brain lesion, a lung nodule, and / or an intravenous tube, and the sketch may be electronically scanned and saved in the element catalog 202). In some cases, the elements stored in the element catalog 202 may be two-dimensional and / or three-dimensional computer-aided design models (e.g., a medical professional may generate a two-dimensional and / or three-dimensional computer-aided design model of a brain lesion, a lung nodule, and / or an intravenous tube on a computer, and the two-dimensional and / or three-dimensional computer-aided design model may be electronically saved and / or stored in the element catalog 202). In various aspects, any other suitable technique may be implemented to generate and / or obtain the elements in the element catalog 202 (e.g., the elements in the element catalog 202 may be cut out from an existing image, etc.).

[0081] In various aspects, as described above, element enhancement component 112 may control and / or manipulate any suitable visual characteristics of elements within element catalog 202. For example, element enhancement component 112 may change / modify any suitable spatial dimension of an element in element catalog 202 (e.g., length, width, height, thickness, color, shading, intensity, etc.). As another example, element enhancement component 112 may change / modify the depicted and / or projected orientation of an element in element catalog 202 (e.g., the element may be depicted facing forward, facing backward, facing upside down, facing sideways, rotated any suitable amount about any suitable axis, reflected about any suitable axis, etc.). In various aspects, modifying / changing dimensions and / or orientations may be more practical and / or more completely facilitated if a computer-aided design model is implemented, as described above.

[0082] In various cases, as described above, the element enhancement component 112 can position the inserted element within the annotated source image 104 at any suitable biologically feasible location (e.g., a gastric gas can be inserted into any portion of the depicted abdominal cavity, but cannot be inserted into any portion of the depicted thoracic cavity). Thus, in various aspects, the element catalog 202 can map and / or correlate different elements to different biologically feasible locations, and the element enhancement component 112 can position the elements during insertion based on the mapping and / or correlation.

[0083] In various aspects, the element enhancement component 112 can insert elements into the annotated source image 104 according to any suitable enhancement strategy and / or scheme.

[0084] In various cases, the element catalog 202 may be considered as a parameterization of a space of possible / potential elements / features that can be inserted into the annotated source image 104. In other words, a space of all possible / potential image elements / features that can be depicted in the annotated source image 104 may be envisioned, and the element catalog 202 may be constructed and / or configured so as to span and / or substantially span that space. In various cases, such a parameterized space may depend on the operating context of the machine learning model 106 (e.g., in a medical context, the space may include possible / potential biological symptom manifestations that may be captured in an image and / or possible / potential medical devices that may be captured in an image).

[0085] In various embodiments, the element enhancement component 112 may update and / or change the element catalog 202 (e.g., may update and / or change the images listed / stored in the element catalog 202 for generating the preliminary training images 204). For example, in some cases, the element catalog 202 may be initialized with a set of existing images of the element of interest and / or a set of existing images of the background elements. However, in various aspects, the element enhancement component 112 may periodically and / or aperiodically query any suitable database and / or data structure accessible to the element enhancement component 112 to check whether new images of the element of interest and / or new images of the background elements are available (e.g., to check whether images that are not yet stored / listed in the element catalog 202 are available for retrieval and / or downloading, so that such new images can be used to generate the preliminary training images 204). If such new images of the element of interest and / or background elements are available in the database and / or data structure, the element enhancement component 112 may retrieve such new images and add them to the element catalog 202, and thus may begin inserting such new images into the annotated source images 104 to generate the preliminary training images 204. As another example, element enhancement component 112 may receive input from an operator including new images of elements of interest and / or background elements that are not already stored / listed in element catalog 202. In various aspects, element enhancement component 112 may thus add the new images to element catalog 202 and may thus begin using the new images to generate preliminary training images 204. In this way, element catalog 202 may be updated, changed, modified, edited, and / or modified as needed to accommodate different operational contexts.

[0086] For example, assume that the element catalog 202 includes an image of a lung nodule, an image of a stomach gas, and an image of a breathing tube. Therefore, the element enhancement component 112 may insert different combinations / arrangements (e.g., with different locations and / or orientations and / or sizes) of the image of the lung nodule, the image of the stomach gas, and the image of the breathing tube into the annotated source image 104 to generate the preliminary training image 204. In various aspects, the element enhancement component 112 may retrieve the image of the pacemaker from any suitable database and / or data structure (and / or may receive the image as input from an operator). Since the image of the pacemaker is not yet stored / listed in the element catalog 202, the element enhancement component 112 may add the image of the pacemaker to the element catalog 202. Therefore, the element enhancement component 112 may begin inserting the image of the pacemaker (e.g., with different locations and / or different orientations and / or different sizes) into the annotated source image 104 to generate the preliminary training image 204. In this way, the element catalog 202 may be updated and / or expanded over time and / or as needed.

[0087] Figure 3 to Figure 4 Block diagrams of exemplary, non-limiting preliminary training images 300 and 400 formed from annotated source images are shown according to one or more embodiments described herein.

[0088] like Figure 3 As shown, preliminary training images 204 may be generated from annotated source images 104. In various cases, there may be N preliminary training images 204 (for any suitable integer N). In other words, element enhancement component 112 may create N electronic copies of annotated source image 104, and may insert any suitable number and / or combination / arrangement of elements from element catalog 202 into each of the N electronic copies of annotated source image 104, thereby generating N preliminary training images 204. As described above, the goal of element insertion may be to increase the variety and / or diversity of features depicted in annotated source image 104. Thus, element enhancement component 112 may insert different numbers of different elements with different orientations and / or different sizes into different locations of different copies of annotated source image 104, thereby generating preliminary training images 204. In other words, a single annotated source image 104 may be converted into N preliminary training images 204.

[0089] As described above, if a particular preliminary training image is formed by inserting only background elements or by inserting no elements at all, the annotations of the particular preliminary training image may be the same as the annotations of the annotated source image 104 (e.g., the background elements may have no effect on the accuracy and / or completeness of the annotations). However, if a particular preliminary training image is formed by inserting any element of interest, the annotations of the particular preliminary training image may be initialized to the annotations of the annotated source image 104 and may be adjusted to reflect the inserted element of interest. In this way, all preliminary training images 204 may have annotations based on the annotated source image 104 and / or annotations based on the inserted elements.

[0090] Figure 4 A real-world example is depicted showing how annotated source images 104 may be used to create preliminary training images 204. As shown, there may be an initial chest X-ray 402. In various cases, the initial chest X-ray 402 may be considered an annotated source image 104. In various aspects, different chest X-rays 404 may be generated by inserting various elements into the initial chest X-ray 402. Although Figure 4 Sixteen different chest X-rays 404 are shown arranged in a 4×4 grid, but this is exemplary and non-limiting. For ease of explanation, assume that the top row is row 1 and the bottom row is row 4, and assume that the leftmost column is column 1 and the rightmost column is column 4. As shown, the (row 1, column 1) image, (row 1, column 4) image, (row 2, column 3) image, and (row 4, column 4) image in different chest X-rays 404 can be formed by inserting various features (e.g., gastric gas, intestinal growth / CIST, colon cancer, digestive dye, etc.) into the depicted abdominal cavity of the initial chest X-ray 402 and / or superimposing them thereon. As shown, the (row 1, column 3) images, (row 2, column 1) images, (row 2, column 2) images, (row 2, column 4) images, (row 3, column 4) images, and (row 4, column 1) images in different chest x-rays 404 may be formed by inserting various features (e.g., intravenous tubes, breathing tubes, ECG wires / leads, pacemakers, etc.) into and / or superimposed on the depicted chest cavity of the initial chest x-ray 402. As shown, the (row 1, column 2) images, (row 3, column 1) images, (row 3, column 2) images, (row 3, column 3) images, and (row 4, column 2) images in different chest x-rays 404 may be formed by inserting various features (e.g., chest growths / nodules / masses, fluid-filled cysts, voids, consolidations, etc.) into and / or superimposed on the depicted chest cavity of the initial chest x-ray 402. Finally, as shown, the (row 4, column 3) images in a different chest X-ray 404 may be created by inserting various features (eg, metal screws, rods, and / or implants) into and / or superimposing them on the initial chest X-ray 402.

[0091] In general, different elements having different sizes may be oriented differently at different locations of initial chest x-ray 402 in order to produce different chest x-rays 404 .

[0092] Figure 5 A block diagram of an exemplary, non-limiting system 500 including modality-based features according to one or more embodiments described herein that can facilitate synthetic training data generation to achieve improved machine learning model generalization capabilities is shown. As shown, in some cases, system 500 can include the same components as system 200, and can also include modality-based features 502 and intermediate training images 504.

[0093] In various aspects, the modality enhancement component 114 may include a list of modality-based characteristics 502 applicable to the preliminary training images 204. In various cases, the modality enhancement component 114 may alter, change, and / or modify any modality-based characteristics 502 of the preliminary training images 204 to generate the intermediate training images 504. As described above, the modality-based characteristics 502 may include any suitable image properties that depend on and / or are associated with the image capture device that generated the annotated source image 104. For example, the modality-based characteristics 502 may include gamma / radiation levels exhibited by and / or depicted in the image, brightness levels exhibited by and / or depicted in the image, contrast levels exhibited by and / or depicted in the image, blur levels exhibited by and / or depicted in the image, noise levels exhibited by and / or depicted in the image, textures exhibited by and / or depicted in the image, field of view exhibited by and / or depicted in the image, resolution exhibited by and / or depicted in the image, device artifacts (e.g., lens scratches, lens dust, lens glare) exhibited by and / or depicted in the image, etc. In various aspects, the modality enhancement component 114 may generate the intermediate training images 504 by varying, changing, and / or modifying any modality-based characteristics 502 of the preliminary training images 204 (e.g., different electronic copies of each preliminary training image 204 may be made, and different modality-based characteristics (e.g., 502) of the different electronic copies may be varied in different ways to generate the intermediate training images 504).

[0094] In various cases, the list of modality-based characteristics 502 can be considered to be a parameterization of a space of possible / potential image attributes that depends on the device modality that generated and / or captured the annotated source image 104. In other words, a space of possible / potential image attributes that can vary between image capture device modalities can be envisioned, and the list of modality-based characteristics 502 can be constructed and / or configured so as to span and / or substantially span that space. In various cases, such a parameterized space can depend on the operating context of the machine learning model 106.

[0095] In various embodiments, the modality enhancement component 114 may update and / or change the list of modality-based characteristics 502 (e.g., may update and / or change the list of modifiable image characteristics / attributes that are related and / or associated with the device modality and used to generate the intermediate training image 504). For example, in some cases, the list of modality-based characteristics 502 may be initialized with a set of existing modifiable image characteristics / attributes related to the device modality. However, in various aspects, the modality enhancement component 114 may periodically and / or aperiodically query any suitable database and / or data structure accessible to the modality enhancement component 114 to check whether new modifiable image characteristics / attributes related to the device modality are available (e.g., to check whether image characteristics / attributes that are dependent on the device modality and are not currently marked / identified as modifiable are known, such that such new image characteristics / attributes can be used to generate the intermediate training image 504). If it is determined that such new modifiable image characteristics / attributes associated with the device modality are available, the modality enhancement component 114 can include such new image characteristics / attributes in the list of modality-based characteristics 502, and can therefore begin to modify such new image characteristics / attributes when generating intermediate training images 504. As another example, the modality enhancement component 114 can receive input from an operator indicating a new image characteristic / attribute that is not yet included in the list of modality-based characteristics 502. In various aspects, the modality enhancement component 114 can therefore add the new image characteristic / attribute to the list of modality-based characteristics 502, and can therefore begin to modify the new characteristic / attribute to generate intermediate training images 504. In this way, the list of modality-based characteristics 502 can be updated, changed, modified, edited, and / or modified as needed to accommodate different operational contexts.

[0096] For example, assume that the list of modality-based characteristics 502 includes image gamma / radiation level, image brightness level, and image contrast level. Therefore, the modality enhancement component 114 may modify and / or change different combinations / permutations of the gamma / radiation level, brightness level, and / or contrast level of the preliminary training image 204 in order to generate the intermediate training image 504. In various aspects, the modality enhancement component 114 may retrieve an indication from any suitable database and / or data structure that the image blur level is now a modifiable image attribute associated with the device modality (and / or may receive the indication as input from an operator). Since the list of modality-based characteristics 502 does not yet include the image blur level, the modality enhancement component 114 may add the image blur level to the list of modality-based characteristics 502. Therefore, the modality enhancement component 114 may begin changing / modifying the image blur level of the preliminary training image 204 in order to generate the intermediate training image 504. In this way, the list of modality-based characteristics 502 may be updated and / or expanded over time and / or as needed.

[0097] Figure 6 A block diagram of an exemplary, non-limiting intermediate training image 600 formed from preliminary training images is shown according to one or more embodiments described herein.

[0098] like Figure 6 As shown, the intermediate training images 504 may be generated from the preliminary training images 204. In various cases, for each preliminary training image 204, there may be M intermediate training images 504, where M is any suitable integer (e.g., intermediate training image 1_1 through intermediate training image 1_M formed from preliminary training image 1; intermediate training image N_1 through intermediate training image N_M formed from preliminary training image N, etc.). In other words, the modality enhancement component 114 may create M electronic copies of each of the N preliminary training images 204, and may alter, change, and / or modify any suitable number of modality-based features (e.g., 502) of each of the M electronic copies of each of the N preliminary training images 204, thereby generating a total of N*M intermediate training images 504. As described above, the goal of the modality-based feature modification may be to increase the variety and / or diversity depicted in the preliminary training images 204. Therefore, the modality enhancement component 114 can differently adjust different combinations / arrangements of different modality-based characteristics of different copies of the preliminary training image 204 to generate the intermediate training images 504. In other words, a single annotated source image 104 can be converted into N*M intermediate training images 504.

[0099] As described above, modifications to any of the modality-based characteristics 502 may have no effect on the accuracy and / or completeness of the annotations.Thus, a particular intermediate training image formed from a particular preliminary training image may have the same annotations as the particular preliminary training image.

[0100] Figure 7 A block diagram of an exemplary, non-limiting system 700 including a geometric transformation according to one or more embodiments described herein that can facilitate synthetic training data generation to achieve improved machine learning model generalization capabilities is shown. As shown, in various cases, system 700 can include the same components as system 500, and can also include a geometric transformation 702 and a deployable training image 704.

[0101] In various aspects, geometry enhancement component 116 may include a list of geometric transformations 702 applicable to intermediate training images 504. In various cases, geometry enhancement component 116 may apply any geometric transformation 702 to the intermediate training images 504 to generate deployable training images 704. As described above, geometric transformation 702 may include any suitable mathematical transformation and / or operation that may spatially change the depicted geometry of an image (e.g., of intermediate training image 504). For example, the list of geometric transformations 702 may include reflecting the image about any suitable axis, rotating the image about any suitable axis by any suitable amount, translating the image in any suitable direction by any suitable amount, tilting the image in any suitable direction by any suitable amount, enlarging and / or reducing any suitable portion of the image by any suitable amount, cropping any suitable portion of the image in any suitable manner, expanding the image in any suitable direction and by any suitable amount, shrinking the image in any suitable direction and by any suitable amount, distorting any suitable portion of the image in any suitable manner and by any suitable amount, harmonizing and / or de-harmonizing the image in any suitable manner and by any suitable amount, applying any suitable affine and / or elastic transformations to the image in any suitable manner, etc. In various aspects, geometry enhancement component 116 may generate deployable training images 704 by applying any of geometric transformations 702 to intermediate training images 504 (e.g., different electronic copies of each intermediate training image 504 may be made, and different geometric transformations (e.g., 702) of the different electronic copies may be applied to generate deployable training images 704).

[0102] In various cases, geometric transformations 702 may be considered as a parameterization of a space of possible / potential mathematical transformations that may be applied to an image. In other words, a space of possible / potential mathematical transformations and / or operations that may be applied to an image may be envisioned, and a list of geometric transformations 702 may be constructed and / or configured so as to span and / or substantially span that space. In various cases, such a parameterized space may depend on the operating context of machine learning model 106.

[0103] In various embodiments, the geometry enhancement component 116 may update and / or change the list of geometric transformations 702 (e.g., the list of mathematical operations / transformations used to generate the deployable training image 704 may be updated and / or changed). For example, in some cases, the list of geometric transformations 702 may be initialized with a set of existing mathematical operations / transformations that can be applied to the image. However, in various aspects, the geometry enhancement component 116 may periodically and / or aperiodically query any suitable database and / or data structure accessible to the geometry enhancement component 116 to check whether new mathematical operations / transformations that can be applied to the image are available (e.g., check whether mathematical operations / transformations that are not currently marked / identified as being applicable to the image are known, so that such new mathematical operations / transformations can be used to generate the deployable training image 704). If it is determined that such new mathematical operations / transformations are available, the geometry enhancement component 116 may include such new mathematical operations / transformations in the list of geometric transformations 702, and thus may begin applying such new mathematical operations / transformations when generating the deployable training image 704. As another example, geometry enhancement component 116 may receive input from an operator indicating a new mathematical operation / transformation that is not already included in the list of geometric transformations 702. In various aspects, geometry enhancement component 116 may therefore add the new mathematical operation / transformation to the list of geometric transformations 702, and may therefore begin applying the new mathematical operation / transformation to generate deployable training images 704. In this way, the list of geometric transformations 702 may be updated, changed, modified, edited, and / or modified as needed to accommodate different operational contexts.

[0104] For example, assume that the list of geometric transformations 702 includes image rotation, image reflection, and image tilt. Therefore, the geometry enhancement component 116 may apply different combinations / permutations of image rotation, image reflection, and / or image tilt to the intermediate training image 504 in order to generate the deployable training image 704. In various aspects, the geometry enhancement component 116 may retrieve an indication from any suitable database and / or data structure that image distortion is now an available geometric transformation (and / or may receive the indication as input from an operator). Since the list of geometric transformations 702 does not include the image distortion, the geometry enhancement component 116 may add the image distortion to the list of geometric transformations 702. Therefore, the geometry enhancement component 116 may begin applying the image distortion to the intermediate training image 504 in order to generate the deployable training image 704. In this way, the list of geometric transformations 702 may be updated and / or expanded over time and / or as needed.

[0105] Figure 8 A block diagram of an exemplary, non-limiting deployable training image 800 formed from intermediate training images according to one or more embodiments described herein is shown.

[0106] like Figure 8As shown, deployable training images 704 may be generated from intermediate training images 504. In various cases, for each intermediate training image 504, there may be P deployable training images 704, where P is any suitable integer (e.g., deployable training image 1_1_1 to deployable training image 1_1_P formed by intermediate training image 1_1; deployable training image N_M_1 to deployable training image N_M_P formed by intermediate training image N_M; etc.). In other words, geometry enhancement component 116 may create P electronic copies of each of the N*M intermediate training images 504, and may apply any suitable number of geometric transformations (e.g., 702) to each of the P electronic copies of each of the N*M intermediate training images 504, thereby generating a total of N*M*P deployable training images 704. As described above, the goal of the geometric transformations may be to increase the variety and / or diversity depicted in the intermediate training images 504. Therefore, the geometry enhancement component 116 may apply different combinations / arrangements of different geometric transformations to different copies of the intermediate training image 504 to generate deployable training images 704. In other words, a single annotated source image 104 may be converted into N*M*P deployable training images 704.

[0107] As described above, application of any of the geometric transformations 702 may have no effect on the accuracy and / or completeness of the annotations.Thus, a particular deployable training image formed from a particular intermediate training image may have the same annotations as the particular intermediate training image.

[0108] As shown, in various aspects, the number of deployable training images 704 may be greater than the number of intermediate training images 504 , and the number of intermediate training images may be greater than the number of preliminary training images 204 .

[0109] Fig. 9 Block diagrams illustrating exemplary, non-limiting variations of modality-based and / or geometric properties according to one or more embodiments described herein.

[0110] In other words, Fig. 9 A real-world example is shown that illustrates how preliminary training images 204 may be used to create intermediate training images 504 and / or deployable training images 704. As shown, an enhanced radiograph 902 may be present. In various cases, modality-based properties of the enhanced radiograph 902 and / or geometric properties of the enhanced radiograph 902 may be manipulated and / or modified as described above to create a further enhanced radiograph 904. Although Fig. 9Sixteen further enhanced X-ray films 904 are shown arranged in a 4×4 grid, but this is exemplary and non-limiting. For ease of explanation, assume that the top row is row 1 and the bottom row is row 4, and assume that the leftmost column is column 1 and the rightmost column is column 4. As shown, the (row 1, column 3) image, (row 2, column 2) image, (row 3, column 2) image, (row 4, column 1) image, and (row 4, column 3) image in the further enhanced X-ray film 904 can be formed by increasing and / or decreasing the brightness / contrast / gamma level of the enhanced X-ray film 902. As shown, the (row 1, column 2) image, (row 2, column 1) image, (row 2, column 3) image, (row 2, column 4) image, (row 3, column 2) image, (row 3, column 3) image, (row 4, column 1) image, (row 4, column 2) image, (row 4, column 3) image, and (row 4, column 4) image in the further enhanced X-ray film 904 can be formed by cropping, scaling, and / or optically distorting the enhanced X-ray film 902. In various cases, any other suitable transformations and / or modifications are possible.

[0111] In various aspects, as described above, the training component 218 can train the machine learning model 106 on the deployable training images 704.

[0112] Fig.10 Exemplary, non-limiting experimental results 1000 are shown according to one or more embodiments described herein.

[0113] In various aspects, the inventors of various embodiments of the subject innovation generated training datasets as described herein (e.g., via element / feature insertion, via modality-based property modification, via geometric transformation), and their experiments showed that lung segmentation machine learning models (e.g., 106) trained on the generated training datasets exhibited significantly improved performance / efficacy compared to training on conventional training datasets. Specifically, as Fig.10As shown, four different experiments were performed, in which it was desired to train a lung segmentation machine learning model on an original dataset of size 138 (e.g., there were 138 training images in the original dataset). A first experiment was performed, in which the machine learning model was trained only on the original dataset. A second experiment was performed, in which a limited degree of element / feature insertion was performed (e.g., represented by a small green annular circle). In this second experiment, the element / feature insertion caused the original dataset to grow from size 138 to size 1600 (e.g., after the element / feature insertion, the dataset included 1600 training images). A third experiment was performed, in which a greater degree of element / feature insertion was performed (e.g., represented by a large green annular circle). In this third experiment, the element / feature insertion caused the original dataset to grow from size 138 to size 3670. Finally, a fourth experiment was performed, in which a greater degree of element / feature insertion was performed, and in which modality-based characteristics were changed and geometric transformations were applied (e.g., represented by a blue annular circle). In this fourth experiment, element / feature insertion, modality-based property variation, and geometric transformation increased the original dataset from size 138 to size 73,640.

[0114] The various performance metrics of the trained machine learning models for each of the four experiments are Fig.10 . The inventors used a test data set of size 1966 to test the performance / efficacy of the trained machine learning model in each experiment. As shown in the figure, in the first experiment (e.g., only training the original data size), the machine learning model achieved a coincidence score (Dicescore) of 0.8063; in the second experiment (e.g., implementing element insertion to a limited extent), the machine learning model achieved a coincidence score of 0.8309; in the third experiment (e.g., implementing a greater degree of element insertion), the machine learning model achieved a coincidence score of 0.8795; and in the fourth experiment (e.g., implementing a greater degree of element insertion and implementing modality-based modifications and geometric transformations), the machine learning model achieved a coincidence score of 0.9135. In other words, when trained on data sets generated by various embodiments of the subject innovation, the machine learning model experiences significant performance / efficacy improvements (e.g., the machine learning model becomes more generalizable / robust, becomes more able to handle difficult and / or unseen test cases, etc.). That is, as Fig.10 As shown, the techniques described herein for increasing dataset variability (e.g., element / feature insertion, modality-based changes, geometric transformations) can improve model performance independently and / or collectively. For at least these reasons, various embodiments of the subject innovation constitute specific and tangible technical improvements (e.g., they can improve the computational performance of machine learning models).

[0115] Figures 11 to 20A block diagram of an exemplary, non-limiting image enhancement according to one or more embodiments described herein is shown. Figures 11 to 20 Real-world examples of various element insertions, various modality-based changes, and / or various geometric transformations that may be implemented in accordance with various embodiments are shown.

[0116] Fig.11 Wires / cables 1102 (e.g., EKG leads, intravenous tubes, breathing tubes, etc.) are depicted, and how the wires / cables 1102 may be inserted into and / or superimposed on various x-ray images 1104 at different locations, in different orientations, in different sizes, in different thicknesses / strengths, in different shapes, etc.

[0117] Fig.12 A mass 1202 (e.g., a fluid-filled cyst, a growth, a CIST, etc.) is depicted, and how the mass 1202 may be inserted into and / or superimposed on various x-ray images 1204 at different locations, in different orientations, in different sizes, in different thicknesses / strengths, in different shapes, etc.

[0118] Fig.13 A mass 1302 (e.g., a tumor, etc.) is depicted, and how the mass 1302 may be inserted into and / or superimposed on various X-ray images 1304 at different locations, in different orientations, in different sizes, in different thicknesses / strengths, in different shapes, etc.

[0119] Fig.14 Gas 1402 is depicted, and how gas 1402 may be inserted into and / or superimposed on various x-ray images 1404 at different locations, in different orientations, in different sizes, in different thicknesses / strengths, in different shapes, etc. As described above, note how gas 1402 may be inserted into the abdominal region of various x-ray images 1404, rather than the chest region of various x-ray images 1404 (e.g., gas cannot form in the chest, but may form in the abdomen).

[0120] Fig.15 Various X-ray images 1500 are shown in which the gamma / radiation level varies continuously. As shown, the gamma / radiation level is highest in the upper left X-ray image and lowest in the lower right X-ray image.

[0121] Fig.16 Various X-ray images with continuously varying Gaussian noise levels are depicted 1600. As shown, the Gaussian noise level is lowest in the upper left X-ray image and highest in the lower right X-ray image.

[0122] Fig.17Various X-ray images with continuously varying levels of Gaussian blur are depicted 1700. As shown, the Gaussian blur level is lowest in the upper left X-ray image and highest in the lower right X-ray image.

[0123] Fig.18 Various X-ray images with continuously varying contrast levels are shown 1800. As shown, the contrast is lowest in the upper left X-ray image and highest in the lower right X-ray image.

[0124] Fig.19 Various X-ray images with continuously varying brightness levels are depicted 1900. As shown, the brightness level is lowest in the upper left X-ray image and highest in the lower right X-ray image.

[0125] Fig. 20 Various x-ray images 2000 showing a continuous variation of exemplary optical distortions are shown. As shown, the optical distortions are most noticeable in the upper left x-ray image and the lower right x-ray image.

[0126] It should be understood that Figures 11 to 20 The enhancements shown in are exemplary and non-limiting. In various embodiments, any other suitable enhancements may be implemented.

[0127] Fig.21 A flowchart of an exemplary, non-limiting computer-implemented method 2100 is shown that can facilitate synthetic training data generation to achieve improved machine learning model generalization capabilities according to one or more embodiments described herein.

[0128] In various embodiments, action 2102 may include generating, by a device operatively coupled to a processor (e.g., 112), a set of annotated preliminary training images (e.g., 204) based on annotated source images (e.g., 104). In various aspects, the annotated preliminary training images may be formed by inserting at least one element of interest or at least one background element (e.g., from 202) into the annotated source images.

[0129] In various cases, action 2104 may include generating, by a device (e.g., 114), a set of annotated intermediate training images based on the set of annotated preliminary training images (e.g., 504). In various cases, the annotated intermediate training images may be formed by changing at least one modality-based characteristic of the annotated preliminary training images (e.g., from 502).

[0130] In various aspects, act 2106 may include generating, by a device (e.g., 116), a set of annotated deployable training images based on the set of annotated intermediate training images (e.g., 704). In various cases, the annotated deployable training images may be formed by changing at least one geometric property of the annotated intermediate training images (e.g., by applying any of 702).

[0131] In various cases, action 2108 may include training, by a device (e.g., 118), a machine learning model (e.g., 106) on the set of annotated deployable training images.

[0132] Fig. 22 A flowchart of an exemplary, non-limiting computer-implemented method 2200 is shown that can facilitate synthetic training data generation to achieve improved machine learning model generalization capabilities according to one or more embodiments described herein.

[0133] As described above, for ease of explanation, the teachings herein regarding how to synthetically increase the variability of training data are described with respect to an imaging context (e.g., a machine learning model 106 configured to analyze one or more images). However, in various aspects, the teachings described herein may be applied to any suitable context utilizing a machine learning model (e.g., a model that analyzes images, a model that analyzes sound recordings, and / or a model that analyzes any other suitable type of data). In such cases, it should be understood that the source data (e.g., 104), preliminary training data (e.g., 204), intermediate training data (e.g., 504), deployable training data (e.g., 704), element catalog 202, modality-based characteristics 502, and / or geometric transformations 702 may depend on the operational context (e.g., if the machine learning model 106 is configured to analyze images, source images / training images may be implemented; if the machine learning model 106 is configured to analyze sound recordings, source / training sound recordings may be implemented; the type and / or format of the insertable elements, modifiable modality-based characteristics, and / or mathematical transformations may depend on the format of the data that the machine learning model 106 is configured to analyze; etc.). Computer-implemented method 2200 demonstrates this generalization capability.

[0134] In various embodiments, action 2202 may include parameterizing, by a device operatively coupled to a processor (e.g., 112), a first space (e.g., 202) of potential data features. In various cases, these features may include features of interest and / or background features that may be inserted into the data segment.

[0135] In various cases, act 2204 may include parameterizing, by the device (e.g., 114), a second space (e.g., 502) of potential modality-based data attributes. In various cases, these attributes may include attributes of the data segment associated with a particular device modality used to capture and / or generate the data segment.

[0136] In various aspects, act 2206 may include parameterizing, by the device (eg, 116), a third space of potential (eg, 702) data transformations. In various cases, these transformations may include mathematical transformations and / or operations that may be applied to the data segments.

[0137] In various embodiments, act 2208 may include receiving, by the device, a source data segment (eg, 104) having an associated annotation.

[0138] In various cases, action 2210 may include generating, by a device (e.g., 112), a set of preliminary training data segments (e.g., 204), wherein the preliminary training data segments may be formed by inserting data features from a first space (e.g., 202) into source data segments. In various cases, this may include sampling first parameters of values / states from the first space, and applying different combinations / permutations of the first parameter sampling to the source data segments to generate the set of preliminary training data segments.

[0139] In various aspects, action 2212 may include generating, by a device (e.g., 114), a set of intermediate training data segments (e.g., 504), wherein the intermediate training data segments may be formed by varying modality-based data attributes from a second space of preliminary training data segments (e.g., 502). In various cases, this may include sampling a second parameter of a value / state from the second space, and applying different combinations / permutations of the second parameter sampling to the set of preliminary training data segments to generate the set of intermediate training data segments.

[0140] In various embodiments, action 2214 may include generating, by a device (e.g., 116), a set of deployable training data segments (e.g., 704), wherein the deployable training data segments may be formed by applying a data transformation (e.g., 702) from a third space to the intermediate training data segments. In various cases, this may include sampling a third parameter of a value / state from the third space, and applying different combinations / permutations of the third parameter sampling to the set of intermediate training data segments to generate the set of deployable training data segments.

[0141] In various cases, a machine learning model can then be trained on the set of deployable training data segments.

[0142] Various embodiments of the subject innovation can achieve their technical benefits by parameterizing the simulation space. Specifically, in various embodiments, it may be desirable to train a machine learning model on a source data segment. In various aspects, a simulation space may be defined, wherein the simulation space may be considered to be a domain of possible input data that may be fed to the machine learning model to be trained (e.g., for a machine learning model designed to analyze chest X-ray images, the simulation space may be a space of all possible chest X-ray images with different anatomical structures / features, different brightness / contrast levels, different distortion levels, different orientations / angles, and / or other different image features that may be fed to the machine learning model; for a machine learning model designed to analyze voice recordings, the simulation space may be a space of all possible voice recordings with different volumes and / or loudness / sound pressure levels, different pitches, different tones, and / or other different sound features that may be fed to the machine learning model). In various cases, a source data segment may be considered to represent only a point within the simulation space (e.g., a specific chest X-ray image in the space of all possible chest X-ray images; a specific voice recording in the space of all possible voice recordings). Various embodiments of the subject innovation may automatically generate multiple deployable training data segments based on source data segments, such that many different points within the simulation space are now represented by the multiple deployable training data segments (e.g., the source data segments may be copied, and the copies may be manipulated and / or modulated so that they have various different arrangements / combinations of features and / or attributes to more fully represent the diversity of the simulation space). Specifically, in various aspects, the simulation space may be parameterized by defining one or more modifiable parameters across the simulation space. As fully explained herein, non-limiting examples of such modifiable parameters may include data elements / features of interest and / or background data elements / features that can be inserted into annotated source data segments, modulated modality-based data characteristics / attributes of source data segments, and / or mathematical transformations that can be applied to source data segments. In various cases, the multiple deployable training data segments may be generated by starting from the source data segments and by applying any suitable combination and / or arrangement of values ​​and / or states to the modifiable parameters. In various cases, the result may be that the multiple deployable training data segments sample the simulation space extensively and / or broadly. In other words, the plurality of deployable training data segments may represent a very large sampling and / or scale of the simulation space (e.g., the plurality of deployable training images may represent, capture, and / or approximate a diversity of values / states in the simulation space). Training the machine learning model on the plurality of deployable training data segments may result in improved performance / efficiency compared to training the machine learning model on the source data segments alone.

[0143] For example, consider the following exemplary parameterized hierarchical structure. First, a simulation space can be defined (e.g., it can be a domain of possible input data segments with different data signatures, which can be received by the machine learning model under consideration). Next, a wide enhanced subspace can be defined in the simulation space, wherein the enhanced subspace contains one or more related enhanced parameters. For example, as explained herein, the first enhanced subspace can be a space in which data elements / features can be inserted, and one or more related enhanced parameters in the space in which data elements / features can be inserted can include the type of the insertable data element / feature (e.g., when the data involved is an image, such type of insertable data element / feature can include an image of a breathing tube, an image of a pacemaker, an image of an implant, an image of a fluid sac, an image of a lung growth, an image of gastric gas), the positioning of the insertable data element / feature (e.g., different insertable images can be inserted into different image positions), the orientation of the insertable data element / feature (e.g., different insertable images can be inserted upside down, backward, sideways), the size / intensity of the insertable data element / feature (e.g., different insertable images can be inserted with different sizes / shapes / thicknesses), and the like. For another example, the second enhancement subspace can be a modifiable modality-based characteristic space, wherein one or more related extensible parameters within the modifiable modality-based characteristic space include any suitable data segment attributes that depend on and / or can be otherwise related to the device modality that generates and / or captures the source data segment under consideration (e.g., image gamma / radiation level, image brightness level, image contrast level, image blur level, image noise level, image texture, device artifacts in the image). For another example, the third enhancement subspace can be a mathematical transformation space, wherein one or more related enhanceable parameters within the mathematical transformation space include any suitable operation that can be applied to the data segment under consideration (e.g., image rotation, image reflection, image translation, image tilt, image scaling, image distortion). In various aspects, each of the one or more related enhanceable parameters within each enhancement subspace can vary within a corresponding continuous parameter range of value and / or state. For example, the gamma / radiation level of the image can vary continuously from a minimum value to a maximum value. Similarly, the contrast level of the image can vary continuously from a minimum value to a maximum value. However, in some cases, the enhanced parameter may have a corresponding discrete range of values ​​and / or states (e.g., a modal artifact parameter may include a state corresponding to no depicted artifacts, a state corresponding to depicted lens glare, a state corresponding to depicted lens scratches, a state corresponding to both depicted lens glare and depicted lens scratches, etc.).In various aspects, parameter sampling may be performed for values / states from a continuous (and / or discrete) parameter range for each enhanceable parameter (e.g., for a given data segment, an enhanceable parameter may have any one of a set of possible values / states, and the parameter sampling for the enhanceable parameter may be any suitable subset of the set of possible values / states). For example, the gamma / radiation level of an image may vary continuously from a minimum value (e.g., 1 unit) to a maximum value (e.g., 1000 units), and the parameter sampling may include a minimum value, a maximum value, and any suitable regular steps and / or increments between the minimum and maximum values ​​(e.g., parameter sampling of gamma level values ​​may start from 1 up to 1000 in steps / increments of 0.1). In various aspects, a source data segment may be converted to multiple deployable training data segments by enhancing the source data segment according to any suitable combination and / or arrangement of such sampled parameter ranges of values / states (e.g., copies of the source data segment may be made, and different copies may be modified to have / exhibit different arrangements / combinations of values / states from the sampled parameter ranges). Thus, the result can be that the multiple deployable training data segments more fully span and / or represent the feature / attribute diversity of the simulation space than individual source data segments, and therefore, training a machine learning model on the multiple deployable training data segments can produce better model performance compared to conventional training techniques.

[0144] Fig.23 A block diagram of an exemplary, non-limiting enhanced spatial hierarchy 2300 is shown that can facilitate synthetic training data generation to achieve improved machine learning model generalization capabilities in accordance with one or more embodiments described herein. Fig.23 It may be helpful to illustrate some of the aspects discussed above.

[0145] In various embodiments, it may be desirable to train a machine learning model (e.g., 106). Thus, in various cases, a simulation space 2302 may be defined. In some cases, simulation space 2302 may be the domain of all possible input data segments that may be received and / or analyzed by the machine learning model to be trained (e.g., if the machine learning model is configured to analyze brain CT scans, simulation space 2302 may be the domain of all possible brain CT scans having different brain shapes, different anatomical features / attributes, different disease states, different pixel values, etc.).

[0146] In various aspects, the simulation space 2302 may be decomposed into a set of enhancement subspaces 2304. In various aspects, as shown, X enhancement subspaces may be defined for the simulation space 2302, where X is any suitable integer (e.g., enhancement subspace 1 to enhancement subspace X). In various cases, each enhancement subspace may be viewed as a space of related and enhanceable parameters that collectively constitute the simulation space 2302. As fully explained above, non-limiting examples of such enhancement subspaces may include element / feature subspaces (e.g., a set of enhanceable parameters associated with data elements / features that can be inserted into a data segment), modality-based subspaces (e.g., a set of enhanceable parameters associated with settings of a device modality that captured / generated a data segment and can be modulated with respect to the data segment), and / or data transformation subspaces (e.g., a set of enhanceable parameters associated with mathematical operations that can be applied to a data segment).

[0147] In various aspects, each enhancement subspace may include a set of enhanceable parameters. As shown, enhancement subspace 1 may include a set of enhanceable parameters 2306 (e.g., enhanceable parameters 1_1 to enhanceable parameters 1_Y, where Y is any suitable integer). Similarly, enhancement subspace X may include a set of enhanceable parameters 2308 (e.g., enhanceable parameters X_1 to enhanceable parameters X_Y, where Y is any suitable integer). Although Fig.23 The group of enhanceable parameters 2306 and the group of enhanceable parameters 2308 are shown as having the same number of parameters (e.g., Y), but this is exemplary and non-limiting. In various aspects, each group of enhanceable parameters may have any suitable number of enhanceable parameters (e.g., some groups have the same number of parameters, some groups have different numbers of parameters, etc.). As fully explained above, non-limiting examples of such enhanceable parameters may include the following: the element / feature enhancement subspace may include element / feature type, element / feature location, element / feature size, element / feature orientation, element / feature intensity, etc. as enhanceable parameters; the modality-based subspace may include brightness level, contrast level, noise / blur level, resolution level, device artifacts, etc. as enhanceable parameters; the data transformation subspace may include reflection operations, rotation operations, translation / tilt / zoom operations, distortion operations, etc. as enhanceable parameters.

[0148] In various aspects, as shown, each enhanceable parameter may have its own parametric range of possible values / states. For example, enhanceable parameter 1_1 may have an associated parametric range of possible values / states of enhanceable parameter 1_1, enhanceable parameter 1_Y may have a parametric range of possible values / states of enhanceable parameter 1_Y, enhanceable parameter X_1 may have a parametric range of possible values / states of enhanceable parameter X_1, enhanceable parameter X_Y may have a parametric range of possible values / states of enhanceable parameter X_Y, etc. As a non-limiting example, the enhanceable parameters in the element / feature subspace may be types, and the corresponding parametric ranges of possible values / states may be all possible types of elements / features that may be inserted into the data segment (e.g., images of pacemakers, images of implants, images of breathing tubes, images of ECG leads, images of lung growths, images of fluid sacs, images of gastric gas, etc.). As another non-limiting example, the enhanceable parameter in the element / feature subspace can be positioning, and the corresponding parameter range of possible values / states can be all possible positions in the data segment where the element / feature can be inserted (e.g., the upper left corner of the image, the lower right corner of the image, the middle of the image, etc.). For another example, the enhanceable parameter in the element / feature subspace can be orientation, and the corresponding parameter range of possible values / states can be all possible orientations that the element / feature can have when inserted into the data segment (e.g., right side up, inverted, backward, sideways, tilted, etc.). For another example, the enhanceable parameter in the modal-based subspace can be brightness, and the corresponding parameter range of possible values / states can be all possible brightness levels that the data segment can have (e.g., a continuous range of values ​​from minimum brightness to maximum brightness). For another example, the enhanceable parameter in the modal-based subspace can be contrast, and the corresponding parameter range of possible values / states can be all possible contrast levels that the data segment can have (e.g., a continuous range of values ​​from minimum contrast to maximum contrast). For another example, the enhanceable parameter in the modality-based subspace may be a device artifact, and the corresponding parameter range of possible values / states may be all possible device artifacts that the data segment may have (e.g., lens glare of different sizes / positions, lens scratches of different sizes / positions, other lens obstructions such as dust / dirt of different sizes / positions, combinations of artifacts, no artifacts, etc.). For another example, the enhanceable parameter in the data transformation subspace may be reflection, and the corresponding parameter range of possible values / states may be all possible reflections applicable to the data segment (e.g., horizontal reflection, vertical reflection, reflection around any other suitable axis, etc.). For another example, the extendable parameter in the data transformation subspace may be rotation, and the corresponding parameter range of possible values / states may be all possible rotations applicable to the data segment (e.g., a continuous range of magnitudes from a minimum angle rotation to a maximum angle rotation).As another example, the scalable parameter in the data transformation subspace can be distortion, and the corresponding parameter range of possible values / states can be all possible distortions that can be applied to the data segment (e.g., barrel distortion of different magnitudes, beard distortion of different magnitudes, pincushion distortion of different magnitudes, combinations of distortions, no distortion, etc.).

[0149] In various aspects, the parameter range of possible values / states of the enhanceable parameter 1_1 and the parameter range of possible values / states of the enhanceable parameter 1_Y can be considered as a set of parameter ranges of possible values / states 2310, which corresponds to the set of enhanceable parameters 2306. In various cases, the parameter range of the set of possible values / states 2310 can be considered to span all possible values / states of the enhancement subspace 1. Similarly, the parameter range of possible values / states of the enhanceable parameter X_1 and the parameter range of possible values / states of the enhanceable parameter X_Y can be considered as a set of parameter ranges of possible values / states 2312, which corresponds to the set of enhanceable parameters 2308. In various cases, the parameter range of the set of possible values / states 2312 can be considered to span all possible values / states of the enhancement subspace X. Therefore, in some cases, the parameter range sets 2310 and 2312 of possible values / states can be collectively considered to span the simulation space 2302.

[0150] In various embodiments, samples of the parameter range of each possible value / state can be obtained. For example, the sampling range of the value / state of the enhanced parameter 1_1 can be obtained from the parameter range of the possible value / state of the enhanced parameter 1_1, the sampling range of the value / state of the enhanced parameter 1_Y can be obtained from the parameter range of the possible value / state of the enhanced parameter 1_Y, the sampling range of the value / state of the enhanced parameter X_1 can be obtained from the parameter range of the possible value / state of the enhanced parameter X_1, the sampling range of the value / state of the enhanced parameter X_Y can be obtained from the parameter range of the possible value / state of the enhanced parameter X_Y, and so on. In various aspects, for a given parameter range of possible value / state, the sampling range of the value / state can be any suitable subset of the given parameter range of possible value / state. In various cases, the sampling range of such value / state can be used to generate a deployable training data segment as described herein. In other words, the attributes / characteristics of the training data segment can be manipulated to adopt any suitable combination / arrangement of the value / state represented in the sampling range of the value / state. In various aspects, the sampling range of values / states of augmentable parameter 1_1 and the sampling range of values / states of augmentable parameter 1_Y may be considered as a set of sampling ranges of values / states 2314, which corresponds to the parametric range of the set of possible values / states 2310. Similarly, the sampling range of values / states of augmentable parameter X_1 and the sampling range of values / states of augmentable parameter X_Y may be considered as a set of sampling ranges of values / states 2316, which corresponds to the parametric range of the set of possible values / states 2312. In some cases, the sets 2314 and 2316 of sampling ranges of values / states may be collectively considered to represent and / or approximate the total set and / or set of values / states of simulation space 2302 (e.g., representing and / or approximating the diversity and / or variability of data features, data attributes, and / or data characteristics within simulation space 2302; in some cases, this may be a coarse approximation and / or a fine approximation of the diversity and / or variability of simulation space 2302, depending on the cardinality, resolution, and / or step size of the sampling ranges).

[0151] As described above, when deployable training data segments are synthetically generated based on sampled range sets 2314 and 2316 of values / states, the deployable training data segments may more completely approximate and / or represent the variability and / or diversity of simulation space 2302. Thus, training a machine learning model of interest on such deployable training data segments may improve model efficacy / performance compared to traditional training techniques.

[0152] In various aspects, embodiments of the subject innovation may be considered to be robust and / or methodical techniques for decomposing a simulation space (e.g., 2302) into enhancement subspaces (e.g., 2304), for decomposing the enhancement subspaces into enhanceable parameters (e.g., 2306, 2308), for defining parametric ranges for possible values / states of these enhanceable parameters (e.g., 2310, 2312), for sampling the parametric ranges for these possible values / states (e.g., 2314, 2316), and for applying these sampled parametric ranges to training data segments such that these training data segments adequately span, represent and / or capture the variability and / or diversity of the entire simulation space.

[0153] As described above, in various aspects, embodiments of the subject innovation may update, change, and / or edit the parameterization of the simulation space 2302 (e.g., by defining and / or creating new and / or different enhancement subspaces within the set of enhancement subspaces 2304, by defining and / or creating new and / or different enhanceable parameters for each enhancement subspace, by changing the parametric range of possible values / states for each enhanceable parameter, and / or by obtaining different samples of the parametric range of possible values / states for each enhanceable parameter).

[0154] Although various embodiments of the subject innovation are described herein as applying image / data enhancement in a particular order (e.g., first element insertion, then modality-based modulation, and finally geometric transformation), this is exemplary, non-limiting, and for ease of explanation. In various aspects, such image / data enhancement may be performed in any suitable order.

[0155] The ability of a machine learning model to generalize can be an important aspect of any artificial intelligence project. However, the ability to generalize can depend on the availability and / or variety of annotated training data. Various embodiments of the subject innovation provide systems and / or techniques that can synthetically generate different training data based on given annotated training data. In various aspects, deterministic data augmentation can be applied as described herein to synthetically generate such different training data. Specifically, element / feature insertion, modality-based modulation, and geometric transformations can be performed in any suitable order to synthetically generate large-volume and diverse training data. As explained herein, training a machine learning model on such synthetically generated training data can result in significant performance / efficacy improvements. This performance / efficacy improvement can be achieved because the disclosed data augmentation can cause the synthetically generated training data to simulate and / or approximate the real-world variability that the machine learning model may encounter during operation.

[0156] In various embodiments, images from source data sets may be selected. In various aspects, any suitable arrangement and / or combination of element / feature insertion, modality-based changes and / or geometric transformations may be performed on the selected images to generate deployable training images. In various aspects, any suitable enhancement strategy / scheme for controlling how to enhance each image may be implemented. In various aspects, each parameter of the enhancement strategy may have its own range of values ​​to be applied (e.g., rotation between 0 and 360 degrees, γ between 50 microwatts and 250 microwatts, etc.). In some cases, various enhancements may have associated execution probabilities (e.g., meaning that enhancements may be performed for less than all images). In various aspects, any suitable enhancement strategy / scheme may be implemented to improve / simulate real-world data variability. In some cases, different enhancement strategies (e.g., different strategies for one-dimensional, two-dimensional, three-dimensional, etc.) may be formulated based on data dimensionality.

[0157] As indicated above, various embodiments of the subject innovation are described with respect to annotated source images 104. Specifically, various embodiments of the subject innovation can quickly and automatically generate a set of deployable training images 704 based on annotated source images 104, wherein the set of deployable training images 704 can be used to facilitate supervised training of the machine learning model 106. However, in various other embodiments, the deployable training images 704 can be generated based on unannotated source images (not shown in the figure). In this case, the deployable training images 704 may not contain annotations / labels and can therefore be used to facilitate unsupervised training and / or reinforcement learning of the machine learning model 106. In other words, one of ordinary skill in the art will appreciate that the teachings herein can be applied to annotated source images as well as unannotated source images.

[0158] To provide additional context for the various embodiments described herein, Fig.24 The following discussion is intended to provide a brief, general description of a suitable computing environment 2400 in which various embodiments of the embodiments described herein may be implemented. Although the embodiments have been described above in the general context of computer-executable instructions that may be executed on one or more computers, those skilled in the art will recognize that the embodiments may also be implemented in conjunction with other program modules and / or as a combination of hardware and software.

[0159] Generally, program modules include routines, programs, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In addition, those skilled in the art will appreciate that the methods of the present invention can be practiced with other computer system configurations, including single-processor or multi-processor computer systems, minicomputers, mainframe computers, Internet of Things (IoT) devices, distributed computing systems, as well as personal computers, handheld computing devices, microprocessor-based or programmable consumer electronics, etc., each of which can be operably coupled to one or more associated devices.

[0160] The illustrated embodiments of the embodiments herein may also be practiced in a distributed computing environment where certain tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices.

[0161] Computing devices typically include various media, which may include computer-readable storage media, machine-readable storage media, and / or communication media, where the two terms are used differently in this article, as described below. Computer-readable storage media or machine-readable storage media can be any available storage media that can be accessed by a computer, and include volatile and non-volatile media, removable and non-removable media. By way of example and not limitation, computer-readable storage media or machine-readable storage media can be implemented in conjunction with any method or technology for storing information such as computer-readable or machine-readable instructions, program modules, structured data, or unstructured data.

[0162] Computer-readable storage media may include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD), Blu-ray disk (BD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, solid-state drives or other solid-state storage devices, or other tangible and / or non-transitory media that can be used to store the desired information. In this regard, the terms "tangible" or "non-transitory" as applied to storage devices, memories, or computer-readable media herein should be understood to exclude only propagating transient signals themselves as a modifier, and do not disclaim all standard storage devices, memories, or computer-readable media that are not only propagating transient signals themselves.

[0163] Computer-readable storage media may be accessed by one or more local or remote computing devices, eg, through access requests, queries, or other data retrieval protocols, to perform various operations on the information stored by the media.

[0164] Communication media typically embodies computer readable instructions, data structures, program modules, or other structured or unstructured data in a data signal, which may be, for example, a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery or transmission media. The term "modulated data signal" or "signal" refers to a signal that has one or more of its characteristics set or changed to encode information in one or more signals. By way of example, and not limitation, communication media include wired media, such as a wired network or direct-wired connection, and wireless media, such as acoustic, RF, infrared and other wireless media.

[0165] Reference again Fig.24 , an exemplary environment 2400 for implementing various embodiments of aspects described herein includes a computer 2402, which includes a processing unit 2404, a system memory 2406, and a system bus 2408. The system bus 2408 couples system components including, but not limited to, the system memory 2406 to the processing unit 2404. The processing unit 2404 can be any of a variety of commercially available processors. Dual microprocessors and other multi-processor architectures can also be used as the processing unit 2404.

[0166] The system bus 2408 may be any of several types of bus structures that may further interconnect to a memory bus (with or without a memory controller), a peripheral bus, and a local bus using any of a variety of commercially available bus architectures. The system memory 2406 includes a ROM 2410 and a RAM 2412. A basic input / output system (BIOS) may be stored in a nonvolatile memory such as a ROM, an erasable programmable read-only memory (EPROM), an EEPROM, wherein the BIOS contains basic routines that help transfer information between elements within the computer 2402, such as during startup. The RAM 2412 may also include a high-speed RAM, such as a static RAM for caching data.

[0167] The computer 2402 also includes an internal hard disk drive (HDD) 2414 (e.g., EIDE, SATA), one or more external storage devices 2416 (e.g., a magnetic floppy disk drive (FDD) 2416, a memory stick or flash drive reader, a memory card reader, etc.), and a drive 2420, such as a solid-state drive, an optical drive, which can read or write from a disk 2422 (such as a CD-ROM disk, a DVD, a BD, etc.). Alternatively, in the case of a solid-state drive, the disk 2422 will not be included unless separated. Although the internal HDD 2414 is shown as being located within the computer 2402, the internal HDD 2414 can also be configured for external use in a suitable infrastructure (not shown). In addition, although not shown in the environment 2400, a solid-state drive (SSD) can be used in addition to or instead of the HDD 2414. The HDD 2414, external storage device 2416, and drive 2420 may be connected to the system bus 2408 via a HDD interface 2424, an external storage interface 2426, and a drive interface 2428, respectively. The interface 2424 for an external drive implementation may include at least one or both of a Universal Serial Bus (USB) and the Institute of Electrical and Electronics Engineers (IEEE) 1394 interface technologies. Other external drive connection technologies are within the contemplation of the embodiments described herein.

[0168] The drives and their associated computer-readable storage media provide non-volatile storage of data, data structures, computer-executable instructions, etc. For the computer 2402, the drives and storage media accommodate the storage of any data in a suitable digital format. Although the above description of computer-readable storage media refers to corresponding types of storage devices, it should be understood by those skilled in the art that other types of computer-readable storage media (whether currently existing or developed in the future) can also be used in the exemplary operating environment, and further, any such storage media can contain computer-executable instructions for performing the methods described herein.

[0169] A number of program modules may be stored in the drives and RAM 2412, including an operating system 2430, one or more application programs 2432, other program modules 2434, and program data 2436. All or portions of the operating system, application programs, modules, and / or data may also be cached in RAM 2412. The systems and methods described herein may be implemented using various commercially available operating systems or combinations of operating systems.

[0170] Computer 2402 may optionally include emulation technology. For example, a hypervisor (not shown) or other intermediary may emulate the hardware environment for operating system 2430, and the emulated hardware may optionally be different from the hardware of the operating system 2430. Fig.24The hardware shown. In such embodiments, the operating system 2430 may include one of a plurality of virtual machines (VMs) hosted at the computer 2402. In addition, the operating system 2430 may provide a runtime environment, such as a Java runtime environment or a .NET framework, for the application 2432. The runtime environment is a conforming execution environment that allows the application 2432 to run on any operating system that includes a runtime environment. Similarly, the operating system 2430 may support containers, and the application 2432 may take the form of a container, which is a lightweight, independent, executable software package that includes, for example, code for the application, a runtime, system tools, system libraries, and settings.

[0171] In addition, the computer 2402 can be enabled with a security module such as a trusted processing module (TPM). For example, in the case of a TPM, the boot component is hashed in the next boot component, and the result is waited for to match the security value before loading the next boot component. This process can occur at any layer in the code execution stack of the computer 2402, such as applied to the application execution level or the operating system (OS) kernel level, thereby achieving security at any code execution level.

[0172] A user may enter commands and information into the computer 2402 through one or more wired / wireless input devices (e.g., a keyboard 2438, a touch screen 2440, and a pointing device such as a mouse 2442). Other input devices (not shown) may include a microphone, an infrared (IR) remote control, a radio frequency (RF) remote control or other remote control, a joystick, a virtual reality controller and / or a virtual reality headset, a game pad, a stylus, an image input device (e.g., a camera), a gesture sensor input device, a visual movement sensor input device, an emotion or facial detection device, a biometric input device (e.g., a fingerprint or iris scanner), etc. These and other input devices are typically connected to the processing unit 2404 through an input device interface 2444, which may be coupled to the system bus 2408, but these and other input devices may be connected through other interfaces, such as a parallel port, an IEEE 1394 serial port, a game port, a USB port, an IR port, a Interfaces, etc.

[0173] A monitor 2446 or other type of display device may also be connected to the system bus 2408 via an interface, such as a video adapter 2448. In addition to the monitor 2446, computers typically include other peripheral output devices (not shown), such as speakers, printers, and the like.

[0174] The computer 2402 can operate in a networked environment using logical connections to one or more remote computers, such as remote computer 2450, via wired and / or wireless communications. The remote computer 2450 can be a workstation, server computer, router, personal computer, portable computer, microprocessor-based entertainment device, peer device, or other public network node, and typically includes many or all of the elements described with respect to the computer 2402, but for the sake of simplicity, only the memory / storage device 2452 is shown. The depicted logical connections include wired / wireless connections to a local area network (LAN) 2454 and / or a larger network, such as a wide area network (WAN) 2456. Such LAN and WAN networking environments are common in offices and companies, and facilitate enterprise-wide computer networks, such as intranets, all of which can be connected to a global communication network, such as the Internet.

[0175] When used in a LAN networking environment, the computer 2402 can be connected to the local network 2454 through a wired and / or wireless communication network interface or adapter 2458. The adapter 2458 can facilitate wired or wireless communication with the LAN 2454, which can also include a wireless access point (AP) disposed thereon for communicating with the adapter 2458 in a wireless mode.

[0176] When used in a WAN networking environment, the computer 2402 may include a modem 2460 or may be connected to a communications server on the WAN 2456 via other means for establishing communications over the WAN 2456, such as through the Internet. The modem 2460, which may be internal or external and a wired or wireless device, may be connected to the system bus 2408 via the input device interface 2444. In a networked environment, program modules shown relative to the computer 2402, or portions thereof, may be stored in the remote memory / storage device 2452. It should be appreciated that the network connections shown are examples and other means of establishing a communications link between the computers may be used.

[0177] When used in a LAN or WAN networking environment, in addition to or as an alternative to the external storage device 2416 described above, the computer 2402 may access a cloud storage system or other network-based storage system, such as, but not limited to, a network virtual machine that provides one or more aspects of information storage or processing. Generally speaking, the connection between the computer 2402 and the cloud storage system may be established through the LAN 2454 or WAN 2456, for example, through an adapter 2458 or a modem 2460, respectively. When the computer 2402 is connected to the associated cloud storage system, the external storage interface 2426 may manage the storage provided by the cloud storage system with the help of the adapter 2458 and / or the modem 2460, as with other types of external storage devices. For example, the external storage interface 2426 may be configured to provide access to cloud storage sources as if those sources were physically connected to the computer 2402.

[0178] The computer 2402 is operable to communicate with any wireless device or entity that is operatively arranged to communicate wirelessly, such as a printer, scanner, desktop and / or portable computer, portable data assistant, communication satellite, any device or location associated with a wirelessly detectable tag (e.g., a self-service machine, a newsstand, a store shelf, etc.). This may include Wireless Fidelity (Wi-Fi) and Wireless technology. Therefore, the communication can be a predefined structure like a conventional network, or just an ad hoc communication between at least two devices.

[0179] Fig.25 2500 is a schematic block diagram of a sample computing environment 2500 with which the disclosed subject matter can interact. The sample computing environment 2500 includes one or more clients 2510. The client 2510 can be hardware and / or software (e.g., a thread, a process, a computing device). The sample computing environment 2500 also includes one or more servers 2530. The server 2530 can also be hardware and / or software (e.g., a thread, a process, a computing device). For example, the server 2530 can accommodate threads to perform conversions by adopting one or more embodiments described herein. A possible communication between the client 2510 and the server 2530 can be in the form of data packets suitable for transmission between two or more computer processes. The sample computing environment 2500 includes a communication framework 2550 that can be used to facilitate communication between the client 2510 and the server 2530. The client 2510 is operably connected to one or more client data repositories 2520, which can be used to store information local to the client 2510. Similarly, server 2530 is operably connected to one or more server data repositories 2540 , which may be used to store information local to server 2530 .

[0180] The present invention can be a system, method, device and / or computer program product at any possible technical detail level of integration. A computer program product can include a computer-readable storage medium (or multiple media) having a computer-readable program instruction thereon for causing a processor to perform aspects of the present invention. A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. An incomplete list of more specific examples of computer-readable storage media can also include the following: a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device (such as a punch card or a raised structure in a groove on which instructions are recorded), and any suitable combination of the above items. As used herein, computer-readable storage media should not be understood as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.

[0181] Computer-readable program instructions described herein can be downloaded to corresponding computing / processing equipment from computer-readable storage media, or downloaded to external computers or external storage devices via a network (e.g., the Internet, local area network, wide area network and / or wireless network). Networks can include copper transmission cables, optical transmission optical fibers, wireless transmissions, routers, firewalls, switches, gateway computers and / or edge servers. Network adapter cards or network interfaces in each computing / processing equipment receive computer-readable program instructions from the network, and forward computer-readable program instructions for storage in computer-readable storage media in corresponding computing / processing equipment. Computer-readable program instructions for performing the operation of the present invention can be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, configuration data of integrated circuits, or source code or object code written in any combination of one or more programming languages ​​(including object-oriented programming languages, such as Smalltalk, C++, etc.) and process programming languages ​​(such as "C" programming languages ​​or similar programming languages). Computer readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (for example, by using the Internet of an Internet service provider). In some embodiments, the electronic circuit including, for example, a programmable logic circuit, a field programmable gate array (FPGA) or a programmable logic array (PLA) can be executed by utilizing the state information of the computer readable program instructions to perform computer readable program instructions with personalized electronic circuits, so as to perform aspects of the present invention.

[0182] Various aspects of the present invention are described herein with reference to the flowchart illustration and / or block diagram of the method, device (system) and computer program product according to embodiments of the present invention.It should be understood that each frame of flowchart illustration and / or block diagram, and the combination of frames in flowchart illustration and / or block diagram can be realized by computer-readable program instructions.These computer-readable program instructions can be provided to the processor of general-purpose computer, special-purpose computer or other programmable data processing device to produce machine, so that the instruction executed by the processor of computer or other programmable data processing device creates the device for realizing the function / action specified in one or more frames of flowchart and / or block diagram.These computer-readable program instructions can also be stored in computer-readable storage medium, which can guide computer, programmable data processing device and / or other equipment to work in a particular way, so that the computer-readable storage medium with instructions stored therein includes products, and the products include the instructions of various aspects of the function / action specified in one or more frames of flowchart and / or block diagram. Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational actions to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0183] The flow charts and block diagrams in the accompanying drawings illustrate the possible specific implementation architecture, functionality and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, fragment or part of an instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative specific implementations, the functions indicated in the box may not occur in the order indicated in the figure. For example, two boxes shown in succession may actually be executed substantially simultaneously, or sometimes these boxes may be executed in reverse order, depending on the functionality involved. It should also be noted that each box illustrated in the block diagram and / or flow chart and the combination of boxes in the block diagram and / or flow chart can be implemented by a system based on dedicated hardware that performs a specific function or action or performs a combination of dedicated hardware and computer instructions.

[0184] Although the subject matter has been described above in the general context of computer executable instructions of a computer program product running on one and / or multiple computers, it will be appreciated by those skilled in the art that the present disclosure may also or may be implemented in combination with other program modules. Typically, program modules include routines, programs, components, data structures, etc. that perform specific tasks and / or implement specific abstract data types. In addition, it will be appreciated by those skilled in the art that the computer implementation method of the present invention may be practiced with other computer system configurations, including single-processor or multi-processor computer systems, small computing devices, large computers, and computers, handheld computing devices (e.g., PDAs, phones), microprocessor-based or programmable consumer or industrial electronic devices, etc. The illustrated aspects may also be practiced in a distributed computing environment in which tasks are performed by a remote processing device linked by a communication network. However, some (if not all) aspects of the present disclosure may be practiced on a stand-alone computer. In a distributed computing environment, program modules may be located in local and remote memory storage devices.

[0185] As used in this application, the terms "component", "system", "platform", "interface", etc. may refer to and / or may include computer-related entities or entities related to an operating machine having one or more specific functions. The entities disclosed herein may be hardware, a combination of hardware and software, software, or software in execution. For example, a component may be, but is not limited to, a process, a processor, an object, an executable file, a thread of execution, a program, and / or a computer running on a processor. By way of example, both an application running on a server and a server may be a component. One or more components may reside within a process and / or a thread of execution, and a component may be located on a computer and / or distributed between two or more computers. In another example, the corresponding component may be executed according to various computer-readable media having various data structures stored thereon. Components may communicate via local and / or remote processes, such as according to signals having one or more data packets (e.g., data from a component that interacts with another component in a local system, a distributed system, and / or a network (such as, via the Internet with other systems via signals). As another example, a component may be a device having a specific functionality provided by a mechanical part operated by an electrical or electronic circuit that is operated by a software or firmware application executed by a processor. In this case, the processor may be internal or external to the device and may execute at least a portion of the software or firmware application. As yet another example, a component may be a device that provides a specific functionality through electronic components rather than mechanical parts, where the electronic components may include a processor or other device for executing software or firmware that at least partially imparts functionality to the electronic components. In one aspect, the component may simulate the electronic component, for example, via a virtual machine within a cloud computing system.

[0186] In addition, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X employs A or B" is intended to mean any natural inclusive permutation. That is, if X employs A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied in any of the foregoing cases. In addition, the articles "a" and "an" used in this specification and the drawings should generally be interpreted as meaning "one or more", unless otherwise specified or clear from the context to refer to the singular form. As used herein, the terms "example" and / or "exemplary" are utilized to mean serving as an example, instance, or illustration. For the avoidance of doubt, the subject matter disclosed herein is not limited to such examples. In addition, any aspect or design described herein as "example" and / or "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects or designs, nor is it meant to exclude equivalent exemplary structures and techniques known to those of ordinary skill in the art.

[0187] As used in this specification, the term "processor" may refer to substantially any computational processing unit or device, including but not limited to a single-core processor; a single processor with software multithreaded execution capability; a multi-core processor; a multi-core processor with software multithreaded execution capability; a multi-core processor with hardware multithreading technology; a parallel platform; and a parallel platform with distributed shared memory. In addition, a processor may refer to an integrated circuit, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic controller (PLC), a complex programmable logic device (CPLD), a discrete gate or transistor logic component, a discrete hardware component, or any combination thereof designed to perform the functions described herein. In addition, the processor may utilize a nanoscale architecture (such as, but not limited to, transistors, switches, and gates based on molecules and quantum dots) in order to optimize space usage or enhance the performance of a user device. The processor may also be implemented as a combination of computational processing units. In the present disclosure, terms such as "storage", "storage device", "data storage", "data storage device", "database", and substantially any other information storage component related to the operation and function of a component are used to refer to a "memory component", an entity embodied in a "memory", or a component including a memory. It should be appreciated that the memory and / or memory components described herein may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. By way of example and not limitation, non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), flash memory, or non-volatile random access memory (RAM) (e.g., ferroelectric RAM (FeRAM)). For example, volatile memory may include RAM, which may act as an external cache memory. By way of example and not limitation, RAM may be provided in a variety of forms, such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), direct Rambus RAM (DRRAM), direct Rambus dynamic RAM (DRDRAM) and Rambus dynamic RAM (RDRAM). In addition, the memory components of the systems or computer-implemented methods disclosed herein are intended to include, but are not limited to, these and any other suitable types of memory.

[0188] What has been described above includes only examples of systems and computer-implemented methods. Of course, it is not possible to describe every conceivable combination of components or computer-implemented methods for the purposes of describing the present disclosure, but one of ordinary skill in the art will recognize that many further combinations and permutations of the present disclosure are possible. In addition, to the extent that the terms "including", "having", "having", etc. are used in the detailed description, claims, appendices, and drawings, such terms are intended to be inclusive in a manner similar to the term "comprising", as interpreted when "including" is used as a transitional word in a claim.

[0189] Descriptions of various embodiments have been given for purposes of illustration, but these descriptions are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications or technical improvements over technologies found on the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

[0190] Further aspects of various embodiments of the claimed subject innovation are provided in the following subject matter:

[0191] 1. A system, comprising: a processor, the processor executing computer executable components stored in a memory, the computer executable components comprising: an element enhancement component, the element enhancement component generating a set of annotated preliminary training images based on annotated source images, wherein the annotated preliminary training images are formed by inserting at least one element of interest or at least one background element into the annotated source images; a modality enhancement component, the modality enhancement component generating a set of annotated intermediate training images based on the set of annotated preliminary training images, wherein the annotated intermediate training images are formed by changing at least one modality-based characteristic of the annotated preliminary training images; and a geometry enhancement component, the geometry enhancement component generating a set of annotated deployable training images based on the set of annotated intermediate training images, wherein the annotated deployable training images are formed by changing at least one geometric characteristic of the annotated intermediate training images.

[0192] 2. A system according to any of the preceding clauses, wherein the computer executable component further comprises: a training component, wherein the training component trains a machine learning model on the set of annotated deployable training images.

[0193] 3. A system according to any of the preceding clauses, wherein the element enhancement component maintains an element catalog, the element catalog listing a set of images of possible elements of interest and a set of images listing possible background elements that can be inserted into the annotated source image, wherein the modality enhancement component maintains a list of modality-based features that can be modified in the preliminary training image, and wherein the geometry enhancement component maintains a list of geometric transformations that can be applied to the intermediate training image.

[0194] 4. A system according to any of the preceding clauses, wherein the element enhancement component updates the element catalog by including new images of elements of interest or new images of background elements in the element catalog, wherein the modality enhancement component updates the list of modality-based characteristics by including new image attributes related to the device modality in the list of modality-based characteristics, and wherein the geometry enhancement component updates the list of geometric transformations by including new operations that can be applied to images in the list of geometric transformations.

[0195] 5. A system according to any of the preceding clauses, wherein the at least one element of interest or the at least one background element is a medical device or a biological symptom manifestation.

[0196] 6. A system according to any preceding clause, wherein the element enhancement component randomly positions the at least one element of interest or the at least one background element within a range of biologically feasible locations within the annotated source image.

[0197] 7. A system according to any of the preceding clauses, wherein the changing of at least one modality-based characteristic includes changing the image gamma level, changing the image blur level, changing the image brightness level, changing the image contrast level, changing the image noise level, changing the image texture, changing the image resolution, changing the image field of view, or applying modality artifacts.

[0198] 8. A system according to any of the preceding clauses, wherein the changing of the at least one geometric characteristic comprises rotation about an image axis, reflection about an image axis, image magnification, image translation, image tilt or image distortion.

[0199] 9. A computer-implemented method, comprising: generating, by a device operatively coupled to a processor, a set of annotated preliminary training images based on annotated source images, wherein the annotated preliminary training images are formed by inserting at least one element of interest or at least one background element into the annotated source images; generating, by the device, a set of annotated intermediate training images based on the set of annotated preliminary training images, wherein the annotated intermediate training images are formed by changing at least one modality-based characteristic of the annotated preliminary training images; and generating, by the device, a set of annotated deployable training images based on the set of annotated intermediate training images, wherein the annotated deployable training images are formed by changing at least one geometric characteristic of the annotated intermediate training images.

[0200] 10. A computer-implemented method according to any of the preceding clauses, further comprising: training a machine learning model by the device on the set of annotated deployable training images.

[0201] 11. A computer-implemented method according to any of the preceding clauses, further comprising: maintaining by the device an element catalog, the element catalog listing a set of images of possible elements of interest and a set of images of possible background elements that can be inserted into the annotated source image; maintaining by the device a list of modality-based features that can be modified in the preliminary training image; and maintaining by the device a list of geometric transformations that can be applied to the intermediate training image.

[0202] 12. A computer-implemented method according to any of the preceding clauses, further comprising: updating, by the device, the element catalog by including in the element catalog a new image of an element of interest or a new image of a background element; updating, by the device, the list of modality-based characteristics by including in the list of modality-based characteristics new image attributes related to the device modality; and updating, by the device, the list of geometric transformations by including in the list of geometric transformations new operations that can be applied to an image.

[0203] 13. A computer-implemented method according to any preceding clause, wherein the at least one element of interest or the at least one background element is a medical device or a biological symptom manifestation.

[0204] 14. A computer-implemented method according to any preceding clause, further comprising: randomly positioning, by the device, the at least one element of interest or the at least one background element within a range of biologically feasible locations within the annotated source image.

[0205] 15. A computer-implemented method according to any of the preceding clauses, wherein said changing at least one modality-based characteristic comprises changing an image gamma level, changing an image blur level, changing an image brightness level, changing an image contrast level, changing an image noise level, changing an image texture, changing an image resolution, changing an image field of view, or applying a modality artifact.

[0206] 16. A computer-implemented method according to any of the preceding clauses, wherein the changing of the at least one geometric characteristic comprises rotation about an image axis, reflection about an image axis, image magnification, image translation, image tilt or image distortion.

[0207] 17. A computer program product for facilitating the generation of synthetic training data to achieve improved machine learning generalization capabilities, the computer program product comprising a computer-readable memory having program instructions embodied therein, the program instructions being executable by a processor to cause the processor to: parameterize a simulation space of a data segment by defining a set of enhanced subspaces, wherein each enhanced subspace comprises a corresponding set of enhanceable parameters, and wherein each enhanceable parameter has a corresponding parameter range of possible values ​​or states; receive a source data segment; for each enhanceable parameter, sample the parameter range corresponding to the possible values ​​or states of the enhanceable parameter, thereby producing a set of sampled ranges representing values ​​or states of the simulation space; and generate a set of training data segments by applying the set of sampled ranges of values ​​or states to a copy of the source data segment.

[0208] 18. A computer program product according to any of the preceding clauses, wherein the program instructions are further executable to cause the processor to: train a machine learning model on the set of training data segments.

[0209] 19. A computer program product according to any preceding clause, wherein the program instructions are further executable to cause the processor to: update the parameterization of the simulation space by defining a new enhancement subspace.

[0210] 20. A computer program product according to any preceding clause, wherein the program instructions are further executable to cause the processor to: update the parameterization of the simulation space by defining new enhanceable parameters within the set of enhancer subspaces.

Claims

1. A system for facilitating generation of synthetic training data, comprising: an element augmentation component that accesses an annotated source image, generates a set of annotated preliminary training images based on the annotated source image, wherein each annotated preliminary training image is formed by inserting a corresponding displacement of a visual object into the annotated source image, wherein such visual object comprises a medical device or a biological symptom; a modality enhancement component for generating a set of annotated intermediate training images based on the set of annotated preliminary training images, wherein each annotated intermediate training image is formed by applying a respective permutation of a modality characteristic change to a respective annotated preliminary training image, wherein such modality characteristic change comprises a change in an image property that depends on settings or parameters of a medical imaging device that captured or generated the annotated source images; and a geometric augmentation component that generates a set of annotated deployable training images based on the set of annotated intermediate training images, wherein each annotated deployable training image is formed by applying a respective permutation of a geometric transformation to a respective annotated intermediate training image, wherein such geometric transformation comprises a spatial transformation of an image pixel grid, wherein the element augmentation component maintains an element catalog that lists a set of images of possible visual objects that can be inserted into the annotated source image, wherein the modality augmentation component maintains a list of modality characteristics that can be modified in the preliminary training image, and wherein the geometry augmentation component maintains a list of geometric transformations that can be applied to the intermediate training image, wherein the element enhancement component updates the element catalog by including new images of visual objects in the element catalog, wherein the modality enhancement component updates the list of modal characteristics by including new image attributes related to the device modality in the list of modal characteristics, and wherein the geometry enhancement component updates the list of geometric transformations by including new spatial operations that can be applied to images in the list of geometric transformations.

2. The system for facilitating generation of synthetic training data according to claim 1, wherein the system for facilitating generation of synthetic training data further comprises: A training component that trains a machine learning model on the set of annotated deployable training images.

3. A system for facilitating generation of synthetic training data according to claim 2, wherein the visual object is an object of interest that the machine learning model is configured to detect.

4. The system for facilitating synthetic training data generation of claim 1 , wherein the element augmentation component randomly positions the visual objects within a range of biologically plausible locations within the annotated source image.

5. A system for facilitating generation of synthetic training data according to claim 1, wherein the corresponding permutation of applying modal characteristic changes includes changing the image gamma level, changing the image blur level, changing the image brightness level, changing the image contrast level, changing the image noise level, changing the image texture, changing the image resolution, changing the image field of view, or applying modal artifacts.

6. The system for facilitating generation of synthetic training data of claim 1, wherein the corresponding permutation to which the geometric change is applied comprises rotation about an image axis, reflection about an image axis, image magnification, image translation, image tilt, or image distortion.

7. A computer-implemented method for facilitating generation of synthetic training data, comprising: accessing, by a device operatively coupled to the processor, an annotated source image; generating, by the device, a set of annotated preliminary training images based on the annotated source images, wherein each annotated preliminary training image is formed by inserting a corresponding displacement of a visual object into the annotated source image, wherein such visual object comprises a medical device or a biological symptom; generating, by the device, a set of annotated intermediate training images based on the set of annotated preliminary training images, wherein each annotated intermediate training image is formed by applying a respective permutation of a modality property change to a respective annotated preliminary training image, wherein such modality property change comprises a change in an image property that depends on settings or parameters of a medical imaging device that captured or generated the annotated source images; and generating, by the device, a set of annotated deployable training images based on the set of annotated intermediate training images, wherein each annotated deployable training image is formed by applying a respective permutation of a geometric transformation to a respective annotated intermediate training image, wherein such geometric transformation comprises a spatial transformation of an image pixel grid, The method further comprises: maintaining, by the device, an element catalog, the element catalog listing a set of images of possible visual objects that can be inserted into the annotated source image, wherein the maintaining comprises: maintaining a list of modality characteristics that can be modified in the preliminary training image, and maintaining a list of geometric transformations that can be applied to the intermediate training image; The device updates the element catalog by including a new image of a visual object in the element catalog, wherein the list of modal characteristics is updated by including a new image attribute related to the device modality in the list of modal characteristics, and wherein the list of geometric transformations is updated by including a new spatial operation that can be applied to the image in the list of geometric transformations.

8. A computer program product for facilitating generation of synthetic training data to achieve improved machine learning generalization capabilities, the computer program product comprising a computer readable memory having program instructions embodied therein, the program instructions being executable by a processor to cause the processor to perform the method according to claim 7.

Citation Information

Patent Citations

  • Image-based tumor phenotyping with machine learning from synthetic data

    CN107492090A

  • Image data expansion method for deep learning model training and learning

    CN109767440A