Coordination of Classification and Resource Allocation for Synthetic Data Tasks

The distributed computing system provides synthetic data as a service, and uses the parameters of synthetic data assets and scenarios to generate training data sets, solving the problems of training data set generation and sharing in the existing technology, and achieving efficient and automated training data set management and improvements in machine learning training processes.

CN112714909BActive Publication Date: 2025-05-27MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980060892.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-11-30
Filing Date
2019-06-28
Publication Date
2025-05-27
Estimated Expiration
2039-06-28

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently generate and share machine learning training data sets for multiple different fields, resulting in limited number and availability of training data sets.

Method used

Synthetic data as a service (SDaaS) is provided through a distributed computing system, and training data sets are generated using the eigen parameters and non-eigen parameters of synthetic data assets and scenarios, and training data sets are automatically developed and refined.

Benefits of technology

It enables efficient generation and management of training data sets without the complexity of manually developing training data sets, supports large-scale production and availability, and improves the machine learning training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112714909B_ABST
    Figure CN112714909B_ABST
Patent Text Reader

Abstract

Describes various techniques for classifying synthetic data tasks and coordinating resource allocation among eligible resource groups for processing synthetic data tasks. The received synthetic data tasks can be classified by identifying the task category and the corresponding eligible resource group (e.g., processor) for processing the synthetic data tasks in the task category. For example, synthetic data tasks can include generation of source assets, ingestion of source assets, identification of change parameters, change of change parameters, and creation of synthetic data. Certain categories of synthetic data tasks can be classified for processing using specific eligible resource groups. For example, the task of ingesting synthetic data assets can be classified for processing only on the CPU, while the task of creating synthetic data assets can be classified for processing only on the GPU. Synthetic data tasks can be queued and routed for processing by eligible resources.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Users rely on different types of technology systems to complete tasks. Technology systems can be improved based on machine learning, which uses statistical techniques to enable a computer to have the ability to gradually improve the performance of a specific task by using data without being explicitly programmed. For example, machine learning can be used in data security, physical security, fraud detection, healthcare, natural language processing, online search and recommendation, financial transactions, and intelligent vehicles. For each of these fields or domains, machine learning models utilize training data sets to train. A training data set is an example data set used to create a framework that matches learning tasks and machine learning applications. For example, a facial recognition system can be trained to compare the unique features of a face with a known set of features of a face to appropriately identify a person. With the increasing use of machine learning in different fields and the importance of properly training machine learning models, improvements to the computational operations of machine learning training systems will provide more efficient performance of machine learning tasks and applications, and also improve the user navigation of the graphical user interface of machine learning training systems. Summary of the Invention

[0002] Embodiments of the present invention relate to methods, systems, and computer storage media for providing a distributed computing system that supports synthetic data as a service. As background, a distributed computing system can operate based on a service-oriented architecture, where services are provided using different service models. At a higher level, a service model can provide an abstraction of the underlying operations associated with providing a corresponding service. Examples of service models include infrastructure as a service, platform as a service, software as a service, and function as a service. Using any of these models, a customer can develop, run, and manage various aspects of a service without having to maintain or develop the operational characteristics abstracted using a service-oriented architecture.

[0003] Turning to machine learning and training data sets, machine learning uses statistical techniques to enable a computer to gradually improve the performance of a specific task by using data without being explicitly programmed. Training data sets are an essential part of the machine learning field. High-quality data sets can help improve machine learning algorithms and computational operations associated with machine learning hardware and software. Creating high-quality training data sets may require a lot of effort. For example, labeling the data in a training data set can be particularly cumbersome, which usually results in an inaccurate labeling process.

[0004] When it comes to democratizing training datasets, in other words, making training datasets generally available for multiple different fields, the conventional methods of finding training datasets are significantly insufficient. For example, limited training dataset generation resources, competition among machine learning system providers, confidentiality, security, and privacy concerns, and other factors may limit the number of training datasets that can be generated or shared, and also limit the different fields in which the training datasets are available. Moreover, the theoretical solutions for developing machine learning training datasets have not been fully defined, implemented, or described because the infrastructure for implementing such solutions is inaccessible or too expensive to enable alternatives to the current technology for developing training datasets. Overall, in conventional machine learning training services, the comprehensive functionality and resources for developing machine learning training datasets are limited.

[0005] The embodiments described in this disclosure are directed to techniques for improving access to machine learning training datasets using a distributed computing system that provides synthetic data as a service (“SDaaS”). SDaaS may refer to a distributed (cloud) computing system service implemented using a service-oriented architecture. The service-oriented architecture provides a machine learning training service while abstracting the underlying operations managed via SDaaS. For example, SDaaS provides a machine learning training system that allows customers to configure, generate, access, manage, and process synthetic data training datasets for machine learning. In particular, SDaaS operates without the complexities typically associated with manually developing training datasets. SDaaS can be delivered in multiple ways based on SDaaS engines, managers, modules, or components, which include an asset assembly engine, a scenario assembly engine, a frameset assembly engine, a frameset package generator, a frameset package repository, a management layer, a feedback loop engine, a crowdsourcing engine, and a machine learning training service. The observable effect of implementing SDaaS on a distributed computing system is to support the mass production and availability of synthetic data assets for generating training datasets. In particular, training datasets are generated by altering the intrinsic and extrinsic parameters of synthetic data assets (e.g., 3D models) or scenarios (e.g., 3D scenes and videos). In other words, training datasets are generated based on changes in intrinsic parameters (e.g., asset shape, position, and material properties) and extrinsic parameters (e.g., changes in environmental and / or scene parameters such as lighting, perspective), where the changes in intrinsic and extrinsic parameters provide a programmable machine learning data representation of the asset and the scenario with multiple assets. Additional specific functionality is provided using the components of SDaaS described in more detail below.

[0006] In this way, the embodiments described herein improve the computational functionality and operations for generating training data sets based on providing synthetic data as a service using a distributed computing system. For example, the computational operations required for the manual development (e.g., labeling and tagging) and refinement (e.g., updating) of training data sets are eliminated based on SDaaS operations. SDaaS operations use synthetic data assets to automatically develop training data sets and automatically refine the training data sets based on training data set reports that indicate additional synthetic data assets or scenarios that will improve the machine learning model in the machine learning training service. In this regard, SDaaS addresses the problems caused by the manual development of machine learning training data sets and improves the existing processes for training machine learning models in a distributed computing system.

[0007] This invention content is provided to introduce some concepts in a simplified form, which will be further described in the following detailed description. This invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The following describes the technology herein in detail with reference to the accompanying drawings, in which:

[0009] Figure 1A and Figure 1B is a block diagram of an example distributed computing for providing synthetic data as a service according to an embodiment of the present invention;

[0010] Figure 2A and Figure 2B is a flowchart illustrating an example implementation of synthetic data as a service in a distributed computing system according to an embodiment of the present invention;

[0011] Figure 3 depicts an example synthetic data as a service interface of a distributed computing system according to an embodiment of the present invention;

[0012] Figure 4 depicts an example synthetic data as a service workflow of a distributed computing system according to an embodiment of the present invention;

[0013] Figure 5 depicts an example synthetic data as a service interface of a distributed computing system according to an embodiment of the present invention;

[0014] Figure 6 is a block diagram illustrating an example synthetic data as a service workflow of a distributed computing system according to an embodiment of the present invention;

[0015] Figure 7is a block diagram illustrating an example distributed computing system synthetic data as a service workflow according to an embodiment of the present invention;

[0016] Figure 8 is a flowchart illustrating an example distributed computing system synthetic data as a service operation according to an embodiment of the present invention;

[0017] Figure 9 is a flowchart illustrating an example distributed computing system synthetic data as a service operation according to an embodiment of the present invention;

[0018] Figure 10 is a flowchart illustrating an example distributed computing system synthetic data as a service operation according to an embodiment of the present invention;

[0019] Figure 11 is a flowchart illustrating an example distributed computing system synthetic data as a service operation according to an embodiment of the present invention;

[0020] Figure 12 is a flowchart illustrating an example distributed computing system synthetic data as a service operation according to an embodiment of the present invention;

[0021] Figure 13 is a flowchart illustrating an example distributed computing system synthetic data as a service operation according to an embodiment of the present invention;

[0022] Figure 14 is a flowchart illustrating an example distributed computing system synthetic data as a service operation according to an embodiment of the present invention;

[0023] Figure 15 is a flowchart illustrating an example method for suggesting change parameters according to an embodiment of the present invention;

[0024] Figure 16 is a flowchart illustrating an example method for pre-packaging a synthetic data set according to an embodiment of the present invention;

[0025] Figure 17 is a flowchart illustrating an example method for coordinating resource allocation according to an embodiment of the present invention;

[0026] Figure 18 is a flowchart illustrating an example method for coordinating resource allocation between processor groups with different architectures according to an embodiment of the present invention;

[0027] Figure 19 is a flowchart illustrating an example method for changing resource allocation according to an embodiment of the present invention;

[0028] Figure 20 is a flowchart illustrating an example method for batch processing synthetic data tasks according to an embodiment of the present invention;

[0029] Figure 21 is a flowchart illustrating an example method for retraining a machine learning model using multiple synthetic data assets;

[0030] Figure 22 is a block diagram of an example distributed computing environment suitable for implementing embodiments of the present invention; and

[0031] Figure 23 is a block diagram of an example computing environment suitable for implementing embodiments of the present invention. DETAILED DESCRIPTION

[0032] Distributed computing systems can be used to provide different types of service-oriented models. As background, service models can provide an abstraction of the underlying operations associated with providing a corresponding service. Examples of service models include infrastructure as a service, platform as a service, software as a service, and function as a service. Using any of these models, a customer can expose, run, and manage aspects of a service without having to maintain or develop the operational characteristics abstracted by the service-oriented architecture.

[0033] Machine learning using statistical techniques enables a computer to use data to gradually improve the performance of a specific task without being explicitly programmed. For example, machine learning can be used in data security, personal safety, fraud detection, healthcare, natural language processing, online search and recommendation, financial transactions, and intelligent vehicles. For each of these domains or fields, machine learning models are trained using a training data set, which is an example data set used to create a framework for matching learning tasks and machine learning applications. The training data set is an indispensable part of the machine learning field. High-quality data sets can help improve machine learning algorithms and computational operations associated with machine learning hardware and software. Machine learning platforms operate based on training data sets that support supervised and semi-supervised machine learning algorithms; however, because labeling data takes a significant amount of time, high-quality training data sets are often difficult to generate and expensive. Machine learning models rely on high-quality labeled training data sets for supervised learning, enabling the models to provide reliable results when predicting, classifying, and analyzing different types of phenomena. Without the correct type of training data set, it may be impossible to develop reliable machine learning models. The training data set includes labeled, tagged, and annotated entries to effectively train machine learning algorithms.

[0034] When it comes to democratizing training datasets, in other words, making training datasets generally available for multiple different domains, the conventional methods of finding training datasets are clearly insufficient. Currently, such limited solutions include outsourcing the labeling function, reusing existing training data and labels, collecting your own training data and labels from free sources, relying on third-party models that have been pre-trained on the labeled data, and leveraging crowdsourcing labeling services. Most of these solutions are either time-consuming, expensive, not suitable for sensitive projects, or clearly not robust enough to handle large-scale machine learning projects. Additionally, the theoretical solutions for developing machine learning training datasets have not been fully defined, enabled, or described because the infrastructure to implement such solutions is inaccessible or too expensive to implement alternatives to the current state of the art for developing training datasets. Overall, in conventional machine learning training services, the comprehensive functionality for developing machine learning training datasets is limited.

[0035] The embodiments described herein provide simple and effective methods and systems for implementing a distributed computing system that provides synthetic data as a service ("SDaaS"). SDaaS can refer to a distributed (cloud) computing system service that is implemented using a service-oriented architecture to provide machine learning training services while abstracting the underlying operations managed via the SDaaS service. For example, SDaaS provides a machine learning training system that allows customers to configure, generate, access, manage, and process synthetic data training datasets for machine learning. In particular, SDaaS operates without the complexities typically associated with manually developing training datasets. SDaaS can be delivered in multiple ways based on SDaaS engines, managers, modules, or components, including an asset assembly engine, a scenario assembly engine, a frame set assembly engine, a frame set package generator, a frame set package repository, a management layer, a feedback loop engine, a crowdsourcing engine, and a machine learning training service. The observable effect of implementing SDaaS as a service on a distributed computing system is to support the mass production and availability of synthetic data assets that generate training datasets. In particular, the training datasets are generated by changing the intrinsic and extrinsic parameters of the synthetic data assets (e.g., 3D models) or scenarios (e.g., 3D scenes and videos). In other words, the training datasets are generated based on intrinsic parameter changes (e.g., asset shape, position, and material properties) and extrinsic parameter changes (e.g., changes in the environment and / or scene parameters such as lighting, perspective), where the intrinsic and extrinsic parameter changes provide a programmable machine learning data representation of the asset and the scenario with multiple assets. Additional specific functionality is provided using the components of SDaaS described in more detail below.

[0036] As background, advancements in computer technology and computer graphics support the generation of virtual environments. For example, a graphics processing unit (GPU) can be used to manipulate computer graphics (e.g., 3D objects) and image processing. Virtual environments can be generated using computer graphics to create a realistic depiction of the real world. A virtual environment can consist entirely of computer graphics or can partially have a combination of computer graphics and real-world objects. A partially virtual environment is sometimes referred to as an augmented reality virtual environment. For example, to generate a virtual environment, a 3D scanner can be used to scan a real-world environment, and the objects in the real-world environment are programmatically identified and assigned attributes or properties. For example, a room and the identified walls of the room and a table in the room can be scanned, and the walls and the table can be assigned attributes such as color, texture, and dimensions. Other different types of environments and objects can be scanned, mapped to attributes, and used to realistically render a virtual representation of the real-world environment and objects. Conventionally, images can be captured using cameras and other image capture devices, while virtual environments are generated using different types of computer graphics processes. Conventionally, the captured images can be used to train machine learning models, and virtual images can be used to train machine learning models. Conventionally captured images can be referred to as non-synthetic data, while virtual images can be referred to as synthetic data.

[0037] Synthetic data has several advantages over non-synthetic data. In particular, since non-synthetic data lacks the inherent properties of synthetic data, synthetic data can be used to improve machine learning (i.e., programming a computer to learn how to perform and gradually improve the performance of a specific task). At a higher level, the program structure of synthetic data allows for the generation of training data sets based on changes in intrinsic parameters and extrinsic parameters, where the changes in intrinsic parameters and extrinsic parameters provide a programmable machine learning data representation of an asset and a scenario having multiple assets. For example, synthetic data can include or be generated based on an asset, and the asset can be manipulated to produce different variations of the same virtual environment (e.g., asset, scenario, and frame set variations). In contrast to virtual assets where intrinsic and extrinsic parameters can be changed as described herein, non-synthetic data (e.g., videos and images) used to train machine learning models with respect to a specific object, feature, or environment are static and generally unchangeable. The ability to create variations in the virtual environment (i.e., asset, scenario, and frame set) allows for the generation of an infinite combination of training data sets that can be used for machine learning. It is contemplated that the present invention can eliminate or limit the amount of manual intervention required to provide privacy and security in the training data set. For example, conventional training data sets can include personal identifying information, faces, or locations that must be manually edited for security and privacy purposes. With synthetic data, the training data set can be generated without any privacy or security concerns, and / or be generated to remove or edit information related to privacy or security while retaining the ability to change the intrinsic and extrinsic parameters of the asset.

[0038] With the programmable structure described above, synthetic data can have a variety of applications, and the programmable structure can address problems faced in machine learning, particularly in obtaining training data sets. As discussed, obtaining training data sets can be difficult for a variety of reasons. A specific example is the difficulty in correctly labeling synthetic data. Machine learning models for face recognition can be trained using a training data set that has been manually labeled, which introduces human error. For example, an image frame including a face can be processed with a bounding box around the face; however, the accuracy and nature associated with the bounding box are subject to human interpretation and can introduce errors that affect the training data set and ultimately the machine learning model generated from the training data set. Synthetic data can be used to improve the labeling process and eliminate human errors associated with the labeling process. In another example, in order to train a machine learning model to understand and infer car collisions at an intersection, the machine learning model would have to include a training data set representing car collisions. The problem here is that there is a limited number of images of car collisions, and even fewer collision images at a particular intersection of interest. In addition, the devices, contexts, and scenarios used for capture introduce specific limitations and challenges in the machine learning training process that cannot be overcome when relying on conventional training data set development methods. For example, if artifacts from the recording device cannot be removed or programmed into the machine learning training process and model, the artifacts may bias the picture quality and the machine learning model developed from the training data set.

[0039] Accordingly, the program structure that defines synthetic data or the assets used to generate synthetic data can be used to improve the fundamental elements of machine learning training itself. The rules, data structures, and processes used in machine learning are conventionally associated with non-synthetic data. In this way, machine learning training can be reorganized to incorporate aspects of synthetic data that would otherwise be lost in non-synthetic data. In particular, machine learning training can include rules, data structures, and processes for managing different variations of assets in a way that improves machine learning training and the resulting machine learning model. Synthetic data assets, scenarios, and sets of frames can be assembled and managed, and sets of frame packages are generated and accessed through SDaaS operations and interfaces. In addition, in a hybrid machine learning model, the machine learning model can be trained and refined using only synthetic data or a combination of synthetic data and non-synthetic data. Machine learning training can be completed within defined phases and durations that utilize aspects of both synthetic data and non-synthetic data. Embodiments of the present invention contemplate other variations and combinations of SDaaS operations based on synthetic data.

[0040] In one embodiment, the source asset can include several different parameters that can be computationally determined based on known techniques in the art. For example, the source asset can refer to a three-dimensional representation of geometric data (e.g., a 3D model). The source asset can be represented as a mesh composed of triangles, where the smoother the triangles and the more detailed the surface of the model, the larger the size of the source. The source asset is accessed and processed to be used as a synthetic data asset. In this regard, the source asset can be represented across a spectrum from high-level polygon models with a large amount of detail to low-level polygon models with less detail. The process of representing the source asset at different levels of detail can be referred to as decimation. The low-level polygon models can be used in different types of processes that are computationally expensive for high-level models. The automatic decimation process can be implemented to store the source asset at different levels of detail. Other types of programmable parameters can be determined and associated with the source asset stored as a synthetic data asset.

[0041] Embodiments of the present invention can operate on a two-tier programmable parameter system, in which a machine learning training service can automatically or based on manual intervention or input, train a model based on accessing and determining first-tier parameters (e.g., asset parameters) and / or second-tier parameters (e.g., scene parameters or frame set parameters), and use these parameters to improve the training data set and, by extension, improve the process of machine learning training and the machine learning training model itself. The machine learning training service can support deep learning and deep learning networks as well as other types of machine learning algorithms and networks. The machine learning training service can also implement a generative adversarial network as a type of unsupervised machine learning. SDaaS can utilize these underlying hierarchical parameters in different ways. For example, the hierarchical parameters can be used to determine how much to charge for a frame set, how to develop different types of frame sets for a specific device (e.g., understanding device parameters and being able to manipulate the parameters to develop a training data set), etc.

[0042] Datasets used to train machine learning models are typically customized for specific machine learning scenarios. To accomplish this customization, specific attributes of the scenario (scenario variation parameters) and / or assets within the scenario (asset variation parameters) can be selected and varied to generate different sets of frames (e.g., rendered images based on a 3D scenario) for a frameset package. The number of potential variation parameters is theoretically infinite. However, not all parameters can be varied to produce a frameset package that will improve the accuracy of the machine learning model for a given scenario. For example, a user can generate a synthetic data scenario (e.g., a 3D scenario) that includes an environment (e.g., a piece of land) and various 3D assets (e.g., cars, people, and trees). In some machine learning scenarios, such as facial recognition, variation parameters such as human body shape, latitude, longitude, sun angle, time of day, camera view, and camera position may be relevant to improving the accuracy of the corresponding machine learning model. In other scenarios, such as using computer vision and optical character recognition (OCR) to read a health insurance card, some of these variation parameters are irrelevant. More specifically, variation parameters such as latitude, longitude, sun angle, and time of day may be irrelevant to an OCR-based scenario, while other parameters such as camera view, camera position, and A-to-Z variability may be relevant to an OCR-based scenario. As used herein, A-to-Z variability refers to differences in the ways different handwriting styles, fonts, etc. represent a specific letter. Generally, some variation parameters will be relevant to some scenarios and irrelevant to others.

[0043] In addition, not all users (e.g., customers, data scientists, etc.) will know which variation parameters may be relevant to a specific scenario and may select invalid variation parameters. In such cases, generating the corresponding frameset package will result in a large amount of unnecessary computational resource consumption. Even if a user can accurately select the relevant variation parameters, the process of selecting a relevant subset can be time-consuming and error-prone when the number of possible variation parameters is large.

[0044] Thus, in some embodiments, the proposed variation parameters for a particular scenario can be presented to support the generation of a frame set package. The relevant variation parameters can be specified by a seeding taxonomy that associates machine learning scenarios with a relevant subset of variation parameters. Generally, a user interface that supports scenario generation and / or frame set generation can be provided. After a particular machine learning scenario is identified, the subset of variation parameters associated with the scenario in the seeding taxonomy can be accessed and presented as the proposed variation parameters. The seeding taxonomy can be predefined and / or adaptable. An adaptable seeding taxonomy can be implemented using a feedback loop that tracks the selected variation parameters and updates the seeding taxonomy based on determining that the selected variation parameters for a frame set generation request differ from the proposed variation parameters in the seeding taxonomy. Thus, the presentation of the proposed variation parameters helps the user to identify and select relevant variation parameters more quickly and effectively. In some cases, the presentation of the proposed variation parameters prevents the selection of invalid variation parameters and the accompanying unnecessary consumption of computing resources.

[0045] In some embodiments, a seeding taxonomy that maps machine learning scenarios to relevant variation parameters for the scenarios can be maintained. For example, the seeding taxonomy can include a list of machine learning scenarios (e.g., OCR, face recognition, video surveillance, spam detection, product recommendation, marketing personalization, fraud detection, data security, physical security screening, financial transactions, computer-aided diagnosis, online search, natural language processing, intelligent vehicle controls, Internet of Things controls, etc.). Generally, a particular machine learning scenario can be associated with a relevant set of variation parameters (e.g., asset variation parameters, scenario variation parameters, etc.) that can be varied to produce a frame set package that can be used to train and improve the corresponding machine learning model for the scenario. The seeding taxonomy can associate each scenario with a relevant subset of variation parameters. For example, unique identifiers can be assigned to the variation parameters for all scenarios, and the seeding taxonomy can maintain a list of unique identifiers indicating the relevant variation parameters for each scenario.

[0046] A user interface that supports scenario generation and / or frame set generation can support the identification of a desired machine learning scenario (e.g., based on user input indicating the desired scenario). A list or other indication of a subset of change parameters associated with the scenarios identified in the seed classification criteria can be accessed via the user interface and presented as the proposed change parameters. For example, the user interface can include a table that accepts input indicating one or more machine learning scenarios. The input can be via a text box, drop-down menu, radio button, check box, interactive list, or other suitable input. The input indicating the machine learning scenario can trigger the lookup and presentation of the relevant change parameters maintained in the seed classification criteria. The relevant change parameters can be presented as a list or other indication of the proposed change parameters for the selected machine learning scenario. The user interface can accept input indicating the selection of one or more of the proposed change parameters and / or one or more additional change parameters. Similar to the input indicating the machine learning scenario, the input indicating the selected change parameters can be via a text box, drop-down menu, radio button, check box, interactive list, or other suitable input. The input indicating the selected change parameters can cause a frame set package to be generated using the selected change parameters, and the frame set package can be obtained, for example, by download or other means.

[0047] Generally, the seed classification criteria can be predefined and / or adaptable. For example, in embodiments including adaptable seed classification criteria, a feedback loop can be implemented that tracks the selected change parameters for each machine learning scenario. The feedback loop can update the seed classification criteria based on a comparison of the selected change parameters for a particular frame set package with the proposed change parameters for the corresponding machine learning scenario. The change parameters can be added or removed from the list or other indication of the relevant change parameters for a particular scenario in the seed classification criteria in a variety of ways. For example, a change parameter can be added or removed based on determining that the selected change parameters for a particular frame set generation request are different from the proposed change parameters for the corresponding machine learning scenario in the seed classification criteria. This update to the seed classification criteria can be performed based on the determined differences for a single request, a threshold number of requests, the majority of requests, or other suitable criteria. Thus, in these embodiments, the seed classification criteria are evolving and self-healing.

[0048] As described above, the datasets used to train machine learning models are typically customized to specific machine learning scenarios. Additionally, especially in heavy (artificial intelligence) AI industries such as government, healthcare, retail, oil and gas, gaming, and finance, some scenarios may occur again. However, recreating synthetic datasets for a particular machine learning scenario every time the dataset is requested would result in a significant consumption of computing resources.

[0049] Thus, in some embodiments, synthetic datasets (e.g., frame set packages) for common or expected machine learning scenarios can be pre-packaged by creating the datasets before a customer or other user requests the datasets. A user interface can be provided that includes a form or other suitable tool that allows the user to specify the desired scenario. For example, the form can allow the user to specify an industry sector and a specific scenario within the industry sector. In response, a representation of the available packages of synthetic data for that industry sector and scenario can be presented, the user can select an available package, and the selected package can be obtained, for example, by download or other means. By pre-packaging synthetic datasets such as these, a substantial consumption of computing resources that would otherwise be required to generate the datasets on demand multiple times can be avoided.

[0050] An example of a common machine learning scenario may occur in a government agency. For example, an agency may wish to use computer vision and optical character recognition (OCR) to read the health insurance cards of veterans. If the agency needs to read 1,000 cards per day, but the current model has an accuracy of 75%, the current model will not be able to correctly read 250 cards per day. To improve the accuracy of the model (e.g., to 85%), the agency may wish for the dataset that can be used to train the model to improve its accuracy. In some embodiments, a representative can use the techniques described herein to create a synthetic dataset for improving OCR. However, in other embodiments, the dataset can be pre-packaged and provided on demand.

[0051] In some embodiments, a list or other indication of pre-packaged synthetic datasets can be presented via a user interface that supports scenario generation and / or frame set generation. For example, the user interface can present a form that accepts input indicating one or more industry sectors. The input can be via a text box, drop-down menu, radio button, check box, interactive list, or other suitable input. The input indicating the industry sector can trigger a list or other indication of the pre-packaged synthetic datasets for the selected (s) industry sector. The user interface can accept input indicating a selection of one or more pre-packaged synthetic datasets. Similar to the input indicating the industry sector, the input indicating the pre-packaged synthetic datasets can be via a text box, drop-down menu, radio button, check box, interactive list, or other suitable input. The input indicating the pre-packaged synthetic datasets can cause the selected pre-packaged synthetic dataset to be obtained, for example, by download or other means.

[0052] Different types of processors can better execute different tasks associated with embodiments of the present invention. For example, a CPU is typically optimized to perform serial processing of various different tasks, while a GPU is optimized to perform parallel processing of heavy computational tasks for applications such as 3D visualization, gaming, image processing, big data, deep machine learning, etc. Since the process of creating synthetic data assets may consume a large amount of visual resources, using a GPU can speed up this process and improve computational efficiency. However, the GPU may have limited availability due to various reasons, including being used for other tasks, cost constraints, or other reasons. Therefore, the CPU can be used additionally or alternatively. Thus, techniques for coordinating resource allocation to process synthetic data tasks are needed.

[0053] In some embodiments, a management layer can be used to classify synthetic data tasks and can coordinate resource allocation such as processors or groups of processors (e.g., CPUs and GPUs) to process the tasks. Generally, the management layer can allocate requests to execute synthetic data tasks to be performed on a CPU or GPU based on the classification of the tasks and / or resource availability. In some embodiments, the management layer can complete this task by classifying incoming tasks according to task categories and / or eligible resources for processing the tasks, queuing the classified tasks, and routing or otherwise allocating the queued tasks to the corresponding processors based on task classification, resource availability, or some other criteria.

[0054] For example, when the management layer receives a request to execute a synthetic data task, the management layer can classify the task by task category. Task categories that support synthetic data creation can include generation or specification of source assets (e.g., 3D models), ingestion of source assets, simulation (e.g., identification and variation of first layer and / or second layer parameters), synthetic data creation (e.g., rendering an image based on a 3D model), a subcategory thereof, etc. Some tasks may be GPU-preferred, such as simulation. Other tasks may be CPU-preferred, such as variation or ingestion. For example, ingestion can include a large amount of idle activity, so it may be necessary to process ingestion tasks on the CPU to avoid idling the GPU. At the same time, some tasks may not support the CPU or GPU. Thus, some such tasks can be classified according to the eligible resources for processing the tasks (e.g., CPU-only, GPU-only, hybrid, etc.), and the classification can be based on task category. The classified tasks can be queued for processing (e.g., in a single queue, in separate queues for each classification, or otherwise).

[0055] Typically, management can use a variety of techniques to monitor the availability of resources (e.g., CPUs and GPUs) and route or otherwise allocate tasks to the resources. For example, tasks can be allocated based on task classification, queuing (e.g., FIFO, LIFO, etc.), prioritization schemes (e.g., assigning priorities to specific accounts or queues), scheduling (e.g., changing the allocation based on time such as daily, weekly, annually, etc.), resource availability, some other criteria, or some combination thereof. Typically, tasks classified for processing on a CPU are allocated to the CPU, tasks classified for processing on a GPU are allocated to the GPU, and hybrid tasks can be allocated to one of them (e.g., based on resource availability, assigning priorities to available GPUs, etc.).

[0056] In some cases, computational efficiency can be improved by keeping the GPUs loaded. In this way, management can monitor one or more GPUs to determine when they are available (e.g., when a GPU enters a standby or suspended mode), and management can route or otherwise allocate tasks (e.g., classified and / or queued tasks) to the available GPUs for execution. For example, an available GPU can be allocated to process specific queued tasks selected based on eligible resources for processing the tasks (e.g., assigning priorities to hybrid tasks or GPU-only tasks), based on task category (e.g., assigning priorities to requests for creating synthetic data), based on the time the request is received (e.g., assigning priorities to the earliest queued queue), based on queue fill (e.g., assigning priorities to the most filled queue), based on the estimated completion time (e.g., assigning priorities to tasks estimated to take the longest or shortest time to execute), etc. In this way, tasks in synthetic data creation can be routed or otherwise allocated to specific resources (e.g., processors) based on various criteria.

[0057] In some cases, creating synthetic data can take a significant amount of time. For example, a client can submit a synthetic data request to create 50k synthetic data assets (e.g., images) from a selected source asset (e.g., a 3D model). Processing such a large request may involve performing multiple tasks and can take several hours. Thus, a portal can be provided that allows the requesting account to actively monitor the progress of the SDaaS in processing the request and / or issue commands to allocate resources to process the task even after the request has started being processed. For example, the portal can support commands for changing or adding resources (e.g., CPU / GPU) on the fly (e.g., during an intermediate request) or otherwise coordinating resource allocation. By changing the resource allocation on the fly (e.g., by allocating the creation of a batch of synthetic data to one or more GPUs), the portal can allow the user to reduce the time taken to process the request (e.g., from 10 hours to 6 hours). Thus, the portal can allow the client or other users to visualize the progress characteristics and / or influence the way the processing is done.

[0058] The portal can provide various types of feedback regarding the processing of the synthetic data request. Among other types of feedback, the portal can present an indication of the progress of processing the request (e.g., 20k out of 50k images have been created), CPU / GPU utilization, the average time taken to create synthetic data on the CPU, the average time taken to create synthetic data on the GPU, the remaining estimated time (e.g., based on existing or estimated creation patterns), the corresponding service cost (e.g., incurred, estimated, or predicted for completing the request processing), or other information. For example, the portal can present an indication of resource consumption or availability (e.g., availability in an existing service level agreement, resources available for additional charges, etc.) and can break it down by task category. In another example, the portal can present information indicating which resources have been allocated to which request and / or task. In yet another example, the portal can present information indicating the resource consumption so far (e.g., for billing purposes). These are only examples, and other types of feedback are possible.

[0059] In some embodiments, the portal may execute commands for allocating resources and may support such commands even after a request has started being processed (mid-request). For example, a command for batch creating synthetic data assets (e.g., 20k to 30k assets from a request to create 50k assets) may be executed and the corresponding tasks may be assigned to a specific processor (e.g., GPU). In this way, the user can manually set or change resource allocation on the fly. Similarly, a command for canceling a processing request (e.g., creating a synthetic asset) may be created. Thus, a command interface may be provided to allow SDaaS (e.g., via automated control) and / or the user (e.g., via manual control, selecting corresponding automated controls, etc.) to reduce and / or minimize processing time, cost (e.g., by assigning tasks to be executed on a relatively inexpensive processor such as a CPU), output quality, etc. Consuming a threshold amount of resources (and / or incurring a threshold cost) may trigger automatically stopping further processing requests, tasks, their specific classifications, etc.

[0060] Embodiments of the present invention may support inherently improving the machine learning training process because synthetic data itself can introduce new challenges and opportunities in machine learning training. For example, synthetic data introduces the possibility of an unlimited amount of training data sets. Storing training data sets can be difficult because computational storage is a limited resource. Rendering assets also requires computational resources. To address these issues, machine learning training using synthetic data may include rendering assets in real time without having to pre-render and store the assets. A machine learning model may be trained using synthetic data (e.g., assets, scenes, and sets of frames) generated on the fly, thus avoiding the storage of training data sets. In this regard, the process of machine learning training has been changed to incorporate the ability to process the real-time generation of training data sets, which is generally not required for non-synthetic data. Machine learning training may include making real-time calls to an API or service to generate relevant training data sets with the identified parameters. Correspondingly, after the training based on the real-time generated synthetic data for machine learning training is completed, the training data sets may be discarded. Since there is no need to store pre-rendered synthetic data, the mechanical process of reading gigabytes from a storage device is also avoided.

[0061] In addition, for example, machine learning using synthetic data can be used to address the problem of machine learning bias. In particular, it has been observed that different machine learning models trained by different institutions have attributes related to how the machine learning models are trained. The type of training data introduces bias in the way the machine learning model operates. Thus, the machine learning training process for an existing machine learning model can include: using a machine-trained model to identify defects (e.g., a classification confusion matrix for classification accuracy) and retraining the model based on generating synthetic image data that addresses the defects to generate an updated machine learning training model that removes the bias of the machine learning model. In one example, the evaluation process for improving the machine learning model can be standardized so that data scientists can access and view the defects in the training dataset through a machine learning training feedback interface. For example, the training dataset can be labeled to help identify clusters, where clusters of asset parameters or scenario parameters can be used to generate a new training dataset to retrain the machine learning model. Retraining on the new training dataset can include identifying irrelevant parameters and discarding them, identifying relevant parameters, and adding new relevant parameters based on analyzing metrics from the machine learning model training results.

[0062] Machine learning models can also be generated, trained, and retrained for specific scenarios. At a higher level, a template machine learning model can be generated for a general problem; however, the same template machine learning model can be retrained in several different ways using multiple different training datasets to address different specific scenarios or variations of the same problem. The template machine learning model addresses two challenges that may not be addressed by using non-synthetic data in machine learning training. First, machine learning models may sometimes be customized for the available training dataset, rather than, conversely, the specific reason for generating the machine learning model determining the training dataset to be generated. Second, machine learning training using only the available training dataset results in an inflexible machine learning model that requires a great deal of effort to reuse. In contrast, using synthetic assets, the machine learning model indicates the type of training dataset to be generated for the template machine learning model or a dedicated machine learning model generated based on the template machine learning model.

[0063] A template machine learning model can refer to a partially constructed or trained machine learning model that can be augmented by leveraging a supplemental training dataset to repurpose the machine learning model for different specialized scenarios. For example, a facial recognition model can be trained as a template but then retrained to accommodate a specific camera that has known artifacts or a specific location or configuration, and the artifacts or specific location or configuration are introduced into the retraining dataset to train a machine learning model specifically for the specific camera. The supplemental training dataset can be specifically generated for training a specialized machine learning model. Machine learning training specific to the template can include an interface for training the template, selecting specific retraining parameters for the specialized training dataset, and training the specialized machine learning model. Other variations and combinations of the template machine learning model and the specialized machine learning model are contemplated using the embodiments described herein.

[0064] Additionally, the training dataset can be generated for specific hardware that can be used to perform machine learning training. For example, a first customer who has access to a GPU processor can request and generate a synthetic data training dataset to take advantage of the ideal processing of the GPU processor. A second customer who has access to a CPU processor can request and generate a synthetic data training dataset for the CPU processor. A third customer who has access to both GPU and CPU processors or is agnostic to GPU or CPU considerations can simply request and generate a synthetic data training dataset accordingly. Currently, the training dataset indicates which processors can be used during the machine learning training process, and for synthetic data, and in some specific cases, the combination of synthetic data and non-synthetic data can affect the type of synthetic data generated for machine learning. The synthetic data training data will correspond to the strength or capabilities of different types of processors to handle different types of training datasets for operations in machine learning model training.

[0065] Embodiments of the present invention also support different types of hybrid-based machine learning models. The machine learning model can be trained using non-synthetic data and then augmented by training using synthetic data to fill in the non-synthetic data assets. The hybrid training can be iterative or combinatorial, as the training can be performed by fully utilizing non-synthetic data and then synthetic data. Alternatively, the training can be performed by simultaneously utilizing a combination of non-synthetic data and synthetic data. For example, an A-to-Z variability machine learning model can be trained using an existing training dataset and then augmented by training using synthetic data. However, the training dataset can be augmented or filled in first before training the machine learning model. Another type of hybrid-based training can refer to training a machine learning model only on synthetic data; however, the hybrid machine learning model is compared with a non-synthetic data machine learning model or is constructed based on a reference non-synthetic data machine learning model. In this regard, the final hybrid-based machine learning model is trained only using synthetic data but can benefit from the previous machine learning model training results generated from an existing non-synthetic data machine learning model.

[0066] For certain types of scenarios where one information set is non-synthetic data associated with a first aspect of a scenario and the second dataset is synthetic data associated with a second aspect of the scenario, the synthetic data can be further used to generate supplementary data. For example, a training dataset may have non-synthetic data for the summer season, making the machine learning model accurate during the summer months but inaccurate during other months of the year. The same machine learning model can be trained using non-synthetic data during the summer while simultaneously training using synthetic data generated for non-summer months. Additionally, with synthetic data, there is an infinite amount of training datasets that can be generated. Thus, embodiments of the present invention can support online training and offline training, where, in a high-level sense, online training refers to continuously training a machine learning model using new training datasets, while offline training can refer to training a machine learning model only periodically. It is envisioned that the training can include periodically feeding a combination of non-synthetic data and synthetic data to the machine learning model.

[0067] Embodiments of the present invention can also use extensive metadata associated with each asset, scene, frame set, or frame set package to improve the training of machine learning models. At a higher level, the machine learning training process can use metadata associated with synthetic data to augment and train. This addresses problems caused by non-synthetic data, which includes limited metadata or no metadata, such that machine learning models cannot benefit from the additional details provided by metadata during the training process. For example, a collision picture at an intersection may show wet conditions but does not further include a specific indication of how long it has been raining, an exact quantification of the visibility level, or other parameters of the specific scene. Non-synthetic data is typically limited to pixel-level data with object boundaries. In contrast, synthetic data can include not only information about the parameters identified above but also any relevant parameters that can help the machine learning model make accurate predictions.

[0068] In addition, synthetic data parameters are accurately labeled compared to manually intervened non-synthetic data labeling. The additional and accurate metadata improves the accuracy of predictions using the trained machine learning model. Moreover, as described above, the additional metadata supports the generation of predictive machine learning models for special scenarios using details from the additional metadata. The machine learning model trained using the additional metadata can also support the improvement of query processing for predictions using the machine learning model and also a novel interface for making such queries. The queries can be based on the additional metadata available for training the machine learning model, and corresponding results can be provided. It is envisioned that user interfaces for accessing, viewing, and interacting with the additional metadata for training, querying, and query results can be provided to users of the SDaaS system.

[0069] Example Operating Environments and Diagrams

[0070] Reference Figure 1A and Figure 1B As shown in FIGS. 10 and 11, the components of the distributed computing system 100 can operate together to provide the functionality of the SDaaS described herein. The distributed computing system 100 supports the processing of synthetic data assets to generate and process training data sets for machine learning. At a higher level, the distributed computing supports a distributed framework for large-scale production of training data sets. In particular, a distributed computing architecture based on file compression, large-scale GPU-enabled hardware, unstructured storage, and / or a distributed backbone network inherently supports the ability to provide SDaaS functionality in a distributed manner, such that multiple users (e.g., designers or data administrators) can access and operate on synthetic data assets simultaneously.

[0071] Figure 1AIt includes client device 130A and interface 128A, client device 130B and interface 128B, and client device 130C and interface 128C. The distributed computing system also includes several components that support SDaaS functionality, including asset assembly engine 110, scenario assembly engine 112, frame set assembly engine 114, frame set package generator 116, frame set package repository 118, feedback loop engine 120, crowdsourcing engine 122, machine learning training service 124, SDaaS repository 126, and seed classification criteria 140. Figure 1B Illustrated are assets 126A and frame sets 126B stored in SDaaS repository 126 and integrated with the machine learning training service to automatically access assets, scenarios, and frame sets as described in more detail below.

[0072] Asset assembly engine 110 can be configured to receive a first source asset from a first distributed synthetic data as a service (SDaaS) upload interface and can receive a second source asset from a second distributed SDaaS upload interface. The first source asset and the second source asset can be ingested, where ingesting the source assets includes automatically calculating the values of the asset change parameters of the source assets. For example, Figure 2A Includes source assets 210 ingested into an asset repository (i.e., asset 220). The asset change parameters can be programmed for machine learning. The asset assembly engine can generate a first synthetic data asset and a second synthetic data asset. The first synthetic data asset includes a first set of values for the asset change parameters, and the second synthetic data asset includes a second set of values for the asset change parameters. The first synthetic data asset and the second synthetic data asset are stored in the synthetic data asset repository.

[0073] The distributed SDaaS upload interfaces (e.g., interfaces 128A, 128B, or 128C) are associated with an SDaaS integrated development environment (IDE). The SDaaS IDE supports identifying other values of the asset change parameters of the source assets. These values are associated with generating a training data set based on intrinsic parameter changes and non-intrinsic parameter changes, where the intrinsic parameter changes and non-intrinsic parameter changes provide a programmable machine learning data representation of assets and scenarios. Ingesting the source assets is based on a machine learning synthetic data standard, which includes a file format and a data set training architecture. The file format can refer to a hard standard, while the data set training architecture can refer to a soft standard, e.g., automated or manual human intervention.

[0074] Reference Figure 2B , ingesting the source assets (e.g., source asset 202) also includes automatically calculating the values of the scenario change parameters of the source assets, where the scenario change parameters can be programmed for machine learning. A synthetic data asset profile can be generated, where the synthetic data asset profile includes the values of the asset change parameters.Figure 2B Additional workpieces such as bounding box 208, thumbnail 210, 3D visualization 212, and optimized asset 214 are illustrated. Refer to Figure 3 , the upload interface 300 is an example SDaaS interface that can support uploading and tagging source assets for ingestion.

[0075] The scene assembly engine 112 can be configured to receive a selection of a first synthetic data asset and a selection of a second synthetic data asset (e.g., a 3D model) from a distributed synthetic data as a service (SDaaS) integrated development environment (IDE). For example, refer to Figure 4 , the assets and asset change parameters 410 at the first layer can be used to generate a scene (e.g., a 3D scene), and the scene change parameters 420 at the second layer are further used to define a set of frames 430 (e.g., images) based on the scene. The synthetic data assets are associated with the asset change parameters and the scene change parameters, which can be used to change various aspects of the synthetic data assets and / or the resulting scene. The asset change parameters and the scene change parameters can be programmed for machine learning. The scene assembly engine can receive values for generating a synthetic data scene, where these values correspond to the asset change parameters or the scene change parameters. Based on these values, the scene assembly engine can use the first synthetic data asset and the second synthetic data asset to generate a synthetic data scene.

[0076] The scene assembly engine client (e.g., client device 130B) can be configured to receive a query for a synthetic data asset, where the query is received via the SDaaS IDE, generate a query result synthetic data asset, and cause a corresponding synthetic data scene generated based on the query result synthetic data asset to be displayed. Generating the synthetic data scene can be based on values for scene generation received from at least two scene assembly engine clients. The synthetic data scene can be generated in association with a scene preview and metadata.

[0077] The frame set assembly engine 114 can be configured to access a synthetic data scene and determine a first set of values of scene change parameters, where the first set of values is automatically determined to generate a frame set of the synthetic data scene (e.g., a set of images generated based on scene changes). The frame set assembly engine can also generate a frame set of the synthetic data scene based on the first set of values, where the frame set of the synthetic data scene includes at least a first frame in the frame set and stores the frame set of the synthetic data scene, and the first frame includes the synthetic data scene updated based on the values of the scene change parameters. A second set of values of the scene change parameters is manually selected to generate a frame set of the synthetic data scene. The second set of values is manually selected using a synthetic data as a service (SDaaS) integrated development environment (IDE) that supports machine learning synthetic data standards, and the machine learning synthetic data standards include file formats and dataset training architectures. Generating a frame set of the synthetic data scene includes iteratively generating frames for the frame set of the synthetic data scene based on: updating the synthetic data scene based on the first set of values.

[0078] The frame set package generator 116 can be configured to access a frame set package generator configuration file, where the frame set package generator configuration file is associated with a first image generation device, and the frame set package generator configuration file includes known device variability parameters associated with the first image generation device. The frame set package is based on the frame set package generator configuration file, where the frame set package generator configuration file includes the values of the known device variability parameters. The frame package includes categories based on at least two synthetic data scenes. Generating the frame set package is based on an expected machine learning algorithm that will be trained using the frame set package, and the expected machine learning algorithm is identified in the frame set package generator configuration file. The frame set package includes assigning a value quantifier to the frame set package. The frame set package is generated based on a synthetic data scene that includes synthetic data assets.

[0079] The frame set package repository 118 can be configured to receive a query for a frame set package from a frame set package query interface, where the frame set query interface includes multiple frame set package categories, identify the query result frame set package based on the frame set package configuration file; and transmit the query result frame set package. At least a portion of the query triggers an auto-suggested frame set package, where the auto-suggested frame set package is associated with the synthetic data scene of the frame set, and the synthetic data scene has synthetic data assets. The frame set package is associated with an image generation device, where the image generation device includes known device variability parameters programmable for machine learning. The query result frame set package is transmitted to an internal machine learning model training service (e.g., machine learning training service 124) or an external machine learning model training service operating on a distributed computing system.

[0080] The feedback loop engine 120 can be configured to access a training dataset report that identifies synthetic data assets having values of asset change parameters, where the synthetic data assets are associated with a set of frames. The feedback loop engine 120 can be configured to update the synthetic data assets with the synthetic data asset change parameters based on the training dataset report; and use the updated synthetic data assets to update the set of frames. Values are identified manually or automatically in the training dataset report for updating the set of frames. An updated set of frames is assigned a value quantifier. The training dataset report is associated with an internal machine learning model training service or an external machine learning model training service operating on a distributed system.

[0081] The crowdsourcing engine 122 can be configured to: receive a source asset from a distributed synthetic data as a service (SDaaS) crowdsourcing interface; receive crowdsourcing labels for the source asset via the distributed SDaaS crowdsourcing interface; ingest the source asset based in part on the crowdsourcing labels, where ingesting the source asset includes automatically calculating a value of an asset change parameter of the source asset, where the asset change parameter is programmable for machine learning; and generate a crowdsourced synthetic data asset that includes the value of the asset change parameter. A value quantifier is assigned to the crowdsourced synthetic data asset. The crowdsourced synthetic data asset profile includes the asset change parameter. Refer to Figure 5 , the crowdsourcing interface 500 is an example SDaaS interface that can support crowdsourcing, uploading, and tagging of source assets for ingestion.

[0082] Automatically suggest change parameters

[0083] Return Figure 1A , in some embodiments, the suggested change parameters for a particular machine learning scenario can be presented to support the generation of a set of frames package. The relevant change parameters can be specified by a seed classification criterion 140 that associates the machine learning scenario with a corresponding relevant subset of change parameters. After a particular machine learning scenario is identified via a user interface (e.g., interface 128B or 128C) that supports scenario generation and / or set of frames generation, the subset of change parameters associated with the scenario identified in the seed classification criterion 140 can be accessed and presented as the suggested change parameters. One or more of the suggested change parameters and / or one or more additional change parameters can be selected, the set of frames assembly engine 114 can generate a set of frames (e.g., an image) by changing the selected change parameters (e.g., asset and / or scenario change parameters), and the set of frames package generator 116 can use the generated set of frames to generate a set of frames package.

[0084] Typically, a user (e.g., a customer, a data scientist, etc.) may seek to generate a frame set package (i.e., a synthetic data set) for training a specific machine learning model. At a higher level, a user interface (e.g., interface 128B or 128C) that supports scenario generation and / or frame set generation may be provided. For example, the user interface may allow the user to design or otherwise specify a 3D scenario (e.g., using 3D modeling software). The user interface may allow the user to specify one or more variation parameters to vary when generating different frame sets (e.g., images based on the 3D scenario) for the frame set package. The number of potential variation parameters is theoretically infinite. However, not all parameters can be varied to produce a frame set package that will improve the accuracy of the machine learning model for a given scenario.

[0085] In some scenarios such as face recognition, variation parameters such as human body shape, latitude, longitude, sun angle, time of day, and camera view and position may be relevant to improving the accuracy of the corresponding machine learning model. In other scenarios (such as using computer vision and optical character recognition (OCR) to read a health insurance card), some of these variation parameters are irrelevant. More specifically, some variation parameters (such as latitude, longitude, sun angle, and time of day) may be irrelevant to an OCR-based scenario, while other parameters such as camera view, camera position, and A-to-Z variability may be relevant. The seed classification criteria 140 may maintain the association between the machine learning scenario and the relevant variation parameters to present the proposed variation parameters for a specific scenario.

[0086] The seed classification criteria 140 associates a machine learning scenario with the relevant variation parameters of the scenario. For example, the seed classification criteria 140 may include a list or other identification of various machine learning scenarios (e.g., OCR, face recognition, video surveillance, spam detection, product recommendation, marketing personalization, fraud detection, data security, physical security screening, financial transactions, computer-aided diagnosis, online search, natural language processing, intelligent vehicle controls, Internet of Things controls, etc.). For any specific machine learning scenario, the associated set of variation parameters (e.g., asset variation parameters, scenario variation parameters, etc.) can be varied to produce a frame set package that can improve the corresponding machine learning model for the scenario. The seed classification criteria 140 may associate one or more machine learning scenarios with the corresponding relevant subset of variation parameters. For example, variation parameters across all scenarios may be assigned unique identifiers, and the seed classification criteria 140 may maintain a list of the unique identifiers indicating the relevant variation parameters for each scenario.

[0087] In some embodiments, a user interface (e.g., interface 128B or 128C) that supports scenario generation and / or frame set generation may support identification of a desired machine learning scenario (e.g., based on user input indicating the desired scenario). A list or other indication of a subset of the variation parameters associated with the identified scenario in the seed classification criteria 140 may be accessed via the user interface and presented as the proposed variation parameters. For example, the user interface may include a table that accepts input indicating one or more machine learning scenarios. The input may be via a text box, drop-down menu, radio button, check box, interactive list, or other suitable input. The input indicating the machine learning scenario may trigger a lookup and presentation of the associated variation parameters maintained in the seed classification criteria 140. The associated variation parameters may be presented as a list or other indication of the proposed variation parameters for the selected machine learning scenario. The user interface may accept input indicating a selection of one or more of the proposed variation parameters and / or one or more additional variation parameters. Similar to the input indicating the machine learning scenario, the input indicating the selected variation parameters may be via a text box, drop-down menu, radio button, check box, interactive list, or other suitable input. The input indicating the selected variation parameters may be used by the frame set assembly engine 114 to generate a frame set (e.g., an image) by changing the selected variation parameters, and the frame set package generator 116 may utilize the generated frame set to generate a frame set package. For example, the frame set package may be obtained by download or other means.

[0088] Generally, the seed classification criteria 140 may be predefined and / or adaptable. In embodiments where the seed classification criteria 140 is adaptable, the feedback loop engine 120 may track the selected variation parameters for each machine learning scenario. The feedback loop engine 120 may update the seed classification criteria 140 based on a comparison of the selected variation parameters for a particular frame set package with the proposed variation parameters for the corresponding machine learning scenario. The variation parameters may be added or removed from the list or other indication of the associated variation parameters for a particular scenario in the seed classification criteria 140 in various ways. For example, a variation parameter may be added or removed based on determining that the selected variation parameters for a particular frame set generation request differ from the proposed variation parameters in the seed classification criteria 140 for the corresponding machine learning scenario. This update to the seed classification criteria 140 may be performed based on the determined differences for a single request, for a threshold number of requests, for most requests, or other suitable criteria. Thus, in these embodiments, the seed classification criteria is evolving and self-healing.

[0089] Pre-packaged synthetic data sets for common machine learning scenarios

[0090] In some embodiments, synthetic datasets (e.g., frame set packages) for common or anticipated machine learning scenarios can be pre-packaged by creating the datasets before a customer or other user requests the datasets. A user (e.g., interface 128B or 128C) can be provided with a user interface that includes a form or other suitable tool that allows the user to specify the desired scenario. For example, the form can allow the user to specify an industry sector and a specific scenario within the sector. In response, a representation of the available packages of synthetic data for that industry and scenario can be presented, the user can select an available package, and the selected package can be obtained, for example, by download or other means.

[0091] Generally, synthetic datasets for any number of machine learning scenarios can be pre-packaged. For example, synthetic datasets can be pre-packaged for some common or anticipated scenarios in heavy AI industries such as government, healthcare, retail, oil and gas, gaming, and finance. By pre-packaging synthetic datasets such as these, a substantial consumption of computing resources can be avoided that would otherwise be used to generate the datasets on demand.

[0092] In some embodiments, a list or other indication of pre-packaged synthetic datasets can be presented via a user interface (e.g., interface 128B or 128C) that supports scenario generation and / or frame set generation. For example, the user interface can present a form that accepts input indicating one or more industry sectors. The input can be via a text box, drop-down menu, radio button, checkbox, interactive list, or other suitable input. Input indicating an industry sector can trigger the presentation of a list or other indication of pre-packaged synthetic datasets for the selected (s) industry sector(s). The user interface can accept input indicating a selection of one or more pre-packaged synthetic datasets. Similar to the input indicating an industry sector, the input indicating a pre-packaged synthetic dataset can be via a text box, drop-down menu, radio button, checkbox, interactive list, or other suitable input. Input indicating a pre-packaged synthetic dataset can cause the selected pre-packaged synthetic dataset to be obtained, for example, by download or other means.

[0093] Coordination of Synthetic Data Task Classification and Resource Allocation

[0094] Now turning to Figure 6 , Figure 6FIG. illustrates an example environment 600 in which requests to perform synthetic data tasks can be processed. The components of environment 600 can operate within a distributed computing system to provide functionality for the SDaaS described herein. Environment 600 supports processing synthetic data tasks such as generation or specification of source assets (e.g., 3D models), ingestion of source assets, simulation (e.g., identification and variation of first layer and / or second layer parameters), creation of synthetic data assets (e.g., drawing images based on 3D models), etc.

[0095] Environment 600 includes a task classifier 610, a CPU queue 620, a GPU queue 622, a hybrid queue 624, a scheduler 630, and a load monitor 660 (collectively referred to as the management layer). At a higher level, the management layer can receive requests to perform synthetic data tasks. The management layer classifies the tasks and routes them to a processor group (e.g., CPU 640 or GPU 650) for execution. In this way, the management layer coordinates resource allocation between the CPU 640 and the GPU 650 to process synthetic data tasks. Environment 600 is only an example environment, and variations are possible. For example, although the embodiments herein refer to coordination of allocation between a CPU and a GPU, environment 600 can include any suitable processor / architecture (e.g., CPU, GPU, ASIC, FPGA, etc.), and the corresponding allocation can be coordinated.

[0096] The task classifier 610 receives incoming requests to perform synthetic data tasks. Any given request can correspond to one or more synthetic data tasks. For example, a request to create a frame set package including 100k synthetic data assets (e.g., images) can be decomposed into any number of synthetic data tasks. In one example, the request can be decomposed into synthetic data tasks that create synthetic data assets in batches (e.g., 1k, 5k, 10k, etc. batches), and the synthetic data tasks can be assigned to a specific processor (or processor group) to generate the corresponding batch of synthetic data assets. More generally, any incoming request can be decomposed into a corresponding set of synthetic data tasks. In this way, the task classifier 610 can generate one or more synthetic data tasks based on the received requests.

[0097] The task classifier 610 classifies synthetic data tasks by task category, eligible resources for processing the tasks, or some other criteria. For example, the task classifier can identify the category of the task and the corresponding eligible resources for processing tasks in that category. In one example, a synthetic data task related to ingesting a synthetic data asset can be classified as being processed only on the CPU (e.g., to avoid idling the GPU), or can be classified as hybrid (e.g., eligible to be processed on the CPU or GPU). In another example, a synthetic data task for creating a synthetic data asset (e.g., rendering an image from a source asset) can be classified as being processed only on the GPU (e.g., to minimize processing time), or as hybrid (e.g., eligible to be processed on the CPU or GPU). Of course, these classifications are merely examples, and any other classification scheme can be implemented. The task category and / or corresponding eligible resources can be predefined, learned, etc.

[0098] In some embodiments, the task classifier 610 can queue synthetic data tasks in one or more queues. In one example, the synthetic data tasks can be filled into a single queue, and the scheduler 630 can select tasks from the queue for routing or assignment for processing on the CPU 640 or GPU 650 (e.g., based on some property or other characteristic of the task that indicates which resource the task should be assigned to). In another example, the synthetic data tasks can be filled into multiple queues, such as a CPU queue 620, a GPU queue 622, and a hybrid queue 624. In the latter example, synthetic data tasks classified as being processed only on the CPU 640 are placed in the CPU queue 620, synthetic data tasks classified as being processed only on the GPU 650 are placed in the GPU queue 622, and synthetic data tasks classified as being processed on the CPU 640 or GPU 650 are placed in the hybrid queue 624. It should be understood that the use of queues and / or task properties is merely an example implementation, and other techniques can also be used to indicate which (which) resources a task should be assigned to.

[0099] In Figure 6In the illustrated embodiment, the scheduler 630 selects synthetic data tasks from the CPU queue 620, the GPU queue 622, and the hybrid queue 624, and routes or otherwise assigns the tasks to resources (e.g., the CPU 640 or the GPU 650) for processing. Various criteria can be used to indicate which resource should be assigned. For example, resources can be assigned based on task classification (e.g., tasks classified as ingestion tasks can be assigned to the CPU 640), based on queuing (e.g., based on FIFO, tasks from the CPU queue 620 can be assigned to the CPU 640), based on a priority scheme (e.g., assigning priority to a particular account or queue), based on scheduling (e.g., changing the assignment according to time such as day, week, year, etc.), based on resource availability, based on some other criteria, or some combination thereof. The CPU 640 and the GPU 650 can process the assigned and / or routed tasks in any suitable manner.

[0100] In some embodiments, the load monitor 660 receives telemetry or some other health or status signal from the CPU 640 and / or the GPU 650, and the telemetry or some other health or status signal provides an indication of resource availability. In this way, the load monitor 660 can determine when the CPU or GPU is available. The load monitor 660 can provide an indication to the scheduler 630 that a particular resource is available, and the scheduler 630 can use this information when determining which resource to assign a particular task. For example, the load monitor 660 can determine when the GPU is available (e.g., when the GPU enters a standby or suspended mode), and can provide a corresponding signal or other indication to the scheduler 630. Thus, the scheduler 630 can route or otherwise assign a particular task (e.g., a classified and / or queued task) to the available GPU for execution. The tasks to be routed or assigned can be selected based on various criteria (e.g., from a queue), and the various criteria include classification of eligible resources for processing the task (e.g., only the GPU), based on queue (e.g., assigning priority to the hybrid queue 624 or the GPU queue 622), based on task category (e.g., assigning priority to requests for creating synthetic data), based on the time the request is received (e.g., assigning priority to the task that has been queued earliest across multiple queues), based on queue fill (e.g., assigning priority to the most filled queue), based on the estimated completion time (e.g., assigning priority to the queued task estimated to take the longest or shortest time to execute), etc. These and other variations for assigning tasks to a particular resource are possible. In this way, tasks in synthetic data creation can be routed or otherwise assigned to a particular resource for processing.

[0101] Progress portal for synthetic data tasks

[0102] Figure 7FIG. illustrates an example environment 700 in which requests to perform synthetic data tasks can be managed. The components of environment 700 can operate within a distributed computing system to provide functionality for the SDaaS described herein. Environment 700 supports processing synthetic data tasks such as the generation or specification of source assets (e.g., 3D models), the ingestion of source assets, simulation (e.g., identification and variation of first layer and / or second layer parameters), the creation of synthetic data assets (e.g., rendering images based on 3D models), etc. The components of environment 700 can correspond to the components of the environment 600 in Figure 6 For example, Figure 7 the task classifier 710, CPU queue 720, GPU queue 722, hybrid queue 724, scheduler 730, and load monitor 760 of Figure 6 can correspond respectively to the task classifier 610, CPU queue 620, GPU queue 622, hybrid queue 624, scheduler 630, and load monitor 660 of

[0103] Although the embodiments herein are described using CPUs and GPUs, environment 700 can include any suitable processor / architecture (e.g., CPU, GPU, ASIC, FPGA, etc.) and corresponding queues.

[0104] In Figure 7In the illustrated embodiment, the progress portal 770 includes a feedback component 772 and a command component 774. The progress portal 770 is communicatively coupled to components in the management layer and / or the CPU 740 and GPU 750. In this way, the feedback component can monitor the progress of a particular synthetic data request and corresponding synthetic data task by means of the management layer and by executing on the CPU 740 and GPU 750. The feedback component 772 can present (or otherwise cause to be presented on the client device 705) information regarding the processing of the synthetic data request. In other types of feedback, the feedback component 772 can present an indication of the overall progress of the request processing (e.g., percentage complete, number of synthetic data assets created, number of synthetic data assets remaining to be created, etc.), CPU / GPU usage, average time taken to create synthetic data on the CPU, average time taken to create synthetic data on the GPU, remaining estimated time (e.g., based on existing or estimated creation patterns), corresponding service cost (e.g., incurred, estimated, or predicted for completing the processing request), or other information. In this way, a user or account that issues a particular synthetic data request can monitor the progress of the request processing.

[0105] In some embodiments, the command component 774 can communicate with a command interface associated with the client device 705. A user can use the command interface to input commands, and the commands can be transmitted to the command component 774 for execution. The command component 774 can accept and execute various commands, including commands for allocating resources to process a synthetic data request, batch creating synthetic assets, or processing some other synthetic data request, changing or auto-aiming for a particular outcome (e.g., processing time, cost, output quality, etc.). The command component 774 can support the execution of commands even after the request processing has started (mid-request).

[0106] For example, the command component 774 can accept and execute commands for allocating resources to process synthetic data requests. As a non-limiting example, processing a request to create a large number of synthetic data assets may take several hours. While the feedback component 772 can provide information about request processing (progress, allocated resources, resource consumption, etc.), the command component 774 can allow the user to influence the way the request is processed. For example, assume that a request to generate 50k assets has started processing and 10k assets have been generated using the CPU 740. The user may prefer to speed up the processing. In this case, the user can use the GPU 750 to issue a command via the command component 774 to generate the remaining synthetic data assets. The command component 774 can execute such a command in various ways. For example, the command component 774 can identify the corresponding synthetic data tasks that have not been routed or allocated to the CPU 740 (e.g., by identifying the synthetic data tasks queued in the CPU queue 720 and / or the hybrid queue 724), and can reclassify the synthetic data tasks to be processed on the GPU 750 (e.g., by moving the synthetic data tasks to the GPU queue 722). In this way, the scheduler 730 can pick up the synthetic data tasks from the GPU queue 722 and route or otherwise allocate them to be processed on the GPU 750. In this manner, the user can manually change the resource allocation (e.g., task classification) during runtime by issuing commands via the progress portal 770.

[0107] In some embodiments, the command component 774 can support commands that change or aim for specific outcomes. For example, the command interface can allow the user to issue commands to the command component 774 to reduce or increase (or minimize or maximize) the resulting processing characteristics, such as processing time, cost, output quality, etc. In this way, the command interface can trigger the command component 774 to reduce and / or minimize the processing time (e.g., by allocating synthetic data tasks to the GPU 750 during runtime), reduce and / or minimize the cost (e.g., by allocating tasks to be executed on a relatively inexpensive processor such as the CPU, by automatically reducing the size or quality of the generated synthetic data assets, etc.), and so on. In some embodiments, the user can issue commands that directly set the output quality (e.g., the size or quality of the generated synthetic data assets). Additionally or alternatively, consuming a threshold amount of resources (and / or incurring a threshold cost) can trigger the automatic stopping of further processing of a request, task, its specific classification, etc. The threshold can be predefined (e.g., by a service level agreement, by an incoming request, etc.) and can be changed during runtime (e.g., by a received command). In some embodiments, commands can be issued to stop further processing of a request, task, its specific classification, etc. These are just a few examples of the possible commands that can be implemented.

[0108] Example Flowchart

[0109] Reference Figures 8 - 14 , a flowchart is provided that illustrates a method for implementing synthetic data as a service in a distributed computing system. The method can be executed using the distributed computing system described herein. In an embodiment, one or more computer storage media have computer-executable instructions embodied thereon that, when executed by one or more processors, can cause the one or more processors to execute the method in the distributed computing system 100.

[0110] Figure 8 is a flowchart that illustrates a process 800 for implementing an asset assembly engine in a distributed computing system. Initially, at block 810, a first source asset is received from a first distributed synthetic data as a service (SDaaS) upload interface. At block 820, a second source asset is received from a second distributed SDaaS upload interface. At block 830, the first source asset and the second source asset are ingested. Ingesting the source assets includes automatically calculating values of asset change parameters for the source assets, where the asset change parameters can be programmed for machine learning. At block 840, a first synthetic data asset including a first set of values of the asset change parameters is generated. At block 850, a second synthetic data asset including a second set of values of the asset change parameters is generated. At block 860, the first synthetic data asset and the second synthetic data asset are stored in a synthetic data asset repository.

[0111] Figure 9 is a flowchart that illustrates a process 900 for implementing a scenario assembly engine in a distributed computing system. Initially, at block 910, selections of a first synthetic data asset and a second synthetic data asset are received from a distributed synthetic data as a service (SDaaS) integrated development environment (IDE). The synthetic data assets are associated with asset change parameters and scenario change parameters, where the asset change parameters and the scenario change parameters can be programmed for machine learning. At block 920, values for generating a synthetic data scenario are generated. The values correspond to the asset change parameters or the scenario change parameters. At block 930, based on the values, a synthetic data scenario is generated using the first synthetic data asset and the second synthetic data asset.

[0112] Figure 10FIG. is a flowchart illustrating a process 1000 for implementing a distributed computing system frame set assembly engine according to an embodiment. Initially, at block 1010, a synthetic data scene is accessed. At block 1020, a first set of values of scene change parameters is determined. The first set of values is automatically determined to generate a synthetic data scene frame set. At block 1030, the synthetic data scene frame set is generated based on the first set of values. The synthetic data scene frame set includes at least a first frame in the frame set, and the first frame includes a synthetic data scene updated based on the value of the scene change parameter. At block 1040, the synthetic data scene frame set is stored.

[0113] Figure 11 FIG. is a flowchart illustrating a process 1100 for implementing a distributed computing frame set package generator according to an embodiment. At block 1110, a frame set package generator configuration file is accessed. The frame set package generator configuration file is associated with a first image generation device. The frame set package generator configuration file includes known device variability parameters associated with the first image generation device. At block 1120, a frame set package is generated based on the frame set package generator configuration file. The frame set package generator configuration file includes values of the known device variability parameters. At block 1130, the frame set package is stored.

[0114] Figure 12 FIG. is a flowchart illustrating a process 1200 for implementing a distributed computing system frame set package repository according to an embodiment. At block 1210, a query for a frame set package is received from a frame set package query interface. The frame set query interface includes a plurality of frame set package categories. At block 1220, a query result frame set package is identified based on the frame set package configuration file. At block 1230, the query result frame set package is transmitted.

[0115] Figure 13 FIG. is a flowchart illustrating a process 1300 for implementing a distributed computing system feedback loop engine according to an embodiment. At block 1310, a training data set report is accessed. The training data set report identifies synthetic data assets having asset change parameter values. The synthetic data assets are associated with a frame set. At block 1320, based on the training data set report, the synthetic data assets with synthetic data asset changes are updated. At block 1330, the frame set is updated using the updated synthetic data assets.

[0116] Figure 14FIG. is a flowchart illustrating a process 1400 for implementing a crowdsourcing engine for a distributed computing system according to an embodiment. At block 1410, a source asset is received from a distributed synthetic data as a service (SDaaS) crowdsourcing interface. At block 1420, crowdsourcing tags for the source asset are received via the distributed SDaaS crowdsourcing interface. At block 1430, the source asset is ingested, partially based on the crowdsourcing tags. Ingesting the source asset includes automatically calculating a value of an asset change parameter for the source asset. The asset change parameter is programmable for machine learning. At block 1440, a crowdsourced synthetic data asset including the asset change parameter value is generated.

[0117] Figure 15 FIG. is a flowchart illustrating a process 1500 for suggesting change parameters according to an embodiment of the present invention. At block 1510, a selection of a machine learning scenario from a plurality of machine learning scenarios is received from a distributed synthetic data as a service (SDaaS) interface. At block 1520, a subset of change parameters associated with the selected machine learning scenario is retrieved from a seed classification criterion that associates the plurality of machine learning scenarios with corresponding subsets of a plurality of change parameters. At block 1530, the subset of change parameters is presented as the suggested change parameters for the selected machine learning scenario on the SDaaS interface. At block 1540, a selection of a change parameter from the plurality of change parameters and a request to generate an associated frame set are received from the SDaaS interface. At block 1550, a corresponding frame set package is generated with frames that have the selected change parameter changed.

[0118] Figure 16 FIG. is a flowchart illustrating a process 1600 for pre-packaging a synthetic data set according to an embodiment of the present invention. At block 1610, a synthetic data set customized for training a first machine learning scenario is pre-packaged. At block 1620, a selection of the first machine learning scenario from a plurality of machine learning scenarios is received from a distributed synthetic data as a service (SDaaS) interface. At block 1630, an indication of the availability of the pre-packaged synthetic data set for the selected machine learning scenario is presented on the SDaaS interface. At block 1640, a selection to download the pre-packaged synthetic data set is received via the SDaaS interface. At block 1650, access to the pre-packaged synthetic data set is provided.

[0119] Figure 17FIG. is a flowchart illustrating a process 1700 for coordinating resource allocation according to an embodiment of the present invention. At block 1710, a request to execute a synthetic data task is received from a distributed synthetic data as a service (SDaaS) interface. At block 1720, resource allocation is coordinated by identifying a corresponding category of eligible resources for processing the synthetic data task and routing the synthetic data task to the first resource in the eligible resource category. At block 1730, the synthetic data task is executed on the first resource.

[0120] Figure 18 FIG. is a flowchart illustrating a process 1800 for coordinating resource allocation between processor groups having different architectures according to an embodiment of the present invention. At block 1810, a request to execute a synthetic data task is received from a distributed synthetic data as a service (SDaaS) interface. At block 1820, resource allocation between processor groups having different architectures is coordinated by identifying corresponding eligible processor groups for executing the synthetic data task from each group, and the corresponding eligible processor groups allocate the synthetic data task to the first processor in the eligible processor group. At block 1830, the synthetic data task is executed on the first processor.

[0121] Figure 19 FIG. is a flowchart illustrating a process 1900 for changing resource allocation according to an embodiment of the present invention. At block 1910, a progress portal is presented, which is configured to monitor the progress of the distributed synthetic data as a service (SDaaS) in processing requests associated with executing synthetic data tasks. At block 1920, a command to change the resource allocation for processing the synthetic data task is received via the command interface of the progress portal. At block 1930, the command is executed by: identifying a subset of synthetic data tasks classified as being processed using the first resource and not yet routed for processing from the synthetic data task, and reclassifying the subset of the synthetic data task to be processed on the second resource.

[0122] Figure 20 FIG. is a flowchart illustrating a process 2000 for batch processing synthetic data tasks according to an embodiment of the present invention. At block 2010, a progress portal is presented, which is configured to monitor the progress of the SDaaS in processing requests for executing synthetic data tasks. At block 2020, a command for batch processing synthetic data tasks is received via the command interface of the progress portal. At block 2030, the command is executed by deriving a plurality of batch-processed synthetic data tasks from the requests for executing synthetic data tasks and classifying the plurality of batch-processed synthetic data tasks for processing.

[0123] Figure 21FIG. 2100 is a flowchart illustrating a process 2100 for retraining a machine learning model according to an embodiment of the present invention. At block 2110, a machine learning model is accessed. At block 2120, a plurality of synthetic data assets are accessed, where the synthetic data assets are associated with asset variation parameters programmable for machine learning. At block 2130, the machine learning model is retrained using the plurality of synthetic data assets.

[0124] Advantageously, the embodiments described herein improve the computational capabilities and operations for generating training datasets using a distributed computing system based on providing synthetic data as a service. In particular, the improvements to the computational capabilities and operations are associated with a distributed infrastructure for large-scale production training datasets based on SDaaS operations. For example, based on SDaaS operations, the computational operations required for manually developing (e.g., labeling and tagging) and refining (e.g., searching) training datasets are avoided, and SDaaS operations use synthetic data assets to automatically develop training datasets and automatically refine training datasets based on training dataset reports that indicate additional synthetic data assets or scenarios that will improve the machine learning model in a machine learning training service.

[0125] In addition, the storage and retrieval of training datasets are improved using an internal machine learning training service operating in the same distributed computing system, thereby reducing computational overhead. SDaaS operations are implemented based on an unconventional arrangement of engines and an unconventional set of rules defined by an ordered combination of steps for an SDaaS system. In this regard, SDaaS addresses the problems caused by manually developing machine learning training datasets and improves the existing process of training machine learning models in a distributed computing system. Overall, these improvements also result in less CPU computation, smaller memory requirements, and greater flexibility in generating and utilizing machine learning training datasets.

[0126] Example embodiments of the present invention

[0127] Accordingly, an example embodiment of the present invention provides a distributed computing system asset assembly engine. The asset assembly engine is configured to receive a first source asset from a first distributed synthetic data as a service (SDaaS) upload interface. The asset assembly engine is further configured to receive a second source asset from a second distributed SDaaS upload interface. The asset assembly engine is further configured to ingest the first source asset and the second source asset. Ingesting the source assets includes automatically calculating values of asset change parameters for the source assets. The asset change parameters are programmable for machine learning. The asset assembly engine is further configured to generate a first synthetic data asset, the first synthetic data asset including a first set of values for the asset change parameters. The asset assembly engine is further configured to generate a second synthetic data asset, the second synthetic data asset including a second set of values for the asset change parameters. The asset assembly engine is further configured to store the first synthetic data asset and the second synthetic data asset in a synthetic data asset repository.

[0128] Another example embodiment of the present invention provides a distributed computing system scenario assembly engine. The scenario assembly engine is configured to receive a selection of a first synthetic data asset and a selection of a second synthetic data asset from a distributed synthetic data as a service (SDaaS) integrated development environment (IDE). The synthetic data assets are associated with asset change parameters and scenario change parameters. The asset change parameters and the scenario change parameters are programmable for machine learning. The engine is further configured to receive values for generating a synthetic data scenario. The values correspond to the asset change parameters or the scenario change parameters. The scenario assembly engine is further configured to generate a synthetic data scenario based on the values, using the first synthetic data asset and the second synthetic data asset.

[0129] Another example embodiment of the present invention provides a distributed computing system frame set assembly engine. The frame set assembly engine is configured to access a synthetic data scenario. The frame set assembly engine is further configured to determine a first set of values for the scenario change parameters. The first set of values is automatically determined to generate a synthetic data scenario frame set. The frame set assembly engine is further configured to generate a synthetic data scenario frame set based on the first set of values. The synthetic data scenario frame set includes at least a first frame in the frame set, the first frame including the synthetic data scenario updated based on the values for the scenario change parameters. The frame set assembly engine is further configured to store the synthetic data scenario frame set.

[0130] Another example embodiment of the present invention provides a distributed computing system frame set package generator. The frame set package generator is configured to access a frame set package generator configuration file. The frame set package generator configuration file is associated with a first image generation device. The frame set package generator configuration file includes known device variability parameters associated with the first image generation device. The frame set package generator is also configured to generate a frame set package based on the frame set package generator configuration file. The frame set package generator configuration file includes values of the known device variability parameters. The frame set package generator is also configured to store the frame set package.

[0131] Another example embodiment of the present invention provides a distributed computing system frame set package repository. The frame set package repository is configured to receive a query for a frame set package from a frame set package query interface. The frame query interface includes a plurality of frame set package categories. The engine is also configured to identify a query result frame set package based on the frame set package configuration file. The engine is also configured to transmit the query result frame set package.

[0132] Another example embodiment of the present invention provides a distributed computing system feedback loop engine. The feedback loop engine is configured to access a training data set report. The training data set report identifies synthetic data assets having asset change parameter values. The synthetic data assets are associated with a frame set. The feedback loop engine is also configured to update the synthetic data assets based on the training data set report, using the synthetic data asset changes. The feedback engine is also configured to update the frame set using the updated synthetic data assets.

[0133] Another example embodiment of the present invention provides a distributed computing system crowdsourcing engine. The crowdsourcing engine is configured to receive source assets from a distributed synthetic data as a service (SDaaS) crowdsourcing interface. The crowdsourcing engine is also configured to receive crowdsourcing tags for the source assets via the distributed SDaaS crowdsourcing interface. The crowdsourcing engine is also configured to ingest the source assets, at least in part, based on the crowdsourcing tags. Ingesting the source assets includes automatically calculating values of asset change parameters for the source assets. The asset change parameters are programmable for machine learning. The crowdsourcing engine is also configured to generate crowdsourced synthetic data assets that include the values of the asset change parameters.

[0134] Example Distributed Computing Environment

[0135] Now refer to Figure 22 , Figure 22 FIG. illustrates an example distributed computing environment 2200 in which implementations of the present disclosure may be employed. In particular, Figure 22Shows the high-level architecture of synthetic data as a service in a distributed computing system in a cloud computing platform 2210, where the system supports seamless modification of software components. It should be understood that this and other arrangements described herein are presented only as examples. For example, as described above, many of the elements described herein can be implemented as discrete components or distributed components or in combination with other components, and implemented in any suitable combination and location. Other arrangements and elements (e.g., machines, interfaces, functions, command and function groupings, etc.) can be used in addition to or instead of the arrangements and elements shown.

[0136] The data center can support a distributed computing environment 2200, which includes a cloud computing platform 2210, racks 2220, and nodes 2230 (e.g., computing devices, processing units, or blades) in the racks 2220. The system can be implemented using a cloud computing platform 2210 that runs cloud services across different data centers and geographical regions. The cloud computing platform 2210 can implement a fabric controller 2240 component to supply and manage resource allocation, deployment, upgrade, and management of cloud services. Generally, the cloud computing platform 2210 is used to store data or run service applications in a distributed manner. The cloud computing infrastructure 2210 in the data center can be configured to host and support the operation of endpoints of specific service applications. The cloud computing infrastructure 2210 can be a public cloud, a private cloud, or a dedicated cloud.

[0137] The nodes 2230 can be configured using a host 2250 (e.g., an operating system or a runtime environment) that runs a defined software stack on the nodes 2230. The nodes 2230 can also be configured to perform specialized functionality (e.g., computing nodes or storage nodes) within the cloud computing platform 2210. The nodes 2230 are assigned to run one or more parts of a tenant's service application. A tenant can refer to a customer who utilizes the resources of the cloud computing platform 2210. The service application components of the cloud computing platform 2210 that support a specific tenant can be referred to as tenant infrastructure or tenancy. The terms service application, application, or service are used interchangeably herein and generally refer to any software or part of software that runs on top of the data center or accesses the storage devices and computing device locations within the data center.

[0138] When node 2230 supports more than one individual service application, node 2230 can be partitioned into virtual machines (e.g., virtual machine 2252 and virtual machine 2254). A physical machine can also run individual service applications simultaneously. The virtual machines or physical machines can be configured as personalized computing environments supported by resources 2260 (e.g., hardware resources and software resources) in cloud computing platform 2210. It is expected that the resources can be configured for a specific service application. Additionally, each service application can be divided into functional parts such that each functional part can run on a separate virtual machine. In cloud computing platform 2210, multiple servers can be used to run service applications and perform data storage operations in a cluster. In particular, the servers can perform data operations independently, but are exposed as a single device called a cluster. Each server in the cluster can be implemented as a node.

[0139] Client device 2280 can be linked to a service application in cloud computing platform 2210. Client device 2280 can be any type of computing device, which can correspond to, for example, computing device 2200 described with reference to Figure 22 Client device 2280 can be configured to issue commands to cloud computing platform 2210. In an embodiment, client device 2280 can communicate with the service application by means of a virtual Internet Protocol (IP) and a load balancer or other means of directing communication requests to a specified endpoint in cloud computing platform 2210. Components of cloud computing platform 2210 can communicate with each other via a network (not shown), which can include but is not limited to one or more local area networks (LANs) and / or wide area networks (WANs).

[0140] Exemplary Computing Environment

[0141] After briefly describing an overview of embodiments of the present invention, the following describes an exemplary operating environment in which embodiments of the present invention can be implemented to provide a general context for various aspects of the present invention. First, with particular reference to Figure 23 , an exemplary operating environment for implementing embodiments of the present invention is shown and generally designated as computing device 2300. Computing device 2300 is only one example of a suitable computing environment and is not intended to imply any limitation as to the scope of use or functionality of the present invention. Computing device 2300 should also not be construed as having any relevance or requirement related to any of the components shown or in combination.

[0142] The present invention may be described in the general context of computer code or machine - usable instructions, which include computer - executable instructions, such as program modules, executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present invention may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general - purpose computers, more specialized computing devices, etc. The present invention may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked by a communications network.

[0143] Reference Figure 23 , computing device 2300 includes a bus 2310 that directly or indirectly couples the following devices: a memory 2312, one or more processors 2314, one or more presentation components 2316, an input / output port 2318, input / output components 2320, and an illustrative power supply 2322. Bus 2310 represents one or more buses, such as an address bus, a data bus, or a combination thereof. For conceptual clarity, Figure 23 the various boxes of [] are shown with lines, and other arrangements of the described components and / or component functionality are also contemplated. For example, a presentation component, such as a display device, may be considered an I / O component. Also, a processor has a memory. We recognize this as being in the nature of the art, and reiterate that Figure 23 the figures of [] merely illustrate one exemplary computing device that may be used in conjunction with one or more embodiments of the present invention. No distinction is made among the categories of "workstation", "server", "laptop computer", "handheld device", etc., as all of these are contemplated within the scope of [] and refer to a "computing device". Figure 23

[0144] Computing device 2300 generally includes a variety of computer - readable media. Computer - readable media can be any available media that can be accessed by computing device 2300 and includes both volatile and non - volatile media, removable and non - removable media. By way of example and not limitation, computer - readable media may include computer - storage media and communication media.

[0145] Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage devices, magnetic tape cartridges, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store the required information and can be accessed by a computing device 2300. Computer storage media does not itself include signals.

[0146] Communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transmission mechanism, and includes any information delivery media. The term "modulated data signal" refers to a signal having one or more sets of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of any of the above are also to be included within the scope of computer-readable media.

[0147] Memory 2312 includes computer storage media in the form of volatile and / or non-volatile memory. The memory can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid state storage devices, hard disk drive devices, optical disk drive devices, etc. Computing device 2300 includes one or more processors that read data from various entities such as memory 2312 or I / O components 2320. The (multiple) presentation components 2316 present data indications to a user or other device. Exemplary presentation components include display devices, speakers, printing components, vibration components, etc.

[0148] I / O port 2318 allows computing device 2300 to be logically coupled to other devices including I / O components 2320, some of which may be built-in. Exemplary components include microphones, joysticks, gamepads, satellite dishes, scanners, printers, wireless devices, etc.

[0149] Referring to distributed computing system synthetic data as a service, the distributed computing system synthetic data as a service component refers to the integrated components for providing synthetic data as a service. The integrated components refer to the hardware architecture and software framework that support the functionality within the system. The hardware architecture refers to the physical components and their interrelationships, and the software framework refers to the software that provides functionality that can be implemented using the hardware embodied on the device.

[0150] A system based on end-to-end software can operate within system components to operate computer hardware to provide system functionality. At a lower level, a hardware processor executes instructions selected from the machine language (also known as machine code or native) instruction set of a given processor. The processor recognizes native instructions and performs corresponding low-level functions related to, for example, logical, control, and memory operations. Low-level software written in machine code can provide more complex functionality for high-level software. As used herein, computer-executable instructions include any software, including low-level software written in machine code, higher-level software such as application software, and any combination thereof. In this regard, system components can manage resources and provide services for system functionality. For embodiments of the present invention, any other variations and combinations thereof can be expected.

[0151] As an example, a distributed computing system synthetic data as a service can include an API library that includes specifications referring to routines, data structures, object classes, and variables that can support the interaction between the hardware architecture of a device and the software framework of the distributed computing system synthetic data as a service. These APIs include configuration specifications for the distributed computing system synthetic data as a service such that different components therein can communicate with each other in the distributed computing system synthetic data as a service as described herein.

[0152] Having identified the various components used herein, it should be understood that within the scope of the present disclosure, any number of components and arrangements can be employed to achieve the desired functionality. For example, for conceptual clarity, the components in the embodiments depicted in the figures are shown by lines. Other arrangements of these components and other components can also be implemented. For example, although some components are depicted as single components, many of the elements described herein can be implemented as discrete components or distributed components or in combination with other components and in any suitable combination and location. Certain elements can be completely omitted. Additionally, as described below, the various functions described herein as being performed by one or more entities can be performed by hardware, firmware, and / or software. For example, the various functions can be performed by a processor executing instructions stored in a memory. Thus, other arrangements and elements (e.g., machines, interfaces, functions, command and function groupings, etc.) can be used in addition to or in place of the arrangements and elements shown.

[0153] The embodiments described in the following paragraphs can be combined with one or more of the specifically described alternatives. In particular, the claimed embodiments can alternatively include references to more than one other embodiment. The claimed embodiments can specify further limitations of the claimed subject matter.

[0154] The subject matter of the embodiments of the present invention is specifically described herein to meet statutory requirements. However, the specification itself is not intended to limit the scope of this patent. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, in combination with other current or future technologies, to include steps or combinations of steps different from those described in this document. Additionally, although the terms "step" and / or "block" may be used herein to represent different elements of the methods employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein unless the order of individual steps is explicitly described.

[0155] For the purposes of this disclosure, the word "including" has the same broad meaning as the word "comprising", and the word "access" includes "receiving", "referencing", or "retrieving". Additionally, the word "communicating" has the same broad meaning as "receiving" or "transmitting" using the communication media described herein by a software- or hardware-based bus, transmitter, or receiver. Further, unless otherwise stated to the contrary, words such as "a" and "an" include the plural as well as the singular. Thus, for example, in the presence of one or more features, the constraint of "feature" is satisfied. Similarly, the term "or" includes conjunctive, disjunctive, and both (thus a or b includes a or b as well as a and b).

[0156] For the purposes of the detailed discussion above, embodiments of the present invention have been described with reference to a distributed computing environment; however, the distributed computing environment described herein is merely exemplary. Components may be configured to perform novel aspects of the embodiments, where the term "configured to" may refer to "programmed to" perform a particular task or implement a particular abstract data type using code. Additionally, although embodiments of the present invention may generally refer to synthetic data as a service and the diagrams described herein in the context of a distributed computing system, it should be understood that the techniques described may be extended to other implementation contexts.

[0157] Embodiments of the present invention have been described with respect to specific embodiments, which are intended in all respects to be illustrative rather than restrictive. Alternative embodiments will become apparent to those of ordinary skill in the art to which the present invention pertains without departing from the scope of the present invention.

[0158] From the foregoing, it can be seen that the present invention is well-suited to achieve all of the purposes and objectives set forth above, as well as other obvious advantages and advantages inherent in the structure.

[0159] It will be understood that certain features and sub-combinations are useful and may be employed without reference to other features or sub-combinations. This is contemplated by the claims and is within the scope of the claims.

Claims

1. A computer system, comprising: one or more hardware processors and a memory, the memory being configured to provide computer program instructions to the one or more hardware processors; a task classifier, configured to use the one or more hardware processors to: receive a request representing a task from an interface of a distributed computing environment, the distributed computing environment being configured to generate a training data set including synthetic images, the task supporting the generation of the training data set; identify a task category of the task from a support category group, the support category group including the generation of a 3D model or a 3D scene including the 3D model or ingestion into the distributed computing environment, the simulation of a plurality of parameters representing different variations of the 3D scene, and the rendering of the synthetic images from the different variations of the 3D scene; identify an eligible processor group by classifying the task as eligible for execution only on the CPU of the distributed computing environment, eligible for execution only on the GPU of the distributed computing environment, or eligible for execution on the CPU or the GPU based on the identified task category; wherein the task is the ingestion of the 3D model into the distributed computing environment, and the task classifier is configured to identify the eligible processor group by classifying the task as eligible for processing only on the CPU based on the identified task category being the ingestion of the 3D model; and a scheduler, configured to use the one or more hardware processors to route the task to a first processor in the eligible processor group.

2. The computer system according to claim 1, wherein the simulation includes the identification of the plurality of parameters or the variation of the plurality of parameters.

3. The computer system according to claim 1, wherein the task classifier is further configured to place the task in a queue corresponding to the eligible processor group, and wherein the scheduler is further configured to select the task from the queue for allocation by assigning a priority to a request for rendering the synthetic image.

4. The computer system according to claim 1, wherein the task classifier is further configured to derive a plurality of tasks for batching the rendering of the synthetic image from the request.

5. The computer system according to claim 1, wherein the scheduler is further configured to route the task based on a monitoring signal indicating the availability of the first processor.

6. One or more computer storage media storing computer-usable instructions that, when used by one or more computing devices, cause the one or more computing devices to perform operations, the operations including: receiving a request representing a task from an interface of a distributed computing environment, the distributed computing environment being configured to use a service-oriented architecture of the distributed computing environment to generate a training data set including synthetic images, the task supporting the generation of the training data set; coordinating resource allocation by the service-oriented architecture by: From a support category group that identifies a task category of the task, the support category group including generation or ingestion of a 3D model or a 3D scene including the 3D model, simulation of a plurality of parameters representing different variations of the 3D scene, and rendering of the synthetic image from the different variations of the 3D scene; Identifying a category of eligible processors in the distributed computing environment by classifying the task as eligible for execution only on a CPU of the distributed computing environment, eligible for execution only on a GPU of the distributed computing environment, or eligible for execution on the CPU or the GPU based on the identified task category, wherein identifying the category of eligible processors includes: Identifying the task category as the ingestion of the 3D model into the distributed computing environment; and Based on the identified task category being the ingestion of the 3D model, classifying the task as eligible for execution only on the CPU of the distributed computing environment; Routing the task to a first processor in the category of eligible processors; and Causing the task to be executed on the first processor.

7. One or more computer storage media according to claim 6, wherein the simulation includes identification of the plurality of parameters or variation of the plurality of parameters.

8. One or more computer storage media according to claim 6, wherein causing the task to be executed on the first processor includes: Queuing the task in a queue corresponding to the category of the eligible processor, and selecting the task from the queue for assignment by assigning a priority to a request to render the synthetic image.

9. One or more computer storage media according to claim 6, the operations further including routing the task based on a monitoring signal indicating availability of the first processor.

10. One or more computer storage media according to claim 6, the operations further include: Deriving from the request a plurality of tasks that batch the rendering of the synthetic image.

11. One or more computer storage media storing computer-usable instructions that, when used by one or more computing devices, cause the one or more computing devices to perform operations, the operations include: Receiving a request representing a task from an interface of a distributed computing environment configured to generate a training data set including a synthetic image using a service-oriented architecture of the distributed computing environment, the task supporting generation of the training data set; Coordinating resource allocation by the service-oriented architecture by: Identifying a task category of the task from a support category group that includes generation or ingestion of a 3D model or a 3D scene including the 3D model, simulation of a plurality of parameters representing different variations of the 3D scene, and rendering of the synthetic image from the different variations of the 3D scene; Identifying a class of eligible processors in the distributed computing environment by classifying the task as eligible for execution only on a CPU of the distributed computing environment, eligible for execution only on a GPU of the distributed computing environment, or eligible for execution on the CPU or the GPU, wherein identifying the class of eligible processors includes: Identifying the task class as the rendering of the synthetic image of the training dataset for the different variations of the 3D scene; and Based on the identified task class being the rendering of the synthetic image of the different variations of the 3D scene, classifying the task as eligible for execution only on the GPU of the distributed computing environment; Routing the task to a first processor in the class of eligible processors; and Causing the task to be executed on the first processor.

12. The one or more computer storage media according to claim 11, wherein the simulation includes the identification of the plurality of parameters, or the variation of the plurality of parameters.

13. The one or more computer storage media according to claim 11, wherein causing the task to be executed on the first processor includes: Queuing the task in a queue corresponding to the class of the eligible processor, and selecting the task from the queue for allocation by assigning a priority to a request for rendering the synthetic image.

14. The one or more computer storage media according to claim 11, the operation further includes routing the task based on a monitoring signal indicating the availability of the first processor.

15. The one or more computer storage media according to claim 11, the operation further includes: Deriving a plurality of tasks for batching the rendering of the synthetic image from the request.

16. A method for coordinating resource allocation, includes: Receiving a request representing a task from an interface of a distributed computing environment, the distributed computing environment being configured to generate a training dataset including synthetic images, the task supporting the generation of the training dataset; Coordinating resource allocation between a group of processors in the distributed computing environment having different architectures by: Identifying a task class of the task from a group of support classes, the group of support classes including the generation or capture of a 3D model or a 3D scene including the 3D model, the simulation of a plurality of parameters representing different variations of the 3D scene, and the rendering of the synthetic image from the different variations of the 3D scene; Identifying a group of eligible processors from the group of processors by classifying the task as eligible for execution only on a first type of processor, eligible for execution only on a second type of processor, or eligible for execution on the first type of processor or the second type of processor based on the identified task class, wherein the task is the rendering of the synthetic image from the different variations of the 3D scene, and identifying the eligible processor group includes classifying the task as eligible for processing only on the GPU based on the identified task category being the rendering of the synthetic image from the different variations of the 3D scene; assigning the task to a first processor in the eligible processor group of the distributed computing environment; and causing the task to be executed on the first processor.

17. The method according to claim 16, wherein the simulation includes the identification of the plurality of parameters, or the variation of the plurality of parameters.

18. The method according to claim 16, wherein the task is the ingestion of the 3D model, and identifying the eligible processor group includes classifying the task as eligible for processing only on the CPU based on the identified task category being the ingestion of the 3D model.

19. The method according to claim 16, wherein the first type of processor is a CPU, the second type of processor is a GPU, and the eligible processor group is selected from a plurality of groups including: a group of only CPUs, a group of only GPUs, and a group of CPUs and GPUs.

20. The method according to claim 16, further including: queuing the task in a queue corresponding to the eligible processor category; and selecting the task from the queue for assignment by assigning a priority to the request to render the synthetic image.

21. The method according to claim 16, further including: deriving a plurality of tasks for batching the rendering of the synthetic image from the request.