Pipeline for efficient training and introduction of machine learning models
The customizable pipeline architecture facilitates efficient training, retraining, and deployment of MLMs, addressing the challenge of domain-specific customization and resource-intensive reconfiguration, thereby improving model adaptability and performance for user-specific applications.
Patent Information
- Application Number
- JP2021147468
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-02-12
- Filing Date
- 2021-09-10
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2041-09-10
AI Technical Summary
Current machine learning model (MLM) pipelines are cumbersome and difficult to customize for different domains, requiring significant development effort and reconfiguration, making them sub-optimal for user-specific needs and resource-intensive to adapt.
A customizable pipeline architecture (CP) that allows users to efficiently train, retrain, and deploy MLMs using a pipeline orchestrator, training engine, and export engine, enabling seamless integration and configuration of pre-trained, retrained, or custom MLMs on user-specific platforms.
Enables efficient management and configuration of MLMs, allowing users to adapt models to their specific needs without advanced developer expertise, reducing resource requirements and enhancing model performance.
Smart Images

Figure 0007702314000001 
Figure 0007702314000002 
Figure 0007702314000003
Abstract
Description
Technical Field
[0001] At least one embodiment relates to processing resources used to perform and facilitate artificial intelligence. For example, at least one embodiment relates to provisioning a pipeline for efficiently training, configuring, deploying, and using machine learning models in a user-specific platform.
Background Art
[0002] Machine learning is often used in office and hospital environments, robotic automation, security applications, autonomous transportation, law enforcement, and many other settings. In particular, machine learning has applications in audio and video processing, such as in voice, speech, and object recognition. One prevalent approach to machine learning involves training a computing system using training data (sounds, images, and / or other data) to identify patterns in the data, which may facilitate data classification, such as the presence of a particular type of object in a training image or a particular word in a training voice. The training may be either supervised or unsupervised. Machine learning models can use various computational algorithms, such as decision tree algorithms (or other rule-based algorithms), artificial neural networks, and the like. During a subsequent deployment stage, also referred to as the "inference stage," new data is input into the trained machine learning model, and various target objects, sounds, or texts of interest can be identified using the patterns and features established during training.
Brief Description of the Drawings
[0003]
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7A
Figure 7B
Figure 8
Figure 9
Figure 10
DETAILED DESCRIPTION OF THE INVENTION
[0004] Machine learning has become a staple in many industries and activities where at least some levels of decision-making may be delegated to a computer system. Currently, machine learning models (MLMs) are being developed for specific target domains and applications. Since the purposes of various machine learning applications can be extremely diverse, MLMs may need to be set up, configured, and trained differently depending on the intended users of the trained MLMs. Models belonging to the same general type, for example, speech recognition models, may nonetheless be set up very differently in different use cases. For example, a speech recognition model designed for automated customer phone support may be different from a model developed for a hospital application, for example, to analyze patient diagnostic data or record a doctor's narrative in response to a patient's request. Moreover, even MLMs operating in the same target domain (e.g., the medical field) may need to be configured differently in different contexts. For example, a model designed to recognize speech in an operating room may need to be trained or configured differently from a model designed for an observation or recovery ward.
[0005] Currently, configuring an MLM for an application in a user-specific domain may require significant development effort. MLM developers may need to design the architecture of the model (e.g., in the case of a neural network MLM, the number of layers and the topology of node connections), train the MLM on relevant domain-specific training data, etc. In many cases, the MLM may be just a part of a larger codebase that involves various additional support stages of computation, such as pre-audio processing, audio artifact removal, filtering, post-processing, spectral Fourier analysis, and the like. Multiple MLMs may exist in the same computational pipeline, and developers may need to incorporate multiple MLMs, each providing distinct functionality, into a single computational pipeline. For example, developers of natural language processing applications may need to integrate feature extraction (using spectral analysis), an acoustic MLM (processing the extracted features), an acoustic post-processing module (removing artifacts, fillers, and stop words, etc.), language pre-processing (performing word tokenization, lemmatization, etc.), a language MLM (identifying topics, speaker intent, intonation of speech, etc.), language post-processing (performing rule-based correction / verification of the language MLM output), etc. In such a pipeline, the output of the acoustic MLM may be input to the language MLM, and the language MLM may further input data to an intent-identification model, and so on.
[0006] Currently, to integrate one or more MLMs with various support stages into a single computational pipeline or workflow, developers must create and manage code that encompasses the entire workflow. For example, developers may create platform-specific code such as a voice recognition MLM pipeline customized to meet the needs of a medical clinic. Such code can make it cumbersome and technically difficult to share and scale MLM applications outside of their original use cases. In particular, such pipelines may not be easily customizable to other domains or computer platforms. Specifically, another developer attempting to reconfigure the MLM of the pipeline for a different domain (e.g., from a stock trading company to an investment brokerage) may not only have to reconfigure the actual MLM, but may also have to re-engineer and overhaul a large portion of the overall code, even though some parts of the code may implement one or more standard modules of the pipeline (e.g., digital audio signal preprocessing). As a result, users who have access to the MLM pipeline but lack advanced developer expertise (e.g., customers) may not be able to customize the pipeline to their specific needs, for example, to reconfigure a natural language programming MLM pipeline for use in a different linguistic domain of interest to the user (e.g., sports broadcasting). Therefore, users may have to use an MLM pipeline that is sub-optimal when it comes to meeting the users' objectives. Alternatively, users may have to incur additional resources and hire specialized developers to configure the pipeline and, in some cases, retrain some or all of the MLMs of the pipeline.
[0007] Aspects and examples of the present disclosure address these and other difficulties of the art by describing methods and systems that enable efficient management and configuration of MLM pipelines. Implementations enable training and retraining of MLMs for user-specific target platforms, changing the parameters and architectures of previously trained MLMs, selecting and configuring previously trained MLMs, adapting the selected MLMs to user-specific needs, further integrating the selected MLMs into customizable workflows, introducing the customized workflows on user and cloud hardware, inputting actual inference data, reading, storing, and managing inference outputs, and so on. System Architecture
[0008] FIG. 1 is a block diagram of an example architecture of a customizable pipeline (CP) 100 that supports the training, configuration, and deployment of one or more machine learning models according to at least some embodiments. As depicted in FIG. 1, CP 100 may be implemented on a computing device 102, although it should be understood that any engine and component of computing device 102 may be implemented (or shared among) any number of computing devices or implemented in the cloud. Computing device 102 may be a desktop computer, laptop computer, smartphone, tablet computer, server, computing device accessing a remote server, computing device utilizing a virtualized computing environment, gaming console, wearable computer, smart TV, and the like. A user of CP 100 may have local or remote access to computing device 102 (e.g., via a network). Computing device 102 may have any number of central processing units (CPUs) and graphical processing units (GPUs) (not shown in FIG. 1), including virtual CPUs and / or virtual GPUs, or any other suitable processing device capable of executing the techniques described herein. Computing device 102 may further have any number of memory devices, network controllers, peripheral devices, and the like (not shown in FIG. 1). Peripheral devices may include a camera (e.g., a video camera) for capturing an image (or sequence of images), a microphone for capturing sound, a scanner, a sensor, or any other device for data data capture.
[0009] In some embodiments, CP100 may include several engines and components for an efficient MLM implementation. A user (customer, end-user, developer, data scientist, etc.) may interact with CP100 via a user interface UI104, which may include a command line, a graphical UI, a web-based interface (e.g., a web browser-accessible interface), a mobile application-based UI, or any combination thereof. UI104 may display menu, table, graph, flowchart, graphical and / or text representations of software, data, and workflows. UI104 may include selectable items that enable a user to input various pipeline settings and provide training / retraining and other data, as will be described in more detail below. User actions input via UI104 may be communicated to the pipeline orchestrator 110 of CP100 via the pipeline API 106. In some embodiments, prior to receiving pipeline data from the pipeline orchestrator 110, the user (or the remote computing device the user is using to access the pipeline) may download an API package to the remote computing device. The downloaded API package may be used to install the pipeline API 106 on the remote computing device so that the user can have two-way communication with the pipeline orchestrator 110 during the setup and use of CP100.
[0010] The pipeline orchestrator 110 may configure and introduce one or more MLMs via the pipeline API 106 and may provide various data to the user for use when using the introduced MLMs for processing (inference) of various input user data. For example, the pipeline orchestrator 110 may provide the user with information about available pre-trained MLMs and may enable retraining of pre-trained MLMs on user-specific data provided by the user or training of new (previously untrained) MLMs. The pipeline orchestrator 110 may then construct the CP100 based on the information received from the user. For example, the pipeline orchestrator 110 may configure the user-selected MLM and introduce the selected MLM along with various other (e.g., pre-processing and post-processing) stages used when implementing the selected MLM. To perform these and other tasks, the pipeline orchestrator 110 may coordinate and manage several engines, each of which implements a part of the overall pipeline functionality.
[0011] In some embodiments, CP100 may have access to one or more previously trained (pre-trained) MLMs, and thus may provide the user with access to at least some of these pre-trained MLMs (e.g., based on the user's subscription). The MLMs may be trained for common tasks in the area of CP specialization. For example, a CP specializing in speech processing may have access to one or more MLMs trained to recognize some general speech, such as customer service requests, general conversations, and the like. CP100 may further include a training engine 120. The training engine 120 may implement retraining (additional training) of the pre-trained MLMs. The retraining may be performed using retraining data tailored for the user-specific domain of use. In some embodiments, the retraining data may be provided by the user. For example, the user may provide retraining data to improve the natural language processing capabilities of one of the pre-trained MLMs to improve the recognition of speech that may be encountered in an investment brokerage environment or a securities trading environment. The data may be provided (e.g., by a technical expert at the user's financial company) in the form of digital recordings of audio in any available (compressed or uncompressed) digital format, such as WAV, WavPack, WMA, MP3, MPEG-4, as the sound track of video recordings, TV programs, and the like.
[0012] The pre-trained MLM 122 may be stored in a trained model repository 124 and may be accessible to the computing device 102 via the network 140. The pre-trained MLM 122 may be trained by a training server 162. The network 140 may be a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), a wireless network, a personal area network (PAN), or a combination thereof. In some embodiments, the training server 162 may be part of the computing device 102. In other embodiments, the training server 162 may be communicatively coupled to the computing device 102 directly or via the network 140. The training server 162 may be a rack-mounted server, a router computer, a personal computer, a laptop computer, a tablet computer, a desktop computer, a media center, or any combination thereof (and / or include the same). The training server 162 may include a training engine 160. The training engine 160 on the training server 162 may be the same as (or similar to) the training server 162 on the computing device 102. In some embodiments, there may be no training engine 120 on the computing device 102, and instead, all training and retraining may be performed by the training engine 160 on the training server 162. In some embodiments, the training engine 160 may perform off-site training of the pre-trained MLM 122, while the training engine 120 on the computing device 102 may perform retraining of the pre-trained MLM 122 and training of a new (custom) MLM 125.
[0013] During training or retraining, the training engine 160(120) may generate and configure one or more MLMs. The MLM may include a regression algorithm, a decision tree, a support vector machine, a k-means clustering model, a neural network, or any other machine learning algorithm. The neural network MLM may include a convolutional, recurrent, fully connected, long short-term memory model, a Hopfield, a Boltzmann, or any other type of neural network. Generating an MLM may include setting up the MLM type (e.g., neural network), architecture, number of layers of neurons, type of connections between the layers (e.g., fully connected, convolutional, deconvolutional, etc.), number of nodes within each layer, type of activation function used at various layers / nodes of the network, type of loss function used during training of the network, etc. Generating an MLM may include setting (e.g., randomly) the initial parameters (weights, biases) of the various nodes of the network. The generated MLM may be trained by the training engine 160 using training data that may include training inputs 165 and corresponding target outputs 167.
[0014] For example, for training the speech recognition MLM122, the training input 165 may include one or more digital recordings having utterances of words, phrases, and / or sentences that the MLM is trained to recognize. The target output 167 may include an indication of whether the target words and phrases are present in the training input 165. The target output 167 may also include, for example, a transcription of the utterance. In some embodiments, the target output 167 may include an identification of the speaker's intent. For example, a customer calling a food delivery service may convey a limited number of intents (such as placing an order, checking on the status of an order, canceling an order, etc.), but may do so in virtually unlimited ways. The particular words and sentences uttered may sometimes be less important, while the determination of intent may sometimes be important. Thus, in such embodiments, the target output 167 may include the correct category of intent. Similarly, for the training input 165 including a client's utterance in a customer service call, the target output 167 may be both a transcription of the utterance and an indication of the client's emotional state (such as anger, worry, satisfaction, etc.). Additionally, the training engine 160 may generate mapping data 166 (such as metadata) that associates the training input 165 with the correct target output 167. During training of the MLM122 (or custom MLM125), the training engine 160 (or 120) may identify patterns in the training input 165 based on the desired target output 167 and train each MLM to perform the desired task. The predictive utility of the identified patterns may then be verified in the inference stage, prior to being used, in future processing of new voices, using additional training input / target output associations. For example, upon receiving a new voice message, the trained MLM122 may be able to identify that the customer desires to check on the status of a previously placed order, confirm the customer's name, order number, etc.
[0015] In some embodiments, the plurality of MLMs may be trained simultaneously or separately. The speech recognition pipeline may involve a plurality of models, such as an acoustic model for audio processing, such as parsing speech into words, a language model for recognizing the parsed words, a model for intent identification, a model for understanding questions, or any other model. In some embodiments, some of the models may be trained independently, and other models may be trained concurrently. For example, the acoustic model may be trained separately from all other models of language processing, and the intent identification model may be trained together with the speech-to-text model, etc.
[0016] In some embodiments, each or some of the MLM122 (and / or MLM125) may be implemented as a deep learning neural network having multiple levels of linear or non-linear operations. For example, each or some of the voice recognition MLMs may be a convolutional neural network, a recurrent neural network (RNN), a fully connected neural network, or the like. In some embodiments, each or some of the MLM122 (and / or MLM125) may include a plurality of neurons, each neuron may receive its input from other neurons or an external source, and may generate an output by applying an activation function to the sum of the (trainable) weighted inputs and a bias value. In some embodiments, each or some of the MLM122 (and / or 125) may include a plurality of neurons arranged in layers, including an input layer, one or more hidden layers, and an output layer. Neurons from adjacent layers may be connected by weighted edges. Initially, the edge weights may be assigned some starting (e.g., random) values. For every training input 165, the training engine 160 may cause each or some of the MLM122 (and / or MLM125) to generate an output. The training engine 137 may then compare the observed output with the desired target output 167. The resulting error or mismatch, e.g., the difference between the desired target output 167 and the actual output of the neural network, may be backpropagated through each respective neural network, and the weights in the neural network may be adjusted to bring the actual output closer to the target output 167. This adjustment may be repeated until the output error for a given training input 165 meets a predetermined condition (e.g., falls below a predetermined value). Thereafter, different training inputs 165 may be selected, new outputs may be generated, and a new series of adjustments may be implemented until each respective neural network is trained to an acceptable accuracy.
[0017] The training engine 120 may include additional components (as compared to the training engine 160) to implement retraining of the previously trained MLM 122 for domain-specific applications. For example, the training engine 120 may include a data augmentation module to augment existing training data (e.g., training input 165) using domain-specific data. For example, existing recordings may be augmented with target words and phrases commonly encountered in the target domain. For example, the data augmentation module may augment existing training inputs with utterances of phrases such as "naked short selling", "capital gains tax", "hedge fund", "economic fundamentals", "initial public offering", etc. The target output 167 may be augmented similarly. For example, the data augmentation module may update the target output 167 using various technical terms having domain-specific meanings, such as "option", "future", etc. The training engine 120 may additionally have a pruning module for reducing the number of nodes, and an evaluation module for determining whether pruning of the nodes reduced the accuracy of the retrained model below a minimum threshold accuracy.
[0018] Figure 2A is a block diagram of an example architecture 200 of a training engine (e.g., training engine 120) of the customizable pipeline 100 of FIG. 1 according to at least some embodiments. As depicted in FIG. 2A, the training engine architecture 200 may include several modules (sub-engines) such as an initial training module 210, an evaluation module 210, a retraining module 230, etc., that perform the operations described above. For example, the initial training module 210 may train the MLM 122 using the initial data 202. The initial training module 210 may also train a custom (user-specific and / or user-provided) MLM 125 using custom data 204. The evaluation module 220 may determine whether the training of the MLM 122 (or custom MLM 125) was successful or whether additional training should be performed. For example, the evaluation module 220 may use a portion of the initial data 202 (or custom data 204) reserved for testing / evaluation. If each MLM does not meet a minimum accuracy or reliability level, the initial training module 210 may provide additional training for the MLM, as depicted using the return arrow. When the MLM successfully passes the evaluation, the MLM may be stored (e.g., in the trained model repository 124 of FIG. 1) for immediate or future use by the user. The retraining module 230 may perform retraining of the stored MLM using new data / tuning data 232 to generate a retrained MLM 123 or a retrained custom MLM 127. For example, a previously trained MLM may be retrained for applications in a different domain. Alternatively, a previously trained MLM may be retrained to account for changed or additional conditions, such as a change in the terminology used in the domain, the hiring of a new employee with different voice characteristics than other employees, and the like. Retraining may be performed similarly until the evaluation criteria (determined, e.g., by the evaluation module 220) are met. The retraining criteria may be different from the initial training criteria.
[0019] Referring back to FIG. 1, during MLM retraining, the user may interact with the pipeline orchestrator 110 via the pipeline API 106 to monitor the retraining process. For example, at the start of retraining (or at any other stage), the user may select a first set of pre-trained MLMs 122 as the models that should be used "as is" without retraining. The user may further select a second set of pre-trained MLMs 122 for retraining by the training engine 120 (or training engine 160) to generate the retrained MLM 123. Additionally, the user may cause the training engine 120 (or training engine 160) to generate a custom (user-trained) MLM 125. The user may select the architecture and network parameters for the custom MLM 125 via the UI 104, and the training engine 120 (or training engine 160) may receive the user-specified parameters via the pipeline orchestrator 110 and execute the training of the custom MLM 125 according to the received parameters. The parameters of the pre-trained MLM 122, the retrained MLM 123, and / or the custom MLM 125 may be stored in a memory device accessible to the pipeline orchestrator 110. The memory device storing the MLMs may be local (e.g., non-volatile) memory on the computing device 102 or remote (e.g., cloud-based) memory accessible by the computing device 102 via the network 140. The user of the CP100 may be provided with a listing of some or all of the MLMs (e.g., MLMs 122, 123, and / or 125) available to the user upon authentication for the user's login / session on the CP100. The user-accessible listing may include the MLMs (123) retrained or (125) user-trained during the current session, as well as the MLMs retrained or user-trained in any of the previous user sessions.Thus, during or after each user session, the newly user-trained custom MLM 125 may be stored in the trained model repository 124 for future use.
[0020] CP100 may further include an export engine 130. The export engine 130 may enable a user to select any number of pre-trained MLMs 122, retrained MLMs 123, or custom MLMs 125 for subsequent introduction during a current (or future) user session. The export engine may export the user-selected MLM using an implementation-independent format. In some embodiments, exporting the MLM may include identifying and retrieving the topology of the MLM, the number and types of neural network layers of the MLM, and the values of the weights determined during training performed by the training engine. In such embodiments, the export engine 130 may generate a representation of the user-selected MLM, causing the generated representation to be displayed on the UI 104. The displayed representation may include graphs, tables, numbers, text inputs, and other objects that characterize the architecture of the selected MLM, such as the number of layers, nodes, edges, and topology of each selected MLM. The displayed representation may further include parameters of the selected MLM, such as weights, biases, activation functions for various nodes, and the like. The metadata loaded by the export engine 130 may also indicate what additional components and modules may be required for the introduction of the user-selected MLM. In speech recognition, such additional components may include a spectral analyzer of the sound of speech, a speech feature extractor for the acoustic model, an acoustic post-processing component that removes speech artifacts, fillers, and stop words, a language pre-processing component that performs word tokenization, lemmatization, etc., a language post-processing component that performs rule-based modification / verification of the language MLM output, and other components.
[0021] CP100 may further include a build engine 150. The build engine 150 may enable a user to construct an exported pre-trained MLM 122, a retrained MLM 123, or a custom MLM 125 prior to introduction on the user platform. The exported representation of the MLM provided by the export engine 130 notifies the user of the MLM's architecture and properties. The representation may show the user which aspects of the MLM are static (parameters) and which aspects (settings) are customizable. For example, the type of neural network (e.g., convolutional vs. fully connected), the number of neuron layers in the neural network, the topology of the edges connecting nodes in the network, the type of activation function used at various nodes, etc. may be fixed parameters. If the user desires to change some of the fixed parameters, the user may have to use the training engine 120 to retrain each neural network model for the new architecture. On the other hand, the user may be able to change the settings of the MLM without retraining the model. Such configurable settings may include the chunk size, e.g., the size of the audio buffer to be processed in a streaming speech recognition application, the alphabet (e.g., Latin, Cyrillic, etc.) to be used to map the output, the language (e.g., English, German, Russian, etc.), the window size for FFT processing of the input voice data, the window overlap (e.g., 25%, 50%, 75%, etc.), the Hamming window parameter, the end-of-utterance detection parameter, the audio buffer size, the latency setting, etc. The build engine 150 may also enable the user to select from available domains (e.g., financial industry domain, medical field domain, etc.), and the selected MLM is trained for that domain (during initial training, retraining, or training on user-specific data).Build engine 150 (or pipeline orchestrator 110) may cause the display of the exported MLM settings on UI 104 along with the exported MLM parameters. The display may be annotated with an indication of which modifications (to the settings) should be handled by build engine 150 and which modifications (to the parameters or settings) should be handled by training engine 120. For example, modifications that call for changes to the language model may be handled by build engine 150 (without calling training engine 120), whereas modifications that call for changes to the acoustic model, e.g., changes to audio data processing in the speech recognition decoder stage, may be handled by training engine 120. In some embodiments, modifications that call for changes to the language model may include, without limitation, adjusting the weights of one or more layers in the model. As a further example, changes to the settings between models, such as from an acoustic model of a certain language trained using adult voice data to an acoustic model of the same language but trained using child voice data, may trigger a call to training engine 120 and retraining of the acoustic model with the new data.
[0022] Some settings may relate to a single exported MLM. For example, alphabet (voice spelling or standard English spelling) settings may affect the language model but may not affect the acoustic model. Some settings may relate to multiple MLMs. For example, configuring a language model for Chinese speech recognition may also call for changes (e.g., automatic or default) to the settings of the acoustic model to adjust for different cadences and tones of the voice. Some settings may affect the interaction of one or more MLMs with various pre-processing and post-processing components of the pipeline. For example, changing the settings of the language model for the use of a model for speech recognition in different linguistic domains may also require modification of the settings of the pre-processing block that removes stop words. Specifically, a language model that should perform speech recognition of a formal presentation in an experts' meeting may use less aggressive stop / filler word removal than the transcription of an informal brainstorming business meeting.
[0023] Figure 2B is a block diagram of an example architecture 200 of the build and deployment stages of the customizable pipeline 100 of FIG. 1 according to at least some embodiments. Build engine 150 enables a user to configure the exported MLM prior to deployment, while deployment engine 170 executes the actual implementation of CP100 on a user-accessible platform. As depicted in FIG. 2B, the MLM exported by export engine 130 may have various model artifacts 252 (e.g., modules, dependencies, metadata, etc.) and one or more configuration files 254 that constitute the actual execution of the exported MLM. The user may input modified configuration settings 256 (e.g., received through UI104 of FIG. 1), which may then be processed by build engine 150. The modified configuration settings may be written back to configuration file 254 using various fields, such as default fields initially provided by training engine 120, and overwritten with user-specified settings. Build engine 150 may process configuration file 254 and model artifacts 252 (e.g., source code, libraries, and other dependencies) for the exported MLM, and generate executable artifacts, configuration files, and various other dependencies for the exported MLM, such as executable code for libraries, pre-processing and post-processing components of the pipeline, and the like. The output of build engine 150 may be an intermediate representation (IR) of the pipeline. In some embodiments, the IR of the MLM pipeline may be packaged as a Docker image 260, or an image for any similar platform for containerized application execution. In some implementations, images in different (non-Docker) formats may be used, e.g., any proprietary format may be used.In some implementations, a suitable archive in a format that bundles together various executable components, libraries, data, and metadata may be used. The content of the IR may be stored (e.g., as an archive) in the memory device of the user's computing device, such as computing device 102, or in a memory device accessible to computing device 102 (e.g., on the cloud). The build engine 150 may be a tool or application implemented in the Python language, the C++ language, the Java (registered trademark) language, or any other programming language.
[0024] The introduction engine 170 of a configurable pipeline (e.g., CP100 of FIG. 1) may implement the pipeline on user-accessible hardware resources (the target platform). The user may have access to a (e.g., local) computing device 102 having several CPUs, GPUs, and memory devices. Alternatively or in addition, the user may have access to one or more cloud computing servers that provide virtualization services. The introduction engine 170 may enable the user to input a description or identification of user-accessible target platform resources 262, which may include identification of available resources such as computing, memory, network, etc. In some embodiments, the introduction engine 170 may collect information regarding the local resources of the computing device 102 using any available metric collection device or driver. Alternatively or in addition, the introduction engine 170 or the pipeline orchestrator 110 may collect information regarding available virtual (cloud) processing resources (e.g., using the authentication service of a remote access server or a remote virtualization server). The collected information may include, but is not limited to, CPU speed, number of CPU cores (physical or virtual), number and type of GPUs (physical or virtual), amount of available system memory and / or GPU memory, number of remote processing nodes available for pipeline introduction, type and version of the operating system installed on the computing device 102 (or type / version of the guest operating system instantiated on a virtualized environment), and the like. The introduction engine 170 enables execution of the MLM pipeline (e.g., CP100) on user-accessible computing resources without reconfiguring the functional characters of the MLM pipeline after the user-selected configuration settings are implemented by the build engine 150.
[0025] In some embodiments, the ingestion engine 170 may access the IR of the MLM pipeline stored (locally or in the cloud) by the build engine 150 and generate an inference ensemble of executable code (e.g., implemented in object code or bytecode), configuration files, libraries, dependencies, and other resources for use by the inference engine 180. In embodiments where the IR of the MLM pipeline is a Docker (or similar) image 260, the ingestion engine 170 may instantiate a pipeline Docker container 270 based on the Docker image 260, for example, using a containerized service of the user-accessible target platform. In some embodiments, the inference ensemble may be a Triton ensemble for Triton Inference Server that facilitates the ingestion of MLMs and enables users to execute MLMs using various available frameworks (e.g., TensorFlow, TersorRT, PyTorch, ONNX Runtime, etc.) or custom user-provided frameworks. The ingestion engine 170 may execute the commands specified in the IR to execute the executable code, libraries, and other dependencies generated by the build engine 150. Additionally, the ingestion engine 170 may perform a mapping of the MLM pipeline configuration generated by the build engine 150 for processing on the computing resources of the specific target platform, including available GPUs and / or CPUs. The configuration files generated by the ingestion engine 170 may be stored using a platform-neutral protocol buffer that may be ASCII-serialized. The use of protocol buffers allows minimizing key input (typographical) and serialization-related errors that may sometimes occur when a user (developer) adds support for an MLM architecture.
[0026] Referring back to FIG. 1, after the introduction engine 170 converts the IR to an inference ensemble (e.g., in a pipeline Docker container) that is ready for execution by the inference engine 180, the MLM pipeline may be ready to process the input user data 182. The user data 182 may be any data to which the configured MLM pipeline may be applied. For example, in the case of audio processing, the user data 182 may include an audio recording, e.g., a digital recording of a conversation, presentation, narration, or any other recording that is to be transcribed into text. In some embodiments, the user data may include a question (or series of questions) to be answered. In some embodiments, the user data 182 may be an image (or sequence of images) with an object to be identified, a pattern of movement to be detected, etc. Any other user data 182 may be input to the inference engine 180 along with the type of user data that depends on the user-specific domain. In embodiments involving natural language processing, the user-specific domain may include customer service support, medical questions, educational environments, courtroom environments, conversations with emergency responders, or any other type of environment.
[0027] The inference engine 180 may process the user data 182 and generate an inference output 184. The inference output 184 may have any suitable type and format. For example, the inference output 184 may be a transcription of speech or conversation, an identification of the speaker's intent or emotion, an answer to a question (e.g., in the form of text or a numerical value), etc. The format of the inference output 184 may be text, numbers, a numerical spreadsheet, an audio file, a video file, or any combination thereof.
[0028] The various engines of CP100 need not be applied in linear progression. In some embodiments, the various engines of CP100 may be applied multiple times. For example, user data 182 may be used as test data, and the obtained inference output 184 may be used as feedback regarding the current state of the pipeline. The feedback may inform the user how to modify the pipeline to improve its performance. In some embodiments, such modifications may be performed iteratively. For example, upon receiving feedback, the user may initiate retraining of some of the MLMs included in the pipeline using the training engine 120. The user may also replace some of the MLMs with other (e.g., pre-trained or user-trained) MLMs and export the newly added MLMs using the export engine 130. The user may change the configuration of some of the old, retrained, or newly trained MLMs using the build engine 150. The build engine 150 may be configured to update only the MLMs of the pipeline and the model artifacts and configuration files of the components (e.g., by generating an updated IR) that have been modified (e.g., via updated configuration settings) without changing the models and components that have remained unchanged for faster installation. In some embodiments, such faster installation may be achieved using the PIP (Package Installation for Python) Wheel build-package format (.whl). Thereafter, the introduction engine 170 may use the updated IR to introduce the updated pipeline onto the target platform. In some embodiments, the user may keep the MLMs and their respective configuration settings but modify the resources available on the target platform (by increasing, decreasing, or otherwise modifying).
[0029] FIG. 3 is a diagram of an example computing device 300 capable of implementing a customizable pipeline that supports training, configuring, and deploying one or more machine learning models, according to at least some embodiments. In some embodiments, computing device 300 may include a customizable pipeline, such as some or all of the engines of CP100 of FIG. 1, including a training engine 120, an export engine 130, a build engine 150, a deployment engine 170, and an inference engine 180. Although FIG. 3 depicts all of the engines as part of the same computing device, in some implementations, any of the engines shown may actually be implemented on different computing devices, including virtual computing devices, cloud-based processing devices, and the like. For example, computing device 300 may include inference engine 180 but may not include other engines of the customizable pipeline. Inference engine 180 (and / or any other engine of the pipeline) may be executed by one or more GPUs 310 to perform speech recognition, object recognition, or any other inference involving machine learning. In some embodiments, GPU 310 includes a plurality of cores 311, each of which is capable of executing a plurality of threads 312. Each core may execute a plurality of threads 312 concurrently (e.g., in parallel). In some embodiments, threads 312 may have access to registers 313. Registers 313 may be thread-specific registers having restricted access to registers for each thread. Additionally, shared registers 314 may be accessed by all threads of a core. In some embodiments, each core 311 may include a scheduler 315 for distributing computational tasks and processes among different threads 312 of core 311. Dispatch unit 316 may implement scheduled tasks on appropriate threads using the correct private registers 313 and shared registers 314.The computing device 300 may include input / output components 334 to facilitate the exchange of information with one or more users or developers.
[0030] In some embodiments, the GPU 310 may have a (high-speed) cache 318, access to which may be shared by multiple cores 311. Additionally, the computing device 300 may include a GPU memory 319, where the GPU 310 may store intermediate and / or final results (outputs) of various calculations performed by the GPU 310. After completion of a particular task, the GPU 310 (or the CPU 330) may move the output to the (main) memory 304. In some embodiments, the CPU 330 may execute a process involving serial calculation tasks (assigned by one of the pipeline engines), while the GPU 310 may execute tasks such as (multiplication by the weights of the inputs of neural nodes and adding biases) that are capable of parallel processing. In some embodiments, each engine of the pipeline (e.g., the build engine 150, the inference engine 180, etc.) may determine which processes managed by each engine should be executed on the GPU 310 and which processes should be executed on the CPU 330. In some embodiments, the CPU 330 may determine which processes should be executed on the GPU 310 and which processes should be executed on the CPU 330.
[0031] Figure 4 is a block diagram of an example customizable pipeline 400 that uses one or more machine learning models for natural language processing of speech, according to at least some embodiments. Some or all of the MLMs of pipeline 400 may be trained neural network models, e.g., pre-trained and / or custom-trained deep learning neural networks. The speech input 402 to pipeline 400 may be an analog signal generated, for example, by a microphone and converted to a digital file in any audio format readable by a processing device. The input speech 402 may undergo audio preprocessing 410, which may include spectral analysis and other processing. For example, the input speech 402 may undergo filtering, upsampling or downsampling, pre-emphasis, windowing (using, for example, a 20 ms window advanced every 10 ms), application of the Mel Frequency Cepstral Coefficient (MFCC) algorithm, and / or other processing. The multi-dimensional vector representing the extracted features may be input to a first MLM 420, which may be an acoustic MLM (e.g., an acoustic neural network model). The acoustic MLM 420 may be a neural network model trained to output the identification of various phonemes (elemental sub-word sounds), each assigned a certain probability. The acoustic postprocessing 430 may include speech decoding (e.g., assigning probabilities to various words), removing audio artifacts, fillers or stop words, or may further include some other type of acoustic postprocessing. The language preprocessing 440 may include word stemming or lemmatization (determining the base form of a word), tokenization (identifying sequences of characters and words), and the like. The output of the language preprocessing 440 may be used as input to a second MLM 450, which may include one or more language MLMs (e.g., language neural network models).The second MLM 450 may be another neural network trained to generate a text output 460, which may be the speech-to-text of the audio input 402. In some embodiments, the second MLM 450 may be implemented via additional output / neuron layers of the same second MLM 450 or as an additional neural network model to generate punctuation detection (450-1), utterance detection (450-2), and speaker intent detection (450-3). The second MLM 450 may include a language understanding neural network model 450-4 and / or a question answering (QA) neural network model 450-5. The output of the QA model 450-5 may be an expression (e.g., a text expression) of an answer to a question included in the audio input 402. In some embodiments, the output of the QA model 450-5 may be provided to a text-to-speech synthesizer, which may be a third MLM 470 trained to output synthesized speech as part of the audio output 472.
[0032] As described above with respect to FIGS. 1 and 2, the first MLM 420, the second MLM 450, and / or the third MLM 470 may be pre-trained or custom-trained by the training engine 120, as schematically depicted using dashed arrows. In some embodiments, some or all of the MLMs 420, 430, and 470 may be re-trained using domain-specific user data. Additionally, as further depicted in FIG. 4, some or all of the MLMs 420, 430, and 470 may be configured (without re-training) using the build engine 150 and based on configuration settings provided by the user.
[0033] FIG. 5 and FIG. 6 are flow diagrams of example methods 500 and 600, respectively, relating to the provisioning of customizable machine learning pipelines according to at least some embodiments. Methods 500 and 600 may be performed to introduce an MLM for use in voice recognition, speech recognition, speech synthesis, object detection, object recognition, motion detection, hazard detection, robotics applications, prediction, and many other contexts and applications where machine learning may be used. In at least one embodiment, methods 500 and 600 may be performed by a processing unit of computing device 102, computing device 300, or some other computing device, or a combination of multiple computing devices. Methods 500 and 600 may be performed by one or more processing units (e.g., a CPU and / or a GPU) that may include (or communicate with) one or more memory devices. In at least one embodiment, methods 500 and 600 may be performed by multiple processing threads (e.g., CPU threads and / or GPU threads), with each thread performing one or more individual functions, routines, subroutines, or operations of the method. In at least one embodiment, the processing threads implementing method 500 (and similarly, method 600) may be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads implementing method 500 (and similarly, method 600) may be executed asynchronously with respect to each other. The various operations of methods 500 and 600 may be performed in an order different from that shown in FIGS. 5 and 6. Some operations of the method may be performed concurrently with other operations. In at least one embodiment, one or more of the operations shown in FIGS. 5 and 6 may not always be performed.
[0034] FIG. 5 is a flow diagram of an example method 500 that provides a customizable pipeline that supports the training, configuration, and deployment of one or more machine learning models, according to at least some embodiments. The customizable pipeline may be a CP100 that may include various engines, modules, and components, as described above with respect to FIGS. 1 and 2A-2B. The processing unit that executes method 500 may access, at block 510, a plurality of trained MLMs, which may be pre-trained MLM122 (e.g., by a provider of pipeline services), or a custom MLM125 (trained by a user of the pipeline). The selected MLM may be pre-trained using a first set of training data, which may have been previously supplied by a provider of pipeline services or by a user. Maintaining a trained MLM may include storing files and data sufficient for the deployment and execution of the MLM, or storing references to such files and data, e.g., links to downloadable files and data stored elsewhere (such as on cloud storage). Some or each of the maintained MLMs may be associated with initial configuration settings, which may be stored in a configuration file, database, etc., for each respective MLM. The configuration settings may also be accessible by the processing unit that executes method 500. In some embodiments, at least some of the selected MLMs may be (or include) neural network models having a plurality of neuron layers. Some of the neural network models may be deep learning neural network models.
[0035] In block 520, the processing unit that executes method 500 may provide a user interface (UI) for receiving user input indicating the selection of one or more of the plurality of trained MLMs. For example, a user may be attempting to set up a machine learning pipeline and may select an MLM from the available MLMs that can solve a particular problem or group of problems for the user. For example, a user attempting to set up a pipeline for speech recognition may select an acoustic model for decoding input speech and a language model for recognizing the decoded speech. To assist in the selection of MLMs for the pipeline, prior to receiving a user selection of one or more MLMs, the processing unit may cause a listing of one or more of the plurality of trained MLMs that may be available to the user to be displayed on the UI. In some embodiments, the listing may be in the form of enumerated items, clickable buttons, icons, menus, or any other prompt. The listing for some or each of the listed MLMs may include a representation of the architecture of each MLM, which may be in the form of a graph, table, layer depiction, description of the MLM topology, etc. In some embodiments, the listing may further include the parameters of each MLM, such as the number of nodes, edges, activation functions, or any other specification of the properties of the MLM.
[0036] In response to viewing the listings of MLMs, the user may decide that some of the selected MLMs should be modified to better match the details of the user's project. In some examples, the modification of the MLM may require, for example, retraining the selected MLM with user-selected data, which may be a second set of training data that is significant enough to be used for retraining the selected MLM for a particular domain to which the MLM pipeline should be applied. In particular, the processing unit executing method 500 may receive a second set of training data for domain-specific training of one or more selected MLMs. For example, the processing unit may receive an instruction for the MLM to be retrained based on user input via the UI. For example, from an acoustic MLM and a language MLM selected for a speech recognition pipeline, the user may indicate that the language MLM (previously trained for English) should be retrained for Japanese. The user may also identify a location (e.g., a cloud storage address) where the retraining data (e.g., the second data) is available. The retraining data may include training inputs (e.g., sound files), target outputs (e.g., the transcription of the speech in the sound file), and mapping data (instructions for the correspondence of the training inputs to the target outputs). In response to receiving the instruction for the MLM to be retrained and the retraining data, the processing unit may cause the execution of a pipeline training engine (e.g., training engine 120) to perform domain-specific training of one or more selected MLMs, as described in more detail above in conjunction with FIG. 1. The retraining may be performed for both the pre-trained MLM 122 and the previously trained custom MLM 125.
[0037] In some instances, the required modifications to the MLM may not be as significant as to require retraining. In some embodiments, the modified configuration settings may include language settings for a language neural network model. Using the previous example, to implement the change from English to Japanese, the user may decide to modify the settings of the acoustic MLM, for example, change the size of the sliding window. In response to receiving user input specifying how the initial configuration settings for one or more MLMs should be modified, at block 530, the processing unit executing method 500 may determine modified configuration settings for one or more selected MLMs based on the user input. In some embodiments, the modified configuration settings for one or more selected MLMs may include the audio buffer size, the utterance end setting, or the latency setting for the acoustic MLM. Other natural language processing MLMs that may be similarly configured may include a text-to-speech MLM, a language understanding MLM, a question answering MLM, or any other MLM.
[0038] At block 540, the processing unit executing method 500 may cause (e.g., by issuing an instruction to start) the execution of a build engine of an MLM pipeline to modify one or more selected MLMs according to the modified configuration settings, as described in more detail above in conjunction with FIG. 1. At block 550, the method may continue by causing the processing unit to execute an introduction engine of a pipeline to introduce one or more modified MLMs, as described in more detail above in conjunction with FIG. 1. At block 560, the processing unit executing method 500 may cause the display on the UI of the representation of one or more introduced MLMs, as described in more detail above in conjunction with FIG. 1. The displayed representation may indicate to the user that the introduced MLM is ready to process user data. A configurable MLM pipeline may be used for training, introduction, and inference using any number of machine learning models of any type.
[0039] FIG. 6 is a flow diagram of an example method 600 of using an introduced customizable machine pipeline machine learning model according to at least some embodiments. In some embodiments, method 600 may be used in conjunction with method 500. Method 600 may be performed after providing a listing of the pre-trained MLM 122 and the custom MLM 125 for presentation to the user and, optionally, after retraining of the MLM selected by the user for retraining by the training engine. In block 610, the processing unit executing method 600 may receive a user selection of one or more MLMs (to be placed in a configurable pipeline) and may cause execution of an export engine of the pipeline to initialize the one or more selected MLMs. In block 620, method 600 may continue to create one or more selected MLMs available for processing user input data (e.g., using the export engine). Additionally, method 600 may include causing execution of a build engine (block 540) and an introduction engine (block 560) as described with respect to method 500 of FIG. 5.
[0040] In block 630, a processing unit that executes method 600 may receive user input data, and in block 640, cause one or more introduced MLMs to be applied to the user input data (e.g., using an inference engine) to generate output data. The output of the application of one or more MLMs may cause, in block 650, a display on the UI of at least one of a representation of the output data or a reference to a stored representation of the output data. For example, if the user input data includes audio, the representation of the output data may be text by speech recognition of the input audio, identification of the speaker's intent, punctuation of the input audio, a voice response to a question asked (e.g., using a synthesized voice), etc. In some embodiments, the representation of the output data may be, for example, text or sound presented directly to the user on the UI. In some embodiments, the representation of the output data may be an indication that the output data is stored (on a local machine or in the cloud).
[0041] Inference and training logic FIG. 7A shows inference and / or training logic 715 used to perform inference and / or training operations with respect to one or more embodiments.
[0042] In at least one embodiment, the inference and / or training logic 715 may include, without limitation, code and / or data storage 701 for storing forward propagation and / or output weights, and / or input / output data, and / or other parameters for constructing neurons or layers of a neural network that are trained and / or used to infer in aspects of one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data storage 701 for storing graph code or other software for controlling timing and / or order, and the code and / or data storage 701 has weight and / or other parameter information loaded therein to configure logic including integer and / or floating point units (collectively arithmetic logic units (ALUs) or simply circuits). In at least one embodiment, code such as graph code loads weight or other parameter information into the processor ALU based on the architecture of the corresponding neural network. In at least one embodiment, the code and / or data storage 701 stores the weight parameters and / or input / output data of each layer of a neural network that is trained or used in conjunction with one or more embodiments while propagating the input / output data and / or weight parameters forward during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 701 may be included with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory.
[0043] In at least one embodiment, any portion of the code and / or data storage 701 may be internal or external to one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 701 may be cache memory, dynamic randomly addressable memory (“DRAM”), static randomly addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 701 is internal or external to, for example, a processor, or the choice of including DRAM, SRAM, flash, or some other type of storage, may be determined according to on-chip versus off-chip available storage, latency requirements of the training and / or inference functions being executed, batch size of data used in neural network inference and / or training, or any combination of these factors.
[0044] In at least one embodiment, the inference and / or training logic 715 may include, without limitation, backward propagation and / or output weights corresponding to neurons or layers of a neural network that are trained and / or used for inference in one or more embodiments, and / or code and / or data storage 705 for storing input / output data. In at least one embodiment, the code and / or data storage 705 stores the weight parameters and / or input / output data of each layer of the neural network that is trained or used in conjunction with one or more embodiments while backpropagating the input / output data and / or weight parameters during training and / or inference using one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data storage 705 for storing graph code or other software for controlling timing and / or order, and the code and / or data storage 705 has weights and / or other parameter information loaded therein to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)).
[0045] In at least one embodiment, a code such as a graph code causes the processor ALU to load weight or other parameter information based on the architecture of the neural network to which such code corresponds. In at least one embodiment, any portion of the code and / or data storage 705 may be included with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory. In at least one embodiment, any portion of the code and / or data storage 705 may be internal or external to one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 705 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the selection of whether the code and / or data storage 705 is internal or external to, for example, the processor, or the selection of whether to include DRAM, SRAM, flash memory, or some other type of storage, may be determined according to on-chip versus off-chip available storage, latency requirements of the training and / or inference functions being executed, the batch size of the data used in neural network inference and / or training, or any combination of these factors.
[0046] In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be separate storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be combined storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be partially combined storage structures and partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 701 and code and / or data storage 705 may be included with L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory.
[0047] In at least one example, the inference and / or training logic 715 includes, without limitation, one or more arithmetic logic units (ALUs) 710 including integer and / or floating point units to perform logical and / or mathematical operations based at least in part on and / or indicated by training and / or inference code (e.g., graph code), the result of which may generate activations (e.g., output values from a layer or neuron within a neural network) stored in activation storage 720, which are a function of code and / or data storage 701 and / or input / output and / or weight parameter data stored in code and / or data storage 705. In at least one example, the activations stored in activation storage 720 are generated according to linear algebra computations or matrix-based computations performed by the ALU 710 in response to executing instructions or other code, where the weight values stored in code and / or data storage 705 and / or data storage 701 are used as operands along with other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 705, or code and / or data storage 701, or another storage on-chip or off-chip.
[0048] In at least one embodiment, the ALU 710 is included within one or more processors, or other hardware logic devices or circuits, but in another embodiment, the ALU 710 may be external to the processors or other hardware logic devices or circuits (such as a coprocessor) that use them. In at least one embodiment, the ALU 710 may be included within the execution unit of a processor, or may be distributed among execution units of processors (such as a central processing unit, a graphics processing unit, a fixed function unit, etc.) that are either within the same processor or different processors of different types, and are accessible within an ALU bank that may be otherwise included in some other way. In at least one embodiment, the code and / or data storage 701, the code and / or data storage 705, and the activation storage 720 may share the same processor or other hardware logic device or circuit, but in another embodiment, they may be in different processors or other hardware logic devices or circuits, or may be in any combination of the same processor or other hardware logic device or circuit and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of the activation storage 720 may be included together with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory. Further, the inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuit, and may be fetched and / or processed using the fetch, decode, scheduling, execution, retirement, and / or other logic circuits of the processor.
[0049] In at least one embodiment, the activation storage 720 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activation storage 720 may be wholly or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, the selection of whether the activation storage 720 is internal or external to, for example, a processor, or the selection of whether to include DRAM, SRAM, flash memory, or some other type of storage, may be determined according to on-chip versus off-chip available storage, latency requirements of the training and / or inference functions being performed, batch size of data used in neural network inference and / or training, or any combination of these factors.
[0050] In at least one embodiment, the inference and / or training logic 715 shown in FIG. 7A may be used in conjunction with an application specific integrated circuit (ASIC) such as Google's TensorFlow® processing unit, Graphcore's inference processing unit (IPU), or Intel Corp's Nervana® (e.g., "Lake Crest") processor. In at least one embodiment, the inference and / or training logic 715 shown in FIG. 7A may be used in conjunction with other hardware such as central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, data processing unit (DPU) hardware, or field programmable gate array (FPGA).
[0051] FIG. 7B shows inference and / or training logic 715 according to at least one embodiment. In at least one embodiment, the inference and / or training logic 715 may include, without limiting the hardware logic, in which computing resources are dedicated to one or more layers of neurons in a neural network for weight values or other information, or are used only in combination with them in other ways. In at least one embodiment, the inference and / or training logic 715 shown in FIG. 7B may be used in conjunction with an application-specific integrated circuit (ASIC) such as a TensorFlow® processing unit from Google, an inference processing unit (IPU) from Graphcore®, or a Nervana® (e.g., "Lake Crest") processor from Intel Corporation. In at least one embodiment, the inference and / or training logic 715 shown in FIG. 7B may be used in conjunction with other hardware such as central processing unit (CPU) hardware, graphics processing unit ("GPU") hardware, data processing unit ("DPU") hardware, or a field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 715 includes, without limitation, code and / or data storage 701 and code and / or data storage 705, which may be used to store code (e.g., graph code), weight values, and / or bias values, gradient information, momentum values, and / or other information including other parameters or hyperparameter information. In at least one embodiment shown in FIG. 7B, each of code and / or data storage 701 and code and / or data storage 705 is associated with dedicated computing resources such as computing hardware 702 and computing hardware 706, respectively. In at least one embodiment, each of computing hardware 702 and computing hardware 706 includes one or more ALUs that execute mathematical functions such as linear algebra functions only on the information stored in code and / or data storage 701 and code and / or data storage 705, respectively, and the results are stored in activation storage 720.
[0052] In at least one embodiment, each of code and / or data storage 701 and 105, and corresponding computing hardware 702 and 706, respectively corresponds to different layers of a neural network, such that the activation resulting from one storage / computation pair 701 / 702 of code and / or data storage 701 and computing hardware 702 is provided as an input to the next storage / computation pair 705 / 706 of code and / or data storage 705 and computing hardware 706 in order to reflect the conceptual organization of the neural network. In at least one embodiment, the storage / computation pairs 701 / 702 and 705 / 706 may correspond to two or more layers of a neural network. In at least one embodiment, additional storage / computation pairs (not shown) may be included in the inference and / or training logic 715 after or in parallel with the storage / computation pairs 701 / 702 and 705 / 706.
[0053] Training and Introduction of Neural Networks FIG. 8 shows the training and deployment of a deep neural network according to at least one embodiment. In at least one embodiment, an untrained neural network 806 is trained using a training data set 802. In at least one embodiment, the training framework 804 is the PyTorch framework, while in other embodiments, the training framework 804 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 804 trains the untrained neural network 806 and enables it to be trained using the processing resources described herein to generate a trained neural network 808. In at least one embodiment, the weights may be randomly selected or selected by pre-training using a deep belief network. In at least one embodiment, the training may be performed in any of a supervised, semi-supervised, or unsupervised manner.
[0054] In at least one embodiment, the untrained neural network 806 is trained using supervised learning, where the training data set 802 includes inputs paired with the desired outputs for the inputs, or the training data set 802 includes inputs with known outputs, and the output of the neural network 806 is scored manually. In at least one embodiment, the untrained neural network 806 is trained in a supervised manner, processes the inputs from the training data set 802, and compares the resulting output to a set of predicted or desired outputs. In at least one embodiment, an error is then backpropagated through the untrained neural network 806. In at least one embodiment, the training framework 804 adjusts the weights that control the untrained neural network 806. In at least one embodiment, the training framework 804 includes a tool that monitors how well the untrained neural network 806 converges towards a model such as a trained neural network 808 that is suitable for generating correct answers in results 814 etc. based on input data such as a new data set 812. In at least one embodiment, the training framework 804 repeatedly trains the untrained neural network 806 while using an adjustment algorithm such as a loss function and stochastic gradient descent to adjust the weights to refine the output of the untrained neural network 806. In at least one embodiment, the training framework 804 trains the untrained neural network 806 until the untrained neural network 806 reaches a desired accuracy. In at least one embodiment, the trained neural network 808 can then be introduced to implement any number of machine learning operations.
[0055] In at least one embodiment, the untrained neural network 806 is trained using unsupervised learning, where the untrained neural network 806 attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised learning training data set 802 includes input data that does not have any associated output data or "ground truth" data. In at least one embodiment, the untrained neural network 806 can learn to group within the training data set 802 and can determine how individual inputs relate to the untrained data set 802. In at least one embodiment, an unsupervised training can be used within the trained neural network 808 that can perform operations useful for reducing the dimensions of the new data set 812 to generate a self-organizing map. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which enables identification of data points within the new data set 812 that deviate from the normal pattern of the new data set 812.
[0056] In at least one embodiment, semi-supervised learning may be used, which is a technique in which labeled data and unlabeled data are mixed in the training data set 802. In at least one embodiment, the training framework 804 may be used to perform incremental learning, such as by transfer learning techniques. In at least one embodiment, incremental learning enables the trained neural network 808 to adapt to the new data set 812 without forgetting the knowledge taught into the trained neural network 808 during initial training.
[0057] Referring to FIG. 9, FIG. 9 is an example data flow diagram for a process 900 of generating and introducing a processing and inference pipeline according to at least one embodiment. In at least one embodiment, the process 900 may be introduced to perform game name recognition analysis and inference on user feedback data at one or more facilities 902, such as a data center.
[0058] In at least one embodiment, process 900 may be executed within training system 904 and / or introduction system 906. In at least one embodiment, training system 904 may be used to train, introduce, and implement a machine learning model (e.g., a neural network, an object detection algorithm, a computer vision algorithm, etc.) for use in introduction system 906. In at least one embodiment, introduction system 906 may be configured to offload processing and compute resources among a distributed computing environment to reduce infrastructure requirements in facility 902. In at least one embodiment, introduction system 906 may provide a streamlined platform for selecting, customizing, and implementing virtual appliances for use with computing devices in facility 902. In at least one embodiment, a virtual appliance may include a software-defined application for performing one or more processing operations on feedback data. In at least one embodiment, one or more applications in the pipeline may use or call services (e.g., inference, visualization, computing, AI, etc.) of introduction system 906 during execution of the application.
[0059] In at least one embodiment, some of the applications used in the advanced processing and inference pipeline may use a machine learning model or other AI to perform one or more processing steps. In at least one embodiment, the machine learning model may be trained in facility 902 using feedback data 908 (such as feedback data stored in facility 902) or feedback data 908 from another one or more facilities, or a combination thereof. In at least one embodiment, an application, service, and / or other resources may be provided using training system 904 to generate a practical and deployable machine learning model for introduction system 906.
[0060] In at least one embodiment, the model registry 924 may be backed up by an object storage that can support version management and object metadata. In at least one embodiment, the object storage may be accessible, for example, from within a cloud platform, via a compatibility application programming interface (API) of cloud storage (e.g., cloud 1026 of FIG. 10). In at least one embodiment, a machine learning model within the model registry 924 may be uploaded, listed, modified, or deleted by a developer or partner of the system interacting with the API. In at least one embodiment, the API may provide access to a way for a user with appropriate credentials to associate a model with an application, thereby enabling the model to be executed as part of running a containerized instance of the application.
[0061] In at least one embodiment, the training pipeline 1004 (FIG. 10) may include a situation where the facility 902 is training its own machine learning model, or a situation where there is an existing machine learning model that needs to be optimized or updated. In at least one embodiment, the feedback data 908 may be received from various channels such as a forum, a web form, or the like. In at least one embodiment, when the feedback data 908 is received, AI-assisted annotation 910 may be used to assist in generating an annotation corresponding to the feedback data 908 to be used as ground truth data for the machine learning model. In at least one embodiment, the AI-assisted annotation 910 may include one or more machine learning models (e.g., a convolutional neural network (CNN)) trained to generate an annotation corresponding to a certain type of feedback data 908 (e.g., from a certain device) and / or a certain type of anomaly in the feedback data 908. In at least one embodiment, the AI-assisted annotation 910 may then be used directly, or may be adjusted or fine-tuned using an annotation tool to generate ground truth data. In at least one embodiment, in some examples, the labeled data 912 may be used as ground truth data for training the machine learning model. In at least one embodiment, the AI-assisted annotation 910, the labeled data 912, or a combination thereof may be used as ground truth data for training the machine learning model. In at least one embodiment, the trained machine learning model may be referred to as the output model 916 and may be used by the introduction system 906 described herein.
[0062] In at least one embodiment, the training pipeline 1004 (FIG. 10) may include a situation where the facility 902 needs a machine learning model to execute one or more processing tasks for one or more applications within the onboarding system 906, but the facility 902 may not currently have such a machine learning model (or may not have a model optimized, efficient, or effective for such purposes). In at least one embodiment, an existing machine learning model may be selected from the model registry 924. In at least one embodiment, the model registry 924 may include machine learning models trained to perform various different inference tasks on imaging data. In at least one embodiment, the machine learning models in the model registry 924 may be trained on imaging data from a facility different from the facility 902 (e.g., a facility in a remote location). In at least one embodiment, the machine learning model may be trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when trained on imaging data from a particular location, the training may be performed at that location, or at least in a manner that protects the confidentiality of the imaging data or restricts the transfer of the imaging data outside the facility (e.g., in accordance with HIPPA regulations, privacy regulations). In at least one embodiment, when a model is trained or partially trained at one location, the machine learning model may be added to the model registry 924. In at least one embodiment, the machine learning model may then be retrained or updated at any number of other facilities, and the retrained or updated model may be made available in the model registry 924. In at least one embodiment, the machine learning model may then be selected from the model registry 924, may be referred to as the output model 916, and may be used in the onboarding system 906 to execute one or more processing tasks for one or more applications of the onboarding system.
[0063] In at least one embodiment, the training pipeline 1004 (FIG. 10) can be used in a scenario that includes a situation where the facility 902 needs a machine learning model to execute one or more processing tasks for one or more applications within the onboarding system 906, but the facility 902 may not currently have such a machine learning model (or may not have a model optimized, efficient, or effective for such purposes). In at least one embodiment, the machine learning model selected from the model registry 924 may not be fine-tuned or optimized for the feedback data 908 generated at the facility 902 because there may be differences in the population, genetic variations, robustness of the training data used to train the machine learning model, diversity of anomalies in the training data, and / or other issues associated with the training data. In at least one embodiment, AI-assisted annotation 910 may be used to assist in generating annotations corresponding to the feedback data 908 that will be used as ground truth data for retraining or updating the machine learning model. In at least one embodiment, the labeled data 912 may be used as ground truth data for training the machine learning model. In at least one embodiment, retraining or updating the machine learning model may be referred to as model training 914. In at least one embodiment, model training 914, such as AI-assisted annotation 910, labeled data 912, or a combination thereof, may be used as ground truth data for retraining or updating the machine learning model.
[0064] In at least one embodiment, the introduction system 906 may include software 918, services 920, hardware 922, and / or other components, features, and functions. In at least one embodiment, the introduction system 906 may include a software “stack,” whereby software 918 may be built on top of services 920, may use services 920 to perform some or all of the processing tasks, and services 920 and software 918 may be built on top of hardware 922 and may use hardware 922 to perform the processing, storage, and / or other computing tasks of the introduction system 906.
[0065] In at least one embodiment, software 918 may include any number of different containers, where each container may perform application instantiation. In at least one embodiment, each application may perform one or more processing tasks of an advanced processing and inference pipeline (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.). In at least one embodiment, for each type of computing device, there may be any number of containers capable of performing data processing tasks on feedback data 908 (or other types of data such as those described herein). In at least one embodiment, the advanced processing and inference pipeline may be defined based on the selection of different containers desired or required to process feedback data 908, in addition to the containers that receive and configure the imaging data used by each container and / or used by facility 902 after processing through the pipeline (to re-convert the output to a usable type of data for storage and display at facility 902). In at least one embodiment, a combination of containers within software 918 (e.g., that make up the pipeline) may be referred to as a virtual device (described in more detail herein), and the virtual device may use services 920 and hardware 922 to perform some or all of the processing tasks of the applications instantiated in the containers.
[0066] In at least one embodiment, the data may be preprocessed as part of a data processing pipeline so that the data is prepared to be processed by one or more applications. In at least one embodiment, post-processing may be performed on the output of one or more inference tasks or other processing tasks of the pipeline so that output data is prepared for the next application and / or so that output data is prepared for transmission and / or for use by a user (e.g., as a response to an inference request). In at least one embodiment, the inference task may be performed by one or more machine learning models, such as a trained or introduced neural network, which may include the output model 916 of the training system 904.
[0067] In at least one embodiment, the tasks of the data processing pipeline may be encapsulated in containers, each of which represents an individual fully functional instantiation of an application and a virtualized computing environment in which a machine learning model can be referenced. In at least one embodiment, the container or application may be issued to a private (e.g., restricted access) area of a container registry (described in more detail herein), and the trained or introduced model may be stored in the model registry 924 and associated with one or more applications. In at least one embodiment, an image of an application (e.g., an image of a container) may be available in the container registry and, when selected by a user from the container registry for introduction into the pipeline, the image may be used to generate a container for instantiating the application for use on the user's system.
[0068] In at least one embodiment, a developer may develop, publish, and store an application (e.g., as a container) to perform processing and / or inference on supplied data. In at least one embodiment, the development, publishing, and / or storage may be performed using a software development kit (SDK) associated with the system (e.g., to ensure that the developed application and / or container complies with or is compatible with the system). In at least one embodiment, the developed application may be tested locally (e.g., at a first facility, for data from the first facility) using an SDK that can support at least a portion of service 920 as a system (e.g., system 1000 of FIG. 10). In at least one embodiment, once verified by system 1000 (e.g., in terms of accuracy, etc.), the application is made available in a container registry for selection and / or implementation by a user (e.g., a hospital, clinic, research institute, healthcare provider, etc.), and one or more processing tasks may be performed on data at the user's facility (e.g., a second facility).
[0069] In at least one embodiment, the developer may then share the application or container through a network so that it can be accessed and used by a user of the system (e.g., system 1000 of FIG. 10). In at least one embodiment, the completed and verified application or container may be stored in a container registry, and the associated machine learning model may be stored in a model registry 924. In at least one embodiment, a requesting entity that issues an inference or image processing request may browse the container registry and / or model registry 924 to find applications, containers, datasets, machine learning models, etc., select a desired combination of elements for inclusion in a data processing pipeline, and send a processing request. In at least one embodiment, the request may include the input data required to execute the request and / or may include the selection of an application and / or machine learning model to be executed when processing the request. In at least one embodiment, the request is then passed to one or more components of an ingress system 906 (e.g., a cloud) to execute the processing of the data processing pipeline. In at least one embodiment, the processing by the ingress system 906 may include referring to elements (e.g., applications, containers, models, etc.) selected from the container registry and / or model registry 924. In at least one embodiment, when a result is generated by the pipeline, the result may be returned to and viewed by the user (e.g., locally, viewed in a viewing application suite running on an in-premises workstation or terminal).
[0070] In at least one embodiment, service 920 may be utilized to assist in the processing or execution of an application or container in a pipeline. In at least one embodiment, service 920 may include a computing service, an artificial intelligence (AI) service, a visualization service, and / or other types of services. In at least one embodiment, service 920 may provide common functionality to one or more applications of software 918, whereby the functionality may be abstracted with respect to services that can be called or utilized by the applications. In at least one embodiment, the functionality provided by service 920 may be executed dynamically and more efficiently, and at the same time, may scale well by enabling an application to process data in parallel (e.g., using parallel computing platform 1030 (FIG. 10)). Instead of requiring each application sharing the same functionality provided by service 920 to have its own instance of service 920, service 920 may be shared among various applications. In at least one embodiment, the service may include, as a non-limiting example, an inference server or engine that may be used to perform tasks of detection or segmentation. In at least one embodiment, a model training service capable of providing functionality for training and / or retraining a machine learning model may be included.
[0071] In at least one embodiment, when service 920 includes an AI service (e.g., an inference service), one or more machine learning models associated with an application for anomaly detection (e.g., tumors, abnormal growths, scarring, etc.) may be executed by calling the inference service (e.g., an inference server) (as an API call) to execute the machine learning model, or its processing, as part of the application execution. In at least one embodiment, when another application includes one or more machine learning models for a segmentation task, the application may call the inference service to execute a machine learning model for executing one or more of the processing operations associated with the segmentation task. In at least one embodiment, software 918 implementing the advanced processing and inference pipeline may be rationalized because each application may call the same inference service to execute one or more inference tasks.
[0072] In at least one embodiment, hardware 922 may include a GPU, a CPU, a DPU, a graphics card, an AI / deep learning system (e.g., an AI supercomputer such as NVIDIA's DGX supercomputer system), a cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 922 may be used to provide efficient and dedicated support for software 918 and service 920 of the introduction system 906. In at least one embodiment, the use of GPU processing for performing processing locally (e.g., at facility 902) may be implemented within an AI / deep learning system, a cloud system, and / or other processing components of the introduction system 906 to improve the efficiency, accuracy, and effectiveness of game name recognition.
[0073] In at least one embodiment, software 918 and / or service 920 may be optimized for GPU processing related to deep learning, machine learning, and / or high-performance computing, by way of non-limiting example. In at least one embodiment, at least a portion of the computing environment of introduction system 906 and / or training system 904 may be executed using GPU-optimized software (e.g., a combination of hardware and software of NVIDIA's DGX system) in one or more supercomputers of a data center, or in a high-performance computing system. In at least one embodiment, hardware 922 may include any number of GPUs, which may be called to perform parallel processing of data as described herein. In at least one embodiment, the cloud platform may further include GPU processing for GPU-optimized execution of deep learning tasks, machine learning tasks, or other computing tasks. In at least one embodiment, the cloud platform (e.g., NVIDIA's NGC) may be executed using an AI / deep learning supercomputer (e.g., provided by NVIDIA's DGX system), and / or GPU-optimized software as a platform for hardware abstraction and scaling. In at least one embodiment, the cloud platform may integrate a plurality of GPU application container clustering systems or orchestration systems (e.g., KUBERNETES) to enable seamless scaling and load balancing.
[0074] FIG. 10 is a system diagram showing an example system 1000 for generating and introducing an imaging introduction pipeline according to at least one embodiment. In at least one embodiment, system 1000 may be used to implement process 900 of FIG. 9 and / or other processes including advanced processing and inference pipelines. In at least one embodiment, system 1000 may include a training system 904 and an introduction system 906. In at least one embodiment, training system 904 and introduction system 906 may be implemented using software 918, services 920, and / or hardware 922 as described herein.
[0075] In at least one embodiment, system 1000 (e.g., training system 904 and / or introduction system 906) may be implemented in a cloud computing environment (e.g., cloud 1026). In at least one embodiment, system 1000 may be implemented locally with respect to a facility or as a combination of cloud and local computing resources. In at least one embodiment, access to the APIs of cloud 1026 may be limited to authorized users via established security measures or protocols. In at least one embodiment, the security protocol may include a web token, which may be signed by an authentication (e.g., AuthN, AuthZ, Gluecon, etc.) service and may have appropriate permissions. In at least one embodiment, the APIs of the virtual appliances (described herein) or other instantiations of system 1000 may be limited to a set of public IPs that have been vetted or approved for communication.
[0076] In at least one embodiment, the various components of system 1000 may communicate with each other using any of a variety of different types of networks including, but not limited to, local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communication between the facility and the components of system 1000 (e.g., to send an inference request, to receive the result of an inference request, etc.) may be communicated via one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (e.g., Ethernet®), etc.
[0077] In at least one embodiment, the training system 904 may execute a training pipeline 1004 similar to that described herein with respect to FIG. 9. In at least one embodiment, if one or more machine learning models are to be used in the introduction pipeline 1010 by the introduction system 906, the training pipeline 1004 may be used to train or retrain one or more (e.g., pre-trained) models and / or one or more of the pre-trained models 1006 may be implemented (e.g., without the need for retraining or updating). In at least one embodiment, as a result of the training pipeline 1004, an output model 916 may be generated. In at least one embodiment, the training pipeline 1004 may include any number of processing steps, namely, AI-assisted annotation 910, labeling or annotation of feedback data 908 to generate labeled data 912, model selection from the model registry, model training 914, training, retraining, or updating of the model, and / or other processing steps. In at least one embodiment, different training pipelines 1004 may be used for different machine learning models used by the introduction system 906. In at least one embodiment, a training pipeline 1004 similar to the first example described with respect to FIG. 9 may be used for the first machine learning model, a training pipeline 1004 similar to the second example described with respect to FIG. 9 may be used for the second machine learning model, and a training pipeline 1004 similar to the third example described with respect to FIG. 9 may be used for the third machine learning model. In at least one embodiment, any combination of tasks within the training system 904 may be used depending on what is required for each respective machine learning model. In at least one embodiment, one or more of the machine learning models may already be trained and ready for introduction, such that the machine learning model may not undergo any processing by the training system 904 and may be implemented by the introduction system 906.
[0078] In at least one embodiment, the output model 916 and / or the pre-trained model 1006 may include any type of machine learning model, depending on the embodiment or embodiments. In at least one embodiment, without limitation, the machine learning model used by the system 1000 may be a linear regression, logistic regression, decision tree, support vector machine (SVM), naive Bayes, k-nearest neighbor (Knn), k-means clustering, random forest, dimensionality reduction algorithm, gradient boosting algorithm, neural network (e.g., auto-encoder, convolutional, recurrent, perceptron, long / short term memory (LSTM), Bi-LSTM, Hopfield, Boltzmann, deep belief, inverse convolutional, adversarial generation, liquid state machine, etc.), and / or other types of machine learning models.
[0079] In at least one embodiment, the training pipeline 1004 may include AI-assisted annotation. In at least one embodiment, the labeled data 912 (e.g., conventional annotation) may be generated by any number of techniques. In at least one embodiment, the labels or other annotations may be generated within a drawing program (e.g., an annotation program), a computer-aided design (CAD) program, a labeling program, an annotation or other type of program suitable for generating ground truth labels, and / or in some instances, may be handwritten. In at least one embodiment, the ground truth data may be generated synthetically (e.g., generated from a computer model or rendering), generated realistically (e.g., designed and generated from real-world data), machine automated (e.g., using feature analysis and learning to extract features from data and then generate labels), human annotated (e.g., a labeler or annotation expert may define the location of the labels), and / or combinations thereof. In at least one embodiment, for each instance of feedback data 908 (or other type of data used by a machine learning model), there may be corresponding ground truth data generated by the training system 904. In at least one embodiment, in addition to or instead of the AI-assisted annotation included in the training pipeline 1004, the AI-assisted annotation may be performed as part of the introduction pipeline 1010. In at least one embodiment, the system 1000 may include a multi-layer platform, which may include a software layer (e.g., software 918) of a diagnostic application (or other type of application) that can perform one or more medical imaging and diagnostic functions.
[0080] In at least one embodiment, the software layer may be implemented as a secure, encrypted, and / or authenticated API through which an application or container may be called (e.g., invoked) from an external environment (e.g., facility 902). In at least one embodiment, the application may then call or execute one or more services 920 to perform computational, AI, or visualization tasks associated with each application, and the software 918 and / or services 920 may utilize the hardware 922 to perform the processing tasks in an efficient and effective manner.
[0081] In at least one embodiment, the ingestion system 906 may execute an ingestion pipeline 1010. In at least one embodiment, the ingestion pipeline 1010 may include any number of applications, which may be applied continuously, discontinuously, or otherwise to feedback data (and / or other types of data), including the AI-assisted annotation described above. In at least one embodiment, as described herein, the ingestion pipeline 1010 for an individual device may be referred to as a virtual appliance for the device. In at least one embodiment, depending on the information required for the data generated by the device, there may be more than one ingestion pipeline 1010 for one device.
[0082] In at least one embodiment, the applications available to the ingestion pipeline 1010 may include any application that can be used to perform processing tasks on feedback data or other data from the device. In at least one embodiment, since various applications may share common image operations, in some embodiments, a data augmentation library (e.g., as one of the services 920) may be used to accelerate these operations. In at least one embodiment, a parallel computing platform 1030 may be used to GPU-accelerate these processing tasks to avoid the bottleneck of conventional processing techniques that rely on CPU processing.
[0083] In at least one embodiment, the deployment system 906 may include a user interface 1014 (e.g., a graphical user interface, a web interface, etc.), and the user interface 1014 is used to select an application for inclusion in the deployment pipeline 1010, to place the application, to modify or change the application or its parameters or structure, to use and interact with the deployment pipeline 1010 during setup and / or deployment, and / or to interact with the deployment system 906 in other ways. In at least one embodiment, although not shown with respect to the training system 904, the user interface 1014 (or a different user interface) may be used to select a model for use in the deployment system 906, to select a model to be trained or retrained in the training system 904, and / or to interact with the training system 904 in other ways.
[0084] In at least one embodiment, in addition to the application orchestration system 1028, a pipeline manager 1012 may be used to manage interactions between the applications or containers of the onboarding pipeline 1010 and the services 920 and / or the hardware 922. In at least one embodiment, the pipeline manager 1012 may be configured to facilitate interactions from application to application, from an application to the service 920, and / or from an application or service to the hardware 922. In at least one embodiment, although illustrated as being included in the software 918, this is not intended to be limiting, and in some instances, the pipeline manager 1012 may be included in the service 920. In at least one embodiment, the application orchestration system 1028 (e.g., Kubernetes, DOCKER, etc.) may include a container orchestration system that can group applications into containers as logical units for coordination, management, scaling, and onboarding. In at least one embodiment, rather than associating applications (e.g., rebuilt applications, segmented applications, etc.) from the onboarding pipeline 1010 with individual containers, each application can execute within a self - contained environment (e.g., kernel level) to improve speed and efficiency.
[0085] In at least one embodiment, each application and / or container (or its image) may be developed, modified, and introduced individually (e.g., a first user or developer may develop, modify, and introduce a first application, and a second user or developer may develop, modify, and introduce a second application separately from the first user or developer), thereby enabling concentration and attention on the tasks of one application and / or container without being interrupted by the tasks of another application or container. In at least one embodiment, communication and cooperation between different containers or applications may be assisted by the pipeline manager 1012 and the application orchestration system 1028. In at least one embodiment, as long as the predicted inputs and / or outputs of each container or application are known to the system (e.g., based on the structure of the application or container), the application orchestration system 1028 and / or the pipeline manager 1012 can facilitate communication between and resource sharing among the respective applications or containers. In at least one embodiment, since one or more of the applications or containers in the introduction pipeline 1010 can share the same services and resources, the application orchestration system 1028 may orchestrate services or resources, perform load balancing, and determine sharing among different applications or containers. In at least one embodiment, a scheduler may be used to track the resource requirements of applications or containers, the current or planned usage of these resources, and the availability of resources. In at least one embodiment, the scheduler may thus allocate resources to different applications and distribute resources among applications considering the system requirements and availability.In some examples, the scheduler (and / or other components of the application orchestration system 1028) may determine resource availability and allocation based on constraints imposed on the system (e.g., user constraints), such as quality of service (QoS), the urgency of data output required (e.g., to determine whether to perform real-time processing or deferred processing).
[0086] In at least one embodiment, the services 920 utilized and shared by the applications or containers of the introduction system 906 may include computing services 1016, AI services 1018, visualization services 1020, and / or other types of services. In at least one embodiment, an application may call (e.g., execute) one or more of the services 920 to perform processing operations for the application. In at least one embodiment, the computing service 1016 may be utilized by an application to perform supercomputing or other high-performance computing (HPC) tasks. In at least one embodiment, parallel processing may be performed using the computing service 1016 (e.g., using the parallel computing platform 1030) to substantially simultaneously process data through one or more of the applications and / or to substantially simultaneously process one or more tasks of one application. In at least one embodiment, the parallel computing platform 1030 (e.g., NVIDIA's CUDA) may enable general-purpose computing on graphics processing units (GPGPUs) (e.g., GPU 1022). In at least one embodiment, the software layer of the parallel computing platform 1030 may provide a virtual instruction set and access to the parallel computing elements of the GPU to execute computing kernels. In at least one embodiment, the parallel computing platform 1030 may include memory, and in some embodiments, the memory may be shared among multiple containers and / or between different processing tasks within one container. In at least one embodiment, inter-process communication (IPC) calls may be generated for multiple containers and / or multiple processes within a container to use the same data from the shared segment of the memory of the parallel computing platform 1030 (e.g., when multiple different stages of an application or multiple applications are processing the same information).In at least one embodiment, rather than creating a copy of the data and moving the data to different locations in memory (e.g., read / write operations), the same data at the same location in memory may be used for any number of processing tasks (e.g., at the same time, different times, etc.). In at least one embodiment, when data is used and new data is generated as a result of processing, this information about the new location of the data may be stored in and shared among various applications. In at least one embodiment, the location of the data and the location of the updated or modified data may be part of the definition of how the payload is understood within the container.
[0087] In at least one embodiment, the AI service 1018 may be utilized to execute an inference service for executing a machine learning model associated with an application (e.g., tasked with performing one or more processing tasks of the application). In at least one embodiment, the AI service 1018 may utilize the AI system 1024 to execute a machine learning model (e.g., a neural network such as a CNN) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, the application of the introduction pipeline 1010 may perform inference on imaging data (e.g., DICOM data, RIS data, CIS data, REST-compliant data, RPC data, raw data, etc.) using one or more of the output model 916 from the training system 904 and / or other models of the application. In at least one embodiment, two or more instances of inference using the application orchestration system 1028 (e.g., a scheduler) may be available. In at least one embodiment, the first category may include a high-priority / low-latency path that can achieve a higher service level agreement, such as for performing inference for emergency requests during an emergency or for a radiologist during a diagnosis. In at least one embodiment, the second category may include a standard-priority path that can be used for non-emergency requests or when the analysis may be performed later. In at least one embodiment, the application orchestration system 1028 may allocate resources (e.g., service 920 and / or hardware 922) based on priority paths for different inference tasks of the AI service 1018.
[0088] In at least one embodiment, the shared storage may be attached to the AI service 1018 within the system 1000. In at least one embodiment, the shared storage may operate as a cache (or other type of storage device) and may be used to process inference requests from applications. In at least one embodiment, when an inference request is sent, the request may be received by a set of API instances of the ingress system 906, and one or more instances may be selected (e.g., for best fit, for load balancing, etc.) for the request to be processed. In at least one embodiment, to process the request, the request may be placed in a database, and the machine learning model may be identified from the model registry 924 if it is not already in the cache, and the verification step may ensure that the appropriate machine learning model is loaded into the cache (e.g., shared storage), and / or a copy of the model may be saved in the cache. In at least one embodiment, if the application has not yet been run or there are not sufficient instances of the application, a scheduler (e.g., the pipeline manager 1012) may be used to start the application referenced in the request. In at least one embodiment, if the inference server for running the model has not yet been started, the inference server may be started. In at least one embodiment, any number of inference servers may be started per model. In at least one embodiment, in a pull model where the inference servers are clustered, the model may be cached whenever load balancing is advantageous. In at least one embodiment, the inference server may be statically loaded onto the corresponding distributed server.
[0089] In at least one embodiment, the inference may be performed using an inference server that runs within a container. In at least one embodiment, an instance of the inference server may be associated with a model (optionally with multiple versions of the model). In at least one embodiment, when a request to perform an inference on a model is received, if no instance of the inference server exists, a new instance may be loaded. In at least one embodiment, when starting the inference server, the model may be passed to the inference server, such that as long as the inference server is running as a different instance, the same container may be used to serve different models.
[0090] In at least one embodiment, during the execution of an application, an inference request may be received for a given application, a container (e.g., hosting an instance of an inference server) may be loaded (if not already loaded), and a start procedure may be called. In at least one embodiment, the preprocessing logic of the container may load, decode, and / or perform any additional preprocessing on the input data (e.g., using a CPU and / or GPU and / or DPU). In at least one embodiment, once the data is prepared for inference, the container may perform the inference on the data as needed. In at least one embodiment, this may include a single inference call for one image (e.g., an X-ray of a hand), or may request inferences for hundreds of images (e.g., a chest CT). In at least one embodiment, the application may summarize the results before completion, which may include, without limitation, generating a single confidence score, pixel-level segmentation, voxel-level segmentation, visualization, or text for summarizing the findings. In at least one embodiment, different models or applications may be assigned different priorities. For example, there may be models with a real-time (TAT less than 1 minute) priority, and there may be models with a low priority (e.g., TAT less than 10 minutes). In at least one embodiment, the model execution time may be measured from the requesting facility or entity and may include the partner network transit time in addition to the execution on the inference service.
[0091] In at least one embodiment, the transfer of requests between the service 920 and the inference application may be hidden behind a software development kit (SDK), and reliable transfer may be provided through a queue. In at least one embodiment, in response to a combination of individual application / tenant IDs, requests are queued via an API, and the SDK retrieves requests from the queue and provides the requests to the application. In at least one embodiment, the name of the queue may be provided in an environment where the SDK picks up requests. In at least one embodiment, asynchronous communication via a queue may be useful because it allows any instance of the application to pick up work when that communication becomes available. In at least one embodiment, results may be returned via the queue to prevent data loss. In at least one embodiment, the highest priority work may proceed to a queue where most instances of the application are connected to the queue, while the lowest priority work may proceed to a queue that processes tasks in the order received, where one instance is connected to the queue, so the queue can also provide a function for segmenting work. In at least one embodiment, the application may execute on a GPU-accelerated instance generated in the cloud 1026, and the inference service may perform inference on the GPU.
[0092] In at least one embodiment, visualization may be generated using visualization service 1020 to view the output of application and / or onboarding pipeline 1010. In at least one embodiment, GPU 1022 may be utilized by visualization service 1020 to generate the visualization. In at least one embodiment, rendering effects such as ray tracing may be implemented by visualization service 1020 to generate higher quality visualizations. In at least one embodiment, the visualization may include, without limitation, rendering of 2D images, rendering of 3D volumes, reconstruction of 3D volumes, 2D tomography slices, virtual reality displays, augmented reality displays, and the like. In at least one embodiment, a virtual interactive display or interactive environment (e.g., a virtual environment) for a user of the system to interact with may be generated using a virtualized environment. In at least one embodiment, visualization service 1020 may include an internal visualizer, cinematics, and / or other rendering or image processing capabilities or functions (e.g., ray tracing, rasterization, internal optics, etc.).
[0093] In at least one embodiment, the hardware 922 may include the GPU 1022, the AI system 1024, the cloud 1026, and / or any other hardware used to execute the training system 904 and / or the introduction system 906. In at least one embodiment, the GPU 1022 (e.g., NVIDIA's TESLA and / or QUADRO GPUs) may include any number of GPUs, which may be used to execute processing tasks for the computing service 1016, the AI service 1018, the visualization service 1020, other services, and / or any features or functions of the software 918. For example, with respect to the AI service 1018, pre-processing may be performed on the imaging data (or other types of data used by the machine learning model) using the GPU 1022, post-processing may be performed on the output of the machine learning model, and / or inference may be performed (e.g., the machine learning model may be executed). In at least one embodiment, the cloud 1026, the AI system 1024, and / or other components of the system 1000 may use the GPU 1022. In at least one embodiment, the cloud 1026 may include a GPU-optimized platform for deep learning tasks. In at least one embodiment, the AI system 1024 may use a GPU, and at least a portion assigned the role of the cloud 1026, or deep learning or inference, may be executed using one or more AI systems 1024. Thus, although the hardware 922 is shown as individual components, this is not intended to be limiting, and any components of the hardware 922 may be combined with and utilized by any other components of the hardware 922.
[0094] In at least one embodiment, the AI system 1024 may include a dedicated computing system (e.g., a supercomputer or HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, the AI system 1024 (e.g., NVIDIA's DGX) may include GPU-optimized software (e.g., a software stack), which may be executed using multiple GPUs 1022 in addition to a CPU, RAM, storage, and / or other components, features, or functions. In at least one embodiment, one or more AI systems 1024 may be implemented in the cloud 1026 (e.g., in a data center) to perform some or all of the AI-based processing tasks of the system 1000.
[0095] In at least one embodiment, cloud 1026 may include a GPU-accelerated infrastructure (e.g., NVIDIA's NGC), which may provide a GPU-optimized platform for executing the processing tasks of system 1000. In at least one embodiment, cloud 1026 may include an AI system 1024 (e.g., as a platform for hardware abstraction and scaling) for executing one or more of the AI-based tasks of system 1000. In at least one embodiment, cloud 1026 may utilize multiple GPUs and be integrated with an application orchestration system 1028 to enable seamless scaling and load balancing between applications and services 920. In at least one embodiment, cloud 1026 may be tasked with executing at least a portion of the services 920 of system 1000, including the computing service 1016, AI service 1018, and / or visualization service 1020 described herein. In at least one embodiment, cloud 1026 may perform large and small batch inferences (e.g., execution of NVIDIA's TensorRT), provide an accelerated parallel computing API and platform 1030 (e.g., NVIDIA's CUDA), execute an application orchestration system 1028 (e.g., KUBERNETES), provide a graphics rendering API and platform (e.g., ray tracing for generating high-quality cinematics, 2D graphics, 3D graphics, and / or other rendering techniques), and / or provide other functions for system 1000.
[0096] In at least one embodiment, to protect patient confidentiality (e.g., where patient data or records are to be used off - site), cloud 1026 may include a registry such as a deep - learning container registry. In at least one embodiment, the registry may store containers for the instantiation of applications that can perform pre - processing, post - processing, or other processing tasks on patient data. In at least one embodiment, cloud 1026 may receive data including patient data as well as sensor data in containers, perform only the processing requested for the sensor data in these containers, and then transfer the resulting output and / or visualization to the appropriate parties and / or devices (e.g., in - house medical devices used for visualization or diagnosis) without the need to extract, store, or otherwise access the patient data at all. In at least one embodiment, the confidentiality of patient data is protected in accordance with HIPAA and / or other data regulations.
[0097] At least one embodiment of the present disclosure can be described in view of the following clauses.
[0098] Clause 1. A method for introducing a customized machine - learning model (MLM), the method comprising: accessing a plurality of trained MLMs, each associated with an initial configuration setting; providing a user interface (UI) for receiving user input indicating a selection of one or more of the plurality of trained MLMs and a modification to the initial configuration setting for the one or more selected MLMs; determining, based on the user input, a modified configuration setting for the one or more selected MLMs; causing the execution of a build engine to modify the one or more selected MLMs according to the modified configuration setting; causing the execution of an introduction engine to introduce the one or more modified MLMs; and causing the display on the UI of the representation of the one or more introduced MLMs.
[0099] The method according to clause 1, wherein in clause 2, one or more selected MLMs are pre-trained using a first set of training data.
[0100] The method according to clause 2, further comprising, in clause 3, receiving a second set of training data for domain-specific training of one or more selected MLMs and triggering the execution of a training engine of a pipeline for performing domain-specific training of one or more selected MLMs.
[0101] The method according to clause 1, further comprising, in clause 4, triggering the display on a UI of a listing of a plurality of pre-trained MLMs before receiving a user input indicating the selection of one or more MLMs, and triggering the execution of an export engine for initializing one or more selected MLMs after receiving a user input indicating the selection of one or more MLMs.
[0102] The method according to clause 4, wherein the step of triggering the execution of the export engine further comprises triggering the display on a UI of at least one representation of the architecture of one or more selected MLMs or the parameters of one or more selected MLMs.
[0103] The method according to clause 4, wherein the step of triggering the execution of the export engine further comprises making one or more selected MLMs available for processing user input data.
[0104] The method according to clause 6, further comprising, in clause 7, receiving user input data, triggering one or more introduced MLMs to be applied to the user input data to generate output data, and triggering the display on a UI of at least one of a representation of the output data or a reference to a stored representation of the output data.
[0105] The method according to clause 1, wherein in clause 8, one or more selected MLMs comprise a neural network model arranged in the pipeline.
[0106] The method according to clause 1, wherein in clause 9, one or more selected MLMs comprise an acoustic neural network model and a language neural network model.
[0107] The method according to clause 9, wherein in clause 10, the modified configuration settings for one or more selected MLMs comprise at least one of the audio buffer size of the acoustic preprocessing stage for the acoustic neural network model, the utterance end setting for the acoustic neural network model, the latency setting for the acoustic neural network model, or the language setting for the language neural network model.
[0108] The method according to clause 8, wherein in clause 11, one or more selected MLMs comprise at least one of a text-to-speech neural network model, a language understanding neural network model, or a question answering neural network model.
[0109] In clause 12, a memory device, and a plurality of trained machine learning models (MLMs) each communicatively coupled to the memory device and each associated with an initial configuration setting, access to the plurality of trained MLMs, receive user input indicating selection of one or more of the plurality of trained MLMs and modification of the initial configuration setting for the one or more MLMs via a user interface (UI), determine modified configuration settings for the one or more selected MLMs based on the user input, trigger execution of a pipeline build engine to modify the one or more selected MLMs according to the modified configuration settings, trigger execution of a pipeline introduction engine to introduce the one or more modified MLMs, and trigger display on the UI of the representation of the one or more introduced MLMs. A system comprising one or more processing devices for performing the above.
[0110] In clause 13, the system according to clause 12, wherein one or more processing devices are further configured to cause display on the UI of a listing of the plurality of trained MLMs prior to receiving user input indicating selection of one or more of the plurality of trained MLMs, and to cause execution of a pipeline export engine to initialize the one or more selected MLMs after receiving user input indicating selection of the one or more MLMs.
[0111] In clause 14, the system according to clause 13, wherein in order to cause execution of the export engine, one or more processing devices are further configured to cause display on the UI of at least one of the architectures of the one or more selected MLMs or a representation of at least one of the parameters of the one or more selected MLMs.
[0112] In clause 15, one or more processing devices are further for receiving user input data, causing one or more introduced MLMs to be applied to the user input data to generate output data, and causing at least one of a presentation of the output data or a reference to a stored presentation of the output data to cause a display on the UI, the system according to clause 12.
[0113] In clause 16, the modified configuration settings for one or more selected MLMs comprise at least one of an audio buffer size of an acoustic preprocessing stage for an acoustic neural network model, an utterance end setting for an acoustic neural network model, a latency setting for an acoustic neural network model, or a language setting for a language neural network model, the system according to clause 12.
[0114] In clause 17, one or more selected MLMs comprise at least one of a speech synthesis neural network model, a language understanding neural network model, or a question answering neural network model, the system according to clause 12.
[0115] In clause 18, a non-transitory computer-readable medium storing instructions thereon, which, when executed by a processing device, cause the processing device to access a plurality of trained machine learning models (MLMs), each associated with an initial configuration setting; provide a user interface (UI) for receiving user input indicating selection of one or more of the plurality of trained MLMs and modification of the initial configuration setting for the one or more MLMs; determine modified configuration settings for the one or more selected MLMs based on the user input; trigger execution of a build engine of a pipeline for modifying the one or more selected MLMs according to the modified configuration settings; trigger execution of an import engine of a pipeline for importing the one or more modified MLMs; and cause display on the UI of the representation of the one or more imported MLMs.
[0116] In clause 19, the instructions are further for causing the processing device to cause display on the UI of a listing of the plurality of trained MLMs before receiving user input indicating selection of one or more MLMs, and to trigger execution of an export engine of a pipeline for initializing the one or more selected MLMs after receiving the user input indicating selection of the one or more MLMs, the computer-readable medium of the clause computer-readable medium according to claim 18.
[0117] In clause 20, the instructions are further for causing the processing device to cause display on the UI of a representation of at least one of the architecture of the one or more selected MLMs or the parameters of the one or more selected MLMs, for triggering execution of the export engine, the computer-readable medium according to clause 19.
[0118] In clause 21, the instructions further cause the processing device to receive user input data and cause one or more introduced MLMs to be applied to the user data to generate output data, and cause at least one of a presentation of the output data or a reference to a stored presentation of the output data to cause a display on the UI. A computer-readable medium according to clause 18 for this purpose.
[0119] In clause 22, the modified configuration settings for one or more selected MLMs comprise at least one of an audio buffer size for an acoustic preprocessing stage for an acoustic neural network model, an utterance end setting for an acoustic neural network model, a latency setting for an acoustic neural network model, or a language setting for a language neural network model. A computer-readable medium according to clause 20 for this purpose.
[0120] In clause 23, one or more selected MLMs comprise at least one of a speech synthesis neural network model, a language understanding neural network model, or a question answering neural network model. A computer-readable medium according to clause 20 for this purpose.
[0121] Other variations are within the scope of the present disclosure. Accordingly, the disclosed techniques are capable of various modifications and alternative configurations, some of which are illustrated in the drawings and described in detail above. However, there is no intention to limit the present disclosure to the specific one or more disclosed forms, and on the contrary, it is intended to cover all modifications, alternative configurations, and equivalents that fall within the spirit and scope of the disclosure as defined in the claims.
[0122] In the context of describing the disclosed embodiments (in particular, in the context of the following claims), the use of the terms "a", "an", and "the", as well as similar indicators, should be construed to cover both the singular and the plural, unless otherwise specified in this specification or clearly contradicted by the context, and should not be construed as a definition of the terms. The terms "comprising", "having", "including", and "containing" should be construed as open-ended terms (meaning "including but not limited to") unless otherwise specified. "Connected" is construed as being partially or fully enclosed within, attached to, or joined to each other, even if there is something intervening, when it refers to a physical connection without modification. The recitation of a range of values herein is merely intended to serve as a concise way of referring individually to each separate value that falls within the range, unless otherwise specified herein or unless each separate value is incorporated into the specification as if it were individually recited herein. In at least one embodiment, the use of the term "set" (e.g., "a set of items") or "subset" should be construed as a non-empty collection comprising one or more members, unless otherwise specified or contradicted by the context. Further, unless otherwise specified or contradicted by the context, the term "subset" of a corresponding set does not necessarily refer to a strict subset of the corresponding set, and the subset and the corresponding set may be equal.
[0123] Conjunctive terms such as "at least one of A, B, and C" or phrases in the form of "at least one of A, B, and C" are understood in the context generally used to indicate that an item, term, etc. is A, B, or C, or a non-empty subset of the set of A, B, and C, unless there is a specific description to the contrary or it is not explicitly negated by the context. For example, in an illustrative example of a set having three members, the conjunctive phrases "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunctive terms do not generally imply that a particular embodiment requires the presence of at least one of each of A, at least one of B, and at least one of C. Further, unless otherwise stated or not explicitly negated by the context, the term "a plurality of" indicates a plural state (e.g., "a plurality of items" indicates multiple items). In at least one embodiment, the number of items that are a plurality is at least two, but may be more if explicitly stated or indicated by the context. Further, unless otherwise stated or not clear from the context to the contrary, the phrase "based on" means "at least partially based on" and does not mean "based only on".
[0124] The operations of the processes described in this specification can be performed in any suitable order, unless otherwise stated in this specification or clearly precluded by the context. In at least one embodiment, a process such as the process described in this specification (or a variation and / or combination thereof) is executed under the control of one or more computer systems configured with executable instructions and is implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or by a combination thereof. In at least one embodiment, the code is stored in a computer-readable storage medium in the form of a computer program comprising a plurality of instructions executable, for example, by one or more processors. In at least one embodiment, the computer-readable storage medium excludes a transient signal (e.g., a propagating transient electrical or electromagnetic transmission), but includes a non-transient computer-readable storage medium including non-transient data storage circuits (e.g., buffers, caches, and queues) within a transceiver of the transient signal. In at least one embodiment, the code (e.g., executable code or source code) is stored in a set of one or more non-transient computer-readable storage mediums, and the storage mediums store executable instructions that, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described in this specification (or have other memory for storing the executable instructions). In at least one embodiment, the set of non-transient computer-readable storage mediums comprises a plurality of non-transient computer-readable storage mediums, and one or more of the individual non-transient storage mediums of the plurality of non-transient computer-readable storage mediums do not have all the code, but the plurality of non-transient computer-readable storage mediums collectively store all the code.In at least one embodiment, executable instructions are executed such that different instructions are executed by different processors. For example, a non-transitory computer-readable storage medium stores the instructions, a main central processing unit (“CPU”) executes some of the instructions, and a graphics processing unit (“GPU”) and / or potentially a data processing unit (“DPU”) in conjunction with the GPU execute other instructions. In at least one embodiment, different components of a computer system have separate processors, and the different processors execute different subsets of instructions.
[0125] Accordingly, in at least one embodiment, a computer system is configured to implement one or more services that perform the operations of the processes described herein, and such computer systems are composed of applicable hardware and / or software that enable the performance of the operations. Further, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment, is a distributed computer system comprising multiple devices operating in different ways, such that the distributed computer system performs the operations described herein without a single device performing all of the operations.
[0126] Any examples provided herein, or the use of exemplary language (e.g., “such as”) are intended merely to clarify embodiments of the present disclosure and do not limit the scope of the present disclosure unless otherwise claimed. No language in this specification should be construed as indicating any non-claimed element as essential to the practice of the present disclosure.
[0127] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference had been individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
[0128] In the specification and claims, the terms "coupled" and "connected" may be used along with their derivatives. It should be understood that these terms may not be intended as synonyms for each other. Rather, in certain instances, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. Also, "coupled" may mean that two or more elements are not in direct contact with each other but still coact or interact with each other.
[0129] Unless otherwise specifically stated, throughout the specification, terms such as "processing", "computing", "calculating", or "determining" refer to the act and / or process of a computer or computing system, or similar electronic computing device, that manipulates and / or transforms data represented as physical, such as electronic, quantities within a register and / or memory of the computing system into other data similarly represented as physical quantities within a memory, register, or other such information storage device, transmission device, or display device of the computing system.
[0130] Similarly, the term "processor" may refer to any device, or portion of a device, that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. By way of non-limiting example, a "processor" may be a CPU or a GPU. A "computing platform" may comprise one or more processors. A "software" process, as used herein, may include software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Also, each process may refer to multiple processes for executing instructions serially or in parallel, continuously or intermittently. In at least one embodiment, the terms "system" and "method" are used interchangeably herein only where one or more methods can be embodied by a system and the method can be considered a system.
[0131] In this specification, it is possible to refer to obtaining, acquiring, receiving, or inputting analog data or digital data into a subsystem, computer system, or computer-implemented machine. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog data or digital data can be realized in various ways, such as receiving the data as a parameter of a function call or a call to an application programming interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog data or digital data can be realized by transferring the data via a serial or parallel interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog data or digital data can be realized by transferring the data via a computer network from an entity providing the data to an entity acquiring the data. Also, in at least one embodiment, it is possible to refer to providing, outputting, transmitting, sending, or presenting analog data or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog data or digital data can be realized by transferring the data as a parameter of an input or output of a function call, an application programming interface, or an inter-process communication mechanism.
[0132] The description in this specification describes exemplary embodiments of the described techniques, but other architectures may be used to implement the described functions, and this other architecture is intended to be within the scope of the present disclosure. Further, for the purpose of explanation, specific assignments of roles may be defined, but various functions and roles may be assigned and divided in different ways depending on the situation.
[0133] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.
Claims
1. A method for introducing a customized machine learning model (MLM), the method comprising: accessing a plurality of trained MLMs, each associated with an initial configuration setting; providing a user interface (UI) for receiving user input indicating selection of one or more of the plurality of trained MLMs and modification to the initial configuration setting for the one or more MLMs; determining, based on the user input, a modified configuration setting for the one or more selected MLMs; causing execution of a build engine to modify the one or more selected MLMs according to the modified configuration setting; causing execution of an introduction engine to introduce the one or more modified MLMs; causing display on the UI of the representation of the one or more introduced MLMs; pre-training the one or more selected MLMs using a first set of training data; receiving a second set of training data for domain-specific training of the one or more selected MLMs; causing execution of a pipeline training engine to perform the domain-specific training of the one or more selected MLMs A method comprising the above steps.
2. Before receiving the user input indicating selection of the one or more MLMs, causing display on the UI of a listing of the plurality of trained MLMs; and after receiving the user input indicating selection of the one or more MLMs, causing execution of an export engine to initialize the one or more selected MLMs The method according to claim 1, further comprising the above steps.
3. The step of causing execution of the export engine further comprises: causing display on the UI of at least one of the architecture of the one or more selected MLMs or a representation of one or more of the parameters of the one or more selected MLMs The method according to claim 2, further comprising the above step.
4. The step of causing execution of the export engine further comprises: making the one or more selected MLMs available for processing user input data The method according to claim 2, further comprising the above step.
5. The step of receiving the user input data; Causing the one or more introduced MLMs to be applied to the user input data to generate output data; Causing at least one of a representation of the output data or a reference to a stored representation of the output data to be displayed on the UI The method according to claim 4, further comprising.
6. The method according to claim 1, wherein the one or more selected MLMs comprise a neural network model arranged in the pipeline.
7. The method according to claim 1, wherein the one or more selected MLMs comprise an acoustic neural network model and a language neural network model.
8. The modified configuration settings for the one or more selected MLMs are The audio buffer size of the acoustic preprocessing stage for the acoustic neural network model, The end-of-utterance setting for the acoustic neural network model, The latency setting for the acoustic neural network model, or The language setting for the language neural network model The method according to claim 7, comprising at least one of.
9. The method according to claim 6, wherein the one or more selected MLMs comprise at least one of a text-to-speech neural network model, a language understanding neural network model, or a question answering neural network model.
10. A memory device; Communicatively coupled to the memory device, Each accessing a plurality of trained machine learning models (MLMs) associated with initial configuration settings; Providing a user interface (UI) for receiving user input indicating selection of one or more of the plurality of trained MLMs and modification of the initial configuration settings for the one or more MLMs; Determining modified configuration settings for the one or more selected MLMs based on the user input; Causing execution of a pipeline build engine to modify the one or more selected MLMs according to the modified configuration settings; causing execution of an introduction engine of the pipeline for introducing the one or more modified MLMs; causing display on the UI of the representation of the one or more introduced MLMs; pre-training the one or more selected MLMs using a first set of training data; receiving a second set of training data for domain-specific training of the one or more selected MLMs; causing execution of a training engine of the pipeline for performing the domain-specific training of the one or more selected MLMs one or more processing devices for performing; A system comprising. [
11. ] The one or more processing devices further causing display on the UI of a listing of the plurality of trained MLMs, prior to receiving the user input indicating selection of the one or more MLMs; causing execution of an export engine of the pipeline for initializing the one or more selected MLMs, after receiving the user input indicating selection of the one or more MLMs The system according to claim 10, for performing. [
12. ] To cause execution of the export engine, the one or more processing devices further cause display on the UI of at least one of an architecture of the one or more selected MLMs, or a representation of at least one of the parameters of the one or more selected MLMs The system according to claim 11, for performing. [
13. ] The one or more processing devices further receiving user input data; causing the one or more introduced MLMs to be applied to the user input data to generate output data; causing display on the UI of at least one of a representation of the output data, or a reference to a stored representation of the output data The system according to claim 10, for performing. [
14. ] The modified configuration settings for the one or more selected MLMs are an audio buffer size of an acoustic preprocessing stage for an acoustic neural network model; an utterance end setting for the acoustic neural network model; The latency setting for the acoustic neural network model, or The language setting for the language neural network model The system according to claim 10, comprising at least one of them.
15. The system according to claim 10, wherein the one or more selected MLMs comprise at least one of a speech synthesis neural network model, a language understanding neural network model, or a question answering neural network model.
16. A non-transitory computer-readable medium storing instructions thereon, which, when executed by a processing device, cause the processing device to Access a plurality of trained machine learning models (MLMs), each associated with an initial configuration setting, Provide a user interface (UI) for receiving user input indicating selection of one or more of the plurality of trained MLMs and modification of the initial configuration settings for the one or more MLMs, Determine modified configuration settings for the one or more selected MLMs based on the user input, Cause execution of a pipeline build engine to modify the one or more selected MLMs according to the modified configuration settings, Cause execution of a pipeline introduction engine to introduce the one or more modified MLMs, Cause display of the representation of the one or more introduced MLMs on the UI, Pre-train the one or more selected MLMs using a first set of training data, Receive a second set of training data for domain-specific training of the one or more selected MLMs, Cause execution of a pipeline training engine to perform the domain-specific training of the one or more selected MLMs A non-transitory computer-readable medium that causes the above to be performed.
17. The instructions further cause the processing device to Cause display of a listing of the plurality of trained MLMs on the UI before receiving the user input indicating selection of the one or more MLMs. After receiving the user input indicating the selection of the one or more MLMs, causing execution of an export engine of the pipeline for initializing the one or more selected MLMs The computer-readable medium according to claim 16, which is for causing the above to be performed
18. In order to cause execution of the export engine, the instructions further cause the processing device to Cause a display on the UI of at least one representation of the architecture of the one or more selected MLMs or of the parameters of the one or more selected MLMs The computer-readable medium according to claim 17, which is for causing the above to be performed
19. The instructions further cause the processing device to Receive user input data, Cause the one or more introduced MLMs to be applied to the user data to generate output data, and Cause a display on the UI of at least one of a representation of the output data or a reference to a stored representation of the output data The computer-readable medium according to claim 16, which is for causing the above to be performed
20. The modified configuration settings for the one or more selected MLMs include The audio buffer size of the acoustic preprocessing stage for the acoustic neural network model, The utterance end setting for the acoustic neural network model, The latency setting for the acoustic neural network model, or The language setting for the language neural network model The computer-readable medium according to claim 18, comprising at least one of the above
21. The computer-readable medium according to claim 18, wherein the one or more selected MLMs comprise at least one of an audio synthesis neural network model, a language understanding neural network model, or a question answering neural network model
Citation Information
Patent Citations
Neural network construction method applied to medical ultrasonic image
CN111047563A
Acoustic model learning device, model learning device, model learning method, and program
JP2018180045A
Machine for development and deployment of analytical models
US20170178027A1
Artificial intelligence model and data collection / development platform
US20180089591A1
Learned model provision method and learned model provision device
WO2018142766A1