Language model assisted system installation, diagnosis and commissioning
The language model assists non-expert users in the installation, troubleshooting and maintenance of complex systems, solves the problems of advanced professional knowledge requirements, and achieves low-cost and efficient system operation and troubleshooting.
Patent Information
- Application Number
- CN202510004962.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-03
- Filing Date
- 2025-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
The installation, troubleshooting and maintenance of complex systems often requires a lot of advanced expertise, resulting in high operating costs and temporary loss of system availability when immediate access to expert help is not available.
Using Language Model (LLM) to provide real-time guidance, assist non-expert users in system installation, troubleshooting and maintenance through natural language queries and responses, and generate concise instructions in combination with system-specific documents and diagnostic tools.
Reduces reliance on expert help, reduces operational costs, and improves system availability and maintenance efficiency in failure situations.
Smart Images

Figure CN120255908A_ABST
Abstract
Description
Technical Field
[0001] At least one embodiment relates to computing resources for providing technical support for complex systems. For example, at least one embodiment relates to systems and techniques for using language models to facilitate the installation, diagnosis, and debugging of complex systems. Background Art
[0002] Well-trained language models (e.g., large language models (LLMs)) are capable of supporting conversations in natural language, understanding the intentions and emotions of speakers, explaining complex topics, generating new text upon receiving appropriate prompts, providing suggestions on topics of interest to users, processing images, audio, and / or other data types, and / or performing other functions. LLMs are typically self-supervised trained on large amounts of text data and / or other data types according to embodiments and learn to predict the next token and / or missing tokens in a phrase / sentence (which may correspond to sub-words, symbols, words, etc.), detect the intentions and / or emotions of human speakers, determine whether two sentences are relevant or not, and / or perform other basic language tasks. After initial training, LLMs are typically subjected to guided (prompt-based) supervised fine-tuning, which enables the LLMs to acquire deeper language proficiency and / or master more specialized tasks. Supervised fine-tuning involves using learning prompts (questions, hints, etc.) that are accompanied by example texts (e.g., answers, example articles, etc.) as training groundtruths. In reinforcement fine-tuning, human evaluators assign grades that indicate the degree of similarity between the generated text and human-generated text. Brief Description of the Drawings
[0003] Figure 1 is a block diagram of an example computer system capable of using a language model to aid non-expert users in the installation, troubleshooting, and / or maintenance of complex systems according to at least one embodiment;
[0004] Figure 2 illustrates an example computing device that supports the deployment of a language model to assist users in the installation, troubleshooting, and / or maintenance of complex systems according to at least one embodiment;
[0005] Figure 3 illustrates an example training of a language model to assist users in the installation, troubleshooting, and / or maintenance of complex systems according to at least one embodiment;
[0006] Figure 4 illustrates an example use of a trained language model in assisting users in the installation, troubleshooting, and / or maintenance of complex systems according to at least one embodiment;
[0007] Figure 5 is a flowchart of an example method for using a language model to assist a user in installing, troubleshooting, and / or maintaining a complex system according to at least one embodiment;
[0008] Figure 6 is a flowchart of an example method for training a language model to assist a user in installing, troubleshooting, and / or maintaining a complex system according to at least one embodiment;
[0009] Figure 7A illustrates inference and / or training logic according to at least one embodiment;
[0010] Figure 7B illustrates inference and / or training logic according to at least one embodiment;
[0011] Figure 8 illustrates training and deployment of a neural network according to at least one embodiment;
[0012] Figure 9 is an example data flow diagram of an advanced computing pipeline according to at least one embodiment; and
[0013] Figure 10 is a system diagram of an example system for training, adapting, instantiating, and deploying a machine learning model in an advanced computing pipeline according to at least one embodiment. DETAILED DESCRIPTION
[0014] Complex systems typically include a combination of multiple high-level hardware components, software code or modules, and / or firmware code or modules. Initial installation, troubleshooting, and debugging of such systems typically require a significant amount of work, advanced expertise, or both, often far beyond the technical knowledge of most users. For example, diagnosis of a complex system can involve the use of complex test equipment capable of generating hundreds or more different error / defect indicators or codes. Individual error codes can indicate a malfunction in one or more hardware units and / or software or firmware blocks. System manuals describing such a large number of error codes can be very long and are written in technical language that is difficult for anyone other than experienced engineers or technicians to understand. Similarly, even routine maintenance, updates, and / or part replacements of complex systems typically require the work of advanced experts. This increases the operating costs of such systems and can result in a temporary loss of system availability when such expert assistance is not immediately available.
[0015] Aspects and embodiments of the present disclosure address these and other technical challenges by providing systems and techniques that utilize the capabilities of language models (LMs), including large language models (LLMs), to provide real-time guidance, instructions, and recommendations for the installation (including assembly), troubleshooting, and / or maintenance of complex systems to non-expert users. In some embodiments, a user performing one of such (or similar) operations may describe a problem to the LM in natural language, such as "The Ethernet network card needs to be replaced or updated", and the LM may respond with instructions on how to replace the card and run diagnostic (e.g., hardware, software, and / or firmware) tools, which may be part of a system installation and / or maintenance kit, built-in tools, etc. The diagnostic tools may test the system and determine whether the new card is working properly by generating a success code or one or more error codes indicating a system failure (e.g., "Error 2576 - driver conflict" or simply "Error" 2576). The code may be provided to the user, e.g., via a suitable user interface (e.g., a general computer monitor or a dedicated diagnostic tool monitor). In the case where the diagnostic tools output one or more error codes, the user may enter the displayed code as another prompt into the LM, such as "The system displays Error 2576". The LM may process the prompt and generate instructions for the user, such as "Download and run the latest driver update from the system's support center". After the user completes the instructions, the user may re-run the diagnostic tools, which may identify any remaining system failures and output additional error codes. The user may use such additional error codes in further prompts entered into the LM to receive further instructions in plain natural language that non-experts can understand. This process may end when the diagnostic tools output a success code or any other indication of no system failures. In some embodiments, the LM may use a complete test log (or any part of the test log) as additional input, which may be automatically appended to the user's prompt. In some embodiments, the LM may use one or more application programming interfaces or plugins to (e.g., recursively) query third-party data sources or applications to retrieve (e.g., via retrieval augmentation) additional information or context to help provide the most accurate, precise, and / or useful answer to the query or prompt.
[0016] The LM can be trained, for example, on the system - specific documents (such as installation and diagnostic manuals (IDM) and system architecture documents (SAD), third - party plugins, and / or API information, etc.) on the support center side. For example, the IDM can be an expert - level manual for system engineers and / or technicians and can include technical descriptions of diagnostic error codes (or some other fault indicators). The SAD can include descriptions of various hardware blocks and components of the system, software modules, firmware modules, interconnections of blocks and modules, data flows that occur during system operation, mappings of system inputs and outputs, mappings of addresses of various blocks, modules, and / or data paths, etc. The LM can be trained using any additional documents published by system developers or otherwise available to expert technicians. The LM can also be trained using multiple training diagnostic logs indicating one or more system failures. Some training diagnostic logs can be logs of actual failures encountered during system installation (including assembly), updates, or operation. Some training diagnostic logs can be logs generated by diagnostic tools for hardware and / or software failures deliberately induced by developers to generate training inputs for the LM. However, some training diagnostic logs can be synthetic logs generated (simulated) by developers or by randomly selecting one or more error codes. The ground truth for training the LM can include sample responses to user prompts prepared by system developers (e.g., in the case of supervised LM training) that the LM attempts to mimic, and a suitable loss function is used to evaluate the difference between the training responses of the LM and the sample (ground truth) responses. In some cases, the training responses generated by the LM can be evaluated (e.g., by developers, expert technicians, or non - professional user) at any suitable ratio (e.g., 1 to N), which indicates the usefulness or effectiveness of the training responses (in the case of reinforcement learning). In still other cases, the training engine performing LM training can access a database of historical user queries and corresponding expert responses (which may have been provided via phone, chat, and / or some other recorded customer service means) and then use such queries / responses for unsupervised training of the LM.
[0017] Advantages of the disclosed technology include, but are not limited to, providing timely and effective technical support to non - expert users of complex systems during initial installation, troubleshooting, and / or maintenance of such systems. This enables users to perform system services directly when such services are needed in many cases without waiting for expert assistance. This reduces the cost of system operation and the amount of time such systems remain offline in case of failures or maintenance issues. Additionally, the technology disclosed herein can allow low - complexity repair facilities to provide more advanced technical support than otherwise possible and minimize the number of referrals to system engineers for high - level and / or on - site assistance.
[0018] Figure 1 is a block diagram of an example computer system 100 capable of using a language model to assist non-expert users in the installation, troubleshooting, and / or maintenance of complex systems according to at least one embodiment. As Figure 1 shown, the computer system 100 may include a computing device 102, a data store 150, and a target system support center 170 connected to a network 140. The network 140 may be a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), a wireless network, a personal area network (PAN), a combination thereof, and / or any other network type. The computing device 102 may be communicatively coupled (e.g., via the network 140 or a local connection) to a diagnostic tool 160 that performs diagnostics on any suitable target system 180.
[0019] The target system 180 may include any computing device, manufacturing device, industrial device, automotive device, household device, entertainment device, and / or any combination of any other system or systems. For example, the target system 180 may include a network server that supports the operation of a local computing network. In another example, the target system 180 may include an automotive entertainment system of a vehicle. In yet another example, the target system 180 may include an industrial robot system. The target system 180 may include any combination of hardware components, such as processing devices, memory devices, peripherals, network controllers, instrument components, robotic device components, engines, actuators, mechanical devices, electronic circuits, displays, lidar, radar, cameras, lighting devices, audio devices, video devices, and / or any other hardware device. The target system 180 may also include any combination of software code, modules, routines, drivers, application programming interfaces (APIs), and / or any other software program, script, instruction, and / or the like that executes on, is associated with, or facilitates the operation of the hardware devices.
[0020] The diagnostic tool 160 can be any device or combination of devices that are designed, assembled, and / or configured to test, troubleshoot, identify faults, confirm compliance with the specifications of the target system 180, and / or otherwise facilitate the operation of the target system 180. The diagnostic tool 160 can have any level of complexity, e.g., ranging from a tool capable of performing a single type of measurement to a sophisticated and elaborate diagnostic machine whose complexity can significantly exceed the complexity of the target system 180 itself. In some embodiments, the diagnostic tool 160 can include a hardware diagnostic 162 that is capable of testing one, multiple, or all of the hardware components of the target system 180. For example, the hardware diagnostic 162 can use any number of sensors that are capable of measuring any relevant environmental conditions (e.g., temperature, pressure, etc.) as well as any number of metrics associated with the performance of the target system 180, such as the speed and accuracy of the operation of the target system 180, network (e.g., Ethernet and / or wireless network) bandwidth, processor speed / utilization, available disk space, energy usage, memory / disk health status, signal strength, signal range, signal quality, response latency, fan speed, voltage, load, and clock speed and / or the like. The software diagnostic 164 can use any number of test program / script sensors that are capable of measuring the effectiveness of the software execution on the test system 180, e.g., queuing time, network throughput, bit error rate, latency, memory usage, processor speed / utilization, etc. It should be understood that in some cases, there may be no clear boundary between the hardware diagnostic 162 and the software diagnostic 164, so a specific diagnostic component can have both hardware diagnostic functions and software diagnostic functions. For example, network latency / throughput / bit error diagnosis can be a combination of hardware and software test tools.
[0021] The diagnostic tool 160 can apply the hardware diagnostic 162 and the software diagnostic 164 to the target system 180 and generate one or more diagnostic logs 166 that represent a record of one or more performed diagnostic operations made in any suitable format. In one example, the diagnostic log 164 can include a list of operations, individual operations indexed by corresponding operation identifiers (IDs), and the results of the operations. Some operations can include binary results (e.g., "Check: Pass" or "Check: Fail"), while other operations can have multi-valued (e.g., continuous) results (e.g., network latency "83 ms"). Some operations can include discrete error / defect fault indicators (e.g., codes / defects), each error or defect indicating a specific type of fault of the target system 180. The number of operations performed and recorded in the diagnostic log 166 by the hardware diagnostic 162 and / or the software diagnostic 164 is not limited and can range from a single operation (for a relatively simple target system 180) to hundreds, thousands, or even more operations. Similarly, individual diagnostic operations can have any number (two or more) of defects, errors, and / or continuous values.
[0022] In some embodiments, the diagnostic tool 160 can be integrated into the target system 180. In some embodiments, the diagnostic tool 160 can be separate from the target system 180 and can engage (e.g., be coupled via one or more interfaces) the target system 180 in response to a failure occurring, a periodic or one-time update, replacement of a worn component, etc. In some embodiments, the diagnostic tool 160 can be initiated by the user 101, e.g., in response to a user indication received by the computing device 102 via the user interface (UI) 106 (e.g., after the user 101 or another user detects a failure). In some embodiments, the diagnostic tool 160 can be automatically initiated by the computing device 102, e.g., in response to the time for periodic maintenance. The computing device 102 can include a desktop computer, a laptop computer, a smartphone, a tablet, a server, a wearable device, a virtual / augmented / mixed reality headset or a head-up display, a digital avatar or a chatbot kiosk, an in-vehicle infotainment computing device, and / or any suitable computing device capable of performing the techniques described herein. The computing device 102 can be configured to communicate with the user 101 via the UI 106. The user 101 can be an individual user (e.g., the owner of a computer, a vehicle, an entertainment device), a collective user (e.g., a business organization, an institution, a government agency, etc.), an agent of a repair agency, etc. In some embodiments, the query (question) generated by the user 101 can include text (e.g., a sequence of one or more typed words), voice (e.g., a sequence of one or more spoken words), or an image (e.g., an image of the output of the target system 180 indicating a failure) and / or some combination thereof. The query can be generated as part of the interaction between the user 101 and the diagnostic engine 120, which interfaces with the LM 122 and facilitates communication with the LM 122, and in some embodiments, can communicate with the diagnostic tool 160. In some embodiments, the LM 122 can be an LLM, e.g., a model having hundreds of millions or billions or more learned parameters.
[0023] The UI 106 can include one or more devices of various modalities, such as a keyboard, a touch screen, a touchpad, a writing tablet, a graphical interface, a mouse, a stylus, and / or any other pointing device capable of selecting words / phrases displayed on the screen, and / or some other suitable device. In some embodiments, the UI 106 can include audio devices, such as a combination of a microphone and a speaker, and video devices, such as a digital camera for capturing an image or a sequence of two or more images (video frames). In some embodiments, the text, voice, and / or video input devices can be integrated together (e.g., integrated into a smartphone, a tablet, a desktop computer, and / or a similar device).
[0024] The computing device 102 may include a memory 104 (e.g., one or more memory devices or units) communicatively coupled to one or more processing devices, such as one or more graphics processing units (GPUs) 110, one or more central processing units (CPUs) 130, one or more data processing units (DPUs), one or more parallel processing units (PPUs), and / or other processing devices (e.g., field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and / or the like). The memory 104 may include read only memory (ROM), flash memory, dynamic random access memory (DRAM) (e.g., synchronous DRAM (SDRAM)), static memory (e.g., static random access memory (SRAM)), and / or some other memory capable of storing digital data. The memory 104 may store a diagnostic engine 120 and an LM 122, as well as an application 124. The application 124 may be any application that runs on, is deployed on, or otherwise uses the target system 180, or may include any application that runs on, is deployed on, or otherwise uses the target system 180.
[0025] In some embodiments, the LM 122 may be a model trained and deployed by a separate entity (e.g., an LM service 190), which may be a cloud service, a subscription service, and / or some combination thereof. In some embodiments, the LM 122 may be trained in multiple stages. Initially, the LM service 190 may train the LM 122 to capture the grammar and semantics of human language, e.g., by predicting the next, previous, and / or missing words in a sequence of words (e.g., one or more sentences of human speech or text). The LM 122 may be further trained using training data that includes a large amount of text (e.g., human conversations, newspaper text, magazine text, book text, web-based text, and / or any other text). Since the ground truth for such training is embedded within the text itself, the LM service 190 may use such text to perform self-supervised training of the LM 122. This teaches the LM 122 how to converse with a user (human user or another computer) in natural language in a manner very similar to conversing with a human speaker, including understanding the user's intent and responding in a manner expected by the user from a conversation partner.
[0026] After initial self-supervised training, the LM training engine 174 (e.g., deployed by the target system support center 170) can perform supervised fine-tuning of the LM 122 to teach the LM 122 the details of the target system 180. Specifically, system-specific documents can be used to train the LM 122, such as the Installation and Diagnostic Manual (IDM) 152, System Architecture Document (SAD) 154, and training logs 156. The IDM 152 can include information such as technical descriptions of diagnostic error codes. The IDM 152 can be written to be understandable by expert technicians or engineers rather than users of the target system 180. The SAD 154 can include the addresses, descriptions, and connections of the hardware components and software modules of the target system 180, and / or other architecture information, such as data paths. The SAD can also include a list of tests available via the diagnostic tool 160 and the mapping of various hardware components and software modules to the listed tests, e.g., descriptions of the components / modules tested by each test. The training logs 156 can include logs of actual faults encountered during the operation of one or more historical instances of the target system 180, faults deliberately induced by developers (specifically, to generate training data for the LM 122), (simulated) synthetic logs generated by developers or via random selection of one or more error codes, and / or error logs generated by any other suitable technique. The LM training engine 174 can use training queries 158, which can be natural language inquiries related to the training logs 156 (e.g., "The diagnostic tool shows error code 8250"), which may be asked by users seeking technical assistance. The training queries 158 can be used to form training prompts that are input into the LM 122. The training prompts can also include the IDM 152 and / or SAD 154 and an indication to the LM122 that can indicate the interpretation of the error code in the training query 158 and that various additional information related to the training query 158 can be found in those (IDM 152 and / or SAD 154) documents. During training, the LM 122 can generate a (training) response to the training query 158, which includes one or more indications to the user on how to resolve the fault mentioned in the user's query. The training query 158 can be mapped to a ground truth 159, which can include a sample response to the training query 158, e.g., a sample response prepared by the system developer. The training response generated by the LM can be evaluated using a suitable loss function or some evaluation scale indicating the effectiveness of the training response (e.g., evaluated by developers, expert technicians, or non-expert user). The LM training engine 174 can also use a reference log 172 (e.g., a diagnostic log corresponding to a fault-free state of the target system 180) to teach the LM 122 how to recognize the desired (target) state of the system.
[0027] In some embodiments, any, some, or all of the IDM 152, SAD 154, training logs 156, training queries 158, and ground truth 159 may be stored in the data store 150, and the target system support center 170 may access the data store 150 directly (e.g., via a bus, interconnect, etc.) or (as Figure 1 shown) via the network 140. The data store 150 may include persistent storage and may be hosted by one or more storage devices, such as main memory, magnetic or optical storage disks, magnetic tape, or hard disk drives, network-attached storage (NAS), storage area network (SAN), etc. Although depicted as separate from the target system support center 170 and / or the computing device 102, in at least some embodiments, the data store 150 may be part of the target system support center 170 and / or the computing device 102. In at least some embodiments, the data store 150 may be a network-attached file server, while in other embodiments, the data store 150 may be some other type of persistent storage, such as an object-oriented database, a relational database, etc., which may be hosted by the target system support center 170 and / or the computing device 102 or one or more different machines coupled to the target system support center 170 and / or the computing device 102 via the network 140.
[0028] In some embodiments, any, some, or all of the IDM 152, SAD 154, and training logs 156 may be stored in the embedding database 151 as an embedding set (vectors in a suitable embedding space). An embedding should be understood as any suitable numerical representation of the input data, e.g., as a vector (string) with any number m of components, which may have integer or floating-point values. An embedding can be regarded as a point in an m-dimensional embedding space. The dimension m of the embedding space may be less than the size of the input data. During training, the embedding generator 176 learns to associate similar input data sets with similar embeddings represented by points that are close together in the embedding space and further learns to associate dissimilar input data sets with points that are far apart in the space.
[0029] In some embodiments, the target system support center 170 may train multiple LMs 122 for multiple types of target systems 170. In some embodiments, any given LM 122 trained by the target system support center 170 is capable of facilitating the maintenance and debugging of multiple types of target systems, e.g., trained using multiple training data sets, such as target system-specific IDM 152, SAD 154, training logs 156, training queries 158, and / or ground truth 159.
[0030] The LM 122 can be implemented using a neural network with a large number (e.g., billions) of artificial neurons. In at least one embodiment, the LM 122 and / or other deployed models can be implemented as a deep learning neural network with multi-level linear and non-linear operations. For example, the LM 122 can include a convolutional neural network, a recurrent neural network, a fully connected neural network, a long short-term memory (LSTM) neural network, a neural network with an attention mechanism (e.g., a transformer neural network, a combination of a convolutional network and one or more transformers (convolutional transformer)), and / or other types of neural networks. In at least one embodiment, the LM 122 can include multiple neurons, where individual neurons receive their inputs from other neurons and / or from an external source and produce an output by applying an activation function to the sum of the weighted (using trainable weights) inputs and (possibly) a bias value. In at least one embodiment, the LM 122 can include multiple neurons arranged in layers, including an input layer, one or more hidden layers, and / or an output layer. Neurons from adjacent layers can be connected by weighted edges.
[0031] Initially, some starting (e.g., random) values can be assigned to the parameters (e.g., edge weights and biases) of one or more LMs 122. For various training inputs, the LM training engine 174 can cause the LM 122 to generate training outputs. Then, the LM training engine 174 can compare these training outputs with the desired target outputs. The resulting error or mismatch (e.g., the difference between the target output and the training output) can be backpropagated through the various neural layers of the LM 122, and the weights and biases of the LM 122 can be adjusted to make the training output closer to the target (true) output. This adjustment can be repeated until the output error for a given training input satisfies a predetermined condition (e.g., is below a predetermined value). Subsequently, different training inputs can be selected, new training outputs can be generated, and a new series of adjustments can be implemented until the LM 122 is trained to the target accuracy or until the LM 122 converges to the limit of the accuracy determined by its architecture.
[0032] The trained LM 122 can be located on the target system support center 170 or the LM server 190. In one embodiment, the target system support center 170 or the LM server 190 can be implemented on a single computing device. The target system support center 170 or the LM server 190 can be (and / or include) a rack server, a router computer, a personal computer, a laptop, a tablet, a desktop computer, a media center, or any combination thereof.
[0033] Figure 2FIG. 200 illustrates an example computing device that supports deploying a language model to assist a user with installation, troubleshooting, and / or maintenance of a complex system, according to at least one embodiment. In at least one embodiment, computing device 200 may be part of computing device 102. In at least one embodiment, computing device 200 may include a diagnostic engine 120 that obtains a user query from UI 106 and provides a response generated by the LM to the user via UI 106. Diagnostic engine 120 may communicate with the LM via an LM application programming interface (API) 202. In some embodiments, the LM may be located on a different computing device / server, such as a cloud-based server of LM service 190 (see Figure 1 ). The LM API 202 may be downloaded and installed on computing device 102 from LM service 190 (or target system support center 170) to facilitate communication with an LM 122 that is remotely provided by LM service 190 or target system support center 170. Computing device 200 may also include a diagnostic tool API 204 that supports communication with diagnostic tool 160, including (but not limited to) sending requests to diagnostic tool 160 to perform one or more specific tests, receiving diagnostic logs 166, etc. Computing device 200 may also support any suitable application 124 that runs on or in association with target system 180 (see Figure 1 ), such as control software for a product production line, camera software for a digital camera, etc. Application 124 may communicate with target system 180 via any suitable target system API 206.
[0034] The operations of the diagnostic engine 120, application 124, LM API 202, diagnostic tool API 204, target system API 206, and / or other software / firmware operating on the computing device 200 may be performed using one or more GPUs 110, one or more CPUs 130, one or more parallel processing units (PPUs) or accelerators (such as deep learning accelerators), data processing units (DPUs), and / or the like. In at least one embodiment, the GPU 110 includes a plurality of cores 211, each core capable of executing multiple threads 212. Each core may concurrently (e.g., in parallel) run multiple threads 212. In at least one embodiment, the threads 212 may access registers 213. The registers 213 may be thread-specific registers, and access to the registers is limited to the corresponding thread. Additionally, shared registers 214 may be accessed by one or more (e.g., all) of the threads of a core. In at least one embodiment, each core 211 may include a scheduler 215 for allocating computational tasks and processes among the different threads 212 of the core 211. A dispatch unit 216 may implement the scheduled tasks on the appropriate threads using the correct private registers 213 and shared registers 214. The computing device 200 may include input / output components 217 for facilitating the exchange of information with one or more users or developers.
[0035] In at least one embodiment, the GPU 110 may have a (high-speed) cache 218, and the multiple cores 211 may share access to the cache. Additionally, the computing device 200 may include GPU memory 219, where the GPU 110 may store intermediate and / or final results (outputs) of various computations performed by the GPU 110. After completing a particular task, the GPU 110 (or CPU 130) may move the output to the (main) memory 104. In at least one embodiment, the CPU 130 may execute processes involving serial computational tasks, while the GPU 110 may execute tasks suitable for parallel processing (e.g., multiplying the inputs of a neural node by weights and adding a bias).
[0036] The systems and methods described herein may be used for various purposes, such as but not limited to machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, data center processing, conversational AI, generative AI, optical transmission simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.
[0037] The disclosed embodiments may be included in various different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, sensing systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems for generating or presenting at least one of augmented reality content, virtual reality content, and mixed reality content, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing generative AI operations, systems for performing conversational AI operations, systems for performing optical transmission simulation, systems for performing collaborative content creation of 3D assets, systems implementing one or more language models (e.g., large language models (LLMs) that can process text, speech, images, and / or other data types to generate outputs in one or more formats), systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0038] Figure 3 An example training 300 of a language model according to at least one embodiment is shown to assist a user in complex system installation, troubleshooting, and / or maintenance. In at least one embodiment, an LM training engine 174 ( Figure 1 ) may train an LM 301. The LM 301 may be initially trained (pre-trained) by an LM service 190 for various general tasks of natural language processing, including but not limited to predicting the next, previous, and / or missing words in a sentence (or some other sequence of words) encountered in any suitable text corpus. The pre-trained LM 301 may be sufficient to talk to a human user in natural language.
[0039] The training 300 may include supervised training, self-supervised training, reinforcement training, unsupervised training, or any combination thereof. The training 300 may be used to train the LM 301 to provide installation, troubleshooting, maintenance, and / or any other technical support for the operation of any suitable system (e.g., Figure 1 the target system 180 therein). The training 300 may use various documents associated with the target system, including but not limited to some or all of the installation and diagnostic manuals 152, system architecture documents 154, training logs 156, and reference logs 172.
[0040] In some embodiments, a training log 156 may be selected for a given instance (epoch) of training 300. The training log 156 may be associated with historical failures encountered during previous operations of the target system 180, deliberately induced failures, synthetic logs generated by developers, randomly selected logs including one or more error codes, and the like. As Figure 3 shown by the dashed arrow in, a training query 158 may be generated for the selected training log 156. The training query 158 may reflect the likely reactions of non-expert users to the training log 156 and may be generated by developers, engineers, or non-expert users. The training query 158 may be phrased in natural language that a potential user might type or say. For example, the training query 158 may include questions about the meaning of error codes in the training log 156, requests for help, requests for step-by-step instructions on how to resolve the error codes, and / or the like.
[0041] The tokenizer 310 may transform the training query 158 into tokens recognizable by the LM 301. A set of tokens may be specific to the LM 301 (e.g., may be different for different models and / or model creators) and is fixed during the pre-training of the LM 301. The set of tokens may include any suitable numerical representation of speech units (e.g., syllables, words, etc.). In one example of GPT-4 tokens, the word "the" may be represented by the token "280", the word "import" may be represented by the token "476", the word "description" may be represented by the token "4097", and so on. In some embodiments, an individual word may be represented by any number of tokens or word transformations. For example, a long word or a word containing multiple characters may be represented by multiple tokens, e.g., one token for the beginning part of the word and another token for the middle or end part of the word. In some cases, even a long word / compound word may be represented by a single token. Thus, tokenization may be performed in any way suitable for input to the LM 301.
[0042] The training query 158 may be used to form a training prompt 320 input into the LM 301. The training prompt 320 may also include (but is not limited to) any, some, or all of the documentation of the target system, such as the IDM 152, SAD 154, training log 156, and / or reference log 172 (logs of the target system operating normally). Various parts of the system documentation may be converted into embeddings using the embedding generator 176 before being used for the training prompt 320.
[0043] The LM 301 can generate a training response 330 to the training query 158. The training response 330 can include an indication to the user that has guidance regarding troubleshooting the fault mentioned in the training query 158. The training response 330 can undergo a response evaluation 340 to determine the degree of applicability / effectiveness of the training response 330 for troubleshooting the fault. In some embodiments, the response evaluation 340 can be performed by developers, expert technicians, non-expert users, etc. For example, if a sample response 350 (e.g., prepared by an expert) is available, a loss function can be used to evaluate the difference between the training response 330 and the sample response 350. This difference can be used to adjust the LM 301, e.g., by directly changing the parameters of the LM 301 (using techniques such as backpropagation, gradient descent, etc.), by exposing the sample response 350 to the LM 301, and / or using various other learning techniques. In the case where the sample response 350 is not available, the response evaluation 340 can grade the training response 330 using a suitable evaluation scheme (e.g., 1 - N) and provide the evaluation grade to the LM 301. After receiving the evaluation grade, the LM 301 can generate a new (updated) training response 330 and evaluate it similarly. This process can continue until an acceptable training response 330 is generated. Multiple training queries 158 associated with the training log 156 can be used to train the LM 301, and these training logs have various error codes (or other fault indicators) that may be encountered during the installation, operation, troubleshooting, and / or maintenance of the target system.
[0044] When the target system undergoes an update (hardware update, software update, firmware update, combination thereof, etc.) or a modification that changes the IDM 152 (e.g., adding a new test to the diagnostic tool 160) or the SAD 154 (e.g., adding a new hardware component and / or software module to the target system 180), the language model (e.g., LM 301) can be retrained using additional training data. In some embodiments, the retraining can be limited to the new features / tests / etc., e.g., by utilizing additional training logs 156 that characterize the faults of the added features and / or using the new tests added to the diagnostic tool 160 and training queries 158 that have questions regarding the faults of the added features and / or the errors associated with the new tests. After retraining, the language model can be made available for in-field use, e.g., as described below in conjunction with Figure 4 as described.
[0045] Figure 4Illustrated is an example use 400 of a trained language model in assisting a user with the installation, troubleshooting, and / or maintenance of a complex system. Use 400 can be implemented by an individual user, an organizational user, a repair agency, or any other entity that may be responsible for system assembly, operation, and / or upkeep. The LM 122 deployed in use 400 can be the LM 301 trained using the example training 300 using Figure 3 As disclosed in connection with Figure 3 During training, the LM 122 can learn various system-specific documents, such as the IDM 152, SAD 154, reference log (RL) 172, and / or any additional knowledge of the target system 180.
[0046] As Figure 4 shown, the diagnostic tool 160 can run a diagnosis on the target system 180 to generate one or more diagnostic logs 166. The diagnosis can be performed in response to a failure of the target system 180, a hardware and / or software update of the target system 180, a scheduled (e.g., periodic) maintenance to be performed on the target system 180, or any other reason. The diagnostic log 166 can be provided to the user 101. After receiving the diagnostic log 166, the user 101 can submit a query 402 to the LM 122 via the UI 106. The query 402 can be phrased in natural language and can include questions about the meaning of error codes in the diagnostic log 166, requests for help, requests for guidance on resolving error codes, and / or the like (e.g., "PCIe port not working", "camera not found", etc.).
[0047] The tokenizer 310 can transform the query 402 into tokens recognizable by the LM 122. The query 402 can be used to form a prompt 404 that is input into the LM 122. The prompt 404 can include the diagnostic log 166 (or any excerpt of the log), and the diagnostic log can be converted into one or more embeddings using the embedding generator 176. The LM 122 can generate a response 406 to the prompt 404. The response 406 can include an indication to the user with guidance on resolving the failure mentioned in the query 402, an indication of how to perform an update, maintenance, and / or any other operation associated with the target system 180.
[0048] The user 101 can cause an action to be performed on the target system 180, such as replacing a component of the target system 180, installing or reinstalling software, etc. Subsequently, the diagnostic tool 160 can perform an additional diagnosis on the target system 170 and can generate additional diagnostic logs. If there is a "no defect" signal in the log, the troubleshooting process can end. If there are additional errors / defects in the log, the user 101 can enter a new query, and the above process can continue for as many cycles as necessary to resolve the remaining failures.
[0049] In some embodiments, the LM may communicate back and forth with one or more (e.g., third-party) plugins or APIs to facilitate communication with the user. For example, when a specific error code exists, the LM may use a third-party plugin associated with the error code to retrieve information about the nature of the error code and any underlying troubleshooting steps or procedures, and may use this additional information to generate an output to the user. This process may be repeated for any number of plugins and / or APIs until communication with the user is complete.
[0050] Figure 5 and Figure 6 Example methods 500 and 600 are shown, which are intended to train and use a trained language model to assist a user in the installation, troubleshooting, and / or maintenance of a complex system. Methods 500 and 600 may be used in the context of providing technical support to users and / or maintenance personnel of any complex system. A system may be considered complex if the installation, troubleshooting, and / or maintenance of the system involves expertise that requires more time, effort, and / or experience than is expected of a typical user / maintenance person. The system may be a hardware system, a software / firmware system, and / or a system that deploys a combination of hardware and software / firmware. In at least one embodiment, methods 500 and / or 600 may be executed using Figure 1 the processing unit of computing device 102 and / or target system support center 170. In at least one embodiment, the processing unit that executes methods 500 and / or 600 may execute instructions stored on a non-transitory computer-readable storage medium. In at least one embodiment, methods 500 and / or 600 may be executed using multiple processing threads (e.g., CPU threads and / or GPU threads), where each thread executes one or more individual functions, routines, subroutines, or operations of the method. In at least one embodiment, the processing threads that implement any one of methods 500 and / or 600 may be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads that implement any one of methods 500 and / or 600 may execute asynchronously relative to each other. Each operation in methods 500 and / or 600 may be Figure 5 and Figure 6 executed in an order different from that shown. Some operations in methods 500 and / or 600 may be executed concurrently with other operations. In at least one embodiment, Figure 5 and Figure 6 one or more of the operations shown may not always be executed.
[0051] Figure 5It is a flowchart of an example method 500 for using a language model to assist a user in installing, troubleshooting, and / or maintaining a complex system according to at least one embodiment. Method 500 can be executed using one or more processing units (e.g., CPU, GPU, accelerator, PPU, DPU, etc.) of computing device 102, and the processing unit includes one or more memory devices (or communicates with one or more memory devices).
[0052] At block 510, method 500 can include: receiving, via a user interface (UI), a natural language (NL) query associated with one or more fault indicators indicating a fault state of the system. The fault indicators can include error / defect codes (e.g., numeric or alphabetic codes), descriptions of the fault state (e.g., using natural language or any predefined set of descriptors), and / or other indicators capable of identifying one or more faults of the system. In some embodiments, one or more fault indicators are received in association with a system diagnosis 501 performed on the system by a diagnostic tool. The NL query can include a text prompt, a voice prompt, an audio prompt, an image prompt, etc., or any combination thereof. For example, a user can type text using a keyboard, speak words using a microphone, or perform some other action. In some embodiments, the NL query is generated in response to a scheduled maintenance of the system, a hardware update of the system, a software update of the system, an inoperable condition of the system, etc., or some combination thereof.
[0053] At block 520, method 500 can include: providing an input to the LM. In some embodiments, the LM can include a large language model. The input can include at least a prompt based on the NL query. The input to the LM can also include one or more test logs associated with the diagnosis performed on the system. In some embodiments, the LM communicates with at least one of an API or a plugin to retrieve additional information related to the NL query. In some embodiments, a document associated with the system can be used to train the LM. For example, the document for training can include diagnostic documents of the system (e.g., installation, assembly, maintenance manuals, etc.). The document for training can also include system architecture documents of the system. The document used in training can also include one or more training diagnostic logs associated with historical or simulated faults of the system.
[0054] At block 530, method 500 can include: receiving, from the LM, a response to the NL query. The response can include one or more indications associated with the resolution of the fault state of the system. At block 540, method 500 can continue to cause the UI to display the response. As Figure 5 shown by the dashed arrow in, the operations of blocks 501-540 can be repeated until the fault state is resolved (block 550).
[0055] In one example, in the case of two or more iterations of method 500, during the initial (first) iteration, an initial input (which is at least based on an initial NL query) can be provided to the LM, and an initial response (with one or more initial indications) to the initial NL query can be received from the LM. After the user's initial attempt (e.g., one or more initial actions) fails to resolve the fault, a new system diagnosis 501 can be performed, generating a fault indicator that reflects the initial attempt to resolve the fault state of the system. These fault indicators can then be used to formulate subsequent NL queries (e.g., for the second iteration of method 500).
[0056] Figure 6 is a flowchart of an example method 600 for training a language model according to at least one embodiment to assist a user in complex system installation, troubleshooting, and / or maintenance. Method 600 can be executed using Figure 1 one or more processing units (e.g., CPU, GPU, accelerator, PPU, DPU, etc.) of the target system support center 170, which can include one or more memory devices (or communicate therewith).
[0057] At block 610, method 600 can include: generating training data. The training data can include natural language (NL) queries (611) associated with one or more fault indicators indicating the fault state of the system. The training data can also include documents associated with the system (612). The training data can also include training diagnostic logs associated with the fault state of the system (613). In some embodiments, the documents associated with the system can include diagnostic documents of the system and / or system architecture documents of the system. In some embodiments, the training diagnostic logs can include diagnostic logs associated with the diagnostics performed on the system, comprehensive diagnostic logs of the system fault state, and / or both.
[0058] At block 620, method 600 can continue to use the training data to train the LM (e.g., a neural network-based large language model) to generate NL responses to the NL queries. The NL responses can include one or more indications associated with resolving the system fault state. In some embodiments, using the training data to train the LM can include Figure 6The operations shown in the callout portion of
[0059] The systems and methods described herein can be used for a variety of purposes, such as but not limited to, performing one or more operations related to machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, optical transmission simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable applications.
[0060] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., in-vehicle infotainment systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing optical transmission simulation, systems for performing collaborative content creation of 3D assets, systems for performing generative AI operations, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0061] Inference and training logic
[0062] Figure 7A Shown is inference and / or training logic 715 for performing inference and / or training operations associated with one or more embodiments.
[0063] In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, code and / or data storage 701 for storing forward and / or output weights and / or input / output data, and / or other parameters that configure neurons or layers of a neural network trained to and / or for inference in aspects of one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data storage 701 for storing graph code or other software to control timing and / or sequencing, where weight and / or other parameter information is loaded to configure the logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs) or simple circuits). In at least one embodiment, the code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 701 stores the weight parameters and / or input / output data of each layer of the neural network used or trained in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 701 may be included within other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0064] In at least one embodiment, any portion of the code and / or data storage 701 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 701 may be cache memory, dynamic random-access memory (“DRAM”), static random-access memory (“SRAM”), non-volatile memory (such as flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 701 is internal or external to the processor, e.g., or consists of DRAM, SRAM, flash memory, or some other storage type, may depend on the available storage space on or off the chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in the inference and / or training of the neural network, or some combination of these factors.
[0065] In at least one embodiment, the inference and / or training logic 715 can include, but is not limited to, code and / or data storage 705 to store the backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained as and / or used for inference in aspects of one or more embodiments. In at least one embodiment, during training and / or inference using aspects of one or more embodiments, the code and / or data storage 705 stores the weight parameters and / or input / output data of each layer of the neural network used or trained in conjunction with one or more embodiments during the backpropagation of the input / output data and / or weight parameters. In at least one embodiment, the training logic 715 can include or be coupled to code and / or data storage 705 for storing graph code or other software to control timing and / or sequencing, where the weights and / or other parameter information are loaded to configure the logic, which includes integer and / or floating-point units (collectively referred to as arithmetic logic units (ALUs)).
[0066] In at least one embodiment, the code (such as graph code) causes the weights or other parameter information to be loaded into the processor ALU based on the architecture of the neural network corresponding to the code. In at least one embodiment, any portion of the code and / or data storage 705 can be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data storage 705 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 705 can be cache memory, DRAM, SRAM, non-volatile memory (such as flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 705 is internal or external to the processor, e.g., whether it consists of DRAM, SRAM, flash memory, or some other storage type, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in the inference and / or training of the neural network, or some combination of these factors.
[0067] In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be separate storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be combined storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be partially combined and partially separated. In at least one embodiment, any portion of code and / or data storage 701 and code and / or data storage 705 may be included with other on-chip or off-chip data storage, including the L1, L2, or L3 cache of the processor or system memory.
[0068] In at least one embodiment, inference and / or training logic 715 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 710 (including integer and / or floating point units) for performing logical and / or mathematical operations at least in part based on training and / or inference code (e.g., graph code) or as indicated thereby, the result of which may produce activations (e.g., output values from a layer or neuron within a neural network) stored in activation storage 720, which are a function of input / output and / or weight parameter data stored in code and / or data storage 701 and / or code and / or data storage 705. In at least one embodiment, the activations are generated by linear algebra and / or matrix-based mathematics performed by ALU 710 in response to execution of instructions or other code, where weight values stored in code and / or data storage 705 and / or data storage 701 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 705 or code and / or data storage 701 or another on-chip or off-chip storage.
[0069] In at least one embodiment, one or more ALUs 710 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 710 may be outside of the processor or other hardware logic devices or circuits that use them (such as a coprocessor). In at least one embodiment, one or more ALUs 710 may be included within the execution unit of a processor or otherwise included in a group of ALUs accessible by the execution unit of a processor, where the execution unit of the processor may be within the same processor or distributed among different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed function unit, etc.). In at least one embodiment, code and / or data storage 701, code and / or data storage 705, and activation storage 720 may share a processor or other hardware logic device or circuit, while in another embodiment, they may be in different processors or other hardware logic devices or circuits or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any part of activation storage 720 may be included together with other on-chip or off-chip data storage, including the L1, L2, or L3 cache of the processor or system memory. Additionally, inference and / or training code may be stored together with other code accessible by the processor or other hardware logic or circuits and may be extracted and / or processed using the fetch, decode, schedule, execute, retire, and / or other logic circuits of the processor.
[0070] In at least one embodiment, activation storage 720 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, activation storage 720 may be entirely or partially inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether activation storage 720 is internal or external to the processor may depend on on-chip or off-chip available storage, the latency requirements for training and / or inference functions, the batch size of data used in an inference and / or training neural network, or some combination of these factors, e.g., or include DRAM, SRAM, flash memory, or some other storage type.
[0071] In at least one embodiment, Figure 7A the inference and / or training logic 715 shown may be used in conjunction with an application specific integrated circuit (“ASIC”), such as the processing unit from Google, the TM inference processing unit (IPU) from Graphcore, or the (e.g., “Lake Crest”) processor from Intel Corp. In at least one embodiment, Figure 7AThe inference and / or training logic 715 shown may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as field programmable gate array (“FPGA”).
[0072] Figure 7B The inference and / or training logic 715 according to at least one embodiment is shown. In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, hardware logic where computing resources are dedicated or otherwise uniquely used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 7B the inference and / or training logic 715 shown in may be used in conjunction with an application specific integrated circuit (ASIC), such as the processing unit from Google, the TM inference processing unit (IPU) from Graphcore or the (e.g., “Lake Crest”) processor from Intel Corp. In at least one embodiment, Figure 7B the inference and / or training logic 715 shown in may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware such as field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 715 includes, but is not limited to, code and / or data storage 701 and code and / or data storage 705, which may be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In Figure 7B at least one embodiment shown, each of code and / or data storage 701 and code and / or data storage 705 is respectively associated with dedicated computing resources (such as computing hardware 702 and computing hardware 706). In at least one embodiment, each of computing hardware 702 and computing hardware 706 includes one or more ALUs that respectively perform mathematical functions (such as linear algebra functions) only on the information stored in code and / or data storage 701 and code and / or data storage 705, and the results of the executed functions are stored in activation storage 720.
[0073] In at least one embodiment, each of code and / or data stores 701 and 705 and corresponding computing hardware 702 and 706 corresponds to a different layer of a neural network such that activations obtained from one storage / computation pair 701 / 702 of code and / or data store 701 and computing hardware 702 are provided as input to the next storage / computation pair 705 / 706 of code and / or data store 705 and computing hardware 706 to reflect the conceptual organization of the neural network. In at least one embodiment, each storage / computation pair 701 / 702 and 705 / 706 may correspond to more than one layer of a neural network. In at least one embodiment, additional storage / computation pairs (not shown) may be included in inference and / or training logic 715 after or in parallel with storage / computation pairs 701 / 702 and 705 / 706.
[0074] Neural Network Training and Deployment
[0075] Figure 8 Illustrated is the training and deployment of a deep neural network according to at least one embodiment. In at least one embodiment, an untrained neural network 806 is trained using a training data set 802. In at least one embodiment, the training framework 804 is the PyTorch framework, while in other embodiments, the training framework 804 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 804 trains the untrained neural network 806 and enables it to be trained using the processing resources described herein to generate a trained neural network 808. In at least one embodiment, the weights may be randomly selected or pre-trained using a deep belief network. In at least one embodiment, training may be performed in a supervised, partially supervised, or unsupervised manner.
[0076] In at least one embodiment, supervised learning is used to train an untrained neural network 806, where the training dataset 802 includes inputs paired with desired outputs for the inputs, or where the training dataset 802 includes inputs with known outputs and outputs for which the neural network 806 is manually graded. In at least one embodiment, the untrained neural network 806 is trained in a supervised manner, and inputs from the training dataset 802 are processed and the resulting outputs are compared to a set of desired or wanted outputs. In at least one embodiment, the error is then propagated back through the untrained neural network 806. In at least one embodiment, the training framework 804 adjusts the weights that control the untrained neural network 806. In at least one embodiment, the training framework 804 includes tools for monitoring the degree to which the untrained neural network 806 converges to a model (e.g., a trained neural network 808), a model adapted to generate correct answers (e.g., results 814) based on input data (e.g., a new dataset 812). In at least one embodiment, the training framework 804 repeatedly trains the untrained neural network 806 while adjusting the weights to improve the output of the untrained neural network 806 using a loss function and an adjustment algorithm (e.g., stochastic gradient descent). In at least one embodiment, the training framework 804 trains the untrained neural network 806 until the untrained neural network 806 reaches a desired accuracy. In at least one embodiment, the trained neural network 808 can then be deployed to perform any number of machine learning operations.
[0077] In at least one embodiment, unsupervised learning is used to train an untrained neural network 806, and the untrained neural network 806 attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised learning training dataset 802 will include input data without any associated output data or “ground truth” data. In at least one embodiment, the untrained neural network 806 can learn groupings within the training dataset 802 and can determine how individual inputs relate to the untrained dataset 802. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in the trained neural network 808, which can perform operations useful for reducing the dimensionality of a new dataset 812. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows identification of data points in a new dataset 812 that deviate from the normal pattern of the new dataset 812.
[0078] In at least one embodiment, semi-supervised learning can be used, which is a technique in which a mixture of labeled data and unlabeled data is included in a training data set 802. In at least one embodiment, a training framework 804 can be used to perform incremental learning, for example, by transfer learning techniques. In at least one embodiment, incremental learning enables the trained neural network 808 to adapt to a new data set 812 without forgetting the knowledge injected into the trained neural network 808 during initial training.
[0079] Referring Figure 9 , Figure 9 FIG. is an example data flow diagram of a process 900 for generating and deploying a processing and inference pipeline according to at least one embodiment. In at least one embodiment, the process 900 can be deployed to perform game name recognition analysis and inference on user feedback data at one or more facilities 902 such as a data center.
[0080] In at least one embodiment, the process 900 can be executed within a training system 904 and / or a deployment system 906. In at least one embodiment, the training system 904 can be used to perform training, deployment, and implementation of a machine learning model (e.g., a neural network, an object detection algorithm, a computer vision algorithm, etc.) for the deployment system 906. In at least one embodiment, the deployment system 906 can be configured to offload processing and computing resources in a distributed computing environment to reduce the infrastructure requirements of the facility 902. In at least one embodiment, the deployment system 906 can provide a pipeline platform for selecting, customizing, and implementing virtual instruments for use with computing devices at the facility 902. In at least one embodiment, a virtual instrument can include a software-defined application for performing one or more processing operations on feedback data. In at least one embodiment, one or more applications in the pipeline can use or invoke services (e.g., inference, visualization, computing, AI, etc.) of the deployment system 906 during application execution.
[0081] In at least one embodiment, some applications used in an advanced processing and inference pipeline can use a machine learning model or other AI to perform one or more processing steps. In at least one embodiment, feedback data 908 (e.g., imaging data) stored at the facility 902 or feedback data 908 from another or more facilities can be used at the facility 902, or a combination thereof, to train a machine learning model. In at least one embodiment, the training system 904 can be used to provide applications, services, and / or other resources to generate a working, deployable machine learning model for the deployment system 906.
[0082] In at least one embodiment, the model registry 924 can be supported by an object store that can support version control and object metadata. In at least one embodiment, the object store can be accessed from within a cloud platform via, for example, an application programming interface (API) compatible with cloud storage (e.g., Figure 10 cloud 1026). In at least one embodiment, machine learning models within the model registry 924 can be uploaded, listed, modified, or deleted by developers or partners of systems that interact with the API. In at least one embodiment, the API can provide access to methods that allow users with appropriate credentials to associate a model with an application such that the model can be executed as part of the execution of a containerized instantiation of the application.
[0083] In at least one embodiment, the training pipeline 1004 ( Figure 10 ) can include scenarios where the facility 902 is training their own machine learning models or has existing machine learning models that need to be optimized or updated. In at least one embodiment, feedback data 908 can be received from various channels such as forums, web forms, etc. In at least one embodiment, once the feedback data 908 is received, AI-assisted annotation 910 can be used to help generate annotations corresponding to the feedback data 908 to be used as ground truth data for the machine learning model. In at least one embodiment, the AI-assisted annotation 910 can include one or more machine learning models (e.g., a convolutional neural network (CNN)) that can be trained to generate annotations corresponding to certain types of feedback data 908 (e.g., from certain devices) and / or certain types of anomalies in the feedback data 908. In at least one embodiment, the AI-assisted annotation 910 can then be used directly or can be adjusted or fine-tuned using an annotation tool to generate the ground truth data. In at least one embodiment, in some examples, the labeled data 912 can be used as the ground truth data for training the machine learning model. In at least one embodiment, the AI-assisted annotation 910, the labeled data 912, or a combination thereof can be used as the ground truth data for training the machine learning model (e.g., via Figures 9 - 10 model training 914 therein). In at least one embodiment, the trained machine learning model can be referred to as the output model 916 and can be used by the deployment system 906 as described herein.
[0084] In at least one embodiment, the training pipeline 1004 ( Figure 10) may include the following scenarios: where the facility 902 requires a machine learning model to perform one or more processing tasks for deploying one or more applications in the deployment system 906, but the facility 902 may not currently have such a machine learning model (or may not have a model that is optimized, efficient, or effective for this purpose). In at least one embodiment, an existing machine learning model can be selected from the model registry 924. In at least one embodiment, the model registry 924 may include machine learning models that are trained to perform various different inference tasks on imaging data. In at least one embodiment, the machine learning models in the model registry 924 can be trained on imaging data from different facilities (e.g., a facility located remotely) rather than the facility 902. In at least one embodiment, the machine learning model may have been trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when trained on imaging data from a specific location (which may be in the form of feedback data 908), the training can be performed at that location or at least in a manner that protects the confidentiality of the imaging data or restricts the transfer of the imaging data off-site (e.g., complying with HIPAA regulations, privacy regulations, etc.). In at least one embodiment, once the model or a portion of the model has been trained at one location, the machine learning model can be added to the model registry 924. In at least one embodiment, the machine learning model can then be retrained or updated at any number of other facilities, and the retrained or updated model can be used in the model registry 924. In at least one embodiment, a machine learning model (referred to as the output model 916) can then be selected from the model registry 924 and used in the deployment system 906 to perform one or more processing tasks for one or more applications of the deployment system.
[0085] In at least one embodiment, the training pipeline 1004( Figure 10)Can be used in scenarios including facility 902, which requires a machine learning model to perform one or more processing tasks for deploying one or more applications in system 906, but facility 902 may not currently have such a machine learning model (or may not have an optimized, efficient, or effective model for this purpose). In at least one embodiment, due to population differences, genetic variations, robustness of the training data for training the machine learning model, diversity of training data anomalies, and / or other issues with the training data, the machine learning model selected from model registry 924 may not be fine-tuned or optimized for the feedback data 908 generated at facility 902. In at least one embodiment, AI-assisted annotation 910 can be used to help generate annotations corresponding to feedback data 908 to be used as ground truth data for retraining or updating the machine learning model. In at least one embodiment, labeled data 912 can be used as ground truth data for training the machine learning model. In at least one embodiment, retraining or updating the machine learning model can be referred to as model training 914. In at least one embodiment, model training 914 (e.g., AI-assisted annotation 910, labeled data 912, or a combination thereof) can be used as ground truth data for retraining or updating the machine learning model.
[0086] In at least one embodiment, deployment system 906 can include software 918, services 920, hardware 922, and / or other components, features, and functions. In at least one embodiment, deployment system 906 can include a software "stack" such that software 918 can be built on top of services 920 and services 920 can be used to perform some or all of the processing tasks, and services 920 and software 918 can be built on top of hardware 922 and hardware 922 can be used to perform the processing, storage, and / or other computing tasks of deployment system 906.
[0087] In at least one embodiment, software 918 may include any number of different containers, where each container may execute an instantiation of an application. In at least one embodiment, each application may perform one or more processing tasks (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.) in an advanced processing and inference pipeline. In at least one embodiment, for each type of computing device, there may be any number of containers that may perform data processing tasks on feedback data 908 (or other data types, such as the data types described herein). In at least one embodiment, in addition to the containers that receive and configure imaging data for use by each container and / or by facility 902 after processing through the pipeline, an advanced processing and inference pipeline may be defined based on the selection of different containers that are desired or required for processing feedback data 908 (e.g., to convert the output back to a usable data type for storage and display at facility 902). In at least one embodiment, a combination of containers within software 918 (e.g., which constitutes a pipeline) may be referred to as a virtual instrument (as described in more detail herein), and the virtual instrument may utilize services 920 and hardware 922 to perform some or all of the processing tasks of the applications instantiated in the containers.
[0088] In at least one embodiment, data may be preprocessed as part of a data processing pipeline to prepare the data for processing by one or more applications. In at least one embodiment, post-processing may be performed on the output of one or more inference tasks or other processing tasks in the pipeline to prepare the output data for the next application and / or to prepare the output data for transmission and / or use by the user (e.g., as a response to an inference request). In at least one embodiment, the inference tasks may be performed by one or more machine learning models, such as trained or deployed neural networks, which may include output model 916 of training system 904.
[0089] In at least one embodiment, the tasks of a data processing pipeline may be encapsulated in one or more containers, where each container represents a discrete, fully functional instantiation of an application and a virtualized computing environment that is capable of referencing a machine learning model. In at least one embodiment, the container or application may be published to a private (e.g., limited access) area of a container registry (described in more detail herein), and the trained or deployed model may be stored in model registry 924 and associated with one or more applications. In at least one embodiment, an image of the application (e.g., a container image) may be used in the container registry, and once the user selects the image from the container registry for deployment in the pipeline, the image may be used to generate a container for instantiation of the application for use by the user's system.
[0090] In at least one embodiment, a developer may develop, publish, and store an application (e.g., as a container) for performing processing and / or inference on provided data. In at least one embodiment, a software development kit (SDK) associated with the system may be used to perform development, publishing, and / or storage (e.g., to ensure that the developed application and / or container conforms to or is compatible with the system). In at least one embodiment, the developed application may be tested locally using the SDK (e.g., at a first facility, on data from the first facility), and the SDK, as the system (e.g., Figure 10 system 1000 in) may support at least some services 920. In at least one embodiment, once verified by system 1000 (e.g., for accuracy, etc.), the application becomes available in the container registry for selection and / or implementation by users (e.g., hospitals, clinics, laboratories, healthcare providers, etc.) to perform one or more processing tasks on data at the users' facilities (e.g., a second facility).
[0091] In at least one embodiment, the developer may then share the application or container over a network for access and use by users of the system (e.g., Figure 10 system 1000). In at least one embodiment, the completed and verified application or container may be stored in the container registry, and the associated machine learning model may be stored in the model registry 924. In at least one embodiment, a requesting entity (which provides an inference or image processing request) may browse the container registry and / or the model registry 924 for applications, containers, data sets, machine learning models, etc., select the desired combination of elements to include in a data processing pipeline, and submit a processing request. In at least one embodiment, the request may include the input data necessary to execute the request, and / or may include a selection of the application and / or machine learning model to be executed when processing the request. In at least one embodiment, the request may then be passed to one or more components (e.g., the cloud) of the deployment system 906 to perform the processing of the data processing pipeline. In at least one embodiment, the processing performed by the deployment system 906 may include referencing elements selected from the container registry and / or the model registry 924 (e.g., applications, containers, models, etc.). In at least one embodiment, once a result is generated by the pipeline, the result may be returned to the user for reference (e.g., for viewing in a viewing application suite executed locally, on a local workstation, or terminal).
[0092] In at least one embodiment, to assist in processing or executing applications or containers in a pipeline, service 920 can be utilized. In at least one embodiment, service 920 can include computing services, collaborative content creation services, simulation services, artificial intelligence (AI) services, visualization services, and / or other service types. In at least one embodiment, service 920 can provide functions common to one or more applications in software 918, and thus functions can be abstracted as services that can be called or utilized by applications. In at least one embodiment, the functions provided by service 920 can run dynamically and more efficiently, while also scaling well by allowing applications to process data in parallel (e.g., using Figure 10 the parallel computing platform 1030 therein). In at least one embodiment, rather than each application that requires the same function provided by shared service 920 having to have a corresponding instance of service 920, service 920 can be shared among and within various applications. In at least one embodiment, by way of non-limiting example, the service can include an inference server or engine that can be used to perform detection or segmentation tasks. In at least one embodiment, a model training service can be included, which can provide machine learning model training and / or retraining capabilities.
[0093] In at least one embodiment, in the case where service 920 includes an AI service (e.g., an inference service), as part of application execution, an inference service (e.g., an inference server) can be invoked (e.g., as an API call) to execute one or more machine learning models or their processing to perform one or more machine learning models associated with an application for anomaly detection (e.g., tumors, growth anomalies, scar formation, etc.). In at least one embodiment, in the case where another application includes one or more machine learning models for a segmentation task, the application can call the inference service to execute the machine learning model for performing one or more processing operations associated with the segmentation task. In at least one embodiment, the software 918 that implements the advanced processing and inference pipeline can be pipelined because each application can call the same inference service to perform one or more inference tasks.
[0094] In at least one embodiment, the hardware 922 can include a GPU, a CPU, a graphics card, an AI / deep learning system (e.g., an AI supercomputer such as NVIDIA's DGX TMa supercomputer system), a cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 922 may be used to provide efficient, specially - built support for software 918 and services 920 in the deployment system 906. In at least one embodiment, GPU processing may be implemented to perform local processing (e.g., at the facility 902) within an AI / deep - learning system, in a cloud system, and / or in other processing components of the deployment system 906 to improve the efficiency, accuracy, and effectiveness of game name recognition.
[0095] In at least one embodiment, by way of non - limiting example, with respect to deep learning, machine learning, and / or high - performance computing, simulation, and visual computing, software 918 and / or services 920 may be optimized for GPU processing. In at least one embodiment, at least some of the computing environments of the deployment system 906 and / or the training system 904 may be executed in a data center with GPU - optimized software (e.g., NVIDIA DGX TM system's hardware and software combination), or in one or more supercomputers or high - performance computer systems. In at least one embodiment, as described herein, the hardware 922 may include any number of GPUs, which may be invoked to perform data processing in parallel. In at least one embodiment, the cloud platform may also include GPU processing for GPU - optimized execution of deep - learning tasks, machine - learning tasks, or other computing tasks. In at least one embodiment, an AI / deep - learning supercomputer and / or GPU - optimized software (e.g., as provided on NVIDIA's DGX TM system) may be used as a hardware abstraction and scaling platform to execute a cloud platform (e.g., NVIDIA's NGC TM ). In at least one embodiment, the cloud platform may integrate an application container cluster system or a coordination system (e.g., KUBERNETES) on multiple GPUs to achieve seamless scaling and load balancing.
[0096] Figure 10 is a system diagram of an example system 1000 for generating and deploying a deployment pipeline according to at least one embodiment. In at least one embodiment, the system 1000 may be used to implement Figure 9 process 900 and / or other processes, including advanced processing and inference pipelines. In at least one embodiment, the system 1000 may include a training system 904 and a deployment system 906. In at least one embodiment, software 918, services 920, and / or hardware 922 may be used to implement the training system 904 and the deployment system 906 as described herein.
[0097] In at least one embodiment, system 1000 (e.g., training system 904 and / or deployment system 906) may be implemented in a cloud computing environment (e.g., using cloud 1026). In at least one embodiment, system 1000 may be implemented locally (with respect to a facility), or as a combination of cloud computing resources and local computing resources. In at least one embodiment, access to the APIs in cloud 1026 may be restricted to authorized users by establishing security measures or protocols. In at least one embodiment, the security protocol may include a network token, which may be signed by an authentication (e.g., AuthN, AuthZ, Gluecon, etc.) service and may carry appropriate authorization. In at least one embodiment, the APIs of the virtual instruments (described herein) or other instances of system 1000 may be restricted to a set of common Internet service providers (ISPs) that have been audited or authorized for interaction.
[0098] In at least one embodiment, the various components of system 1000 may communicate with each other and among themselves using any of a variety of different network types, including but not limited to local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communication between the facilities and components of system 1000 (e.g., for sending inference requests, for receiving results of inference requests, etc.) may be conveyed via one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (e.g., Ethernet), etc.
[0099] In at least one embodiment, similar to what is described herein with respect to Figure 9 the training system 904 may execute a training pipeline 1004. In at least one embodiment, where the deployment system 906 will use one or more machine learning models in a deployment pipeline 1010, the training pipeline 1004 may be used to train or retrain one or more (e.g., pre-trained) models, and / or implement one or more pre-trained models 1006 (e.g., without retraining or updating). In at least one embodiment, an output model 916 may be generated as a result of the training pipeline 1004. In at least one embodiment, the training pipeline 1004 may include any number of processing steps, such as AI-assisted annotation 910, labeling or annotation of feedback data 908 to generate labeled data 912, selection of a model from a model registry, model training 914, training, retraining, or updating of a model, and / or other processing steps. In at least one embodiment, different training pipelines 1004 may be used for different machine learning models used by the deployment system 906. In at least one embodiment, similar to the training pipeline 1004 described in the first example with respect to Figure 9 the first machine learning model, and similar to what is described with respect to Figure 9The training pipeline 1004 of the second example described can be used for a second machine learning model, similar to the Figure 9 The training pipeline 1004 of the third example described can be used for a third machine learning model. In at least one embodiment, any combination of tasks within the training system 904 can be used according to the requirements of each respective machine learning model. In at least one embodiment, one or more machine learning models may have been trained and be ready for deployment, so the training system 904 may not perform any processing on the machine learning models, and the machine learning models can be implemented by the deployment system 906.
[0100] In at least one embodiment, according to the embodiment, one or more output models 916 and / or pre-trained models 1006 can include any type of machine learning model. In at least one embodiment and without limitation, the machine learning models used by the system 1000 can include linear regression, logistic regression, decision trees, support vector machines (SVMs), naive Bayes, k-nearest neighbors (Knn), k-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptron, long / short-term memory (LSTM), Bi-LSTM, Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid state machines, etc.), and / or other types of machine learning models.
[0101] In at least one embodiment, the training pipeline 1004 can include AI-assisted annotation. In at least one embodiment, labeled data 912 (e.g., traditional annotation) can be generated by any number of techniques. In at least one embodiment, in some examples, labels or other annotations can be generated in a drawing program (e.g., an annotation program), a computer-aided design (CAD) program, a markup program, another type of application suitable for generating ground truth annotations or labels, and / or can be hand-drawn. In at least one embodiment, ground truth data can be synthetically generated (e.g., generated from a computer model or rendering), real-world generated (e.g., designed and generated from real-world data), machine-generated automatically (e.g., using feature analysis and learning to extract features from data and then generate labels), manually annotated (e.g., by a tagger or annotation expert to define the location of the label), and / or combinations thereof. In at least one embodiment, for each instance of feedback data 908 (or other data types used by the machine learning model), there can be corresponding ground truth data generated by the training system 904. In at least one embodiment, AI-assisted annotation can be performed as part of the deployment pipeline 1010; supplementing or replacing the AI-assisted annotation included in the training pipeline 1004. In at least one embodiment, the system 1000 can include a multi-layer platform, and the multi-layer platform can include a software layer (e.g., software 918) of a diagnostic application (or other application type), which can perform one or more medical imaging and diagnostic functions.
[0102] In at least one embodiment, the software layer can be implemented as a secure, encrypted, and / or certified API through which an application or container can be invoked (e.g., called) from an external environment (e.g., facility 902). In at least one embodiment, the application can then call or execute one or more services 920 to perform computing, AI, or visualization tasks associated with the respective application, and the software 918 and / or the services 920 can utilize the hardware 922 to perform processing tasks in an effective and efficient manner.
[0103] In at least one embodiment, the deployment system 906 can execute the deployment pipeline 1010. In at least one embodiment, the deployment pipeline 1010 can include any number of applications, which can be sequential, non-sequential, or otherwise applied to the feedback data (and / or other data types) - including AI-assisted annotation, as described above. In at least one embodiment, as described herein, the deployment pipeline 1010 for an individual device can be referred to as a virtual instrument for the device. In at least one embodiment, for a single device, there can be more than one deployment pipeline 1010, depending on the information desired from the data generated by the device.
[0104] In at least one embodiment, the applications that can be used to deploy pipeline 1010 can include any application that can be used to perform processing tasks on feedback data or other data from a device. In at least one embodiment, since various applications can share common image operations, in some embodiments, a data augmentation library (e.g., as one of services 920) can be used to accelerate these operations. In at least one embodiment, to avoid the bottleneck of traditional processing methods that rely on CPU processing, parallel computing platform 1030 can be used for GPU acceleration of these processing tasks.
[0105] In at least one embodiment, deployment system 906 can include a user interface 1014 (e.g., a graphical user interface, a web interface, etc.), which can be used to select the applications to be included in deployment pipeline 1010, arrange the applications, modify or change the applications or their parameters or configurations, use and interact with deployment pipeline 1010 during setup and / or deployment, and / or otherwise interact with deployment system 906. In at least one embodiment, although not shown with respect to training system 904, UI 1014 (or a different user interface) can be used to select the models to be used in deployment system 906, to select the models for training or retraining in training system 904, and / or to otherwise interact with training system 904. In at least one embodiment, training system 904 and deployment system 906 can include DICOM adapters 1002A and 1002B.
[0106] In at least one embodiment, in addition to application coordination system 1028, pipeline manager 1012 can also be used to manage the interaction between the applications or containers of deployment pipeline 1010 and services 920 and / or hardware 922. In at least one embodiment, pipeline manager 1012 can be configured to facilitate interaction from application to application, from application to service 920, and / or from application or service to hardware 922. In at least one embodiment, although shown as being included in software 918, this is not intended to be limiting, and in some examples, pipeline manager 1012 can be included in service 920. In at least one embodiment, application coordination system 1028 (e.g., Kubernetes, DOCKER, etc.) can include a container coordination system, which can group applications into containers as logical units for coordination, management, scaling, and deployment. In at least one embodiment, by associating the applications from deployment pipeline 1010 (e.g., reconstruction applications, segmentation applications, etc.) with individual containers, each application can execute in a self - contained environment (e.g., at the kernel level) to improve speed and efficiency.
[0107] In at least one embodiment, each application and / or container (or its image) can be developed, modified, and deployed separately (e.g., a first user or developer can develop, modify, and deploy a first application, and a second user or developer can develop, modify, and deploy a second application separate from the first user or developer), which can allow for focusing on and attending to the tasks of a single application and / or container without being hindered by the tasks of other applications or containers. In at least one embodiment, the pipeline manager 1012 and the application coordination system 1028 can assist in the communication and collaboration between different containers or applications. In at least one embodiment, as long as the expected inputs and / or outputs of each container or application are known to the system (e.g., based on the construction of the application or container), the application coordination system 1028 and / or the pipeline manager 1012 can facilitate communication and resource sharing between and among each application or container. In at least one embodiment, since one or more applications or containers in the deployment pipeline 1010 can share the same services and resources, the application coordination system 1028 can coordinate, perform load balancing, and determine the sharing of services or resources between and among the various applications or containers. In at least one embodiment, a scheduler can be used to track the resource requirements of applications or containers, the current or planned usage of these resources, and the resource availability. Thus, in at least one embodiment, the scheduler can allocate resources to different applications, taking into account the system's requirements and availability, and distribute resources between and among applications. In some examples, the scheduler (and / or other components of the application coordination system 1028) can determine resource availability and distribution based on constraints imposed on the system (e.g., user constraints), such as quality of service (QoS), the urgency of data output (e.g., to determine whether to perform real-time processing or deferred processing), etc.
[0108] In at least one embodiment, the services 920 utilized and shared by applications or containers in the deployment system 906 can include computing services 1016, collaborative content creation services 1017, AI services 1018, simulation services 1019, visualization services 1020, and / or other service types. In at least one embodiment, an application can invoke (e.g., execute) one or more services 920 to perform processing operations for the application. In at least one embodiment, an application can utilize the computing services 1016 to perform supercomputing or other high-performance computing (HPC) tasks. In at least one embodiment, one or more computing services 1016 can be utilized to perform parallel processing (e.g., using the parallel computing platform 1030) to process data substantially simultaneously by one or more applications and / or one or more tasks of a single application. In at least one embodiment, the parallel computing platform 1030 (e.g., NVIDIA's ) General computing can be implemented on a GPU (GPGPU) (e.g., GPU 1022). In at least one embodiment, the software layer of the parallel computing platform 1030 can provide access to the virtual instruction set and parallel computing elements of the GPU to execute compute kernels. In at least one embodiment, the parallel computing platform 1030 can include memory, and in some embodiments, the memory can be shared between and among multiple containers and / or between and among different processing tasks within a single container. In at least one embodiment, inter - process communication (IPC) calls can be generated for multiple containers and / or multiple processes within a container to use the same data from a shared memory segment of the parallel computing platform 1030 (e.g., where multiple different stages of one application or multiple applications are processing the same information). In at least one embodiment, rather than copying data and moving it to different locations in memory (e.g., read / write operations), the same data in the same location in memory can be used for any number of processing tasks (e.g., at the same time, different times, etc.). In at least one embodiment, since data is used to generate new data as a result of processing, this information about the new location of the data can be stored and shared among various applications. In at least one embodiment, the location of the data and the location of the updated or modified data can be part of the definition of how to understand the payload in a container.
[0109] In at least one embodiment, an AI service 1018 can be utilized to perform an inference service for executing a machine learning model associated with an application (e.g., the task is to perform one or more processing tasks of the application). In at least one embodiment, the AI service 1018 can utilize an AI system 1024 to execute a machine learning model (e.g., a neural network such as a CNN) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, the application of the deployment pipeline 1010 can use one or more output models 916 from the training system 904 and / or other models of the application to perform inference on imaging data (e.g., DICOM data, RIS data, CIS data, REST-compliant data, RPC data, raw data, etc.). In at least one embodiment, two or more examples of performing inference using an application coordination system 1028 (e.g., a scheduler) can be available. In at least one embodiment, the first category can include a high-priority / low-latency path, which can implement a higher service level agreement, such as for performing inference on an emergency request in an emergency situation or for a radiologist during a diagnostic process. In at least one embodiment, the second category can include a standard-priority path, which can be used for requests that may not be urgent or for situations where analysis can be performed at a later time. In at least one embodiment, the application coordination system 1028 can allocate resources (e.g., service 920 and / or hardware 922) based on the priority path for different inference tasks of the AI service 1018.
[0110] In at least one embodiment, a shared memory can be installed in the AI service 1018 in the system 1000. In at least one embodiment, the shared memory can operate as a cache (or other storage device type) and can be used to process inference requests from applications. In at least one embodiment, when an inference request is submitted, a set of API instances of the deployment system 906 can receive the request and can select one or more instances (e.g., for best fit, for load balancing, etc.) to process the request. In at least one embodiment, to process the request, the request can be input into a database, and if not already in the cache, the machine learning model can be located from the model registry 924. A validation step can ensure that the appropriate machine learning model is loaded into the cache (e.g., shared storage), and / or a copy of the model can be saved to the cache. In at least one embodiment, if the application is not already running or there are not enough instances of the application, a scheduler (e.g., the scheduler of the pipeline manager 1012) can be used to start the application referenced in the request. In at least one embodiment, if an inference server has not been started to execute the model, the inference server can be started. In at least one embodiment, each model can start any number of inference servers. In at least one embodiment, in a pull model where inference servers are clustered, the model can be cached whenever load balancing is beneficial. In at least one embodiment, the inference server can be statically loaded into the corresponding distributed server.
[0111] In at least one embodiment, an inference server running in a container can be used to perform inference. In at least one embodiment, an instance of the inference server can be associated with a model (and optionally with multiple versions of the model). In at least one embodiment, if an instance of the inference server does not exist when a request to perform inference on a model is received, a new instance can be loaded. In at least one embodiment, when starting the inference server, the model can be passed to the inference server so that the same container can be used to serve different models as long as the inference server runs as different instances.
[0112] In at least one embodiment, during application execution, an inference request for a given application may be received, and a container (e.g., an instance of a hosted inference server) may be loaded (if not already loaded), and a launcher may be invoked. In at least one embodiment, preprocessing logic in the container may load, decode, and / or perform any additional preprocessing on the incoming data (e.g., using a CPU and / or GPU). In at least one embodiment, once the data is ready for inference, the container may perform inference on the data as needed. In at least one embodiment, this may include a single inference call on an image (e.g., a hand X-ray), or may require inference on hundreds of images (e.g., chest CTs). In at least one embodiment, the application may summarize the results before completion, which may include but is not limited to a single confidence score, pixel-level segmentation, voxel-level segmentation, generating a visualization, or generating text to summarize the results. In at least one embodiment, different priorities may be assigned to different models or applications. For example, some models may have real-time (turnaround time less than 1 minute) priority, while other models may have a lower priority (e.g., turnaround time less than 10 minutes). In at least one embodiment, the model execution time may be measured from the requesting agency or entity, and may include the cooperative network traversal time as well as the execution time of the inference service.
[0113] In at least one embodiment, the transfer of requests between the service 920 and the inference application may be hidden behind a software development kit (SDK), and a robust transfer may be provided via a queue. In at least one embodiment, requests are placed in the queue via an API for an individual application / tenant ID combination, and the SDK pulls requests from the queue and provides the requests to the application. In at least one embodiment, the name of the queue may be provided in the environment from which the SDK picks up the requests. In at least one embodiment, asynchronous communication via the queue may be useful because it may allow any instance of the application to pick up work when it is available. In at least one embodiment, results may be transferred back via the queue to ensure no data loss. In at least one embodiment, the queue may also provide the ability to split the work, as the highest priority work may go into the queue connected to most instances of the application, while the lowest priority work may go into the queue connected to a single instance that processes tasks in the order received. In at least one embodiment, the application may run on a GPU-accelerated instance that is generated in the cloud 1026, and the inference service may perform inference on the GPU.
[0114] In at least one embodiment, a visualization service 1020 can be utilized to generate visualizations for viewing the output of an application and / or a deployment pipeline 1010. In at least one embodiment, the visualization service 1020 can utilize a GPU 1022 to generate visualizations. In at least one embodiment, the visualization service 1020 can implement rendering effects such as ray tracing or other light transport simulation techniques to generate higher quality visualizations. In at least one embodiment, the visualizations can include, but are not limited to, 2D image rendering, 3D volume rendering, 3D volume reconstruction, 2D tomographic slices, virtual reality displays, augmented reality displays, etc. In at least one embodiment, a virtualization environment can be used to generate a virtual interactive display or environment (e.g., a virtual environment) for interaction by system users (e.g., doctors, nurses, radiologists, etc.). In at least one embodiment, the visualization service 1020 can include an internal visualizer, movie and / or other rendering or image processing capabilities or functions (e.g., ray tracing, rasterization, internal optics, etc.).
[0115] In at least one embodiment, the hardware 922 can include a GPU 1022, an AI system 1024, a cloud 1026, and / or any other hardware for executing the training system 904 and / or the deployment system 906. In at least one embodiment, the GPU 1022 (e.g., NVIDIA's and / or QUADRO GPUs) can include any number of GPUs with any features or functions that can be used to perform processing tasks for computing services 1016, collaborative content creation services 1017, AI services 1018, simulation services 1019, visualization services 1020, other services, and / or software 918. For example, for the AI service 1018, the GPU 1022 can be used to perform preprocessing on imaging data (or other data types used by machine learning models), perform postprocessing on the output of machine learning models, and / or perform inference (e.g., to execute a machine learning model). In at least one embodiment, the cloud 1026, the AI system 1024, and / or other components of the system 1000 can use the GPU 1022. In at least one embodiment, the cloud 1026 can include a GPU-optimized platform for deep learning tasks. In at least one embodiment, the AI system 1024 can use GPUs, and one or more AI systems 1024 can be used to execute the cloud 1026 (or at least part of the tasks for deep learning or inference). Similarly, although the hardware 922 is shown as discrete components, this is not intended to be limiting, and any component of the hardware 922 can be combined with or utilized by any other component of the hardware 922.
[0116] In at least one embodiment, the AI system 1024 may include a specially constructed computing system (e.g., a supercomputer or HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, in addition to the CPU, RAM, memory, and / or other components, features, or functions, the AI system 1024 (e.g., NVIDIA's DGX TM ) may also include software (e.g., a software stack) that can use multiple GPUs 1022 to perform GPU-optimized processing. In at least one embodiment, one or more AI systems 1024 may be implemented in the cloud 1026 (e.g., in a data center) to perform some or all of the AI-based processing tasks of the system 1000.
[0117] In at least one embodiment, the cloud 1026 may include GPU-accelerated infrastructure (e.g., NVIDIA's NGC TM ) that can provide a GPU-optimized platform for performing the processing tasks of the system 1000. In at least one embodiment, the cloud 1026 may include an AI system 1024 for performing one or more AI-based tasks of the system 1000 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, the cloud 1026 may be integrated with the application coordination system 1028 that utilizes multiple GPUs to achieve seamless scaling and load balancing between and within the applications and services 920. In at least one embodiment, as described herein, the cloud 1026 may be responsible for performing at least some of the services 920 of the system 1000, including the computing service 1016, the AI service 1018, and / or the visualization service 1020. In at least one embodiment, the cloud 1026 may perform inference for large and small batches (e.g., execute NVIDIA's TensorRT TM ), provide an accelerated parallel computing API and platform 1030 (e.g., NVIDIA's ), execute the application coordination system 1028 (e.g., KUBERNETES), provide a graphics rendering API and platform (e.g., for ray tracing, 2D graphics, 3D graphics, and / or other rendering techniques to produce higher-quality movie effects), and / or may provide other functions for the system 1000.
[0118] In at least one embodiment, to protect patient confidentiality (e.g., in the case of off-site use of patient data or records), cloud 1026 can include a registry - such as a deep learning container registry. In at least one embodiment, the registry can store containers for instantiating applications that can perform pre-processing, post-processing, or other processing tasks on patient data. In at least one embodiment, cloud 1026 can receive data that includes patient data as well as sensor data in the containers, perform the requested processing only on the sensor data in those containers, and then forward the result output and / or visualization to appropriate parties and / or devices (e.g., local medical devices for visualization or diagnosis) without extracting, storing, or otherwise accessing the patient data. In at least one embodiment, patient data confidentiality is maintained in accordance with HIPAA and / or other data regulations.
[0119] Other variations are within the spirit of the present disclosure. Thus, although the disclosed techniques are susceptible to various modifications and alternative configurations, certain of its illustrated embodiments are shown in the drawings and have been described in detail above. However, it is to be understood that the disclosure is not intended to be limited to the one or more specific forms disclosed, but on the contrary, is intended to cover all modifications, alternative configurations, and equivalents falling within the spirit and scope of the present disclosure as defined by the appended claims.
[0120] Unless otherwise stated or clearly contradicted by the context, in the context of describing the disclosed embodiments (especially in the context of the appended claims), the use of the terms "a" and "an" and "the" and similar referents should be construed to cover both the singular and the plural, rather than as a definition of the terms. Unless otherwise stated, the terms "comprising", "having", "including", and "containing" should be construed as open-ended terms (meaning "including but not limited to"). The term "connected" (when unmodified refers to a physical connection) should be construed to mean partly or wholly included within, attached to, or joined together, even with some intervening elements. Unless otherwise indicated herein, references to numerical ranges in this document are only intended as a shorthand method of referring separately to each individual value falling within the range, and each individual value is incorporated into the specification as if it were individually recited herein. In at least one embodiment, unless otherwise indicated or contradicted by the context, the use of the term "set" (e.g., "set of items") or "subset" should be construed to mean a non-empty set including one or more members. Further, unless otherwise indicated or contradicted by the context, a "subset" of a corresponding set does not necessarily mean a proper subset of the corresponding set, but rather the subset and the corresponding set can be equal.
[0121] Unless otherwise expressly indicated or clearly contradicted by context, a conjunctive phrase such as "at least one of A, B, and C" or "at least one of A, B and C" is understood in context to mean generally that the items, clauses, etc. can be A or B or C, or any non-empty subset of the set A, B, and C. For example, in an illustrative example of a set having three members, the conjunctive phrases "at least one of A, B, and C" and "at least one of A, B and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunctive language is generally not intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise stated or contradicted by context, the term "plurality" denotes a plural state (e.g., "a plurality of items" means multiple items). In at least one embodiment, the number of items in a plurality of items is at least two, but can be more if expressly indicated or indicated by context. Further, unless otherwise stated or clear from context, the phrase "based on" means "at least partially based on" rather than "solely based on".
[0122] Unless otherwise indicated herein or clearly contradicted by context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that execute jointly on one or more processors by hardware or a combination thereof. In at least one embodiment, the code is stored, for example, in the form of a computer program on a computer-readable storage medium, the computer program including a plurality of instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transitory signals (e.g., propagated transient electrical or electromagnetic transmissions), but includes non-transitory data storage circuits (e.g., buffers, caches, and queues) within a transitory signal transceiver. In at least one embodiment, the code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which the executable instructions are stored, which when executed by one or more processors of a computer system (i.e., as a result of being executed) cause the computer system to perform the operations described herein. In at least one embodiment, a set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the plurality of non-transitory computer-readable storage media lack all of the code, but the plurality of non-transitory computer-readable storage media together store all of the code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors, e.g., the non-transitory computer-readable storage medium stores instructions and a main central processing unit (“CPU”) executes some instructions while a graphics processing unit (“GPU”) executes other instructions. In at least one embodiment, different components of a computer system have separate processors and different processors execute different subsets of the instructions.
[0123] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or jointly perform the operations of the processes described herein, and such a computer system is configured with suitable hardware and / or software that enables the implementation of the operations. Additionally, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system that includes a plurality of devices operating in different ways such that the distributed computer system performs the operations described herein and such that a single device does not perform all of the operations.
[0124] The use of any and all examples or exemplary language (e.g., "such as") provided herein is for illustrative purposes only to better clarify the embodiments of the present disclosure and does not limit the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating that any non-claimed element is essential for the practice of the disclosure.
[0125] All references cited herein, including publications, patent applications, and patents, are incorporated herein by reference to the extent that each reference is specifically and individually indicated to be incorporated by reference and the entire content thereof is set forth herein.
[0126] In the specification and claims, the terms "coupled" and "connected" and their derivatives may be used. It should be understood that these terms are not intended as synonyms for each other. Instead, in a particular example, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other.
[0127] Unless otherwise explicitly stated, it is understood that throughout the specification, terms such as "process", "compute", "calculate", "determine", etc., refer to actions and / or processes of a computer or computing system or similar electronic computing device that process and / or transform data represented as physical quantities (e.g., electrons) in the registers and / or memories of the computing system into other data similarly represented as physical quantities in the memories, registers, or other such information storage, transmission, or display devices of the computing system.
[0128] In a similar manner, the term "processor" may refer to any device or part of a device that processes electronic data from registers and / or memories and converts that electronic data into other electronic data that can be stored in registers and / or memories. As a non-limiting example, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, a "software" process may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process may refer to multiple processes that execute instructions sequentially or in parallel, continuously or intermittently. In at least one embodiment, the terms "system" and "method" may be used interchangeably herein, provided that a system can embody one or more methods and a method can be considered a system.
[0129] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog and digital data may be accomplished in a variety of ways, such as by receiving data as a parameter of a function call or a call to an application programming interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be accomplished by transmitting data via a serial or parallel interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. In at least one embodiment, reference may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data may be implemented by transmitting the data as an input or output parameter of a function call, an application programming interface, or a parameter of an interprocess communication mechanism.
[0130] Although the description herein sets forth example embodiments of the described techniques, other architectures may be used to implement the described functionality and are intended to fall within the scope of the present disclosure. Additionally, although specific assignments of responsibilities are defined above for purposes of description, the various functions and responsibilities may be assigned and partitioned in different ways depending on the circumstances.
[0131] Moreover, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter claimed in the appended claims need not be limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.
Claims
1. A method, comprising: Receiving, via a user interface UI, a natural language NL query associated with one or more fault indicators indicating a fault state of a system; Providing an input to a language model LM trained using documents associated with the system, the input including a prompt based at least on the NL query; Receiving, from the LM, a response to the NL query, the response including one or more indications associated with resolving the fault state of the system; And Causing the UI to display the response.
2. The method according to claim 1, wherein the one or more fault indicators are received from a diagnostic tool in association with a diagnosis performed on the system.
3. The method according to claim 2, wherein the input further includes one or more test logs associated with the diagnosis performed on the system.
4. The method according to claim 1 further comprises: Before receiving the NL query, Providing an initial input to the LM based at least on an initial NL query; And Receiving, from the LM, an initial response to the initial NL query, the initial response including one or more initial indications; And wherein the one or more fault indicators indicate an initial attempt to resolve the fault state of the system.
5. The method according to claim 1, wherein the LM includes a large language model.
6. The method according to claim 1, wherein the documents associated with the system include one or more of the following: Diagnostic documents of the system, System architecture documents of the system, or One or more training diagnostic logs associated with historical or simulated faults of the system.
7. The method according to claim 1, wherein the NL query is generated in response to at least one of the following: Regular maintenance of the system, Hardware updates of the system, Software updates of the system, Firmware updates of the system, or Inoperable conditions of the system.
8. The method according to claim 1, wherein the query includes at least one of a text prompt, a voice prompt, an audio prompt, or an image prompt.
9. A method, comprising: Generating training data, the training data including: A natural language NL query associated with one or more fault indicators indicating a fault state of a system, Documents associated with the system, and Training diagnostic logs associated with the fault state of the system; and Using the training data to train a language model LM to generate an NL response to the NL query, the NL response including one or more indications associated with resolving the fault state of the system.
10. The method according to claim 9, wherein the documents associated with the system include one or more of the following: Diagnostic documents of the system, or System architecture documents of the system.
11. The method according to claim 9, wherein the training diagnostic logs include at least one of the following: Diagnostic logs associated with a diagnosis performed on the system, or Comprehensive diagnostic logs for the fault state of the system.
12. The method according to claim 9, wherein the training data further includes: a sample NL response to the NL query, and where training the LM using the training data includes: applying the training data to the LM to obtain a trained NL response; and causing one or more parameters of the LM to be modified using a difference between the trained NL response and the sample NL response.
13. The method according to claim 9, wherein training the LM using the training data includes: applying the training data to the LM to obtain a trained NL response; obtaining an evaluation score characterizing the effectiveness of the trained NL response in resolving the fault state of the system; and and causing one or more parameters of the LM to be modified using the evaluation score.
14. The method according to claim 9, wherein the LM includes a neural network-based large language model.
15. A system, comprising: one or more processing units for: receiving, via a user interface UI, a natural language NL query associated with one or more fault indicators indicating a fault state of the system; providing an input to a language model LM trained using documents associated with the system, the input including at least a prompt based on the NL query; receiving, from the LM, a response to the NL query, the response including one or more indications associated with resolving the fault state of the system; and and causing the UI to display the response.
16. The system according to claim 15, wherein the one or more fault indicators are received from a diagnostic tool in association with a diagnosis performed on the system.
17. The system according to claim 16, wherein the input further includes one or more test logs associated with the diagnosis performed on the system.
18. The system according to claim 15, wherein the documents associated with the system include one or more of the following: diagnostic documents of the system, system architecture documents of the system, or one or more training diagnostic logs associated with historical or simulated faults of the system.
19. The system according to claim 15, wherein the LM communicates with at least one of an application programming interface API or a plugin to retrieve additional information related to the NL query.
20. The system according to claim 15, wherein the system is included in at least one of the following: an in-vehicle infotainment system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using edge devices; a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content; a system implemented using robots; a system for performing one or more conversational AI operations; a system implementing one or more large language models LLM; a system implementing one or more language models; A system for performing one or more generative AI operations; A system for generating synthetic data; A system comprising one or more virtual machines (VMs); A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.