Workload management engine in artificial intelligence system
By dynamically switching neural network models through a workload management engine, the problem of insufficient flexibility in artificial intelligence systems when handling diverse workloads on processing units is solved, thereby improving system performance and resource utilization, and reducing latency and resource contention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MICROSOFT TECHNOLOGY LICENSING LLC
- Filing Date
- 2024-11-14
- Publication Date
- 2026-06-23
Smart Images

Figure CN122270767A_ABST
Abstract
Description
Background Technology
[0001] Users rely on computing environments with applications and services to complete computational tasks. Users can interact with different types of applications and services powered by artificial intelligence (AI) systems. Specifically, neural networks, leveraging their ability to learn from data for prediction and decision-making tasks, have become a general-purpose tool across numerous applications. A neural network can refer to a computational model associated with machine learning and artificial intelligence. Neural networks consist of interconnected nodes organized in a layered manner (e.g., input layers, hidden layers, and output layers). For example, neural networks can support image and pattern recognition, assisting tasks such as object detection, facial recognition, and image classification, impacting a wide range of applications from security systems to photo tagging. Furthermore, in the field of natural language processing (NLP), neural networks power language translation, sentiment analysis, and chatbot interactions, thereby facilitating human-like communication between machines and users. Summary of the Invention
[0002] The technical aspects described herein generally relate to systems, methods, and computer storage media for providing workload management engines that utilize artificial intelligence systems for workload management, etc. The workload management engine supports dynamic switching between different neural network models for optimized performance. Specifically, workload management includes adaptive strategies that adjust the neural network model employed by a processing unit (e.g., a neural processing unit, "NPU") based on the dynamic nature of the workload, workload management factors, and workload management logic. The workload management engine includes neural network models that support strategic decisions optimized for processing units, workload management factors, and workload management logic.
[0003] The workload management engine operates by dynamically switching neural network models to optimize processing unit utilization to meet overall system requirements. Neural network models (e.g., full and simplified models) can be trained offline. Simplified models can be generated using different simplification strategies (e.g., quantization, pruning, or network architecture selection "NAS"). Neural network models are deployed to support dynamic selection based on workload management factors (e.g., physical environment conditions, operating mode, power mode, power supply mode, and NPU capacity). These models are also associated with workload management logic that indicates which model to use based on identified workload management factors. Furthermore, neural network models and associated tasks can be associated with priority identifiers, which are incorporated into the logic and decision-making factors for switching between models.
[0004] Typically, artificial intelligence systems lack comprehensive computing logic and infrastructure to effectively manage workloads for their processing units (e.g., NPUs). Unoptimized workload management for processing units can lead to several drawbacks that can impact the overall performance, efficiency, and effectiveness of neural network processing. For example, an intelligent surveillance system might be deployed on a single device equipped with an NPU. This device handles video feeds from multiple cameras and uses different neural networks to perform various computer vision tasks. The simultaneous operation of these neural networks (e.g., object detection neural networks, face recognition neural networks, and anomaly detection neural networks) on a shared NPU can cause contention for computing resources. An unoptimized NPU may lack the flexibility to adapt to diverse workloads, hindering its ability to effectively handle a wide range of applications, resulting in suboptimal inference speeds and latency in real-time processing tasks. Scalability challenges, inefficient memory management, and integration difficulties further exacerbate these limitations.
[0005] One technical solution—addressing the limitations of conventional artificial intelligence systems—may include the challenges of implementing a workload management engine that supports an adaptive policy framework; and the challenges of providing workload management operations and interfaces via the workload management engine in an AI system. This adaptive policy framework supports dynamically switching the neural network models adopted and evaluated to solve various problems in dynamic AI system environments with diverse processing unit workloads. Therefore, AI systems can be improved based on workload management operations that effectively provide NPU workload management.
[0006] During operation, multiple states of workload management factors are identified. Tasks associated with the workload processing unit are identified. Based on the task, the multiple states of the workload management factors, and the workload management logic, a neural network model is selected from multiple neural network models. The workload management logic supports dynamic switching between multiple neural network models. These models include full neural network models and simplified neural network models. The task is then executed using the selected neural network model.
[0007] This summary is provided to introduce some concepts in a simplified form, which will be further described in the detailed description below. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter. Attached Figure Description
[0008] The techniques described herein are described in detail below with reference to the accompanying drawings, wherein:
[0009] Figure 1AThis is a block diagram of an exemplary artificial intelligence system based on various aspects of the technology described herein, the exemplary artificial intelligence system including a workload management engine supporting a first artificial intelligence client (edge device) and a second artificial intelligence client (cloud device);
[0010] Figure 1B This is a block diagram of an exemplary artificial intelligence client (edge device) based on various aspects of the technology described herein, which has a workload manager with a workload management engine;
[0011] Figure 2 This is a block diagram illustrating exemplary schematics associated with a workload management engine based on various aspects of the technology described herein;
[0012] Figure 3 This document provides a first exemplary method for providing workload management using a workload management engine based on the techniques described herein.
[0013] Figure 4 A second exemplary method for providing workload management using a workload management engine based on the techniques described herein is provided.
[0014] Figure 5 A third exemplary method for providing workload management using a workload management engine based on the techniques described herein is provided;
[0015] Figure 6 A block diagram of an exemplary distributed computing environment suitable for use in implementing various aspects of the techniques described herein is provided; and
[0016] Figure 7 This is a block diagram of an exemplary computing environment suitable for use in implementing various aspects of the techniques described herein. Detailed Implementation Overview
[0017] An artificial intelligence (AI) system refers to an AI computing environment or architecture that includes the infrastructure and components supporting the development, training, and deployment of AI models. It provides the necessary hardware, software, and frameworks for developers to create and run AI applications. AI systems can be cloud-based AI solutions that utilize cloud computing infrastructure to develop, train, deploy, and manage AI models and applications. AI models can specifically refer to neural networks, which are computational models associated with machine learning and artificial intelligence. Neural networks consist of interconnected nodes organized in a layered manner (e.g., input layers, hidden layers, and output layers).
[0018] Artificial intelligence systems can include neural networks that leverage their ability to learn from data for prediction and decision-making tasks, becoming a general-purpose tool across numerous applications. For example, neural networks can support image and pattern recognition, assisting tasks such as object detection, facial recognition, and image classification, impacting a wide range of applications from security systems to photo tagging. In the field of natural language processing (NLP), neural networks power language translation, sentiment analysis, and chatbot interactions, thereby facilitating human-like communication between machines and users. Speech recognition systems rely on neural networks for transcription and voice control functions, found in virtual assistants and voice-operated devices. From healthcare applications (where neural networks assist in medical diagnosis by analyzing images and detecting anomalies) to the financial sector (where neural networks help with fraud detection and stock price prediction), neural networks are driving innovation and automation.
[0019] Computing power is becoming a critical resource, especially as artificial intelligence (i.e., AI models) is employed to perform diverse tasks, some of which were not previously performed using AI models. Edge devices, such as monitoring systems, can include several AI models supporting monitoring capabilities. For example, AI models can be associated with subsystems of the edge device (e.g., cameras, sensors); however, AI models are performance-optimized and lack limitations on AI parameters—this can lead to AI tasks being executed serially rather than simultaneously, impacting their speed. Given these physical limitations (e.g., memory, battery status, processor capacity) and market demand for more diverse AI implementations, edge devices require performance improvements and management (e.g., processing unit optimization) to support the different AI models on these devices. Specifically, as the use of AI models on processors older than these models increases, providing the management and flexibility to run these AI models on those processors is crucial.
[0020] Typically, artificial intelligence systems are not configured with comprehensive computing logic and infrastructure to effectively manage the workload of their processing units (e.g., NPUs). Unoptimized workload management can lead to several drawbacks for the processing units, potentially impacting the overall performance, efficiency, and effectiveness of neural network processing. For example, an intelligent surveillance system might be deployed on a single device equipped with an NPU. This device processes video feeds from multiple cameras and uses different neural networks to perform various computer vision tasks. The simultaneous operation of these neural networks (e.g., object detection neural networks, face recognition neural networks, and anomaly detection neural networks) on a shared NPU can cause contention for computing resources.
[0021] As illustrated, all three neural networks can compete for and share the NPU's processing resources. The object detection network requires real-time processing to track and identify objects, the face recognition network requires precise computation to accurately match faces, and the anomaly detection network needs to continuously analyze the video stream for any anomalous patterns. Competition for NPU resources can lead to challenges such as increased inference latency, reduced overall throughput, and potential delays in responding to real-time events. An unoptimized NPU may lack the flexibility to adapt to diverse workloads, hindering its ability to effectively handle a wide range of applications, resulting in suboptimal inference speeds and thus latency in real-time processing tasks. Therefore, a more comprehensive AI system—with alternative bases for performing workload management operations—can improve the computational operations and interfaces associated with the processing units that manage the workloads of the neural network models.
[0022] Embodiments of this technical solution relate to systems, methods, and computer storage media for providing workload management using a workload management engine of an artificial intelligence system, etc. The workload management engine supports dynamic switching between different neural network models for optimized performance. Specifically, workload management includes adaptive strategies that adjust the neural network model employed by a processing unit (e.g., a neural processing unit "NPU") based on the dynamic nature of the workload, workload management factors, and workload management logic. The workload management engine includes a neural network model supporting strategic decisions optimized for the processing unit, workload management factors, and workload management logic.
[0023] The workload management engine operates by dynamically switching neural network models to optimize resource consumption of processing units. Neural network models (e.g., full neural network models and simplified neural network models) can be trained offline. Simplified neural network models can be generated using different simplification strategies (e.g., quantization, pruning, or network architecture selection "NAS"). Neural network models are deployed to support dynamic selection based on workload management factors (e.g., physical environment conditions, operating mode, power mode, power supply mode, and NPU capacity). Neural network models are also associated with workload management logic that indicates which neural network model to use based on identified workload management factors. Furthermore, neural network models and the tasks associated with them can be associated with priority identifiers, which are incorporated into the logic and decision-making factors for switching between neural network models.
[0024] Workload management is provided using a workload management engine that is operationally integrated into the artificial intelligence system. This AI system supports a workload management framework for computing components associated with dynamically switching between different neural network models for optimized performance. The use of an NPU is exemplary; it is conceivable that other types of processing units (e.g., GPUs / TPUs / CPUs) could also be associated with this implementation.
[0025] At a high level, neural networks can have predefined resource footprints—typically optimized for performance. Neural network resource footprints can refer to the resource and performance characteristics of a neural network. However, due to the different operating factors of the processing units supporting these neural networks (e.g., NPU / GPU / TPU), not all neural networks can be optimized for performance. Utilizing performance trade-offs to balance factors such as memory usage, computational efficiency, power consumption, and computational power can support the optimization of processing unit operations. Specifically, different simplification strategies (e.g., quantization, pruning, and network architecture selection) can be used to generate simplified neural network models with corresponding full neural network models; and based on the problems experienced by the processing unit (e.g., runtime bottlenecks, the need for reduced power consumption), optimization logic is used to determine which simplified neural network model should be used to perform the task.
[0026] Processing units (or workload processing units), encompassing neural processing units (NPUs), graphics processing units (GPUs), and tensor processing units (TPUs), are specialized hardware components designed to accelerate computational tasks in electronic devices. NPUs are tailored for neural network operations, efficiently performing parallelized tasks such as matrix multiplication and tensor computation. GPUs, initially developed for graphics rendering, have evolved into powerful parallel processors adept at handling the large-scale mathematical computations indispensable in neural network training and inference. They operate by processing multiple data points concurrently, making them particularly effective for parallelizable deep learning models. TPUs support tensor computation and are optimized for training and inference, offering high throughput and energy efficiency. The operational efficiency of these processing units is crucial for accelerating neural network workloads, where the choice depends on factors such as the nature of the task, its scale, and overall performance requirements.
[0027] In operation, the specific application employing the first neural network model no longer primarily determines how to execute that first neural network model on the processor; instead, a workload manager is provided to determine, for example, based on trade-offs and parameters, how to execute the first neural network model relative to multiple other neural network models, or even for a specific user. For example, a neural network model for a cellular phone could be executed via an optimized NPU based on evaluating workload management factors (such as battery state, power mode, and physical conditions) and employing workload management logic associated with the neural network model to determine how to handle the tasks associated with those neural network models.
[0028] The high-level process may include a preparation phase and a runtime phase for workload management of the processing unit. The preparation phase may include training a full neural network model and a simplified neural network model. A full neural network refers to the complete architecture of the model, encompassing all layers, nodes, and parameters originally designed for a specific task. For example, this may include complex architectures such as deep neural networks, convolutional neural networks (CNNs), or recurrent neural networks (RNNs). Embodiments of this technical solution also envision other types of neural networks.
[0029] On the other hand, simplified neural networks represent models that have undergone simplification or optimization processes to reduce their size, complexity, or computational requirements. Simplification techniques can involve pruning connections or neurons, quantizing weights, employing knowledge distillation, or applying model compression methods. The choice between full and simplified neural networks depends on the specific requirements of the application, where full networks offer high expressiveness and accuracy, while simplified networks offer advantages in computational efficiency and adaptability to resource-constrained environments. This decision involves weighing the pros and cons between model complexity, speed, and task performance.
[0030] Neural network models can be stored and deployed to support dynamic switching between them. A neural network model can include a first complete neural network model associated with at least two corresponding simplified neural network models, and a second complete neural network model associated with at least two corresponding simplified neural network models. Neural network models can be adopted in different types of scenarios, including edge devices with limited computing power and memory that require higher energy efficiency. Thus, simplified neural network models can be used if edge devices need to reduce power consumption or if the NPU is operating at full computing capacity.
[0031] The preparation phase may further include associating workload management factors (such as physical environmental conditions, operating modes, power modes, battery status, power modes, processor activity capacity, task priorities, and neural network model priorities) and workload management logic with different neural network models. This allows the workload management logic to determine the appropriate neural network model for each processing unit to perform the task, based on the workload management factors. For example, a first complete neural network model associated with audio noise cancellation might be a low-priority neural network model; therefore, the logic could instruct that the first complete neural network model should be moved to a first simplified neural network model before a second complete neural network model, which is a high-priority neural network model associated with monitoring and detection, is moved.
[0032] The runtime phase can include a startup mechanism (e.g., event triggering, time loop, state change) used to invoke trigger inputs associated with starting the runtime phase. The runtime phase supports dynamic switching from the current neural network model to a selected subsequent neural network model. Triggered interrupt startup uses logic associated with specific events and priorities (i.e., workload management factors) to select the subsequent neural network. For example, an intelligent monitoring system might associate this with physical environmental conditions (e.g., day or night), operating modes (e.g., active, idle, power-saving), battery status (e.g., low, medium, high), power supply (e.g., power connected, power disconnected), and NPU capacity (e.g., low, medium, high capacity). Workload management factors can be predefined factors (e.g., always on), which are established in advance based on known standards, specifications, or predefined rules; and measurement factors (e.g., power supply), which are dynamically determined during the processing of real-time data or observation. The subsequent neural network model can be used to perform tasks associated with the NPU—ideally optimizing the operation of the NPU.
[0033] This allows for the re-evaluation of initiation mechanisms (e.g., event triggering, time loops, state changes) to assess subsequent neural network models in the current activity. The re-evaluation mechanism determines whether the subsequent neural network model performs as expected (e.g., meets predefined performance thresholds). In other words, the performance of the subsequent neural network model is continuously or periodically monitored based on quantified metrics to determine if it meets predefined performance thresholds. If the subsequent neural network model reaches or exceeds these thresholds, it is retained.
[0034] A predefined performance threshold is a predetermined level or standard used as a benchmark or standard for evaluating the performance of a processor, neural network model, or task. This threshold is established based on specific requirements, expectations, or industry standards, and it represents the minimum acceptable performance level. It can be defined across various metrics, such as accuracy, speed, efficiency, or any other relevant performance indicator depending on the context. The purpose of setting predefined performance thresholds is to establish clear standards for evaluating the effectiveness, reliability, and suitability of a processor, neural network model, or task, thereby ensuring that it meets expected standards and objectives. Exceeding or meeting these predefined thresholds indicates successful and satisfactory performance, while falling below these thresholds may require further optimization or improvement efforts.
[0035] This explanation demonstrates how, for image classification, full neural network models, particularly convolutional neural networks (CNNs), can be evaluated based on accuracy, inference time, model size, and FLOPs (floating-point operations). Subsequently, simplified neural network models derived through quantization—where parameter precision is reduced—are evaluated using the same metrics. Comparing these metrics allows for a comprehensive assessment of the trade-offs between the full and simplified models. This evaluation framework facilitates the selection of optimized models tailored to the specific requirements of the deployment scenario, thus balancing computational efficiency with task performance.
[0036] If the performance improvement of a subsequent neural network model is deemed insufficient, a decision can be made to revert to the previous neural network model. Conversely, if reverting to the previous model is deemed not beneficial, a second subsequent neural network model can be selected and implemented. This dynamic decision-making process ensures adaptability, allowing for continuous optimization of NPU performance based on real-time observations.
[0037] Dynamic decision-making processes can be executed based on workload management logic. The workload management logic for processing units (e.g., NPUs) involves observing different workload management factors (e.g., physical environment conditions, power patterns, NPU capacity) associated with various neural network models (e.g., full neural network models, first, second, and third simplified neural network models). Each neural network model has unique characteristics—the first simplified neural network model is simplified based on quantization, the second simplified neural network model is simplified based on quantization, and the third simplified neural network model is simplified based on NAS.
[0038] By way of explanation, workload management logic can include a decision matrix that correlates observed workload management factors with the characteristics of each neural network model. Example use cases include the following:
[0039] For the "always-on device requiring power saving (efficiency)" use case, the logic is prompted to select a simplified neural network that can support basic detection but does not support tracking;
[0040] For the "low battery, power saving mode, power not connected" use case, the logic is prompted to select a simplified neural network model that is optimized for power and may have lower performance; and
[0041] In response to "the NPU is nearing full computing capacity and the NPU has a request to add one or more networks", the logic is prompted to select one or more simplified neural network models that have been quantized, pruned and optimized for memory.
[0042] The workload management logic incorporates adaptive thresholds and dynamic adjustments, allowing it to respond to real-time workload changes and refine decisions based on historical data. A feedback loop facilitates continuous learning, while a fallback mechanism ensures resilience by providing alternatives in case any neural network model becomes unavailable or encounters problems. Overall, this adaptive workload management logic optimizes resource utilization, enhances system performance, and adapts to evolving workload patterns.
[0043] Workload management logic can be priority-aware logic for tasks associated with the NPU. This logic considers both workload management factors and the priority level associated with each task (e.g., critical, high, medium, low priority). For example, the previously described decision matrix can be extended to include task priorities, establishing priority-aware decision rules to guide the selection process. For instance, for "low battery, power saving mode, device not connected to power," the logic is prompted to select a simplified neural network model optimized for power and potentially with lower performance—if the associated task has a critical priority. This logic incorporates dynamic adjustments to task priorities, ensuring adaptability to changing environments. Integration of task queues and scheduling mechanisms allows for efficient execution based on task priority and component availability. Fallback mechanisms are strengthened, providing redundancy for critical tasks to prevent interruptions. This priority-aware workload management logic optimizes resource allocation, adapts to changing priorities, and enhances the overall efficiency of the system.
[0044] Therefore, processing unit optimization can be based on different workload management factors. For each use case that meets specific workload management factors, workload management logic tailored to the neural network model is applied to determine what to do. Specifically, based on the state of the workload management factors and the workload management logic, the current neural network model can be switched to a subsequent neural network model (e.g., a simplified neural network model). If the subsequent neural network model operates within a predefined performance threshold, then that subsequent neural network model is used—or one can try moving to another subsequent neural network (e.g., another simplified neural network model smaller than the simplified neural network model).
[0045] The processing unit can be configured to support multiple neural network models, which are prioritized for optimization evaluation. Therefore, while optimizing the first neural network model, the next one—based on priority—can be optimized next. It is envisioned that the NPU can return to normal operating mode based on several predefined conditions (e.g., timeouts or changes in one or more workload management conditions). For example, predefined conditions could include environmental changes (e.g., device movement using the inertial measurement unit (IMU), changes in lighting conditions, changes in noise levels) or user changes (e.g., a detection report indicating a user has left, or a change in the number of people in the scene).
[0046] Processing unit optimization can also be associated with additional optional extensions. Administrator users can be given priority control, including the ability to adjust the order or importance of various tasks, processes, or components of processing unit optimization to tailor the optimization to their unique preferences and requirements. User priority control can also include prioritizing specific user types. For example, user priority control can prioritize accessibility users. Users can identify themselves as accessibility users, and based on this parameter, workload management logic can instruct the neural network model to prioritize accessibility associated with that user. Thus, workload management logic can define a first set of decision rules for accessibility users and a second set of decision rules for non-accessibility users.
[0047] Processing unit optimization may also include running full neural network models and simplified neural network models (e.g., quantized simplified neural network models) in parallel and comparing the outputs using a predefined comparison framework. A predefined comparison framework is used because different networks can have different criteria or parameters for evaluating strengths, weaknesses, features, and other relevant attributes. Parallel evaluation of full and simplified neural network models is used because, for example, in edge devices, full neural network models can support different user sets; however, simplified neural network models may be more effective when identifying specific users in a given scenario. Furthermore, processing unit optimization may include continuously evaluating the performance of the active neural network model to make a dynamic switch to a subsequent neural network model. Processing unit optimization may also include interchangeably running neural networks of different sizes to continuously test differences, or interchangeably running neural networks of different sizes to reduce errors and provide more accurate results. Embodiments of this technical solution also envision other variations and combinations with optional extensions.
[0048] Advantageously, embodiments of this technical solution include several inventive features (e.g., operations, systems, engines, and components) associated with an artificial intelligence system having a workload management engine. This workload management engine supports workload management operations used to achieve dynamic switching between different neural network models for optimized performance—and provides the AI system operations and interfaces via the workload management engine within the AI system. These workload management operations are solutions to specific problems in AI systems (e.g., the limited flexibility of the NPU's capacity to adapt to diverse workloads, hindering its ability to effectively handle a wide range of applications, causing suboptimal inference speeds, and thus causing latency in real-time processing tasks). The workload management engine provides an ordered combination of operations, including adaptive strategies that adjust the neural network model adopted by the NPU based on the dynamic properties of the workload and workload management factors—which improves computational operations within the AI system.
[0049] In this way, the workload management engine provides technical improvements in AI, particularly in prioritizing and managing neural network models. The workload management engine involves the integration of neural network models with processor optimizations, providing technical solutions that go beyond abstract concepts—for example, the workload management logic provides priority ranking algorithms and decision rules for performing processor optimizations. Processor optimization results in efficiency gains in managing multiple neural networks, including the efficient allocation of resources, including processing power and memory, for improved tasks. Furthermore, the optimization techniques are particularly well-suited to the unique characteristics of neural network models (e.g., full and simplified neural networks; different simplification strategies; different tasks and task priorities) to address inherent processor challenges, distinguishing it from general optimization methods. Example systems and operations
[0050] Various aspects of this technical solution can be seen through examples and references. Figures 1A to 1B To be described. Figure 1A The diagram illustrates a cloud computing system (environment) 100, which includes an artificial intelligence system 100A; a network 100B; a workload management engine 110, workload management operations 112, a workload manager 120 (including workload management elements 122 and workload management logic 124); an artificial intelligence client 130 (e.g., an edge device) (including a workload manager 130A, a processing unit 130B, and a neural network model 130C); an artificial intelligence client 140 (e.g., a cloud device) (including a workload manager 140A, a processing unit 140B, and a neural network model 140C); and a machine learning engine 150, including a full neural network model 152 and a simplified neural network model 154.
[0051] Cloud computing environment 100 provides computing system resources for different types of managed computing environments. For example, cloud computing environment 100 supports the delivery of computing services—including servers, storage, databases, networks, software-integrated applications and services (collectively, “(multiple) services)”), and artificial intelligence systems (e.g., artificial intelligence system 100A). Multiple artificial intelligence clients (e.g., artificial intelligence client 130) include hardware or software that accesses resources in cloud computing environment 100. Artificial intelligence client 130 may include applications or services that support client-side functionality associated with cloud computing environment 100. Multiple artificial intelligence clients can access the computing components of cloud computing environment 100 via a network (e.g., network 100B) to perform computing operations.
[0052] Artificial intelligence system 100A is responsible for providing an artificial intelligence computing environment or architecture, including the infrastructure and components that support the development, training, and deployment of artificial intelligence models. Artificial intelligence system 100A is also responsible for providing workload management associated with workload management engine 110. Artificial intelligence system 100A operates to support the generation of inference for machine learning models.
[0053] The artificial intelligence system 100A provides a workload management engine 110, which supports the dynamic selection of different neural network models (e.g., multiple neural network models) to provide workload management on a processing unit (e.g., an NPU). The workload management engine 110 can support a preparation phase for preparing and deploying multiple neural network models, as well as a runtime phase for dynamically switching between multiple neural network models on the processing unit. For example, multiple neural network models can be deployed in artificial intelligence clients (e.g., artificial intelligence client 130 and artificial intelligence device 140) to support runtime operations via workload managers (e.g., workload manager 130A and workload manager 140A) on the corresponding artificial intelligence clients. The workload management engine 110 can provide a machine learning engine (e.g., machine learning engine 150) that supports providing multiple neural networks (e.g., a complete neural network model 152 and a simplified neural network model).
[0054] Machine Learning Engine 150 is a machine learning framework or library that operates as a tool providing the infrastructure, algorithms, and capabilities for designing, training, and deploying machine learning models. Machine Learning Engine 150 may include pre-built features and APIs that enable the building and application of machine learning techniques. Machine Learning Engine 150 can provide a machine learning workflow from data processing and feature extraction to model training, evaluation, and deployment.
[0055] Machine Learning Engine 150 trains multiple neural network models (e.g., multiple neural network models), including different versions of full neural network models and different versions of simplified neural network models. Multiple neural network models can be trained using different neural network model optimizations (e.g., neural network simplification strategies). For example, multiple neural networks can be simplified based on quantization, pruning, or network architecture selection. Quantization (a model compression technique) reduces the precision of numerical representations in a neural network, reducing them to a lower bit width, such as 16-bit or 8-bit integers. This results in a more compact model, reduced memory requirements, and faster inference, with a graceful decrease in precision. Pruning involves selectively removing less important connections or neurons during training, producing a sparser model with fewer parameters, reduced memory footprint, and potentially faster inference time. Network architecture selection encompasses customizing or selecting a neural network architecture that strikes a balance between complexity and task performance, often favoring simplicity for efficient resource use and simplified deployment. Transfer learning involves fine-tuning a pre-trained model for a specific task, representing another aspect of architecture selection. Machine Learning Engine 150 provides multiple models that can be deployed to different types of AI clients, such as edge devices or cloud devices.
[0056] The preparation phase may further include the workload management engine 110 associating workload management factors (e.g., workload management factor 122) and workload management logic (e.g., workload management logic 124). Workload management factors may include operational factors that determine processing unit functionality and performance optimization. For example, workload management factors may include physical environment conditions, operating modes, battery life, power modes, power supply modes, processor activity capacity, task priorities, and neural network model priorities. Workload management logic refers to a set of rules, algorithms, and strategies implemented to optimize the processing unit. For example, workload management logic is used to determine what to do for active neural network models and subsequent neural network models—as discussed in more detail herein. Therefore, the workload management engine accesses multiple neural network models and associates these multiple neural network models with corresponding workload management factors 122 and workload management logic 124.
[0057] The workload management engine deploys multiple neural network models to support workload management, including dynamic switching between these models based on workload management factors 122 and workload management logic 124. These multiple neural network models can be deployed to different operational scenarios (e.g., edge devices or cloud computing applications). Convolutional Neural Networks (CNNs) are well-suited for vision tasks such as image classification, object detection in autonomous driving, and medical image analysis. They excel at capturing spatial patterns. Recurrent Neural Networks (RNNs), on the other hand, are effective in sequence-based tasks. They have applications in natural language processing such as language modeling and sentiment analysis, in finance for time series prediction, and in human-computer interaction for speech recognition and gesture recognition. CNNs focus on visual data and spatial relationships, while RNNs specialize in handling sequential and temporal information, demonstrating their versatility in various machine learning applications. Embodiments of this technical solution also envision other variations of neural network models and scenarios.
[0058] The workload management engine 110 provides a workload manager 120, which manages the execution of neural network tasks. The workload manager 120 is responsible for optimizing the utilization of processing unit computing resources, thereby ensuring that neural network workloads are handled effectively. Specifically, the workload manager 120 enables tasks associated with the AI client to be executed using a selected neural network model. The workload manager 120 may include workload management factors 122 and workload management logic 124, supporting workload management on the AI client.
[0059] As shown in the figure, workload manager 120 is deployed in the cloud; however, workload manager 120 can be deployed to different types of devices to support the functionality described herein. Specifically, workload manager 120 can be deployed on different devices and applications to provide the workload management functionality described herein. For example, AI client 130 can be an edge device, including workload manager 130A, processing unit 130B (e.g., NPU), and neural network model 130C (e.g., a neural network model from machine learning engine 150). AI client 140 can be a remote or local cloud device or application, including workload manager 140, processing unit 140B (e.g., GPU), and neural network model 140C (e.g., a neural network model from machine learning engine 150). Workload manager 140A provides similar functionality to workload manager 130A for AI client 140 (e.g., cloud device).
[0060] Figure 1B The illustration depicts an AI client 130 with a workload manager 130A, which includes a workload management element 132, workload management logic 134, a processing unit 130B, and a neural network model 130C. By way of example, the AI client 130 can refer to an exemplary edge device supporting the implementation of multiple neural networks. This edge device could be a smart security camera designed for multifaceted functionality. Equipped with a neural processing unit (NPU) (e.g., processing unit 130B), the edge device can simultaneously perform machine learning tasks via the neural network model 130C. One neural network model can be dedicated to object detection for identifying and classifying various entities within the camera's range. Simultaneously, another neural network model is dedicated to facial recognition, enabling the camera to identify and verify individuals. Additionally, a third neural network model focuses on anomaly detection to identify anomalous patterns or unexpected events for enhanced security.
[0061] The workload manager 130A is associated with the processing unit 130B to support dynamic switching between multiple neural network models based on workload management factors 132 and workload management logic 134. The processing unit 130 (or workload processing unit) can be a neural processing unit (NPU) or other processing units (such as a graphics processing unit (GPU) or tensor processing unit (TPU)), which are dedicated hardware components designed to accelerate computational tasks in electronic devices. NPUs are tailored for neural network operations, efficiently performing parallelized tasks such as matrix multiplication and tensor computation.
[0062] Workload management factors 132 may include operational factors that determine the functionality and performance optimization of the processing unit. For example, workload management factors 132 may include physical environmental conditions, operating mode, battery life, power mode, power supply mode, processor activity capacity, task priority, and neural network model priority. Workload management logic 134 refers to a set of rules, algorithms, and strategies implemented to optimize the processing unit. Applying workload management logic 134 to the processing unit 130B involves observing different workload management factors 132 associated with various neural network models (e.g., a complete neural network model, a first simplified neural network model, a second simplified neural network model, and a third simplified neural network model). Workload management logic 134 may be associated with a first priority type (e.g., a set of priority identifiers: critical, high, medium, low) associated with a task and a second priority type (e.g., a set of priority identifiers: high, medium, low) associated with multiple neural network models. Thus, both priority types can be used to determine which neural network model to select.
[0063] The AI client 130 includes neural network models 130C. Each neural network model has unique characteristics—for example, a first simplified neural network model is simplified based on quantization, a second simplified neural network model is simplified based on quantization, and a third simplified neural network model is simplified based on NAS. Neural networks are selectively chosen based on workload management factors 132 and workload management logic 134 to optimize the performance of the processing unit 130B.
[0064] Therefore, in operation, the workload manager 130A identifies multiple states of workload management factors. The workload manager 130A identifies a task associated with the processing unit 130B. Based on the task and the multiple states of the workload management factors, the workload manager 130A selects a neural network model from the neural network model 130C. The workload manager 130A causes the task to be executed on the processing unit 130B using the selected neural network model.
[0065] refer to Figure 2 , Figure 2 The diagram illustrates an example implementation of a workload management engine for NPU workload optimization 200. The workload management engine provides offline model training 212, supporting the training of multiple neural network models. Model training 202 can support full-size model training 214 for a full-size model, training using quantization 216 for a simplified model 0, and training using pruning for a simplified model 1.
[0066] As discussed in this paper, multiple neural network models can be associated with workload management factors and workload management logic. The workload management engine can also support online workload optimization, such as using a workload manager and workload scheduler. During the preparation phase, offline training of the neural network models is performed. Training can include training both a full neural network model and a simplified neural network model—using simplification strategies. Quantization involves training a simplified neural network model that approximates the original performance of the full neural network model as closely as possible, saving memory and power consumption (e.g., fewer cycles for inference). Pruning involves training the simplified neural network model using several different techniques (e.g., reducing weights, removing arcs from the graph, employing different degrees of pruning). The network architecture can include simplified neural network sizes that leverage different configurations of the full neural network model, where different simplified neural network sizes can be selectively implemented. It is envisioned that different types of simplification strategies can be combined to support the training of simplified neural network models.
[0067] The trained model can be associated with workload management logic that provides decision rules for selecting among neural network models. For example, the decision rules can be associated with workload management factors such as latency, power consumption, and power mode. The workload management logic can vary, for example, by being associated with the physical conditions of the processing unit (e.g., user presence for face detection), software features (e.g., always-on mode), or hardware characteristics (e.g., NPU memory capacity). The workload management logic can prioritize different workload management factors, including priorities for neural network type, task type, and factor type. For example, a first priority set can be used when power is connected, and a second priority set can be applied when power is disconnected.
[0068] During the runtime phase, the runtime phase may include a startup mechanism (e.g., event triggering, time loop, state change) used to invoke trigger inputs associated with starting the runtime phase. The runtime phase supports dynamic switching from the current neural network model to a selected subsequent neural network model. System request 222, such as a system request associated with observing workload management factors, may trigger workload optimization 224. Workload optimization 224 may include accessing workload management factors and employing workload management logic to select one or more neural network models. Model set execution 226 may include a processor unit (e.g., an NPU) executing one or more neural network models. Results 228 (e.g., performance metrics) associated with the execution of one or more neural networks may be generated. It is envisioned that a feedback mechanism 230 may be implemented so that results 228 are communicated, and based on the evaluation results 228 and subsequent workload management factors, another system request 234 may be generated and communicated 236 to dynamically switch the active neural network model to a subsequent neural network model.
[0069] Workload optimization 240 may include an LLM path 250 and a human presence detection path 270. The LLM path 250 may be associated with a first set of neural network models (e.g., a first complete neural network model and a first simplified neural network model) for the LLM task. The human presence detection path 270 may be associated with a second set of neural network models (e.g., a second complete neural network model and a second simplified neural network model) for the human presence detection task.
[0070] In box 252, LLM path 250 includes running an LLM model (e.g., performing LLM tasks on the LLM model); and in box 254, a request is sent to the NPU scheduler. The NPU scheduler can evaluate whether a simplified neural network model should be used. For example, the request may be associated with a high-priority task that can tolerate only a limited amount of latency (e.g., comparing Word document autocomplete, which should occur in real time, with a ChatGPT response that can tolerate some latency). In box 256, a determination is made (e.g., using workload management logic) as to whether the latency requirement can be met. And in box 268, if the latency requirement can be met, a first full neural network is employed; and in box 260, a first simplified neural network model is employed.
[0071] In box 272, the human presence detection path includes running human presence detection; then in box 274, it is determined whether the device power is low; in box 276, if the device power is determined to be low, a second complete neural network model is used; and in box 278, if the device power is determined to be not low, a second simplified neural network model is used. In box 280, the output neural network model from LLM path 250 and human presence detection path 270 is executed at the NPU.
[0072] Examples and references have been provided. Figure 1A , Figure 1B and Figure 2 This describes various aspects of the technical solution. Figure 1A This is a block diagram of an exemplary technical solution environment, based on the reference shown. Figure 6 and Figure 7 The described example environment is for implementing embodiments of this technical solution. Generally, this technical solution environment includes a technical solution system suitable for providing an exemplary artificial intelligence system 100A, wherein the methods of this disclosure can be employed. Among other engines, managers, generators, selectors, or components (collectively referred to herein as "components") not shown, the technical solution environment of the artificial intelligence system 100A corresponds to... Figure 1A and Figure 1B . Figure 2 Illustration A illustrates offline model training associated with AI system 100A and online optimization associated with AI client 130 (e.g., edge device), including two example use cases. Example Method
[0073] refer to Figure 3 , Figure 4 and Figure 5 The document provides flowcharts illustrating methods for using a workload management engine to provide workload management in an artificial intelligence system. These methods can be executed using the artificial intelligence system described herein. In embodiments, one or more computer storage media having computer-executable or computer-usable instructions embodied thereon, which, when executed by one or more processors, can cause one or more processors to perform methods (e.g., computer-implemented methods) within the artificial intelligence system (e.g., a computerized system or computing system).
[0074] Go to Figure 3The document provides a flowchart illustrating a method 300 for providing workload management using a workload management engine in an artificial intelligence system. In box 302, the workload management engine identifies multiple states of workload management factors. Workload management factors are predefined operational factors that support performance optimization for managing workload processing units. In box 304, the workload management engine identifies the task associated with the workload processing unit and the complete neural network model. In box 306, based on the task, the multiple states of the workload management factors, and the complete neural network model, the workload management engine selects a simplified neural network model from multiple neural network models. In box 308, the workload management engine initiates the execution of the task using the simplified neural network model. In box 310, the workload management engine determines whether the simplified neural network model meets a predefined performance threshold. In box 312, based on the determination that the simplified neural network model meets the predefined performance threshold, the workload management engine maintains the execution of the task using the simplified neural network model; or based on the determination that the simplified neural network model does not meet the predefined performance threshold, the workload management engine selects another neural network model.
[0075] Go to Figure 4 The document provides a flowchart illustrating a method 400 for providing workload management using a workload management engine in an artificial intelligence system. In box 402, the workload management engine trains multiple neural network models, including full neural network models and simplified neural network models. In box 404, the workload management engine associates each of the multiple neural network models with corresponding workload management logic and workload management factors. In box 406, the workload management engine identifies the deployment of multiple neural network models to support workload management, which includes dynamically switching between multiple neural network models based on workload management logic and workload management factors.
[0076] Go to Figure 5 The diagram provides a flowchart illustrating a method 500 for providing workload management using a workload management engine in an artificial intelligence system. In box 502, the workload management engine identifies multiple states of workload management factors. Workload management factors are predefined operational factors that support performance optimization of the workload processing unit. In box 504, the workload management engine identifies a task associated with the workload processing unit. In box 506, based on the task and the multiple states of the workload management factors, the workload management engine selects a neural network model from multiple neural network models. In box 508, the workload management engine causes the task to be executed using the selected neural network model. Other embodiments
[0077] In some embodiments, a system, such as the computerized system described in any of the above embodiments, includes at least one computer processor and a computer storage medium storing computer-usable instructions, which, when used by the at least one computer processor, cause the system to perform operations. The operations include multiple states identifying workload management factors. Workload management factors are predefined operational factors that support performance optimization of a workload processing unit. The operations also include identifying tasks associated with the workload processing unit and full neural network models from a plurality of neural network models. The operations include selecting a simplified neural network model from the plurality of neural network models based on the task, the multiple states of the workload management factors, and the full neural network model. The operations also include inducing the execution of a task using the simplified neural network model. The operations further include determining whether the simplified neural network model meets a predefined performance threshold. And, the operations include maintaining the execution of the task utilizing the simplified neural network model based on determining that the simplified neural network meets the predefined performance threshold; or selecting another neural network model from the plurality of neural network models based on determining that the simplified neural network does not meet the predefined performance threshold.
[0078] In any combination of embodiments of the above systems, the workload manager is associated with the workload processing unit to support dynamic switching between multiple neural network models based on workload management factors and workload management logic.
[0079] In any combination of embodiments of the above system, the workload management logic is associated with a first priority type associated with a task and a second priority type associated with multiple neural network models.
[0080] In any combination of embodiments of the above system, selecting a neural network model from multiple neural network models includes: selecting another simplified neural network model or reverting to the full neural network model.
[0081] In any combination of embodiments of the above system, the operation further includes identifying a second complete neural network model for optimization by the workload processing unit, wherein the workload processing unit supports both the complete neural network model and the second complete neural network model, the complete neural network model having a higher optimization priority than the second complete neural network model.
[0082] In any combination of embodiments of the above system, the plurality of neural network models include a first complete neural network model associated with at least two corresponding simplified neural network models, and a second complete neural network model associated with at least two corresponding simplified neural network models.
[0083] In any combination of embodiments of the above systems, the workload processing unit is associated with one of the following: an edge device or a cloud computing application.
[0084] In any combination of embodiments of the above system, the workload processing unit corresponds to one of the following: a neural processing unit (NPU), a graphics processing unit (GPU), or a tensor processing unit (TPU).
[0085] In any combination of embodiments of the above system, the operation further includes training multiple neural network models, including full neural network models and simplified neural network models. Each of the multiple neural network models is associated with a corresponding workload management logic and workload management factors. The multiple neural network models are deployed to support workload management, which includes dynamically switching between the multiple neural network models based on the corresponding workload management logic and workload management factors.
[0086] In any combination of embodiments of the above system, training multiple neural network models includes training a full neural network model and training multiple simplified neural network models, wherein the simplified neural network models are based on one of the following: quantization, pruning, or network architecture selection (NAS).
[0087] In some embodiments, one or more computer storage media are embodied thereon with computer-executable instructions that, when executed by a computing system having a processor and memory, cause the processor to perform operations. The operations include training multiple neural network models, including full neural network models and simplified neural network models. The operations also include associating each of the multiple neural network models with corresponding workload management logic and workload management factors, wherein the workload management factors are predefined operational factors that support performance optimization of management workload processing units. Furthermore, the operations include deploying the multiple neural network models to support workload management, which includes dynamically switching between the multiple neural network models based on corresponding workload management logic and workload management factors.
[0088] In any combination of the embodiments of the above media, training multiple neural network models includes training a full neural network model and training multiple simplified neural network models, wherein the simplified neural network models are based on one of the following: quantization, pruning, or network architecture selection (NAS).
[0089] In any combination of the embodiments of the above-described medium, the plurality of neural network models include a first complete neural network model associated with at least two corresponding simplified neural network models, and a second complete neural network model associated with at least two corresponding simplified neural network models.
[0090] In any combination of the embodiments of the above-described medium, each of the at least two corresponding simplified neural network models of the first complete neural network model is associated with a different simplification strategy.
[0091] In any combination of embodiments of the above-described medium, the operation includes: identifying multiple states of workload management factors; identifying tasks associated with workload processing units; selecting a neural network model from multiple neural network models based on the tasks and multiple states of workload management factors; and causing the task to be executed using the selected neural network model.
[0092] In some embodiments, a computer-implemented method is provided. The method includes identifying multiple states of workload management factors, wherein the workload management factors are predefined operational factors that support performance optimization of a workload processing unit. The method also includes identifying a task associated with the workload processing unit. The method further includes selecting a neural network model from multiple neural network models based on the task and the multiple states of the workload management factors. And, the method includes causing the task to be executed using the selected neural network model.
[0093] In any combination of embodiments of the above methods, the workload manager is associated with the workload processing unit to support dynamic switching between multiple neural network models based on workload management factors and workload management logic.
[0094] In any combination of embodiments of the above methods, the method further includes determining whether the neural network model meets a predefined performance threshold; and maintaining the execution of the task utilizing the simplified neural network model based on the determination that the neural network model meets the predefined performance threshold.
[0095] In any combination of embodiments of the above methods, the method includes determining whether a neural network model meets a predefined performance threshold; and selecting another neural network model from a plurality of neural network models based on the determination that the neural network model does not meet the predefined performance threshold.
[0096] In any combination of embodiments of the above methods, the method includes a simplified neural network; and selecting another neural network model from a plurality of neural network models includes selecting another simplified neural network model. Technological improvements
[0097] Embodiments of this technical solution have been described with reference to several inventive features (e.g., operations, systems, engines, and components) associated with artificial intelligence systems. The described inventive features include: operations, interfaces, data structures, and the arrangement of computing resources associated with the functionality described herein in relation to the workload management engine. The functionality of embodiments of this technical solution is also described through implementation methods and anecdotal examples—to demonstrate operations (e.g., dynamically switching between different neural network models for optimized performance). The workload management engine is a solution to specific problems in artificial intelligence technologies (e.g., the limited flexibility of NPU / GPU / TPU capacity to adapt to diverse neural network architectures and workloads, hindering their ability to effectively handle a wide range of applications, causing suboptimal inference speeds, and thus causing latency in real-time processing tasks). The adaptive policy framework improves the computational operations associated with providing workload management using the workload management engine of an artificial intelligence system. Overall, these improvements result in optimized computation, memory, and increased flexibility in the artificial intelligence system compared to previous conventional artificial intelligence system operations performed for similar functionality. Additional support for specific implementations Example Distributed Computing System Environment
[0098] Now for reference Figure 6 , Figure 6 The illustration depicts an example distributed computing environment 600 in which an implementation of this disclosure may be employed. Specifically, Figure 6 A high-level architecture of an example cloud computing platform 610 that can host a technology solution environment or a portion thereof (e.g., a data trustee environment) is shown. It should be understood that such and other arrangements described herein are illustrated by way of example only. For example, as stated above, many of the elements described herein can be implemented as discrete or distributed components, or combined with other components, and in any suitable combination and location. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, and functional groupings) may be used in addition to or instead of those shown.
[0099] The data center can support a distributed computing environment 600, including a cloud computing platform 610, racks 620, and nodes 630 (e.g., computing devices, processing units, or blades) within the racks 620. The technical solution environment can be implemented using the cloud computing platform 610, which runs cloud services across different data centers and geographical regions. The cloud computing platform 610 can implement a fabrication controller 640 component for provisioning and managing the allocation, deployment, upgrades, and management of cloud services. Typically, the cloud computing platform 610 stores data or runs service applications in a distributed manner. The cloud computing platform 610 in the data center can be configured to host and support the operation of endpoints for specific service applications. The cloud computing platform 610 can be a public cloud, a private cloud, or a dedicated cloud.
[0100] Node 630 may provide a host 650 (e.g., an operating system or runtime environment) on which a defined software stack runs. Node 630 may also be configured to perform specialized functions (e.g., compute nodes or storage nodes) within the cloud computing platform 610. Node 630 is assigned to run one or more parts of a tenant's service application. A tenant may refer to a customer that utilizes the resources of the cloud computing platform 610. The service application components of the cloud computing platform 610 that support a particular tenant may be referred to as multi-tenant infrastructure or lease. The terms “service application,” “application,” or “service” are used interchangeably herein and broadly refer to any software or part of software that runs on top of or accesses storage and compute equipment locations within a data center.
[0101] When node 630 supports more than one individual service application, node 630 can be partitioned into virtual machines (e.g., virtual machine 652 and virtual machine 654). Physical machines can also run individual service applications simultaneously. Virtual machines or physical machines can be configured as personalized computing environments supported by resources 660 (e.g., hardware and software resources) in the cloud computing platform 610. It is envisioned that resources can be configured for specific service applications. Furthermore, each service application can be divided into functional parts, so that each functional part can run on a separate virtual machine. In the cloud computing platform 610, multiple servers can be used to run service applications and perform data storage operations in a cluster. Specifically, servers can perform data operations independently but are exposed as a single device called a cluster. Each server in the cluster can be implemented as a node.
[0102] Client device 680 can be linked to service applications in cloud computing platform 610. Client device 680 can be any type of computing device, which can correspond to the reference... Figure 7 The described computing device 700, for example, client device 680, can be configured to issue commands to cloud computing platform 610. In embodiments, client device 680 can communicate with service applications via Virtual Internet Protocol (IP) and load balancers or other means of directing communication requests to designated endpoints in cloud computing platform 610. Components of cloud computing platform 610 can communicate with each other via a network (not shown), which may include, but is not limited to, one or more local area networks (LANs) and / or wide area networks (WANs). Example computing environment
[0103] Having briefly described an overview of embodiments of this technical solution, the following describes an example operating environment in which embodiments of this technical solution may be implemented, in order to provide a general context for various aspects of this technical solution. First, refer to... Figure 7An example operating environment for implementing embodiments of this technical solution is shown and is generally designated as computing device 700. Computing device 700 is merely an example of a suitable computing environment and is not intended to impose any limitation on the scope of the purpose or functionality of this technical solution. Nor should computing device 700 be construed as having any dependency or requirement on any of the components or combinations illustrated.
[0104] This technical solution can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules, which are executed by a computer or other machine (such as a personal data assistant or other handheld device). Typically, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs a specific task or implements a specific abstract data type. This technical solution can be implemented in various system configurations, including handheld devices, consumer electronics, general-purpose computers, and more specialized computing devices. This technical solution can also be implemented in distributed computing environments, where tasks are performed by remote processing devices linked through a communication network.
[0105] refer to Figure 7 The computing device 700 includes a bus 710 that is directly or indirectly coupled to the following devices: a memory 712, one or more processors 714, one or more presentation components 716, an input / output port 718, an input / output component 720, and an illustrative power supply 722. The bus 710 can represent one or more buses (such as an address bus, a data bus, or a combination thereof). For clarity of concept, Figure 7 The various boxes are shown using lines, and other arrangements of the described components and / or component functions are also envisioned. For example, presentation components (such as display devices) can be considered as I / O components. Furthermore, the processor has memory. It is recognized that this is the nature of the art, and is reiterated... Figure 7 The figures are merely illustrative examples of computing devices, which may be used in conjunction with one or more embodiments of this technical solution. No distinction is made between categories such as "workstation," "server," "laptop," and "handheld device," as all of these are considered within the scope of this technical solution. Figure 7 Within the scope and reference to "computing device".
[0106] Computing device 700 typically includes a variety of computer-readable media. Computer-readable media can be any available medium that can be accessed by computing device 700, and includes volatile and non-volatile media, removable and non-removable media. By way of example, and not limitation, computer-readable media can include computer storage media and communication media.
[0107] Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disc storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by the computing device 700. Computer storage media itself does not include the signal itself.
[0108] Communication media typically embody computer-readable instructions, data structures, program modules, or other data in the form of modulated data signals (such as carrier waves or other transmission mechanisms), and include any information delivery medium. The term "modulated data signal" refers to a signal whose characteristics are set or altered in a manner that encodes information within the signal. By way of example, and not limitation, communication media include wired media (such as wired networks or direct wired connections) and wireless media, such as acoustic, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.
[0109] Memory 712 includes computer storage media in the form of volatile and / or non-volatile memory. The memory can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. Computing device 700 includes one or more processors that read data from various entities such as memory 712 or I / O components 720. Multiple presentation components 716 present data indications to a user or other device. Exemplary presentation components include display devices, speakers, printing components, vibration components, etc.
[0110] I / O port 718 allows computing device 700 to be logically coupled to other devices including I / O components 720, some of which may be built-in. Illustrative components include microphones, joysticks, gamepads, satellite antennas, scanners, printers, wireless devices, etc. Additional structural and functional features of embodiments of this technical solution
[0111] Various components utilized herein have been identified, and it should be understood that any number of components and arrangements can be employed to achieve the desired functionality within the scope of this disclosure. For example, components in the embodiments depicted in the figures are shown using lines for clarity of concept. Other arrangements of these and other components can also be implemented. For example, although some components are depicted as single components, many of the elements described herein can be implemented as discrete or distributed components or combined with other components, and in any suitable combination and location. Some elements can be omitted entirely. Furthermore, the various functions described herein as being performed by one or more entities can be performed by hardware, firmware, and / or software, as described below. For example, various functions can be performed by a processor executing instructions stored in memory. Therefore, other arrangements and elements (e.g., machines, interfaces, functions, sequences, and functional groups) can be used in addition to or instead of those shown.
[0112] The embodiments described in the following paragraphs can be combined with one or more of the specifically described alternatives. Specifically, the claimed embodiments may include references to more than one other embodiment in an alternative manner. The claimed embodiments may specify further limitations on the claimed subject matter.
[0113] This document specifically describes embodiments of the technical solutions to meet legal requirements. However, this description itself is not intended to limit the scope of this patent. Rather, the inventors have envisioned that the claimed subject matter may also be embodied in other ways, in combination with other current or future technologies, including different steps or combinations of steps similar to those described in this document. Furthermore, although the terms “step” and / or “box” may be used herein to mean different elements of the method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless and except where the order of the steps is explicitly described.
[0114] For the purposes of this disclosure, the word "including" has the same broad meaning as the word "comprising," and the word "access" includes "receiving," "quoting," or "retrieval." Furthermore, the word "communication" has the same broad meaning as the words "receiving" or "transmitting," facilitated by a software or hardware-based bus, receiver, or transmitter using the communication medium described herein. Moreover, unless otherwise indicated, words such as "a" and "a" include both plural and singular forms. Thus, for example, the constraint "one feature" is satisfied when one or more features are present. Furthermore, the term "or" includes connectivity, separation, and both (a or b therefore includes a or b, and a and b).
[0115] For the purposes of the detailed discussion above, embodiments of the present technical solution have been described with reference to a distributed computing environment; however, the distributed computing environment depicted herein is merely exemplary. Components can be configured to perform novel aspects of the embodiments, wherein the term "configured for" can mean "programmed to" perform a specific task or implement a specific abstract data type using code. Furthermore, while embodiments of the present technical solution may generally refer to the technical solution environment and schematic diagrams described herein, it should be understood that the described technology can be extended to other implementation contexts.
[0116] Embodiments of the present invention have been described with reference to specific examples, which are intended in all respects to be illustrative and not restrictive. Alternative embodiments will be apparent to those skilled in the art without departing from the scope of the present invention.
[0117] As can be seen from the foregoing, this technical solution is very suitable for achieving all the goals and objectives mentioned above, as well as other obvious advantages inherent in the structure itself.
[0118] It will be understood that certain features and sub-combinations are practical and can be adopted without reference to other features or sub-combinations. This is conceivable and within the scope of the claims.
Claims
1. A computerized system, comprising: One or more computer processors; as well as A computer memory storing computer-usable instructions, which, when used by the one or more computer processors, cause the one or more computer processors to perform operations, the operations including: Identify (302) multiple states of workload management factors, wherein the workload management factors are predefined operational factors that support performance optimization of the management workload processing unit; Identifier (304) is the task associated with the workload processing unit and the complete neural network model from multiple neural network models; Based on the task, the multiple states of the workload management factors, and the complete neural network model, a simplified neural network model is selected from the multiple neural network models (306); This causes (308) the execution of the task using the simplified neural network model; Determine whether the simplified neural network model described in (310) meets a predefined performance threshold; and Based on determining that the simplified neural network meets the predefined performance threshold, maintain (312) the execution of the task using the simplified neural network model; or Based on the determination that the simplified neural network does not meet the predefined performance threshold, another neural network model is selected from the plurality of neural network models (312).
2. The system of claim 1, wherein the workload manager is associated with the workload processing unit to support dynamic switching between the plurality of neural network models based on the workload management factors and workload management logic.
3. The system of claim 1, wherein the workload management logic is associated with a first priority type of the task and a second priority type of the plurality of neural network models.
4. The system of claim 1, wherein selecting the neural network model from the plurality of neural network models comprises: Choose another simplified neural network model or revert to the complete neural network model.
5. The system according to claim 1, further comprising: A second complete neural network model is identified for optimization of the workload processing unit, wherein the workload processing unit supports both the complete neural network model and the second complete neural network model, and the complete neural network model has a higher optimization priority than the second complete neural network model.
6. The system according to claim 1, wherein the operation includes: The plurality of neural network models are trained, including the complete neural network model and the simplified neural network model; Associate each of the plurality of neural network models with the corresponding workload management logic and workload management factors; as well as The multiple neural network models are deployed to support workload management, which includes dynamically switching between the multiple neural network models based on corresponding workload management logic and workload management factors.
7. The system of claim 1, wherein training the plurality of neural network models includes training the full neural network model and training a plurality of simplified neural network models, wherein the simplification of the simplified neural network model is based on one of the following: quantization, pruning, or network architecture selection (NAS).
8. One or more computer storage media having computer-executable instructions embodied thereon, which, when executed by a computing system having a processor and a memory, cause the processor to perform operations, the operations including: Train (402) multiple neural network models, including full neural network models and simplified neural network models; Each of the plurality of neural network models is associated with a corresponding workload management logic and workload management factors (404), wherein the workload management factors are predefined operational factors that support performance optimization of the workload processing unit; and Deploy (406) the plurality of neural network models to support workload management, the workload management including dynamically switching between the plurality of neural network models based on the corresponding workload management logic and workload management factors of the plurality of neural network models.
9. The medium of claim 8, wherein training the plurality of neural network models includes training the full neural network model and training a plurality of simplified neural network models, wherein the simplification of the simplified neural network model is based on one of the following: quantization, pruning, or network architecture selection (NAS).
10. The medium of claim 8, wherein the plurality of neural network models includes a first complete neural network model associated with at least two corresponding simplified neural network models, and a second complete neural network model associated with at least two corresponding simplified neural network models.
11. The medium of claim 10, wherein each of the at least two corresponding simplified neural network models of the first complete neural network model is associated with a different simplification strategy.
12. The medium according to claim 8, wherein the operation further comprises: Identify multiple states of workload management factors; Identify the tasks associated with the workload processing unit; Based on the multiple states of the task and the workload management factors, a neural network model is selected from multiple neural network models; and The task is then performed using the selected neural network model.
13. A computer-implemented method, the method comprising: Identify (502) multiple states of workload management factors, wherein the workload management factors are predefined operational factors that support performance optimization of the management workload processing unit; Identifier (504) is the task associated with the workload processing unit; Based on the multiple states of the task and the workload management factors, a neural network model (506) is selected from multiple neural network models; and The task described in (508) is performed using the selected neural network model.
14. The method of claim 13, further comprising: Determine whether the neural network model meets the predefined performance threshold; as well as Based on determining that the neural network model meets the predefined performance threshold, the execution of the task utilizing the simplified neural network model is maintained.
15. The method of claim 13, further comprising: Determine whether the neural network model meets the predefined performance threshold; as well as Based on the determination that the neural network model does not meet the predefined performance threshold, another neural network model is selected from the plurality of neural network models.