Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

50 results about "Service level objective" patented technology

A service-level objective (SLO) is a key element of a service-level agreement (SLA) between a service provider and a customer. SLOs are agreed upon as a means of measuring the performance of the Service Provider and are outlined as a way of avoiding disputes between the two parties based on misunderstanding.

Machine learning techniques for predictive anomaly detection

As described herein, various embodiments of the present invention improve computational efficiency of performing reliable persistent monitoring of a computer system. Persistent monitoring of the operations of a computer system has various operational reliability benefits for various computer systems and is an important part of service level objectives for highly maintenance-critical computer systems. However, doing the noted persistent monitoring operations in a reliable manner requires access to large amounts of labeled training data that are not always available for more customized computer systems with unique operational / behaviorial pattern signatures. In response, various embodiments of the present invention address the noted challenges by generating training data for an anomalous operational state detection machine learning model whose training data may be generated using a ground-truth validation criterion that is defined based at least in part on an inferred outlier score for a decomposed residual component of a given operational monitoring timeseries trend.
Owner:LIBERTY MUTUAL INSURANCE CO

Server-free workflow dynamic resource configuration system based on resource decoupling

The invention discloses a server-free workflow dynamic resource configuration system based on resource decoupling, and relates to the field of cloud computing. Comprising a workflow sampling scheduling module, a configuration decoupling search module and a dynamic configuration module. The dynamic configuration module shunts the workload in combination with the input characteristic data, the scheduling workflow sampling scheduling module obtains the decoupled optimal resource configuration, and a mapping relation between the input characteristic and the decoupled resource configuration is established by training a random forest model; the workflow sampling scheduling module analyzes a workflow structure through input, generates a weighted directed acyclic graph and identifies a key path, preferentially searches for optimal configuration for the key path, and then iteratively optimizes a sub-path on the premise of not violating the consistency of the key path; the configuration decoupling search module allocates decoupled CPU / memory resources based on a critical path priority policy. According to the method, the cost is minimized while the service level target is met through an automatic searching method for decoupling the memory and the CPU.
Owner:SHANGHAI JIAOTONG UNIV +1

System and method for coordinated resource scaling in microservice-based and serverless applications

A computer-implemented method for trace-driven call-graph-aware proactive coordinated autoscaling of component microservices in an application includes generating performance-resource elasticity models of endpoints of the component microservices of the application. Workload levels of the endpoint of the component microservices is predicted based on user traffic observed at a front end service. A trace-level performance of the application is predicted for different microservice replica scaling based on the performance-resource elasticity models at end points, the ends points on the trace call graph and the predicted workload levels. A microservice replica scaling is recommended for each of the component microservices to meet predefined trace-level user service level objectives.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Trace-driven call dependency-set aware proactive coordinated distributed auto-scaling for resource management

A computer-implemented method for trace-driven dependency-set-aware proactive coordinated autoscaling of component microservices in an application includes generating performance-resource elasticity models at a trace-level for traces of the application using dependency set of microservices for each trace. The method predicts workload levels of each of the traces, and also predicts a trace-level performance of the application for different microservice replica scaling based on the dependency set of microservices for each trace, performance-resource elasticity models and the predicted workload levels. The method uses distributed computing to recommend a microservice replica scaling for each of the component microservices to meet one or more predefined trace-level user service level objectives.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Model input request scheduling method and device, storage medium and program product

Embodiments of the invention provide a model input request scheduling method and device, a storage medium and a program product. The method comprises the steps of receiving an input request of a preset language model; determining prediction time consumption of the input request in a pre-filling stage in a preset language model processing process; determining the latest execution time of the pre-filling stage of the input request according to the receiving time of the input request, the predicted consumed time of the pre-filling stage and the target first lexical element delay; and according to the latest execution time of the pre-filling stage, sending the input request to a service node of a preset language model for processing. According to the embodiment of the invention, the latest execution time of the pre-filling stage is determined based on the predicted time consumption of the pre-filling stage of the input request to schedule the input request, and the system throughput and the calculation load balance are optimized while the service level target is met, so that the resource utilization efficiency and the service quality guarantee capability of the service node of the preset language model can be improved.
Owner:BYTEDANCE TECHNOLOGY CO LTD +1

Intent-based orchestration in heterogenous compute platforms

Various systems and methods for implementing intent-based orchestration in heterogenous compute platforms are described herein. An orchestration system is configured to: receive, at the orchestration system, a workload request for a workload, the workload request including an intent-based service level objective (SLO); generate rules for resource allocation based on the workload request; generate a deployment plan using the rules for resource allocation and the intent-based SLO; deploy the workload using the deployment plan; monitor performance of the workload using real-time telemetry; and modify the rules for resource allocation and the deployment plan based on the real-time telemetry.
Owner:INTEL CORP

Multi-target execution plan optimization method, system and product for intelligent application

The invention discloses a multi-target execution plan optimization method, system and product oriented to intelligent application, and belongs to the technical field of intelligent application. A logic application diagram and context information are obtained, and the context information comprises a service level target and a load feature; the method comprises the steps that a logic application diagram, a service level target, load characteristics and a pre-constructed operator performance model are integrated to obtain a constraint optimization problem, a multi-target constraint optimization solver is used for solving the constraint optimization problem, an optimal physical execution plan is generated, and an execution engine is used for executing the optimal physical execution plan. According to the method, the multi-dimensional service target can be understood and adjusted in real time by obtaining the service level target set by the user, the hysteresis is avoided by generating the optimal physical execution plan before plan execution, the self-adaptive capability and the quality of the generated result are improved by using the pre-constructed operator performance model and the multi-target constraint optimization solver, and the method has the advantages of being high in practicability and the like. And the scientificity, the globality and the reproducibility of a scheduling decision are ensured.
Owner:北京文聿科技有限公司

Improvements in data product development

This disclosure relates to methods, devices, and computer-readable media for use in developing data products. One such method comprises receiving an existing build of the data product, identifying a data product version associated with the existing build, receiving a user-specified modification for the data product, in response to a user input, automatically determining a compatibility result for the modification with the identified data product version, based on the existing build of the data product, and in response to the determined compatibility result being a negative compatibility result, triggering a failure event in relation to the identified data product version. The user specified modification may comprise one or more of: modification of a schema or addition modification or deletion of one or more tables, columns, service level objectives, service level indicators, constraints. Determination of the compatibility result may comprise determining a negative compatibility result by identifying that a table, column, service level objective, constraint or other object present in the existing build is removed by the modification.
Owner:DATAOPS SOFTWARE LTD

A method and system for hybrid deployment of deep learning tasks

The present invention relates to the technical field of computer deep learning, and particularly relates to a method and system for hybrid deployment of deep learning tasks. The method includes: S1. A user submits a task through the native interface of Kubernetes; S2. A task scheduling optimizer analyzes the task according to resource requirements and service levels and allocates it to a suitable node; S3. An SLO manager tracks the real-time data of various performance indicators of the system and compares the real-time data with the pre-set service level objectives; S4. The Koordlet component performs specific resource allocation operations on the target node; S5. A traffic security monitor monitors the data flow between nodes in real time. By analyzing the periodic laws of resource usage of deep learning tasks, the present invention proposes a hybrid deployment strategy for different types of resources, realizes dynamic resource sharing between online and offline tasks, and at the same time ensures that the system can still operate stably and efficiently under high load conditions.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Method and apparatus for using microslices to control network performance of an enterprise wireless communication network

A method and apparatus for utilizing microslices in an enterprise wireless communication network to manage and control network performance. The microslice instances are monitored during network operation, for example at communication nodes, functional blocks, and end-to-end to provide Key Performance Indicators (KPIs). These KPI are compared with performance objectives, which may be Service Level Objectives (SLOs). If the performance objectives are not met by the KPIs, then one or more of the microslice instances may be dynamically adjusted until the performance objectives are sufficiently met. Alternatively, the lower priority microslice instances may be dropped (i.e., terminated) until the performance objectives are sufficiently met.
Owner:CELONA INC

Power oversubscription in LLM cloud providers

Systems and methods for implementing power oversubscription in graphic processing unit (GPU) servers are provided. An increase to a quantity of servers allocated to a group of GPU servers in an inference cluster is applied. Based on the power consumption of the group of GPU servers exceeding a first threshold, a frequency of low priority inference workloads is capped, and based on the power consumption of the group of GPU servers exceeding a second threshold, the frequency of the low priority inference workloads are capped and a frequency of high priority inference workloads are capped, enabling an increase in allocated server capacity in the existing inference clusters while maintaining service level objectives (SLOs).
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Service level objective-based regulator

Techniques are disclosed that enable a self-regulating process to meet a service level objective (SLO). In some embodiments, a self-regulating process is a background process comprising a regulator that receives background job requests and historical information related to the background process for evaluation to determine actions (e.g., speed up, slow down, or maintain the same speed), enabling the background process to adjust its pace gradually and smoothly even when encountering unexpected big changes in load.
Owner:ORACLE INT CORP

Service level objective (SLO) based continuous integration / continuous development (CICD) framework for chaos engineering

A Service Level Objective (SLO)-based Continuous Integration / Continuous Development (CICD) framework is disclosed, which enables SLO-based Canary deployments, SLO-based chaos engineering, and / or SLO-based performance testing. The SLO-based CICD framework enhances and simplifies CICD pipelines and brings the pipelines in line with Site Reliability Engineering (SRE) best practices. By relying on SLOs rather than SLIs, the SLO-based CICD framework avoids the use of threshold-based alerts, which can be noisy and error prone, and the thresholds associated with threshold-based alerts can change over time due to changes in production volume. The SLO-based CICD framework uses multi-window burn rates when evaluating SLOs, which removes the reliance on a static, predefined thresholds of SLI threshold-based alerting, and removes sensitivity of such evaluations to changes in production volume.
Owner:UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)

Network slicing including modeling, distribution, traffic engineering and maintenance

The techniques describe network slicing in computer networks. For example, a node receives a slice policy definition. The slice policy definition comprising a slice selector to identify packets belonging to one or more network slices, referred to as a “slice aggregate,” and one or more network resource requirements for the slice aggregate to meet one or more Service Level Objectives (SLOs). The node configures, based on the slice policy definition, a path for the slice aggregate that complies with the one or more network resource requirements. In response to receiving a packet, the node determines whether the packet is associated with the slice aggregate and, in response to determining that the packet is associated with the slice aggregate, forwards the packet along the path for the slice aggregate.
Owner:JUNIPER NETWORKS INC

Service level objective (SLO) based continuous integration / continuous development (CICD) framework for performance testing

A Service Level Objective (SLO)-based Continuous Integration / Continuous Development (CICD) framework is disclosed, which enables SLO-based Canary deployments, SLO-based chaos engineering, and / or SLO-based performance testing. The SLO-based CICD framework enhances and simplifies CICD pipelines and brings the pipelines in line with Site Reliability Engineering (SRE) best practices. By relying on SLOs rather than SLIs, the SLO-based CICD framework avoids the use of threshold-based alerts, which can be noisy and error prone, and the thresholds associated with threshold-based alerts can change over time due to changes in production volume. The SLO-based CICD framework uses multi-window burn rates when evaluating SLOs, which removes the reliance on a static, predefined thresholds of SLI threshold-based alerting, and removes sensitivity of such evaluations to changes in production volume.
Owner:UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)

Method and system for performing anomaly detection in a distributed multi-tiered computing environment

Techniques described herein relate to a method for managing a distributed multi-tiered computing (DMC) environment. The method includes obtaining, by a local controller associated with a DMC domain, service level objective (SLO) metrics; applying the SLO metrics to a predictive anomaly detection transformer to perform anomaly detection; making a first determination that an anomaly is detected; in response to the first determination: attempting basic remediation to resolve the anomaly; making a second determination that the basic remediation is unsuccessful; in response to the second determination: making a third determination that the anomaly is associated with a silent failure; and in response to the third determination: performing service impairment isolation to obtain a collection of services correlated to the anomaly; and performing root cause analysis to identify causal services.
Owner:DELL PROD LP

Deep reinforcement learning cloud platform resource adjustment method based on prediction of user workload

The invention relates to the technical field of cloud computing, and provides a deep reinforcement learning cloud platform resource adjustment method based on user workload prediction. The method aims at solving the problem of adjustment hysteresis caused by the fact that load changes cannot be predicted in a traditional response type resource adjustment strategy, so that the rule violation rate of a service level target (SLO) is effectively reduced, and the resource utilization rate is improved. According to the main scheme, the method comprises the following steps: firstly, predicting a future workload trend through a long short-term memory (LSTM) network; combining a prediction result with a current system state (such as response delay and copy number) to serve as input of a deep reinforcement learning agent; the intelligent agent outputs a resource adjustment action through a strategy network; and finally, optimizing the decision model according to a composite revenue function integrating the service quality, the resource cost and the adjustment stability. The method can be widely applied to cloud native application under a micro-service architecture, and automatic and prospective elastic resource scaling management is realized.
Owner:INFORMATION & TELECOMM COMPANY SICHUAN ELECTRIC POWER

Service level breach prediction

A method is provided, comprising: obtaining a plurality of measured response times for a given storage group, each of the measured response times corresponding to a different instance of a same time window; classifying the plurality of measured response times with a machine learning model to obtain a first predicted response time and a second predicted response time, the first predicted response time corresponding to a first future instance of the time window, and the second predicted response time corresponding to a second future instance of the time window; detecting whether the first predicted response time and the second predicted response time satisfies a service level objective for the given storage group; and generating a notification when at least one of the first predicted response time and the second predicted response time fails to satisfy the service level objective.
Owner:DELL PROD LP

Computing power scheduling method and device, communication equipment, medium and product

The invention provides a computing power scheduling method and device, communication equipment, a medium and a product, and the method comprises the steps: carrying out the data preprocessing and topology awareness of a computing task, and obtaining a computing graph of the computing task; performing sub-graph division on the calculation graph to obtain calculation sub-graphs, and packaging the calculation sub-graphs into schedulable units; the calculation benefit of the schedulable unit is obtained by determining the execution delay, the power consumption and the heterogeneous coefficient of the calculation nodes which distribute the schedulable unit to different areas; constructing a task scheduling model meeting a service level target according to a maximum and minimum algorithm and the calculation benefit; the optimal allocation of the schedulable units on the computing nodes is determined through the task scheduling model, the schedulable units are scheduled to the computing nodes in parallel according to the optimal allocation, large computing tasks can be scheduled to a proper area according to delay and power consumption, the problem of hardware resource heterogeneity in the scheduling process can be solved, and the scheduling efficiency is improved. And the total power consumption of the whole task is reduced through refined scheduling.
Owner:ZUNYI BRANCH OF CHINA MOBILE GRP GUIZHOU COMPANY +1

Data arrangement method and device oriented to job intention

The invention provides a job intention-oriented data arrangement method and device, and the method comprises the steps: compiling a job intention, and obtaining a multi-level arrangement strategy; the multi-level arrangement strategy comprises at least two of an application layer arrangement strategy, a transmission layer arrangement strategy and a system storage layer arrangement strategy; and executing the data transmission operation based on the multi-level arrangement strategy. According to the method provided by the invention, the operation intention containing the global demand and the constraint is acquired and compiled into the multi-level arrangement strategy covering multiple levels of application, transmission, storage and the like, and finally the data transmission operation is executed based on the strategy, so that the full-link resource configuration can be flexibly adjusted according to the service intention; the problem that downstream data receiving is blocked due to the fact that the network is not matched with the storage performance is effectively solved, and the comprehensive performance and the resource utilization rate of data transmission are remarkably improved while the service level target is guaranteed and complex constraints are met.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Service level target violation diagnosis method and device, equipment and medium

The invention provides a service level target violation diagnosis method, a service level target violation diagnosis device, service level target violation diagnosis equipment and a medium, and the method realizes the service level target violation diagnosis method in a data plane, and comprises the following steps: carrying out compression grouping on equipment on a network path based on path information, and collecting location-based diagnosis information in a probabilistic manner, the location-based diagnostic information includes state information of a specific point on the network path; probabilistically detecting packet loss port information of all devices on the network path based on the compression grouping result to obtain detail-based diagnosis information which refers to specific data required by diagnosis of the suspicious area; and performing native path coding according to the media access control code of each device in the network to determine a target device corresponding to a violation event contained in the location-based diagnostic information and / or the detail-based diagnostic information. The method does not depend on the control plane capability, alleviates the problem of high violation detection overhead, and has complete knowability.
Owner:TSINGHUA UNIVERSITY

Systems and methods for edge resource demand load scheduling

Managing the resource demand load for edge systems is significantly more complex than for other systems, such as cloud environments. Edge resource demand load scheduling systems and methods are disclosed that can ensure that edge systems operate smoothly and efficiently while balancing multiple scheduling objectives. Scheduling techniques disclosed herein may utilize heuristic rules for candidate edge system selection (e.g., utilizing ARMA / ARIMA averages and / or service level objectives) and modified best fit decreasing (mBFD) assignment / allocation techniques.
Owner:DELL PROD LP

System to convert abstract storage solution input data into a cluster topology for deployment and updates of a clustered file system

The technology described herein is directed towards deploying a node cluster, e.g., in a cloud environment, based on abstract user storage solution inputs / service level objective input data, including user-specified storage capacity and performance characteristics. An intelligent system converts the input data into cluster topology data (e.g., virtual machines, storage devices / volumes and other information) that can be used for initial cluster deployment or infrastructure updates to existing node clusters in multi-cloud environments. The input data can be directed to an initial deployment of a new node cluster or an update to an existing node cluster. Further, the system can take customer rules as input data, monitor cluster conditions corresponding to those rules, and update the cluster based on a condition in the rules being met, by taking an action associated with the rule's condition to automatically generate and handle a cluster update request.
Owner:DELL PROD LP

Real-time analysis and intelligent optimization system for software operation data based on machine learning

ActiveCN121302884BDynamically reflect fluctuation characteristicsCharacterizing average performanceBiological modelsDesign optimisation/simulationData packResource utilization
This invention provides a real-time analysis and intelligent optimization system for software operation data based on machine learning, comprising: a data acquisition module for acquiring real-time operational data of the software operation system, including request arrival logs, service processing logs, and resource usage information; a mechanism model construction module for establishing a queuing control mechanism model based on the real-time operational data and defining a set of operational operations within the mechanism model; an operational state estimation module for estimating the real-time operational data based on the mechanism model to obtain request arrival rate intervals, service capacity intervals, and burst traffic intervals, forming a set of uncertainty intervals; and a service level target constraint module for selecting tail performance indicators as service level targets based on the set of uncertainty intervals. This invention can continuously guarantee service level targets under dynamic fluctuations and uncertain environments, improving the stability and resource utilization of the software operation system.
Owner:南昌职业大学

Software operation data real-time analysis and intelligent optimization system based on machine learning

The invention provides a software operation data real-time analysis and intelligent optimization system based on machine learning, and the system comprises a data collection module which is used for obtaining real-time operation data of a software operation system, and the operation data comprises a request arrival log, a service processing log and resource occupation information; the mechanism model building module is used for building a queuing control mechanism model based on the real-time operation data and defining an operation set in the mechanism model; the operation state estimation module is used for estimating real-time operation data based on the mechanism model to obtain a request arrival rate interval, a service capability interval and a burst flow interval to form an uncertainty interval set; and selecting a tail performance index as a service grade target on the basis of the uncertainty interval set. According to the method, the service level target can be continuously guaranteed in a dynamic fluctuation and uncertain environment, and the stability and the resource utilization rate of a software operation system are improved.
Owner:南昌职业大学

Query admission control for online data systems based on response time objectives

Exemplary embodiments include tracking processing time metrics of multiple server queries. A processing time per query type is estimated using the tracked processing time metrics. A current queue wait time is estimated based on a number of queries currently in the queue and the estimated processing times of query types for each of the queries currently in the queue. Upon receiving a current server query from a client, the current query type is mapped to an estimated processing time determined using the tracked processing time metrics. An estimated response time is determined using the current queue wait time and the estimated processing time. The server query is rejected from being added to the queue in response to determining the estimated response time does not satisfy a service level objective and an error message is sent to the client indicating the rejection of the server query.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Service level objective based priority processing of control path operations

A VASA provider is configured to include a control channel that is able to provide differentiated service to control path operations. The VASA provider maps a first set of data path service level objectives (SLOs) available on a storage system to a second set of control path SLOs. The VASA provider maintains a set of control path operation queues, such that a separate control path operation queue is maintained for each of the control path SLOs. Virtual machines are assigned control path SLOs based on their data path SLOs. Control path operations from virtual machines are mapped using the control path SLOs to the control path operation queues. The VASA provider processes control operations from the set of control path operation queues in a differentiated manner, to thereby provide different quality of service levels to the control operations of the different virtual machines.
Owner:DELL PROD LP

Automated management of network resources based on customer quality of experience

PCT designated stage expiredWO2025151893A1TransmissionComputer networkQuality of experience
Described herein are systems and methods that automate management of network resources based on an objective function that determines compliance with service level objectives (SLOs). The disclosed technologies use measurements with fine granularity to estimate customer quality of experience (QoE) and to establish SLOs. The disclosed technologies score SLO performance using these fine-grained measurements, such as passive speed measurements. The disclosed technologies use these measurements to improve or optimize SLO performance across a network in an automated fashion to thereby improve customer QoE.
Owner:VIASAT INC

Methods, apparatus, and articles of manufacture to schedule algorithms based on characteristics of data

Systems, apparatus, articles of manufacture, and methods are disclosed to schedule algorithms based on characteristics of data. An example compute device includes circuitry to determine at least one characteristic of a data object to be processed and adjust metadata associated with the data object to indicate that the data object has the at least one characteristic. Additionally, the example compute device includes machine-readable instructions and at least one programmable circuit to be programmed by the machine-readable instructions to select at least one of two or more programmable circuits to process the data object based on (a) at least one service level objective associated with the data object, (b) the metadata associated with the data object, and (c) telemetry data associated with the two or more programmable circuits.
Owner:OPENCHIP & SOFTWARE TECHNOLOGIES SL

Quick and quality-aware cloud collaborative big language model reasoning method and system

The invention discloses a fast and quality-aware cloud collaborative large language model reasoning method and system (CEC-LM), and belongs to the field of computer networks and artificial intelligence, and the method comprises the steps: employing a task execution selector based on a lightweight model, evaluating the task complexity in real time according to a prompt word requested by a user, and carrying out the real-time evaluation of the task complexity according to the prompt word; the complex tasks are routed to a cloud LLM to ensure response quality, and meanwhile the simple tasks are distributed to a local small language model to be executed. And when sensing that the local SLM is high in load, dynamically unloading the complexity prediction model from the GPU to the CPU, and realizing zero-copy parameter transmission by utilizing a unified memory architecture. According to the method, a task monitoring and priority scheduling mechanism during operation is introduced, the task routing is adaptively adjusted by continuously monitoring the network condition and the system load, and the emergency task is preferentially processed, so that the achievement rate of the service level target is ensured, the reasoning quality, delay and cost are effectively balanced, and the system overhead is remarkably reduced.
Owner:WUHAN UNIV