Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

39 results about "Service level objective" patented technology

A service-level objective (SLO) is a key element of a service-level agreement (SLA) between a service provider and a customer. SLOs are agreed upon as a means of measuring the performance of the Service Provider and are outlined as a way of avoiding disputes between the two parties based on misunderstanding.

System and method for coordinated resource scaling in microservice-based and serverless applications

A computer-implemented method for trace-driven call-graph-aware proactive coordinated autoscaling of component microservices in an application includes generating performance-resource elasticity models of endpoints of the component microservices of the application. Workload levels of the endpoint of the component microservices is predicted based on user traffic observed at a front end service. A trace-level performance of the application is predicted for different microservice replica scaling based on the performance-resource elasticity models at end points, the ends points on the trace call graph and the predicted workload levels. A microservice replica scaling is recommended for each of the component microservices to meet predefined trace-level user service level objectives.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Trace-driven call dependency-set aware proactive coordinated distributed auto-scaling for resource management

A computer-implemented method for trace-driven dependency-set-aware proactive coordinated autoscaling of component microservices in an application includes generating performance-resource elasticity models at a trace-level for traces of the application using dependency set of microservices for each trace. The method predicts workload levels of each of the traces, and also predicts a trace-level performance of the application for different microservice replica scaling based on the dependency set of microservices for each trace, performance-resource elasticity models and the predicted workload levels. The method uses distributed computing to recommend a microservice replica scaling for each of the component microservices to meet one or more predefined trace-level user service level objectives.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Model input request scheduling method and device, storage medium and program product

Embodiments of the invention provide a model input request scheduling method and device, a storage medium and a program product. The method comprises the steps of receiving an input request of a preset language model; determining prediction time consumption of the input request in a pre-filling stage in a preset language model processing process; determining the latest execution time of the pre-filling stage of the input request according to the receiving time of the input request, the predicted consumed time of the pre-filling stage and the target first lexical element delay; and according to the latest execution time of the pre-filling stage, sending the input request to a service node of a preset language model for processing. According to the embodiment of the invention, the latest execution time of the pre-filling stage is determined based on the predicted time consumption of the pre-filling stage of the input request to schedule the input request, and the system throughput and the calculation load balance are optimized while the service level target is met, so that the resource utilization efficiency and the service quality guarantee capability of the service node of the preset language model can be improved.
Owner:BYTEDANCE TECHNOLOGY CO LTD +1

Intent-based orchestration in heterogenous compute platforms

Various systems and methods for implementing intent-based orchestration in heterogenous compute platforms are described herein. An orchestration system is configured to: receive, at the orchestration system, a workload request for a workload, the workload request including an intent-based service level objective (SLO); generate rules for resource allocation based on the workload request; generate a deployment plan using the rules for resource allocation and the intent-based SLO; deploy the workload using the deployment plan; monitor performance of the workload using real-time telemetry; and modify the rules for resource allocation and the deployment plan based on the real-time telemetry.
Owner:INTEL CORP

Multi-target execution plan optimization method, system and product for intelligent application

The invention discloses a multi-target execution plan optimization method, system and product oriented to intelligent application, and belongs to the technical field of intelligent application. A logic application diagram and context information are obtained, and the context information comprises a service level target and a load feature; the method comprises the steps that a logic application diagram, a service level target, load characteristics and a pre-constructed operator performance model are integrated to obtain a constraint optimization problem, a multi-target constraint optimization solver is used for solving the constraint optimization problem, an optimal physical execution plan is generated, and an execution engine is used for executing the optimal physical execution plan. According to the method, the multi-dimensional service target can be understood and adjusted in real time by obtaining the service level target set by the user, the hysteresis is avoided by generating the optimal physical execution plan before plan execution, the self-adaptive capability and the quality of the generated result are improved by using the pre-constructed operator performance model and the multi-target constraint optimization solver, and the method has the advantages of being high in practicability and the like. And the scientificity, the globality and the reproducibility of a scheduling decision are ensured.
Owner:北京文聿科技有限公司

Improvements in data product development

This disclosure relates to methods, devices, and computer-readable media for use in developing data products. One such method comprises receiving an existing build of the data product, identifying a data product version associated with the existing build, receiving a user-specified modification for the data product, in response to a user input, automatically determining a compatibility result for the modification with the identified data product version, based on the existing build of the data product, and in response to the determined compatibility result being a negative compatibility result, triggering a failure event in relation to the identified data product version. The user specified modification may comprise one or more of: modification of a schema or addition modification or deletion of one or more tables, columns, service level objectives, service level indicators, constraints. Determination of the compatibility result may comprise determining a negative compatibility result by identifying that a table, column, service level objective, constraint or other object present in the existing build is removed by the modification.
Owner:DATAOPS SOFTWARE LTD

Method and apparatus for using microslices to control network performance of an enterprise wireless communication network

A method and apparatus for utilizing microslices in an enterprise wireless communication network to manage and control network performance. The microslice instances are monitored during network operation, for example at communication nodes, functional blocks, and end-to-end to provide Key Performance Indicators (KPIs). These KPI are compared with performance objectives, which may be Service Level Objectives (SLOs). If the performance objectives are not met by the KPIs, then one or more of the microslice instances may be dynamically adjusted until the performance objectives are sufficiently met. Alternatively, the lower priority microslice instances may be dropped (i.e., terminated) until the performance objectives are sufficiently met.
Owner:CELONA INC

Power oversubscription in LLM cloud providers

Systems and methods for implementing power oversubscription in graphic processing unit (GPU) servers are provided. An increase to a quantity of servers allocated to a group of GPU servers in an inference cluster is applied. Based on the power consumption of the group of GPU servers exceeding a first threshold, a frequency of low priority inference workloads is capped, and based on the power consumption of the group of GPU servers exceeding a second threshold, the frequency of the low priority inference workloads are capped and a frequency of high priority inference workloads are capped, enabling an increase in allocated server capacity in the existing inference clusters while maintaining service level objectives (SLOs).
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Service level objective-based regulator

Techniques are disclosed that enable a self-regulating process to meet a service level objective (SLO). In some embodiments, a self-regulating process is a background process comprising a regulator that receives background job requests and historical information related to the background process for evaluation to determine actions (e.g., speed up, slow down, or maintain the same speed), enabling the background process to adjust its pace gradually and smoothly even when encountering unexpected big changes in load.
Owner:ORACLE INT CORP

Service level objective (SLO) based continuous integration / continuous development (CICD) framework for chaos engineering

A Service Level Objective (SLO)-based Continuous Integration / Continuous Development (CICD) framework is disclosed, which enables SLO-based Canary deployments, SLO-based chaos engineering, and / or SLO-based performance testing. The SLO-based CICD framework enhances and simplifies CICD pipelines and brings the pipelines in line with Site Reliability Engineering (SRE) best practices. By relying on SLOs rather than SLIs, the SLO-based CICD framework avoids the use of threshold-based alerts, which can be noisy and error prone, and the thresholds associated with threshold-based alerts can change over time due to changes in production volume. The SLO-based CICD framework uses multi-window burn rates when evaluating SLOs, which removes the reliance on a static, predefined thresholds of SLI threshold-based alerting, and removes sensitivity of such evaluations to changes in production volume.
Owner:UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)

Service level objective (SLO) based continuous integration / continuous development (CICD) framework for performance testing

A Service Level Objective (SLO)-based Continuous Integration / Continuous Development (CICD) framework is disclosed, which enables SLO-based Canary deployments, SLO-based chaos engineering, and / or SLO-based performance testing. The SLO-based CICD framework enhances and simplifies CICD pipelines and brings the pipelines in line with Site Reliability Engineering (SRE) best practices. By relying on SLOs rather than SLIs, the SLO-based CICD framework avoids the use of threshold-based alerts, which can be noisy and error prone, and the thresholds associated with threshold-based alerts can change over time due to changes in production volume. The SLO-based CICD framework uses multi-window burn rates when evaluating SLOs, which removes the reliance on a static, predefined thresholds of SLI threshold-based alerting, and removes sensitivity of such evaluations to changes in production volume.
Owner:UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)

Method and system for performing anomaly detection in a distributed multi-tiered computing environment

Techniques described herein relate to a method for managing a distributed multi-tiered computing (DMC) environment. The method includes obtaining, by a local controller associated with a DMC domain, service level objective (SLO) metrics; applying the SLO metrics to a predictive anomaly detection transformer to perform anomaly detection; making a first determination that an anomaly is detected; in response to the first determination: attempting basic remediation to resolve the anomaly; making a second determination that the basic remediation is unsuccessful; in response to the second determination: making a third determination that the anomaly is associated with a silent failure; and in response to the third determination: performing service impairment isolation to obtain a collection of services correlated to the anomaly; and performing root cause analysis to identify causal services.
Owner:DELL PROD LP

Deep reinforcement learning cloud platform resource adjustment method based on prediction of user workload

The invention relates to the technical field of cloud computing, and provides a deep reinforcement learning cloud platform resource adjustment method based on user workload prediction. The method aims at solving the problem of adjustment hysteresis caused by the fact that load changes cannot be predicted in a traditional response type resource adjustment strategy, so that the rule violation rate of a service level target (SLO) is effectively reduced, and the resource utilization rate is improved. According to the main scheme, the method comprises the following steps: firstly, predicting a future workload trend through a long short-term memory (LSTM) network; combining a prediction result with a current system state (such as response delay and copy number) to serve as input of a deep reinforcement learning agent; the intelligent agent outputs a resource adjustment action through a strategy network; and finally, optimizing the decision model according to a composite revenue function integrating the service quality, the resource cost and the adjustment stability. The method can be widely applied to cloud native application under a micro-service architecture, and automatic and prospective elastic resource scaling management is realized.
Owner:INFORMATION & TELECOMM COMPANY SICHUAN ELECTRIC POWER

Service level breach prediction

A method is provided, comprising: obtaining a plurality of measured response times for a given storage group, each of the measured response times corresponding to a different instance of a same time window; classifying the plurality of measured response times with a machine learning model to obtain a first predicted response time and a second predicted response time, the first predicted response time corresponding to a first future instance of the time window, and the second predicted response time corresponding to a second future instance of the time window; detecting whether the first predicted response time and the second predicted response time satisfies a service level objective for the given storage group; and generating a notification when at least one of the first predicted response time and the second predicted response time fails to satisfy the service level objective.
Owner:DELL PROD LP

Computing power scheduling method and device, communication equipment, medium and product

The invention provides a computing power scheduling method and device, communication equipment, a medium and a product, and the method comprises the steps: carrying out the data preprocessing and topology awareness of a computing task, and obtaining a computing graph of the computing task; performing sub-graph division on the calculation graph to obtain calculation sub-graphs, and packaging the calculation sub-graphs into schedulable units; the calculation benefit of the schedulable unit is obtained by determining the execution delay, the power consumption and the heterogeneous coefficient of the calculation nodes which distribute the schedulable unit to different areas; constructing a task scheduling model meeting a service level target according to a maximum and minimum algorithm and the calculation benefit; the optimal allocation of the schedulable units on the computing nodes is determined through the task scheduling model, the schedulable units are scheduled to the computing nodes in parallel according to the optimal allocation, large computing tasks can be scheduled to a proper area according to delay and power consumption, the problem of hardware resource heterogeneity in the scheduling process can be solved, and the scheduling efficiency is improved. And the total power consumption of the whole task is reduced through refined scheduling.
Owner:ZUNYI BRANCH OF CHINA MOBILE GRP GUIZHOU COMPANY +1

Data arrangement method and device oriented to job intention

The invention provides a job intention-oriented data arrangement method and device, and the method comprises the steps: compiling a job intention, and obtaining a multi-level arrangement strategy; the multi-level arrangement strategy comprises at least two of an application layer arrangement strategy, a transmission layer arrangement strategy and a system storage layer arrangement strategy; and executing the data transmission operation based on the multi-level arrangement strategy. According to the method provided by the invention, the operation intention containing the global demand and the constraint is acquired and compiled into the multi-level arrangement strategy covering multiple levels of application, transmission, storage and the like, and finally the data transmission operation is executed based on the strategy, so that the full-link resource configuration can be flexibly adjusted according to the service intention; the problem that downstream data receiving is blocked due to the fact that the network is not matched with the storage performance is effectively solved, and the comprehensive performance and the resource utilization rate of data transmission are remarkably improved while the service level target is guaranteed and complex constraints are met.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Systems and methods for edge resource demand load scheduling

Managing the resource demand load for edge systems is significantly more complex than for other systems, such as cloud environments. Edge resource demand load scheduling systems and methods are disclosed that can ensure that edge systems operate smoothly and efficiently while balancing multiple scheduling objectives. Scheduling techniques disclosed herein may utilize heuristic rules for candidate edge system selection (e.g., utilizing ARMA / ARIMA averages and / or service level objectives) and modified best fit decreasing (mBFD) assignment / allocation techniques.
Owner:DELL PROD LP

Real-time analysis and intelligent optimization system for software operation data based on machine learning

ActiveCN121302884BDynamically reflect fluctuation characteristicsCharacterizing average performanceBiological modelsDesign optimisation/simulationData packResource utilization
This invention provides a real-time analysis and intelligent optimization system for software operation data based on machine learning, comprising: a data acquisition module for acquiring real-time operational data of the software operation system, including request arrival logs, service processing logs, and resource usage information; a mechanism model construction module for establishing a queuing control mechanism model based on the real-time operational data and defining a set of operational operations within the mechanism model; an operational state estimation module for estimating the real-time operational data based on the mechanism model to obtain request arrival rate intervals, service capacity intervals, and burst traffic intervals, forming a set of uncertainty intervals; and a service level target constraint module for selecting tail performance indicators as service level targets based on the set of uncertainty intervals. This invention can continuously guarantee service level targets under dynamic fluctuations and uncertain environments, improving the stability and resource utilization of the software operation system.
Owner:南昌职业大学

Software operation data real-time analysis and intelligent optimization system based on machine learning

The invention provides a software operation data real-time analysis and intelligent optimization system based on machine learning, and the system comprises a data collection module which is used for obtaining real-time operation data of a software operation system, and the operation data comprises a request arrival log, a service processing log and resource occupation information; the mechanism model building module is used for building a queuing control mechanism model based on the real-time operation data and defining an operation set in the mechanism model; the operation state estimation module is used for estimating real-time operation data based on the mechanism model to obtain a request arrival rate interval, a service capability interval and a burst flow interval to form an uncertainty interval set; and selecting a tail performance index as a service grade target on the basis of the uncertainty interval set. According to the method, the service level target can be continuously guaranteed in a dynamic fluctuation and uncertain environment, and the stability and the resource utilization rate of a software operation system are improved.
Owner:南昌职业大学

Methods, apparatus, and articles of manufacture to schedule algorithms based on characteristics of data

Systems, apparatus, articles of manufacture, and methods are disclosed to schedule algorithms based on characteristics of data. An example compute device includes circuitry to determine at least one characteristic of a data object to be processed and adjust metadata associated with the data object to indicate that the data object has the at least one characteristic. Additionally, the example compute device includes machine-readable instructions and at least one programmable circuit to be programmed by the machine-readable instructions to select at least one of two or more programmable circuits to process the data object based on (a) at least one service level objective associated with the data object, (b) the metadata associated with the data object, and (c) telemetry data associated with the two or more programmable circuits.
Owner:OPENCHIP & SOFTWARE TECHNOLOGIES SL

Quick and quality-aware cloud collaborative big language model reasoning method and system

The invention discloses a fast and quality-aware cloud collaborative large language model reasoning method and system (CEC-LM), and belongs to the field of computer networks and artificial intelligence, and the method comprises the steps: employing a task execution selector based on a lightweight model, evaluating the task complexity in real time according to a prompt word requested by a user, and carrying out the real-time evaluation of the task complexity according to the prompt word; the complex tasks are routed to a cloud LLM to ensure response quality, and meanwhile the simple tasks are distributed to a local small language model to be executed. And when sensing that the local SLM is high in load, dynamically unloading the complexity prediction model from the GPU to the CPU, and realizing zero-copy parameter transmission by utilizing a unified memory architecture. According to the method, a task monitoring and priority scheduling mechanism during operation is introduced, the task routing is adaptively adjusted by continuously monitoring the network condition and the system load, and the emergency task is preferentially processed, so that the achievement rate of the service level target is ensured, the reasoning quality, delay and cost are effectively balanced, and the system overhead is remarkably reduced.
Owner:WUHAN UNIV

Trace-driven call_graph_aware proactive coordinated autoscaling for resource management

A computer-implemented method for trace-driven call-graph-aware proactive coordinated autoscaling of component microservices in an application includes generating performance-resource elasticity models of endpoints of the component microservices of the application. Workload levels of the endpoint of the component microservices is predicted based on user traffic observed at a front end service. A trace-level performance of the application is predicted for different microservice replica scaling based on the performance-resource elasticity models at end points, the ends points on the trace call graph and the predicted workload levels. A microservice replica scaling is recommended for each of the component microservices to meet predefined trace-level user service level objectives.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION +1

Management of system health by customer-defined service level agreements and service level objectives

A method for managing system of a production environment includes obtaining, from a client device, a request for remediation of a user-defined parameter of the production environment, determining a set of potential remediation actions for the user-defined parameter, applying each of the set of potential remediation actions to a simulated production environment to obtain a set of remediation results, updating a system health display using the set of remediation results to obtain an updated system health display, wherein the updated system health display comprises a display of the user-defined parameter and an impact of a remediation action of the set of remediation actions on the user-defined parameter, and implementing the remediation action of the set of remediation actions to service the request for remediation.
Owner:DELL PROD LP

Elastic service system based on MOE model

An elastic service system based on an MOE model comprises an off-line reconstruction engine and an on-line scheduling engine, the off-line reconstruction engine carries out one-time structure conversion without retraining on a pre-trained model, an original single expert in the model is deconstructed into a plurality of fine-grained sub experts, and finer-grained calculation configuration options are obtained; an online scheduling engine dynamically decides the number and combination of sub-experts to be activated for each request according to real-time workloads, heterogeneous user requests and system resource conditions to achieve flexible services. According to the method, the single experts with coarse granularity are deconstructed into the sub experts with fine granularity, the online scheduling strategy matched with the sub experts is constructed, and the rigid structure of the model is converted into smooth and elastic service capability, so that the cost-quality accurate control is realized to adapt to diversified service level targets, and the resource utilization efficiency of the system is improved.
Owner:SHANGHAI JIAOTONG UNIV

Prioritized storage array rebuild

One or more aspects of the present disclosure relate to optimizing the rebuild process of a persistent storage device in a storage array is disclosed. The embodiments detect rebuild events, identify affected back-end slices, and prioritize the rebuild order based on a calculated priority score for each slice. This score is derived from service level objectives (SLO) and input / output (IO) statistics of corresponding front-end logical tracks. The embodiments can generate SLO slice objects representing back-end slices, group them in a shared memory database, and update scores during write operations. Rebuild job queues with different priority levels are established, and back-end slices are queued based on their priority scores. This approach ensures efficient rebuilding of critical data, considering both SLOs and real-time IO statistics, thus minimizing performance degradation and enhancing overall system reliability.
Owner:DELL PROD LP

Managing power for serverless computing

Embodiments dynamically measure latency for a plurality of functions with a plurality of corresponding frequency levels within a serverless computing cluster; measure a transition latency from an idle state to an active state for the plurality of functions; determine whether a target response time to perform a service level objective (SLO) within the serverless computing cluster is going to be missed; dynamically reallocate at least one core and changing a frequency level across the plurality of functions by scaling down in response to a determination that the target response time to perform the SLO within the serverless computing cluster is going to be met; and dynamically reallocate the at least one core and changing the frequency level across the plurality of functions by scaling up in response to a determination that the target response time to perform the SLO within the serverless computing cluster is going to be missed.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Adaptive load balancing method for maximizing session affinity and routing device

The invention belongs to the technical field of server cluster load balancing, and relates to a self-adaptive load balancing method for maximizing session affinity and a routing device. The method comprises the following steps: presetting a consistency candidate instance list based on multiple Hash for each session, monitoring an SLO index of each instance in real time, and converting the SLO index into a dynamic access probability through a feedback controller; during routing, traversing the candidate list according to priorities, performing probability judgment in combination with the real-time access probability of each instance, and selecting the first instance suitable for access; and when all the preset paths are unavailable, degrading to an instance with the lowest global load to ensure the final availability of the system. According to the method, the problems that in traditional load balancing, the inherent contradiction between session affinity maintaining and dynamic load balancing is achieved, and a single strategy cannot guarantee the service quality SLO and the capacity expansion and contraction stability of the system at the same time are effectively solved; and adaptive shunting of hotspot traffic and maximum utilization of session locality are realized.
Owner:BEIJING SILICON MOBILE TECHNOLOGY CO LTD

A time-series data-based method for SLO lifecycle management and analysis

PendingCN122309292AAnalysis dataData mining
This invention proposes a method for SLO (Service Level Requirement) lifecycle management and analysis based on time-series data. The method includes: creating Service Level Requirement (SLO) rules and establishing the association between Service Level Indicators (SLIs) and SLOs; starting a first scheduled task, which calls a cold start function upon initial startup to calculate and store daily granularity achievement rate data for SLIs within twice the configuration period; starting a second scheduled task, which queries the daily granularity achievement rate data of all SLIs belonging to the same SLO from the previous day at a second preset time each day, and performs a weighted sum based on their respective weights; and starting a third scheduled task, which executes according to a preset period, obtaining historical daily granularity achievement rate data and combining it with the real-time calculated daily SLO achievement rate data to calculate SLO analysis data each time. This invention solves the technical challenges of long SLI data gaps and interference in achievement rate calculation during non-fault periods in traditional solutions through a data-ready mechanism.
Owner:JIANGXI TONGRUI INFORMATION TECH CO LTD

System and method for proactive customer engagement experiences

ActiveUS12664559B1FinanceAlarm indicatorsPersonalizationCustomer engagement
A method and system of providing customer-specific information and guidance to insurance policyholders during claims processing. The information is created based on outputs from a claims complexity model, customer segmentation model, and claims service level objective (SLO) model. The resulting information is tailored to the individual policyholder and their claim, and includes a summary of their claim activity, guidance regarding their next steps, and predicted wait times for each stage in their claim's processing. In some embodiments, the information is presented as messaging delivered over the policyholder's preferred communication channel(s) that can differ based on which stage of the claims process the claim is in.
Owner:UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)

Service level target management methods, systems, electronic devices and storage media

This application proposes a service level target management method, system, electronic device, and storage medium. The service level target management method includes: acquiring user-defined service level target configuration information for each microservice, as well as monitoring indicator data and workload context information of the currently running workload; a workload refers to the specific application implementing the microservice; workload context information refers to the environment and status information related to the workload during execution; generating corresponding service level target rule data based on the service level target configuration information and workload context information; the service level target rule data includes service level target metadata, service level indicator rules, and service alarm rules; and managing the service level targets of the microservices based on the monitoring indicator data and service level target rule data. This application can determine different service level indicator types and service level targets for different workloads.
Owner:ALIBABA (CHINA) CO LTD