Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

19 results about "Service level objective" patented technology

A service-level objective (SLO) is a key element of a service-level agreement (SLA) between a service provider and a customer. SLOs are agreed upon as a means of measuring the performance of the Service Provider and are outlined as a way of avoiding disputes between the two parties based on misunderstanding.

System and method for coordinated resource scaling in microservice-based and serverless applications

A computer-implemented method for trace-driven call-graph-aware proactive coordinated autoscaling of component microservices in an application includes generating performance-resource elasticity models of endpoints of the component microservices of the application. Workload levels of the endpoint of the component microservices is predicted based on user traffic observed at a front end service. A trace-level performance of the application is predicted for different microservice replica scaling based on the performance-resource elasticity models at end points, the ends points on the trace call graph and the predicted workload levels. A microservice replica scaling is recommended for each of the component microservices to meet predefined trace-level user service level objectives.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Service level objective (SLO) based continuous integration / continuous development (CICD) framework for chaos engineering

A Service Level Objective (SLO)-based Continuous Integration / Continuous Development (CICD) framework is disclosed, which enables SLO-based Canary deployments, SLO-based chaos engineering, and / or SLO-based performance testing. The SLO-based CICD framework enhances and simplifies CICD pipelines and brings the pipelines in line with Site Reliability Engineering (SRE) best practices. By relying on SLOs rather than SLIs, the SLO-based CICD framework avoids the use of threshold-based alerts, which can be noisy and error prone, and the thresholds associated with threshold-based alerts can change over time due to changes in production volume. The SLO-based CICD framework uses multi-window burn rates when evaluating SLOs, which removes the reliance on a static, predefined thresholds of SLI threshold-based alerting, and removes sensitivity of such evaluations to changes in production volume.
Owner:UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)

Service level objective (SLO) based continuous integration / continuous development (CICD) framework for performance testing

A Service Level Objective (SLO)-based Continuous Integration / Continuous Development (CICD) framework is disclosed, which enables SLO-based Canary deployments, SLO-based chaos engineering, and / or SLO-based performance testing. The SLO-based CICD framework enhances and simplifies CICD pipelines and brings the pipelines in line with Site Reliability Engineering (SRE) best practices. By relying on SLOs rather than SLIs, the SLO-based CICD framework avoids the use of threshold-based alerts, which can be noisy and error prone, and the thresholds associated with threshold-based alerts can change over time due to changes in production volume. The SLO-based CICD framework uses multi-window burn rates when evaluating SLOs, which removes the reliance on a static, predefined thresholds of SLI threshold-based alerting, and removes sensitivity of such evaluations to changes in production volume.
Owner:UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)

Service level breach prediction

A method is provided, comprising: obtaining a plurality of measured response times for a given storage group, each of the measured response times corresponding to a different instance of a same time window; classifying the plurality of measured response times with a machine learning model to obtain a first predicted response time and a second predicted response time, the first predicted response time corresponding to a first future instance of the time window, and the second predicted response time corresponding to a second future instance of the time window; detecting whether the first predicted response time and the second predicted response time satisfies a service level objective for the given storage group; and generating a notification when at least one of the first predicted response time and the second predicted response time fails to satisfy the service level objective.
Owner:DELL PROD LP

Data arrangement method and device oriented to job intention

The invention provides a job intention-oriented data arrangement method and device, and the method comprises the steps: compiling a job intention, and obtaining a multi-level arrangement strategy; the multi-level arrangement strategy comprises at least two of an application layer arrangement strategy, a transmission layer arrangement strategy and a system storage layer arrangement strategy; and executing the data transmission operation based on the multi-level arrangement strategy. According to the method provided by the invention, the operation intention containing the global demand and the constraint is acquired and compiled into the multi-level arrangement strategy covering multiple levels of application, transmission, storage and the like, and finally the data transmission operation is executed based on the strategy, so that the full-link resource configuration can be flexibly adjusted according to the service intention; the problem that downstream data receiving is blocked due to the fact that the network is not matched with the storage performance is effectively solved, and the comprehensive performance and the resource utilization rate of data transmission are remarkably improved while the service level target is guaranteed and complex constraints are met.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Systems and methods for edge resource demand load scheduling

Managing the resource demand load for edge systems is significantly more complex than for other systems, such as cloud environments. Edge resource demand load scheduling systems and methods are disclosed that can ensure that edge systems operate smoothly and efficiently while balancing multiple scheduling objectives. Scheduling techniques disclosed herein may utilize heuristic rules for candidate edge system selection (e.g., utilizing ARMA / ARIMA averages and / or service level objectives) and modified best fit decreasing (mBFD) assignment / allocation techniques.
Owner:DELL PROD LP

Real-time analysis and intelligent optimization system for software operation data based on machine learning

ActiveCN121302884BDynamically reflect fluctuation characteristicsCharacterizing average performanceBiological modelsDesign optimisation/simulationData packResource utilization
This invention provides a real-time analysis and intelligent optimization system for software operation data based on machine learning, comprising: a data acquisition module for acquiring real-time operational data of the software operation system, including request arrival logs, service processing logs, and resource usage information; a mechanism model construction module for establishing a queuing control mechanism model based on the real-time operational data and defining a set of operational operations within the mechanism model; an operational state estimation module for estimating the real-time operational data based on the mechanism model to obtain request arrival rate intervals, service capacity intervals, and burst traffic intervals, forming a set of uncertainty intervals; and a service level target constraint module for selecting tail performance indicators as service level targets based on the set of uncertainty intervals. This invention can continuously guarantee service level targets under dynamic fluctuations and uncertain environments, improving the stability and resource utilization of the software operation system.
Owner:南昌职业大学

Software operation data real-time analysis and intelligent optimization system based on machine learning

The invention provides a software operation data real-time analysis and intelligent optimization system based on machine learning, and the system comprises a data collection module which is used for obtaining real-time operation data of a software operation system, and the operation data comprises a request arrival log, a service processing log and resource occupation information; the mechanism model building module is used for building a queuing control mechanism model based on the real-time operation data and defining an operation set in the mechanism model; the operation state estimation module is used for estimating real-time operation data based on the mechanism model to obtain a request arrival rate interval, a service capability interval and a burst flow interval to form an uncertainty interval set; and selecting a tail performance index as a service grade target on the basis of the uncertainty interval set. According to the method, the service level target can be continuously guaranteed in a dynamic fluctuation and uncertain environment, and the stability and the resource utilization rate of a software operation system are improved.
Owner:南昌职业大学

Quick and quality-aware cloud collaborative big language model reasoning method and system

The invention discloses a fast and quality-aware cloud collaborative large language model reasoning method and system (CEC-LM), and belongs to the field of computer networks and artificial intelligence, and the method comprises the steps: employing a task execution selector based on a lightweight model, evaluating the task complexity in real time according to a prompt word requested by a user, and carrying out the real-time evaluation of the task complexity according to the prompt word; the complex tasks are routed to a cloud LLM to ensure response quality, and meanwhile the simple tasks are distributed to a local small language model to be executed. And when sensing that the local SLM is high in load, dynamically unloading the complexity prediction model from the GPU to the CPU, and realizing zero-copy parameter transmission by utilizing a unified memory architecture. According to the method, a task monitoring and priority scheduling mechanism during operation is introduced, the task routing is adaptively adjusted by continuously monitoring the network condition and the system load, and the emergency task is preferentially processed, so that the achievement rate of the service level target is ensured, the reasoning quality, delay and cost are effectively balanced, and the system overhead is remarkably reduced.
Owner:WUHAN UNIV

Elastic service system based on MOE model

An elastic service system based on an MOE model comprises an off-line reconstruction engine and an on-line scheduling engine, the off-line reconstruction engine carries out one-time structure conversion without retraining on a pre-trained model, an original single expert in the model is deconstructed into a plurality of fine-grained sub experts, and finer-grained calculation configuration options are obtained; an online scheduling engine dynamically decides the number and combination of sub-experts to be activated for each request according to real-time workloads, heterogeneous user requests and system resource conditions to achieve flexible services. According to the method, the single experts with coarse granularity are deconstructed into the sub experts with fine granularity, the online scheduling strategy matched with the sub experts is constructed, and the rigid structure of the model is converted into smooth and elastic service capability, so that the cost-quality accurate control is realized to adapt to diversified service level targets, and the resource utilization efficiency of the system is improved.
Owner:SHANGHAI JIAOTONG UNIV

Prioritized storage array rebuild

One or more aspects of the present disclosure relate to optimizing the rebuild process of a persistent storage device in a storage array is disclosed. The embodiments detect rebuild events, identify affected back-end slices, and prioritize the rebuild order based on a calculated priority score for each slice. This score is derived from service level objectives (SLO) and input / output (IO) statistics of corresponding front-end logical tracks. The embodiments can generate SLO slice objects representing back-end slices, group them in a shared memory database, and update scores during write operations. Rebuild job queues with different priority levels are established, and back-end slices are queued based on their priority scores. This approach ensures efficient rebuilding of critical data, considering both SLOs and real-time IO statistics, thus minimizing performance degradation and enhancing overall system reliability.
Owner:DELL PROD LP

Adaptive load balancing method for maximizing session affinity and routing device

The invention belongs to the technical field of server cluster load balancing, and relates to a self-adaptive load balancing method for maximizing session affinity and a routing device. The method comprises the following steps: presetting a consistency candidate instance list based on multiple Hash for each session, monitoring an SLO index of each instance in real time, and converting the SLO index into a dynamic access probability through a feedback controller; during routing, traversing the candidate list according to priorities, performing probability judgment in combination with the real-time access probability of each instance, and selecting the first instance suitable for access; and when all the preset paths are unavailable, degrading to an instance with the lowest global load to ensure the final availability of the system. According to the method, the problems that in traditional load balancing, the inherent contradiction between session affinity maintaining and dynamic load balancing is achieved, and a single strategy cannot guarantee the service quality SLO and the capacity expansion and contraction stability of the system at the same time are effectively solved; and adaptive shunting of hotspot traffic and maximum utilization of session locality are realized.
Owner:BEIJING SILICON MOBILE TECHNOLOGY CO LTD

A time-series data-based method for SLO lifecycle management and analysis

PendingCN122309292AAnalysis dataData mining
This invention proposes a method for SLO (Service Level Requirement) lifecycle management and analysis based on time-series data. The method includes: creating Service Level Requirement (SLO) rules and establishing the association between Service Level Indicators (SLIs) and SLOs; starting a first scheduled task, which calls a cold start function upon initial startup to calculate and store daily granularity achievement rate data for SLIs within twice the configuration period; starting a second scheduled task, which queries the daily granularity achievement rate data of all SLIs belonging to the same SLO from the previous day at a second preset time each day, and performs a weighted sum based on their respective weights; and starting a third scheduled task, which executes according to a preset period, obtaining historical daily granularity achievement rate data and combining it with the real-time calculated daily SLO achievement rate data to calculate SLO analysis data each time. This invention solves the technical challenges of long SLI data gaps and interference in achievement rate calculation during non-fault periods in traditional solutions through a data-ready mechanism.
Owner:JIANGXI TONGRUI INFORMATION TECH CO LTD

System and method for proactive customer engagement experiences

ActiveUS12664559B1FinanceAlarm indicatorsPersonalizationCustomer engagement
A method and system of providing customer-specific information and guidance to insurance policyholders during claims processing. The information is created based on outputs from a claims complexity model, customer segmentation model, and claims service level objective (SLO) model. The resulting information is tailored to the individual policyholder and their claim, and includes a summary of their claim activity, guidance regarding their next steps, and predicted wait times for each stage in their claim's processing. In some embodiments, the information is presented as messaging delivered over the policyholder's preferred communication channel(s) that can differ based on which stage of the claims process the claim is in.
Owner:UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)

Service level target management methods, systems, electronic devices and storage media

This application proposes a service level target management method, system, electronic device, and storage medium. The service level target management method includes: acquiring user-defined service level target configuration information for each microservice, as well as monitoring indicator data and workload context information of the currently running workload; a workload refers to the specific application implementing the microservice; workload context information refers to the environment and status information related to the workload during execution; generating corresponding service level target rule data based on the service level target configuration information and workload context information; the service level target rule data includes service level target metadata, service level indicator rules, and service alarm rules; and managing the service level targets of the microservices based on the monitoring indicator data and service level target rule data. This application can determine different service level indicator types and service level targets for different workloads.
Owner:ALIBABA (CHINA) CO LTD

Service level objective (SLO) based continuous integration / continuous development (CICD) framework for canary release integration

A Service Level Objective (SLO)-based Continuous Integration / Continuous Development (CICD) framework is disclosed, which enables SLO-based Canary deployments, SLO-based chaos engineering, and / or SLO-based performance testing. The SLO-based CICD framework enhances and simplifies CICD pipelines and brings the pipelines in line with Site Reliability Engineering (SRE) best practices. By relying on SLOs rather than SLIs, the SLO-based CICD framework avoids the use of threshold-based alerts, which can be noisy and error prone, and the thresholds associated with threshold-based alerts can change over time due to changes in production volume. The SLO-based CICD framework uses multi-window burn rates when evaluating SLOs, which removes the reliance on a static, predefined thresholds of SLI threshold-based alerting, and removes sensitivity of such evaluations to changes in production volume.
Owner:UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)

Large language model reasoning effective throughput optimization method, system, equipment and medium

The invention discloses a large language model reasoning effective throughput optimization method, system and device and a medium, which are corresponding schemes, in the scheme, lexical element level segmentation is performed on a reasoning request, an optimal segmentation point is dynamically determined, and in the process, the lexical element level segmentation is performed on the lexical element level segmentation point; simulated queueing and scheduling can be carried out on the segmented micro-requests under the condition that reasoning calculation is not actually executed, so that delay constraints and load balancing constraints are verified; moreover, the positions of the segmentation points are adjusted and searched step by step, so that the predicted calculation time of the two segmented micro requests is more balanced on the premise of meeting the service level target, waiting and idling caused by uneven resource matching are reduced, the utilization rate of calculation resources is improved, and the calculation efficiency is improved. Therefore, the effective throughput is improved, and the end-to-end service delay is reduced.
Owner:UNIV OF SCI & TECH OF CHINA

Model reasoning service judgment system and method based on intention engine

The invention discloses a model reasoning service judgment system and method based on an intention engine, which are used for sequentially completing service mode selection, model instance selection and routing execution, result output and closed-loop updating of strategy parameters by taking request characteristics and a system resource state as input for each reasoning request. According to the method, mechanisms such as request-level dual-mode switching, semantic and resource joint judgment, multi-dimensional result judgment and metadata-based closed-loop learning are integrated in the same edge inference system, so that the edge multi-model inference system can consider inference quality and resource utilization efficiency while ensuring a service level target.
Owner:SICHUAN JIEHONG INTELLIGENT TECH CO LTD

Electronic device comprising framework operating on basis of model for converting service level objectives (SLO) of network bandwidth into optimal CPU allocation, and operating method thereof

According to various embodiments, an electronic device comprising a Tasador framework operating on the basis of a model for converting service level objectives (SLO) of a network bandwidth into an optimal CPU allocation value, comprises: a Tasador manager; a Tasador collector; a communication interface; and a processor, wherein the processor is configured to: allow the Tasador manager to receive user request information through the communication interface; predict the optimal CPU allocation value corresponding to the SLO of the network bandwidth by applying the user request information to a CPU allocation prediction model as input data of the CPU allocation prediction model; and enforce CPU allocation corresponding to the optimal CPU allocation value by using the Tasador manager, wherein the CPU allocation prediction model may be trained on the basis of a training dataset received from the Tasador collector through the communication interface.
Owner:KYUNGPOOK NAT UNIV IND ACADEMIC COOP FOUND