Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

21 results about "Job scheduler" patented technology

A job scheduler is a computer application for controlling unattended background program execution of jobs. This is commonly called batch scheduling, as execution of non-interactive jobs is often called batch processing, though traditional job and batch are distinguished and contrasted; see that page for details. Other synonyms include batch system, distributed resource management system (DRMS), distributed resource manager (DRM), and, commonly today, workload automation (WLA). The data structure of jobs to run is known as the job queue.

Distributed pipeline-parallel LLM fine-tuning method for heterogeneous GPU

This application relates to the technical field of natural language processing, and provides a distributed pipeline-parallel LLM fine-tuning method for heterogeneous GPUs. A plurality of LoRA models are fine-tuned simultaneously based on a multi-job fine-tuning system; each LoRA model is partitioned into a plurality of parts distributed on a corresponding number of GPUs, and the GPUs are sorted. A job configuration module generates a plurality of jobs according to a user request, and divides each job into a plurality of training batches; a dynamic job scheduler generates a scheduling scheme based on a training batch sequence of each job and a dynamic scheduling strategy; and the scheduling scheme is sent to a multi-job training module on each corresponding GPU according to a positive sequence of the GPUs, to train all the LoRA models.
Owner:ZHEJIANG NEW INTERNET EXCHANGE CENT CO LTD +1

Job scheduling for a data management systems based on job groups

Methods, systems, and devices for data management are described. A data management system (DMS) may provide backup and recovery services for customer computing systems or databases, which may involve scheduling jobs to perform the backup and recovery services. Each job may define a set of semaphores which may be acquired prior to execution of the job. The semaphores may be representative of an availability of computing resources associated with the DMS. Jobs may be grouped into job groups based on the semaphores associated with each job. Each job group may be scheduled independently by separate dispatchers or job schedulers. Within a job group, the associated dispatcher(s) may schedule jobs if the semaphore(s) associated with the job group are available. Within a job group, the associated dispatchers may refrain from dispatching additional jobs if the semaphore(s) for the job group are full until the semaphore(s) are available.
Owner:RUBRIK INC

System and method for software monitoring and maintenance

A system and method may manage a computer by executing, by a job scheduler on the computer, a framework script written in a scripting language which discovers installed sub-task scripts and executes discovered sub-task scripts. Sub-task scripts may be stored in directories corresponding to accounts in which a sub-task script is to be run. Sub-task scripts may perform maintenance or monitoring tasks. A system and method may assign resource access or permissions for a set of computers, by, for each computer of a set of computers, determining the organizational unit and sub-organizational unit of the computer; and for each computer of the set of computers assigning the computer to a security group, based on the organizational unit and sub-organizational unit of the computer. For each computer, the security group the computer is assigned to may be associated with permissions enforced for each computer in the group.
Owner:MORGAN STANLEY SERVICES GROUP INC

An effective reinforcement learning and graph neural network fusion method for solving flexible shop floor scheduling

This invention discloses an effective reinforcement learning and graph neural network fusion method for solving flexible job scheduling, comprising the following steps: formalizing the Flexible Job Scheduler (FJSP) into a Multi-Level Product (MDP) by setting states, actions, state transitions, and rewards, while describing the scheduling process as dynamically allocating ready operations to compatible and idle machines; during the scheduling process, the scheduling state is first transformed into a heterogeneous graph structure; a heterogeneous graph attention network is used for two-stage embedding to extract feature embeddings of jobs and machines from the heterogeneous graph; a decision network uses the feature embeddings to generate an action probability distribution, from which scheduling actions are sampled; and a deep reinforcement learning model is trained using a proximal policy optimization algorithm. Compared with traditional PDR and SOTA DRL methods, the algorithm achieves significantly high-quality solutions on two synthetic datasets with different distributions, improving performance and generalization ability. It also exhibits good generalization ability because the model is trained on a small number of instances on a large number of or out-of-distribution instances.
Owner:YUNNAN UNIV

Job scheduling and resource allocation system in high-performance computing cluster

The invention discloses a job scheduling and resource allocation system in a high-performance computing cluster, and relates to the technical field of high-performance computing. The system comprises a job scheduler, a resource manager, a newly-added job DAG (directed acyclic graph) analyzer, a cluster data distribution monitor, a cache preheating priority calculator, a collaborative decision engine, a distributed cache preheating actuator and a performance optimization module, and the job DAG analyzer constructs a job dependence directed acyclic graph and analyzes a critical path; the cache preheating priority calculator adopts a dynamic priority drift algorithm to determine a data preheating priority, the collaborative decision engine coordinates scheduling and preheating resource allocation, the performance optimization module dynamically adjusts system parameters, the system realizes deep collaboration of job scheduling and cache preheating, the job access delay is reduced, and the system performance is improved. The method improves the cluster computing efficiency and the resource utilization rate, and is suitable for processing a high-performance computing cluster of complex operation workflow.
Owner:SHANXI WENHUI WEIYE TECHNOLOGY CO LTD

Real-Time Monitoring For Ransomware Attacks Using Exception-Level Transition Metrics

Aspects of the disclosure include a dynamic cloud workload reallocation based on an active ransomware attack. An example method includes receiving a first message that a computing instance is potentially infected by ransomware. The method further includes receiving a security state-based metric related to the computing instance based at least in part on the first message. The method further includes comparing the security state-based metric to a threshold metric. The method further incudes determining a likelihood of a ransomware attack based at least in part on the comparison. The method further includes transmitting second message to a job scheduler to reschedule workloads directed toward the computing instance based at least in part on the determination.
Owner:ORACLE INT CORP

Elastic job pulling for a scheduling system

Disclosed herein are a system, method, and computer program product embodiments for elastic job pulling. For example, a first indication of a number of jobs at the job scheduler stage that are yet to be executed is obtained. This indication may be obtained from a job scheduler. A second indication of a total number of available job executors communicatively coupled to the job scheduler may be obtained from the job scheduler. A pull request for a job may be issued to the job scheduler based on at least one of the first indication, the second indication, or a computing resource status of a job executor. The job executor may obtain and execute the job based on the pull request. A status indication of the job may be transmitted, where the status indication indicates whether the job executed successfully.
Owner:SAP SE

Job processing system and method based on cloud service

A job processing method based on a cloud service includes: receiving a job request and obtaining an identification string of the job according to the job request by a job scheduler, sending the identification string to a memory cache and generating a worker by the job scheduler, obtains the identification string from the memory cache and obtaining data from a database according to the identification string by the worker, and executing the job based on the data and sending a result of the job to the database by the worker.
Owner:WISTRON CORP

Real-time monitoring for ransomware attacks using exception-level transition metrics

Aspects of the disclosure include a dynamic cloud workload reallocation based on an active ransomware attack. An example method includes receiving a first message that a computing instance is potentially infected by ransomware. The method further includes receiving a security state-based metric related to the computing instance based at least in part on the first message. The method further includes comparing the security state-based metric to a threshold metric. The method further incudes determining a likelihood of a ransomware attack based at least in part on the comparison. The method further includes transmitting second message to a job scheduler to reschedule workloads directed toward the computing instance based at least in part on the determination.
Owner:ORACLE INT CORP

Sublimation NPU-based distributed training system and method

The invention relates to a distributed training system and method based on mercuric chloride NPU, and the method comprises the steps: a batch job scheduler is configured to receive a batch job configuration list, and creates and manages a head node Pod and a plurality of work nodes Pod in batches on nodes with idle mercuric chloride NPU in a cluster managed by a container arrangement platform, the head node Pod and the working node Pod are automatically networked based on a unified distributed computing framework to form a distributed computing cluster special for tasks; the main control service module is further configured to monitor the networking state of the distributed computing cluster, and submit a stand-alone training starting command to the head node Pod after confirming that networking is completed; in addition, the head node Pod is further configured to take the received starting command as a distributed task, and the distributed task is scheduled to each working node Pod for parallel execution through a unified distributed computing framework, so that the use threshold is greatly reduced, and the development efficiency and the system reliability are improved.
Owner:SUPCON TECH CO LTD

Data output timeliness measurement method and device of big data platform and storage medium

The invention discloses a data output timeliness measurement method and device for a big data platform and a storage medium, and the method comprises the steps: obtaining the actual completion timeliness of each data processing operation in the big data platform and the corresponding reference timeliness, and calculating the individual timeliness index of each data processing operation; acquiring preset attribute information and a data dependency relationship of each data processing job; determining a dynamic factor according to preset attribute information and the data dependency relationship, and determining an aggregation weight of the data processing job in each evaluation dimension based on the dynamic factor; calculating a dimension timeliness index through the individual timeliness index and the aggregation weight, and generating a comprehensive timeliness index of the big data platform based on the dimension timeliness index; and inputting the comprehensive timeliness index into a job scheduler of the big data platform to monitor the scheduling state of the big data platform. According to the method and the device, the technical effects of prospective scheduling and resource allocation and global aging landslide caused by local delay are realized by quantifying the aging risk conduction dynamic factors of the data dependency relationship between the operations.
Owner:CHINA MERCHANTS BANK

Multi-actuator device with shared read channel

A multi-actuator storage device includes a first actuator supporting a first read element coupled to a shared read channel, a second actuator supporting a second read element coupled to the shared read channel, and a job scheduler that employs a lowest-access time selection methodology when selecting each next job to schedule on the first actuator. When the lowest-access-time methodology results in the selection of a read job for execution by the first actuator, the job scheduler identifies a pending job for execution by the second actuator that can be executed concurrent to the first read job without concurrently accessing the read channel and schedules the pending job on the second actuator concurrent to the first read job.
Owner:SEAGATE TECH LLC

Facilitating distributed job execution

Methods and systems are described herein for facilitating distributed job execution without a central job scheduler. The system may cause a container to, prior to executing job execution code for a job associated with a job data record, update a record instance of the job data record to indicate an updated status for the job and attempt to update the job data record at a database based on the record instance of the job data record. If the container successfully updates the job data record, the container may execute the job execution code for the job. If the container fails to update the job data record, the container may refrain from executing the job execution code for the job. The system may then update a first job data record associated with a first job at the database based on execution of the first job by a first container.
Owner:CAPITAL ONE SERVICES LLC

Multi-actuator device with shared read channel

A multi-actuator storage device includes a first actuator supporting a first read element coupled to a shared read channel, a second actuator supporting a second read element coupled to the shared read channel, and a job scheduler that employs a lowest-access time selection methodology when selecting each next job to schedule on the first actuator. When the lowest-access-time methodology results in the selection of a read job for execution by the first actuator, the job scheduler identifies a pending job for execution by the second actuator that can be executed concurrent to the first read job without concurrently accessing the read channel and schedules the pending job on the second actuator concurrent to the first read job.
Owner:SEAGATE TECH LLC

Runtime scheduler queue introspection

A system includes storage of job data describing a computing job in a job data location, creation of a queue entry associated with the computing job, the queue entry comprising a pointer to the job data location and a subset of the job data, storage of the queue entry in a job scheduler queue of a lock-free skiplist at a position based on a priority of the job, and reading of data from the queue entry and from a plurality of other queue entries of the job scheduler queue.
Owner:SAP SE

Background job processing framework

The described technology relates to scheduling jobs of a plurality of types in an enterprise web application. A processing system configures a job database having a plurality of job entries, and concurrently executes a plurality of job schedulers independently of each other. Each job scheduler is configured to schedule for execution jobs in the jobs database that are of a type different from types of jobs others of the plurality of job schedulers are configured to schedule. The processing system also causes performance of jobs scheduled for execution by any of the plurality of schedulers. Method and computer readable medium embodiments are also provided.
Owner:NASDAQ INC

System and method for network service enterprise computer group permissioning

A system and method may manage a computer by executing, by a job scheduler on the computer, a framework script written in a scripting language which discovers installed sub-task scripts and executes discovered sub-task scripts. Sub-task scripts may be stored in directories corresponding to accounts in which a sub-task script is to be run. Sub-task scripts may perform maintenance or monitoring tasks. A system and method may assign resource access or permissions for a set of computers, by, for each computer of a set of computers, determining the organizational unit and sub-organizational unit of the computer; and for each computer of the set of computers assigning the computer to a security group, based on the organizational unit and sub-organizational unit of the computer. For each computer, the security group the computer is assigned to may be associated with permissions enforced for each computer in the group.
Owner:MORGAN STANLEY SERVICES GROUP INC

A job scheduling and resource allocation system in a high-performance computing cluster

The application discloses a kind of job scheduling and resource allocation system in high-performance computing cluster, it is related to high-performance computing technical field, the system includes job scheduler, resource manager and newly added job DAG parser, cluster data distribution monitor, cache preheating priority calculator, collaborative decision engine, distributed cache preheating executor and performance optimization module, job DAG parser constructs job dependent directed acyclic graph and analyzes critical path, cache preheating priority calculator determines data preheating priority using dynamic priority drift algorithm, collaborative decision engine coordinates the resource allocation of scheduling and preheating, and performance optimization module dynamically adjusts system parameters, the system realizes job scheduling and cache preheating depth collaboration, reduces job access delay, improves cluster computing efficiency and resource utilization, and is suitable for processing complex job workflow high-performance computing cluster.
Owner:SHANXI WENHUI WEIYE TECHNOLOGY CO LTD

Facilitating distributed job execution

Methods and systems are described herein for facilitating distributed job execution without a central job scheduler. The system may cause a container to, prior to executing job execution code for a job associated with a job data record, update a record instance of the job data record to indicate an updated status for the job and attempt to update the job data record at a database based on the record instance of the job data record. If the container successfully updates the job data record, the container may execute the job execution code for the job. If the container fails to update the job data record, the container may refrain from executing the job execution code for the job. The system may then update a first job data record associated with a first job at the database based on execution of the first job by a first container.
Owner:CAPITAL ONE SERVICES LLC

Job scheduling in cloud systems to reduce resource fragmentation

Methods, systems, and computer-readable storage media for a job scheduler system that determines a balance matrix for each group of N jobs to N job workers and a bipartite graph is generated using the balance matrix. The bipartite graph is processed using a maximum matching algorithm to ensure that the N jobs match the N job workers with a maximum value of the sum of a degree of balance of remaining resources of each of the N job workers. Each job is assigned to a respective job worker using the result of the maximum matching algorithm.
Owner:SAP SE

Job scheduler for multi-tenant fairness

Techniques are described for determining whether to process a job request. An example, method can include a device receiving a first message from a first stream, the first message comprising a job request from a tenant and a tenant identifier. The device can detect a base number of units permissible to be processed for the tenant over a unit of time. The device can detect a processing speed of a downstream processor of an asynchronous pipeline. The device can detect a number of messages in a second stream, the downstream processor configured to receive messages from the second stream. The device can determine a target throughput and a historical throughput for the tenant. The device can compare the target throughput with the historical throughput to determine whether to process the job request. The device can schedule the job request for processing based at least in part on the comparison.
Owner:ORACLE INT CORP