Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1267 results about "Scheduling (computing)" patented technology

In computing, scheduling is the method by which work is assigned to resources that complete the work. The work may be virtual computation elements such as threads, processes or data flows, which are in turn scheduled onto hardware resources such as processors, network links or expansion cards.

Heterogeneous resource computing power intelligent scheduling method and system

The invention relates to the technical field of computing power scheduling, and discloses a heterogeneous resource computing power intelligent scheduling method and system. According to the method, real-time state monitoring is conducted on heterogeneous computing resources, and resource state parameters such as the computing unit utilization rate and the memory occupancy rate are obtained; task attributes and user request parameters of the task queue are collected, historical task data are processed based on the genetic algorithm optimization model to execute task demand prediction, and predicted demand parameters are generated. A dependency graph containing resource unit nodes and communication link roadsides is constructed through a resource topology analysis tool, predicted demand parameters are input into a scheduling priority classifier trained by a graph neural network, and an actual scheduling priority is identified. And executing resource conflict prediction based on the priority, inputting task feature vectors into a conflict resolution module of a fuzzy logic decision maker, outputting actual conflict resolution parameters, and finally integrating to generate a scheduling scheme containing a resource allocation sequence and an execution time table.
Owner:BEIJING WEICHENG TECHNOLOGY CO LTD

Computing power scheduling method and system based on dynamic load prediction and resource priority ranking

The invention discloses a computing power scheduling method and system based on dynamic load prediction and resource priority ranking. The computing power scheduling method comprises the following steps: collecting historical load data, task submission data and resource state data of each node in a computing power cluster; on the basis of the preprocessed multi-dimensional load feature data set, constructing an improved hybrid prediction model, optimizing model parameters through training, and predicting the load change trend of each computing power node in a future preset time period by using the trained model to obtain a node load prediction result; extracting a service level protocol parameter, a resource demand type and historical execution efficiency data of a to-be-scheduled task, and establishing a multi-dimensional resource priority evaluation index system; according to the computing power scheduling method, the problems of low resource utilization rate and high task response delay caused by low load prediction precision and mismatching of resource allocation and task priority in a traditional computing power scheduling method are solved, and the overall operation efficiency and service quality of a computing power cluster are improved.
Owner:SHAOGUAN DATA IND RESEARCH INSTITUTE

Large language model reasoning calculation service energy consumption optimization scheduling method based on task length prediction

The invention discloses a task length prediction-based large language model reasoning calculation service energy consumption optimization scheduling method, which comprises the following steps of: firstly, reasoning by taking an Alpaca-52k instruction data set as an input source and a large language model (such as Llama3-8B), counting the number of output tokens of the large language model, and labeling each piece of input data; then, a Qwen2-1. 5B large language model is finely tuned by using an Alpaca-52k instruction data set and the response length of Llama3-8B reasoning, and computing resources of the prediction method are reduced on the premise that the prediction performance is guaranteed; then, the response length of the Llama3-8B reasoning task is predicted through the fine-tuned Qwen2-1. 5B model, task balanced sorting scheduling is carried out according to the response length so as to improve the large language model reasoning speed, and finally, a deep reinforcement learning power selection algorithm is used to reduce the calculation power as much as possible on the premise that the large language model reasoning task time delay is met so as to improve the large language model reasoning efficiency. Therefore, the energy consumption of large language model reasoning calculation is reduced.
Owner:SOUTHEAST UNIV

Solid state disk management system and data processing method

The invention discloses a solid state disk management system and a data processing method, and relates to the technical field of solid state disk management. Flexible task scheduling and resource management are provided through a lightweight operating system module, and a computing task program defined by a user is supported to be dynamically loaded; edge computing or machine learning operators are efficiently executed in combination with a programmable hardware processing unit and a DMA channel of the computing acceleration engine module, the data preloading module is utilized to predict and preload data to DDR based on LBA access history, access delay is reduced, an NVMe protocol is expanded by means of the task unloading interface module, host task issuing and result returning are achieved, and the data processing efficiency is improved. And hardware-level memory protection is ensured through the security isolation module, so that the data calculation processing capacity of the solid state disk is remarkably improved, localized calculation tasks such as edge calculation and machine learning are supported, mass data transmission is effectively reduced, and bus and network loads are relieved.
Owner:HUIJU ELECTRONICS (DONGGUAN) IND CO LTD

Interaction method and platform system based on artificial intelligence equipment

The invention provides an interaction method based on artificial intelligence equipment, belongs to the technical field of artificial intelligence, and remarkably optimizes the efficiency of an AI equipment interaction system through distributed resource dynamic scheduling, multi-framework heterogeneous collaboration and full-link intelligent monitoring. Based on an elastic capacity expansion and contraction and priority scheduling strategy of Kubernetes, GPU resource islands are eliminated, and the computing power distribution efficiency is improved; a Flink checkpoint mechanism and Spark RDD blood relationship tracking are integrated, training task breakpoint continuous calculation and millisecond-level fault recovery are achieved, and repeated data processing is avoided; gPU node performance indexes are collected in real time through Prometheus Exporter, and in combination with Grafana visual early warning, the anomaly detection response speed is shortened to the second level; seamless access of frames such as TensorFlow / PyTorch is supported, heterogeneous device protocol conversion is achieved through a unified API gateway, and the cross-platform migration cost is reduced.
Owner:ZHONGKE NUOXIN BEIJING HI TECH

Energy-efficient task scheduling method for edge computing system

The invention discloses an energy-efficient task scheduling method for an edge computing system, which comprises the following steps of: acquiring a task state, a computing node resource state, a link state and an energy consumption state according to a unified time slot under a computing power network control domain, and constructing a system state vector; performing priority evaluation on the to-be-scheduled task based on the residual delay budget, the candidate node energy efficiency coefficient and the queue position to obtain a target task set; inputting a system state vector and a target task set into an energy efficiency perception deep reinforcement learning scheduling model, outputting a task-node allocation decision under the constraint of computing node resources and task time delay, and introducing a system-level energy consumption ratio, self-adaptive energy consumption penalty and exploration bias facing high-energy-efficiency nodes into rewards; and scheduling tasks according to the allocation decision, recording state transition and instant rewards, updating a double-commentator and actor network, and performing iterative execution in continuous time slots. According to the method, task success rate, time delay, load balancing and energy-saving performance are considered, and system energy consumption is reduced.
Owner:JIANGSU MARITIME INST +2

Large model reasoning acceleration method, device and equipment

The invention provides a large model reasoning acceleration method, device and equipment, a computing service device is configured with a plurality of computing nodes, stores a prefix index structure and is used for indicating a mapping relation between a token sequence prefix and the computing nodes storing cache computing results of the token sequence prefix; the method comprises the following steps: executing a global scheduling mechanism, querying the prefix index structure based on a token sequence of a reasoning request to perform prefix matching so as to determine one or more candidate computing nodes, and selecting a target computing node from the candidate computing nodes according to a real-time load state; executing a local scheduling mechanism, and allocating an execution priority for the reasoning request to perform scheduling processing according to the prefix matching degree of the reasoning request on the target computing node; and loading a cache calculation result corresponding to the matched prefix, and only calling a large model for a non-prefix part of the reasoning request to execute reasoning calculation.
Owner:ZHEJIANG LAB

Task scheduling method applicable to large language model (LLM), and computing device, storage medium and computer program product

Provided in the present disclosure are a task scheduling method applicable to a large language model (LLM), and a computing device, a storage medium and a computer program product. The method comprises: when determining that a target task processing node currently executing a target task meets a resource reporting condition, a target second task scheduler sending a target running resource of the target task processing node to a first task scheduler; when determining that the target running resource meets a task migration condition, the first task scheduler determining a migration task processing node from among a plurality of task processing nodes, and sending node information of the migration task processing node to the target second task scheduler; and the target second task scheduler determining the migration task processing node on the basis of the node information, determining the target task currently executed by the target task processing node, as well as task data of the target task, and migrating the target task and the task data in batches to the migration task processing node for execution, so as to obtain a task processing result of the target task.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Parallel task scheduling algorithm for heterogeneous multi-core processor

The invention relates to the technical field of computer architecture and parallel computing, and discloses a parallel task scheduling algorithm for a heterogeneous multi-core processor, which comprises the steps of task modeling, resource mapping, dynamic load balancing, communication optimization, task scheduling decision and execution monitoring. Task allocation is adjusted in real time through dynamic load balancing, cross-core communication delay is reduced in combination with communication optimization, and an efficient task allocation sequence is generated by using an improved genetic algorithm. According to the method, the resource utilization rate and the task execution efficiency of the heterogeneous multi-core processor in a high-performance computing scene can be improved, meanwhile, the robustness and adaptability of an algorithm are enhanced, and the task allocation problem in a complex computing scene is effectively solved.
Owner:SUZHOU DUXUEKEZHENG INTELLIGENT TECH CO LTD

Resource awareness and task migration method and system for industrial edge node

The invention discloses a resource awareness and task migration method and system for industrial edge nodes. The method comprises the following steps: constructing an edge node resource dynamic monitoring system, and sensing, calculating, storing and network resource states in real time; establishing a node health degree evaluation model, and predicting a potential fault risk; designing an intelligent task migration decision-making mechanism based on a resource state; the guarantee of data consistency and service continuity in the task migration process is realized; and constructing a distributed task scheduling optimization framework. The system comprises a resource monitoring module, a health assessment module, a migration decision module, a data synchronization module and a scheduling optimization module. According to the method, the problem of unstable task execution caused by dynamic resource change in an industrial edge computing environment is solved, and intelligent resource management and efficient task migration of the edge nodes are realized.
Owner:XIAMEN SIGGANG ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Elastic computing power scheduling method and device based on space-time prediction and medium

The invention provides an elastic computing power scheduling method and device based on space-time prediction and a medium, and belongs to the technical field of cloud computing and edge computing. The method comprises the following steps: acquiring a historical load, a satellite cloud picture, meteorological data, a burst flow mark, an energy price and a green electricity proportion in real time; constructing a fusion prediction model which comprises a multi-modal feature coding module, a space-time attention fusion module and a prediction output layer; a multi-objective optimization function is constructed, cost, delay and resource fluctuation are minimized, and green rewards are maximized; solving a task allocation variable and a resource adjustment amount by adopting mixed integer programming; the actual green electricity utilization rate is monitored, and the green reward coefficient is dynamically adjusted to correct prediction and scheduling deviations. According to the method, the computing power demand fluctuation of different areas in the future time period can be predicted in advance, so that resource allocation is dynamically adjusted, resource idleness and waste are reduced, and the utilization rate of computing power resources is remarkably improved.
Owner:SHANDONG FUTURE NETWORK RES INST (PURPLE MOUNTAIN LAB IND INTERNET INNOVATION APPL BASE)

Intelligent edge computing cooperative processing system based on integrated circuit

The invention relates to the technical field of edge computing, and discloses an intelligent edge computing co-processing system based on an integrated circuit, which comprises a heterogeneous computing cluster module, a hardware computing unit set, a multi-core control processor based on RISC-V, a programmable pulse tensor computing array and a reconfigurable engine oriented to streaming processing, according to the intelligent edge computing cooperative processing system based on the integrated circuit, zero-delay data exchange is realized through silicon intermediate layer integration of the heterogeneous computing cluster module and a snakelike data channel, and traditional bus arbitration delay is eliminated through real-time operation code analysis and optimal optical communication path mapping of the hardware task routing matrix (TRF) module; the priority path distribution of the task scheduling subsystem is synchronously coordinated with the time-sensitive task, so that the collaborative efficiency of the computing unit is improved in multiple dimensions, the non-blocking transmission of the high-priority task is ensured, and the effect of enhancing the overall processing capability and response speed of the intelligent edge computing is achieved.
Owner:QIQIHAR QISAN MACHINE TOOL

Network card data local preprocessing system fused with edge computing

The invention discloses a network card data local preprocessing system fused with edge computing, and relates to the technical field of edge computing and artificial intelligence collaborative optimization. Comprising an edge computing unit, a hierarchical collaborative architecture, a model hot switching and generative fragmentation module, an intention recognition and adaptive scheduling module, a delay energy consumption optimization scheduling module, a CXL zero-copy sharing module, an edge computing unit integrated processor, an FPGA or ASIC and a neuromorphic computing unit. According to the method, an FPGA, an ASIC and a neuromorphic computing unit are integrated in an intelligent network card, microsecond-level dynamic connection reconfiguration and adaptive generative model fragmentation execution are realized through a reconfigurable Mesh interconnection matrix, an attention layer and a feed-forward layer of a Transform class model are fragmented and allocated to different computing units for parallel execution, and cross-card streamlined processing is realized in cooperation with a zero-copy shared memory. And the intention recognition module is deeply coupled with the model hot switching module, so that dynamic model switching and fragmentation strategy optimization based on service priorities and system loads are realized.
Owner:ZHUHAI SHININGDA TECH CO LTD

Heterogeneous GPU resource management scheduling method, computer device, medium and product

The invention discloses a heterogeneous GPU resource management scheduling method, a computer device, a medium and a product. The method comprises the following steps: acquiring computing power resources of each node, including a GPU model, a GPU video memory, a GPU number, a computing power segmentation scheme and a GPU use condition; automatically segmenting the heterogeneous GPU of each node according to the computing power segmentation scheme of each node to obtain resource segmentation information; obtaining an expected computing power and an expected video memory of a to-be-executed task; screening out target nodes meeting the expected computing power and the expected video memory according to a node analysis strategy; and according to the sub-resource analysis strategy, screening out sub-resources meeting the expected computing power and the expected video memory, recording the sub-resources as target sub-resources, and allocating the to-be-executed task to the target sub-resources. According to the method, GPU resource fragmentation is effectively reduced through global and node resource conjoint analysis scheduling, meanwhile, pooling management, dynamic configuration and intelligent scheduling of heterogeneous computing equipment can be achieved, and the resource utilization rate is effectively increased.
Owner:BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD

Method and system for testing operation stability of heterogeneous computing system

The invention relates to the field of computer systems, in particular to an operation stability testing method and system oriented to a heterogeneous computing system, and the method comprises the steps: constructing a system overall operation logic model comprising a resource topological structure, a task scheduling rule and a communication link matrix; identifying a key communication path according to the communication link matrix, analyzing a resource competition hotspot in combination with a task scheduling rule, and establishing an interference model library; calling a specific model instance in the interference model library according to a test target, configuring injection parameters, and dynamically applying interference in a system operation process; key performance indexes in the system operation process are collected in real time; based on the real-time monitoring data, a multi-index weighted scoring algorithm is adopted to calculate an overall stability coefficient, a robustness grade and a self-healing capability index of the system, and a stability evaluation result is generated; and generating a visual report according to the topological structure of the operation logic model, and supporting comparison and analysis of historical versions. Unified centralized management is realized, and test efficiency and evaluation accuracy are improved.
Owner:SHANDONG CHAOYUE DATA CONTROL ELECTRONICS CO LTD

Multi-data center computing power-electric power cooperative scheduling and demand response declaration capacity optimization method

The invention provides a multi-data center computing power-electric power cooperative scheduling and demand response declaration capacity optimization method, which comprises the following steps of: firstly, acquiring multi-dimensional historical time sequence data of a data center; a machine learning algorithm is used to predict the workload of each data center in each time period of a future single day in a future scheduling period and the power market price of the place where each data center is located, and then a space-time coupling-oriented multi-data center computing power-power cooperation model considering uncertainty is constructed; the collaborative model comprises a joint optimization framework of computing power scheduling and power scheduling, and takes actual profit maximization as a target, then the established collaborative model is converted into a mixed integer linear programming problem and solved, and the optimal declaration capacity of each data center participating in demand response is obtained; and each data center carries out scheduling according to the task load of space migration and time migration, the charging and discharging power of an energy storage system and the power generation amount of renewable energy sources, so that overall optimization of demand response declaration capacity of the multiple data centers is realized.
Owner:XIAMEN UNIV

Distributed computing power dynamic scheduling method, equipment and medium

The invention discloses a distributed computing power dynamic scheduling method and device and a medium, and the method comprises the steps: packaging heterogeneous computing power resources of a home terminal and an edge cloud node through a lightweight containerization technology, and collecting the hardware resource state data of the home terminal and the load data of the edge cloud node in real time; generating a dynamic scheduling strategy according to the task calculation type and the real-time network state of the to-be-calculated task, and splitting the to-be-calculated task into a plurality of sub-tasks according to the dynamic scheduling strategy; distributing the sub-tasks to home terminals or edge cloud nodes according with a dynamic scheduling strategy, and aggregating calculation results of the sub-tasks; and performing end-to-end encryption and fragmentation verification on the cross-domain transmitted data stream, and dynamically adjusting the task allocation permission of the home terminal based on the equipment security score.
Owner:INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD

Distributed training scheduling and communication optimization method and system of multi-modal large model on domestic computing power platform

The invention discloses a distributed training scheduling and communication optimization method and system of a multi-modal large model on a domestic computing power platform. The method comprises the following steps: virtualizing a heterogeneous computing unit of a preset platform into a virtual device pool, and fusing first-order gradient of a multi-modal sample and Hessian matrix information based on quantitative perception training to generate a sample sensitivity grading atlas; virtual device pool attributes and the sensitivity grading atlas are used as input, an optimal hybrid parallel configuration scheme is automatically generated through a configuration search algorithm, and a parallel combination mode, resource mapping and a high-sensitivity sample scheduling strategy are defined; a distributed training code of an integrated communication optimization strategy is automatically generated according to a configuration scheme, pipeline parallel communication and data parallel gradient synchronization constraint are executed in a topology adjacent equipment subset, and a hierarchical aggregation mechanism is adopted; and dynamically screening a core training set and scheduling a calculation task to complete distributed training. According to the method, efficient cooperative training of the multi-modal large model on the domestic computing power platform is realized.
Owner:GUANGXI POWER GRID CORP

Task scheduling method based on predictable resource state graph modeling

The invention discloses a task scheduling method based on predictable resource state atlas modeling, a platform is oriented to a heterogeneous computing environment, a unified resource state atlas is constructed by collecting multi-dimensional resource state parameters of computing nodes, and performance characteristics and communication topological relations among the nodes are comprehensively described. On the basis, a bidirectional time sequence model and an attention mechanism are fused, and the load trend of each node in a future short time is predicted. The platform constructs a multi-factor scheduling scoring function based on a task feature vector and resource state prediction map, integrates parameters such as resource matching degree, prediction load, communication delay and energy consumption cost, dynamically evaluates the adaptability of tasks and resources, and realizes adaptive scheduling and optimal resource allocation of the tasks. Compared with the prior art, the method has the advantages of being high in resource state predictability, high in task allocation intelligence degree, outstanding in platform evolution capability and the like, and is suitable for intelligent task scheduling application in a large-scale heterogeneous resource environment.
Owner:NANJING NORTH OPTICAL ELECTRONICS

Power consumption and performance balanced scheduling method and system for heterogeneous multi-core processor

The invention discloses a power consumption and performance balanced scheduling method and system for a heterogeneous multi-core processor, and belongs to the technical field of computers, and the method comprises the steps: obtaining historical load data and current system state parameters, generating a load sequence and carrying out load prediction analysis, generating a predicted load value for feature fusion and quantification, and generating a task feature vector; constructing a core capability portrait, calculating a matching degree between the task characteristic vector and the core capability portrait, and generating a scheduling strategy; and dynamically combining the physical cores into a logic calculation unit according to a scheduling strategy, configuring a power consumption state, and executing task allocation and migration. According to the method, the technical means of load prediction analysis, quantitative matching of tasks and cores and dynamic combination and configuration of physical cores are adopted, and pre-judgment and prospective distribution of system resources can be achieved, so that performance requirements are met, meanwhile, the system energy efficiency is effectively optimized, and deep balance of power consumption and performance of a processor is achieved.
Owner:BING TANG INTELLIGENT TECH (SHANGHAI) CO LTD

Computing power resource fusion method based on distributed flow pipeline

The invention relates to the technical field of distributed computing and computing power scheduling, in particular to a computing power resource fusion method based on a distributed flow pipeline, which comprises the following steps of: disassembling a user computing power request into a flow pipeline unit for packaging computing logic, an input / output interface and a resource demand label; meanwhile, computing power types, real-time load rates, memory occupancy rates, network round-trip delays, geographic positioning and energy consumption data of cloud edge end nodes are collected, a hierarchical topology network framework is constructed based on the collected data, node computing power available values are calculated, an inter-node transmission cost matrix is generated, a fault probability prediction model is constructed, and a fault probability prediction model is constructed. And finally, analyzing a data dependency relationship of the pipeline unit through a four-dimensional joint decision engine, executing dynamic mapping, and preferentially mapping the high-computing-power demand unit to a GPU cluster node, so as to realize non-interruption reconstruction during pipeline topology operation. And the global computing power resource utilization rate, the task operation efficiency and the service stability are improved.
Owner:LANZHOU YUNFAN ZHILIAN TECH CO LTD +2

System for dynamic scheduling and optimisation of diagnostic tasks

A system is provided for dynamic scheduling and optimisation of diagnostic tasks in a networked computing environment. The system associates issue tickets with a diagnostic task matrix comprising probable causes, diagnostic tasks, probability values, outcome expectations, and resource parameters. A task scheduling controller generates optimised task sequences based on task success likelihoods, cost, technician availability, and evidentiary sufficiency. As tasks are completed, outcomes are used to update the diagnostic model, enabling automatic self-improvement. A statistical learning model, such as aBayesian or neural network, refines diagnostic probabilities using historical data. Integration with calendaring systems allows real-time rescheduling based on personnel availability. A graphical interface supports live drag-and-drop reconfiguration of task associations, with immediate propagation of updates to task probabilities and cost metrics. The system thereby enhances resolution speed, accuracy, and resource efficiency across evolving operational contexts.
Owner:NOVUM GLOBAL GROUP PTY LTD

Computing power resource allocation method, system and product based on multi-dimensional dynamic evaluation

The invention relates to the technical field of computing power resource allocation, and particularly discloses a computing power resource allocation method and system based on multi-dimensional dynamic evaluation and a product. The method comprises the steps of obtaining evaluation index data of a to-be-scheduled task in multiple dimensions; dynamically configuring a weight value of each evaluation index according to a task attribute and a system state; the evaluation index data is fuzzified by using a preset membership function, a fuzzy evaluation matrix is constructed in combination with the weight value, and comprehensive calculation is performed through a fuzzy inference rule to obtain fuzzy comprehensive evaluation data of the task, so that an accurate evaluation result is generated; modeling a task execution process by adopting a multi-layer perceptron model, and predicting the execution performance of the task; and based on the evaluation result and the prediction result, dynamically selecting a proper strategy from a plurality of predefined task scheduling strategies, and scheduling the tasks to optimize computing power resource allocation. According to the method, the resource utilization efficiency is remarkably improved through multi-dimensional evaluation and a dynamic scheduling mechanism.
Owner:DIGITAL CHONGQING BIG DATA APPL DEV CO LTD

Storage and calculation separation method and device based on multi-level cache and intelligent scheduling, and server

The invention discloses a storage and calculation separation method and device based on multi-level cache and intelligent scheduling and a server, and belongs to the technical field of data processing, and the method comprises the following steps: collecting operation behavior data of a user, carrying out distributed storage, and dynamically distributing and calculating node cache space; performing classification marking to form marked user data; performing intelligent scheduling layer analysis tasks on the marked user data, and screening and distributing adaptive computing nodes; the computing node receives the analyzed task to check the local cache, obtains corresponding data from the local or request remote storage according to the validity of the local corresponding data, computes the corresponding data, and caches the computed data to the local or remote storage according to the user access probability; and carrying out dynamic scheduling and hierarchical storage on the calculated data between local cache and remote storage according to access frequency and heat factors. According to the invention, the problems of low transmission efficiency and unbalanced resource utilization of the existing storage and calculation separation architecture are solved.
Owner:SHENZHEN COOCAA NETWORK TECH CO LTD

Heterogeneous reasoning computing power scheduling method and device and storage medium

The invention provides a heterogeneous reasoning computing power scheduling method and device and a storage medium, and the method comprises the steps: receiving a reasoning request, splitting the reasoning request, and obtaining at least one reasoning task; according to the coding length of each reasoning task, scheduling the reasoning tasks to a long task queue or a short task queue; when it is monitored that reasoning tasks with to-be-processed task states exist in the long task queue and the short task queue, the working state of the reasoning service of the first processor is judged based on the number of tasks in processing of the first processor in the long task queue and the short task queue and the maximum concurrency of the first processor; the maximum concurrency of the first processor is obtained based on the current conditions of the long task queue and the short task queue; and on the basis of the working state of the first processor reasoning service, the reasoning task to be processed is scheduled to the first processor reasoning service or the second processor reasoning service so as to maximize and reasonably utilize the mixed computing power resource of the heterogeneous processor.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Deployment method, device and equipment of large model agent and medium

The invention discloses a deployment method and device of a large model agent, equipment and a medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: determining agent configuration information of a to-be-deployed large model agent after an agent deployment request is received, and determining a target large model and model metadata; determining an available resource condition of a preset computing resource pool, and performing query matching according to the available resource condition and the model metadata to determine a plurality of hardware-compatible candidate computing resource nodes; performing load prediction operation on the to-be-deployed large model agent after the large model agent is online to obtain a predicted load, and generating a target deployment decision scheme based on the predicted load, a preset scheduling strategy, the available resource condition and the candidate computing resource node; and adjusting the runtime parameters of the target large model according to the resource characteristics of the target computing resource node, and loading the adjusted target large model on the target computing resource node so as to deploy and start the agent instance of the to-be-deployed large model agent.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Heterogeneous computing power resource pooling scheduling platform

The invention provides a heterogeneous computing power resource pooling scheduling platform which comprises a task and resource tensor construction module, a scheduling decision module, a breakpoint state generation module and a rescheduling and recovery module. According to the platform, task requirements and node resources are modeled into a tensor structure in a unified mode, and accurate initial scheduling is achieved through a structured scoring function containing conditional penalty terms. In the task running process, interruption judgment is carried out based on the node dynamic state and the scheduling score, and a breakpoint state tensor containing an execution state is generated. When interruption occurs, a proper node is selected through compatibility screening, and task seamless migration and execution recovery are realized by utilizing the stored state. According to the method, the technical problems of non-uniform resource expression, lack of scene adaptability and incapability of realizing uninterrupted migration in a heterogeneous computing power environment are solved, and the resource utilization rate and the service continuity are remarkably improved.
Owner:SHANGYANG TECH CO LTD

Task scheduling method and device based on heterogeneous computing, storage medium and equipment

The invention relates to a heterogeneous computing-based task scheduling method and device, a storage medium and equipment. The method comprises the following steps of: obtaining an inference task of a large language model; in the execution process of the reasoning task, at least based on the memory access intensity and the current reasoning stage, memory access intensive operators and calculation intensive operators in the reasoning task are recognized; allocating the memory access intensive operator to a first computing unit for executing a memory intensive task, and allocating a computing intensive operator to a second computing unit for executing a computing intensive task; obtaining estimated execution time of the two calculation units for executing the corresponding tasks and data transmission time between the two calculation units; according to the pre-estimated execution time and the data transmission time, the task starting moments of the first calculation unit and the second calculation unit are determined with the purpose of minimizing the overall execution delay of the reasoning task; and controlling the first calculation unit and the second calculation unit to asynchronously execute the corresponding tasks in parallel according to the task starting time.
Owner:GUANGDONG UCAP INTERNET INFORMATION TECH

Distributed component task scheduling strategy supporting cross-platform collaboration

The invention discloses a distributed component task scheduling strategy supporting cross-platform collaboration. The distributed component task scheduling strategy comprises the following steps: step 1, collecting computing power, load conditions and network states of different platforms; step 2, carrying out cleaning and standardization treatment; step 3, optimizing task scheduling by using a deep Q network algorithm; 4, deploying the task scheduling model to a monitoring system; 5, adjusting rules according to the priorities and requirements of the tasks and platform load conditions, and dynamically adjusting a task scheduling strategy; step 6, performing anomaly detection on the task scheduling behavior by using a rule engine; 7, regularly analyzing task scheduling results, and optimizing the scheduling model by combining with newly added data; and step 8, optimizing a platform resource allocation strategy by adjusting scheduling algorithm parameters, and updating the model regularly. Through real-time resource perception, dynamic task allocation and an exception handling mechanism, efficient cooperation across multiple heterogeneous platforms is realized, and the resource utilization rate and the task execution efficiency are improved.
Owner:CHENGDU HAIQING TECH CO LTD

Intelligent computing power scheduling method and system for heterogeneous computing power cluster

The invention relates to the technical field of computing power intelligent scheduling, and discloses a computing power intelligent scheduling method and system for a heterogeneous computing power cluster, and the method comprises the steps: firstly analyzing task features from task description submitted by a user, and generating an internal task object; and then, constructing a resource portrait by collecting real-time resource monitoring data and static configuration information of each node in the heterogeneous computing power cluster. Based on the information, the execution performance of different tasks on each node is predicted by using a machine learning model, and a performance prediction mapping table is formed. Then, candidate resources are screened and sorted according to the mapping table, and it is ensured that the optimal node is selected to be bound with the task; therefore, not only is the matching degree of the task characteristics and the hardware attributes considered, but also the dynamic state of the hardware resources is combined, so that a more accurate task scheduling strategy is realized, and the overall operation efficiency and the resource utilization rate are effectively improved.
Owner:SHANGHAI YUANLU JIAJIA INFORMATION SCI & TECH CO LTD