Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

145 results about "Heterogeneous cluster" patented technology

Anatomical cluster which consists of all or some members of one or more organ subclasses and one or more organ part subclasses which are grouped together according to some shared attributes. Examples: joint, internal ear, pharynx.

Reinforcement learning calculation simulation method and device, electronic equipment and storage medium

The invention discloses a reinforcement learning calculation simulation method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence computing, and the method comprises the steps: inputting the determined current model parameter configuration, current hardware configuration and current working load into a target simulation system to obtain a plurality of parallel grouping combinations, determining a target simulation system according to the current hardware configuration, determining an effective parallel packet combination from the plurality of parallel packet combinations based on a preset Monte Carlo method, inputting the effective parallel packet combination into a simulator of a preset neural network model, and performing delay time calculation according to the effective parallel packet combination through the simulator to obtain a delay time sequence; and the combination corresponding to the shortest delay time is used as a target parallel grouping combination, so that the technical problems of mismatching of simulation scenes, insufficient precision and lack of effective support for heterogeneous clusters are solved, reliable performance prediction and optimal parallel strategy suggestions are provided through high-precision performance modeling and automatic exploration, and the method is suitable for large-scale popularization and application. Therefore, the resource consumption of large-scale GRPO training is reduced.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Ground-air heterogeneous cluster formation control method and system based on adjustable funnel algorithm

The invention discloses a ground-to-ground heterogeneous cluster formation control method and system based on an adjustable funnel algorithm, and relates to the technical field of distributed formation maneuvering control, and the method comprises the steps: building a dynamic model of a ground-to-ground heterogeneous cluster, and building an undirected graph representing an interaction topological structure of the ground-to-ground heterogeneous cluster; modeling a distributed safety critical formation maneuvering control problem of the air-ground heterogeneous cluster; a fully distributed dynamic compensator facing each follower and a formation maneuvering controller based on an adjustable funnel algorithm are designed, and then a safety-critical control framework based on a disturbance observer is adopted to construct a composite controller; and acquiring a virtual navigator signal and inputting the virtual navigator signal to the constructed various controllers to obtain an actual formation control signal of each follower, and executing distributed safety-critical formation maneuvering control on the followers. Through a distributed safety critical formation maneuvering strategy based on an adjustable funnel method, collision / obstacle avoidance and disturbance suppression of an air-ground heterogeneous cluster can be realized.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Heterogeneous cluster-oriented resource allocation method and device and storage medium

The invention relates to a heterogeneous cluster-oriented resource allocation method and device and a storage medium, and the method comprises the steps: executing a plurality of iterations of an initial model on a single node of each heterogeneous cluster, and obtaining the performance characteristics of the model; generating a plurality of parallel strategy combinations comprising the first parallelism degree and the second parallelism degree; generating a plurality of assembly line parallel combinations based on the first degree of parallelism, calculating the load imbalance rate of each assembly line parallel combination, and screening out N combinations with the minimum load imbalance rate from the assembly line parallel combinations; traversing the second parallelism degree and the micro-processing batch in each stage corresponding to each parallel combination to generate a plurality of stage strategy combinations; and calculating video memory occupation and execution time under each stage strategy combination, screening out K stage strategy combinations of which the video memories are smaller than or equal to a video memory threshold value and the execution time is shortest, and carrying out resource allocation on the heterogeneous cluster based on the stage strategy combinations. According to the invention, the problems of low training efficiency and low resource utilization rate are solved.
Owner:ZHEJIANG LAB

Heterogeneous cluster hybrid attack defense control method and system oriented to urban confrontation environment, terminal equipment and medium

The invention discloses a hybrid attack defense control method and system for a heterogeneous cluster in an urban confrontation environment, terminal equipment and a medium, and relates to the technical field of unmanned platform cluster control, and the method comprises the steps: constructing an ideal system model of an unmanned platform cluster comprising a leader and a plurality of followers, constructing an information physical hybrid attack model based on the model; using the attack model to simulate an attack behavior of an attacker on an ideal system to obtain an attacked system model; based on an attacked system model, constructing a distributed elastic security estimator with a compensation mechanism, compensating attack influence for each follower and estimating an expected position; and based on the expected position of the follower, constructing a distributed elastic safety controller with a compensation mechanism, and controlling the follower. By designing the estimator and the controller, attack interference is counteracted, control input is generated, and safe collaboration of the heterogeneous nonlinear unmanned cluster under the cyber-physical hybrid attack is guaranteed.
Owner:BEIJING INST OF TECH

A system for energy-conscious LLM-based workflow keying with dynamic resource allocation

An energy-conscious workflow planning system based on LLM with dynamic resource allocation, consisting of: a workflow input interface configured to receive workflow-directed acyclic graphs (DAGs), energy budget constraints, system performance constraints, and natural language requests from human operators; a large language model (LLM) logic module connected to the workflow input interface and configured to analyze the workflow specifications and system constraints in natural language, generate energy-conscious planning recommendations based on the analyzed workflow specifications, and provide explainable planning rationales in natural language; a reinforcement learning-based scheduling unit connected to the LLM reasoning module and configured to: receive scheduling recommendations from the LLM reasoning agent, fine-tune task-resource assignments by dynamically adapting to runtime variations, and perform online resource redistribution under runtime variability; an energy monitoring unit configured to: continuously monitor CPU and GPU utilization in heterogeneous clusters, track power consumption and thermal limits per node, and generate energy profiles for system components; a multi-objective optimization engine configured to: perform a Pareto-optimal scheduling analysis that balances energy consumption, lead time and reliability, apply statistical and AI-supported trade-off analyses and ensure optimal resource allocation based on Pareto frontier analysis; a dynamic resource allocation unit configured to: use predictive models that incorporate LLM inferences and feedback from reinforcement learning, reassign tasks between nodes and clusters while minimizing energy consumption and improving system throughput based on the predictive models; a performance optimization module configured to optimize scheduling decisions using multi-criteria optimization analysis; and a user interface that allows human operators to override and refine planning strategies in real time based on verifiable planning reasons.
Owner:BENEDICT SHAJULIN DR KANYAKUMARI +1

Depth learning task node allocation method and system for executing time-aware computing power network heterogeneous GPU (Graphics Processing Unit) cluster

The invention discloses an execution time aware computing power network heterogeneous GPU cluster deep learning task node allocation method and system. The method comprises the following steps: firstly, based on a deep learning task, extracting and preprocessing task features and available node features; secondly, a sampler equally divides new tasks without historical data to available nodes, and each node performs mixed sampling on the tasks until all the tasks estimate execution time data; taking execution time data as a training set, taking the task features and the node features as a test set, and using a regression decision tree model to predict the execution time of the task on each node; performing task allocation on each node by using a cost search algorithm and a short job total JCT priority strategy; and finally, periodically monitoring node resources released in the cluster to obtain an optimal node allocation result. According to the method, task delay and total task JCT are remarkably reduced, cluster node resource changes are monitored in real time, and the resource utilization rate is increased.
Owner:HANGZHOU DIANZI UNIV +1

Rapid system recovery method and system based on system mirror image management

The invention discloses a system quick recovery method and system based on system mirror image management, and the method comprises the steps: initializing a boot partition of a functional blade in a heterogeneous cluster system and a boot flag bit of the boot partition, the boot flag bit being used for BIOS reading to select a corresponding boot partition when the functional blade is started; the method comprises the following steps: triggering a job operating system or a boot operating system of a target function blade to restart through a soft shutdown interface, and modifying a start boot flag bit in a BMC (Baseboard Management Controller) of the target function blade, so that the target function blade is sequentially restarted and switched into the boot operating system from the job operating system; and executing system quick recovery or backup for the target system partition under the boot operating system, and then restarting and switching to run under the specified job operating system so as to complete system quick recovery or backup. According to the method, the characteristics of the heterogeneous cluster system are combined, and automatic and flexible system mirror image fast rollback recovery is achieved when the heterogeneous cluster system has a serious fault.
Owner:NAT UNIV OF DEFENSE TECH

Heterogeneous cluster-oriented large model module-level training strategy optimization method and device and computer equipment

The invention relates to a large model module level training strategy optimization method and device for a heterogeneous cluster and computer equipment, and the method comprises the steps: obtaining the calculation information of a large model module level and the communication information of an operator level for the heterogeneous cluster; the calculation information comprises calculation time information and module video memory information of modules under different data scales under different distributed strategies and on different chips; the communication information is communication time delays of communication operators under different data scales; determining an initial distributed training strategy of each assembly line based on the module video memory information; based on the calculation time information and the communication information, calculation time of different stages of the assembly line is determined; and optimizing each initial distributed training strategy according to the calculation time and a video memory threshold value for bearing the large model equipment to obtain a target distributed training strategy. Through the method and the device, the problem of relatively low resource utilization rate of a large model oriented to heterogeneous clusters is solved.
Owner:ZHEJIANG LAB

Image processing unit (GPU) cluster communication method and electronic device

The embodiment of the invention provides an image processing unit (GPU) cluster communication method and an electronic device, and the method comprises the steps: splitting a GPU cluster, and obtaining at least one algorithm isomorphic group and at least one algorithm heterogeneous group; the protocol distribution communication and the full collection communication are performed through the isomorphic GPUs in each algorithm isomorphic group, and the full protocol communication is performed through the heterogeneous GPUs in each algorithm heterogeneous group. Therefore, through the embodiment of the invention, the problems that the communication traffic between heterogeneous GPUs is increased and extra data copy overhead is introduced due to the difference of communication libraries between different GPUs in a heterogeneous GPU cluster environment can be solved.
Owner:ZTE CORP

Asymmetric segmentation scheduling system and method in heterogeneous GPU cluster

The invention provides an asymmetric segmentation scheduling system and method in a heterogeneous GPU cluster, and the method comprises the steps: S1, checking the features of an incoming request, and grouping the incoming request into different request buckets according to the token length; s2, in a heterogeneous GPU cluster environment, optimizing a large language model reasoning instance by adopting a double-layer strategy; and S3, calculating the matching degree of the request buckets and the big language model reasoning instances, and scheduling different request buckets to the big language model reasoning instance with the highest matching degree according to a calculation result. According to the method provided by the invention, the model layer can be asymmetrically segmented according to the computing power and the video memory capacity of each GPU on the premise of satisfying the model parallelism degree constraint, the dynamic balance of the execution duration between stages is realized, assembly line cavitation bubbles are fundamentally reduced, and the overall throughput rate is improved.
Owner:SHANGHAI JIAOTONG UNIV

Agricultural unmanned aerial vehicle heterogeneous cluster cooperative task allocation and path planning method

The invention relates to a heterogeneous cluster cooperative task allocation and path planning method for an agricultural unmanned aerial vehicle, and the method comprises the following steps: S1, carrying out the system modeling and initialization, and carrying out the modeling of a cooperative data collection task of an unmanned aerial vehicle mechanism cluster in an agricultural environment as a decentralized partially observable Markov decision process; s2, defining a multi-dimensional space, wherein the decentralized part observes a global state space, a joint action space and a local observation space of a Markov decision process; s3, designing a mixed reward function; s4, performing cluster strategy training; and S5, performing online distributed execution. The method effectively solves the problems that an existing method is poor in environment dynamic change adaptability, does not fully consider the heterogeneous characteristics of the unmanned aerial vehicles and does not fully consider decision-making and execution separation, and remarkably improves the operation efficiency, safety and collaborative intelligence level of an agricultural unmanned aerial vehicle cluster in a complex unstructured environment.
Owner:QINGDAO AGRI UNIV

Large model reasoning method for operator-level distributed scheduling

The invention discloses a large model reasoning method for operator-level distributed scheduling. The large model reasoning method comprises the steps of S1, operator decoupling and dependency modeling; s2, carrying out heterogeneous hardware capability portraying; and S3, dynamically adjusting the strategy. Operator-level fine-grained scheduling is realized in large model reasoning, and the computing power utilization rate of the heterogeneous GPU cluster is maximized.
Owner:BEIJING MOMENT UNLIMITED TECHNOLOGY CO LTD

Machine learning workload orchestration in heterogeneous clusters

Systems and methods are described herein to orchestrate the execution of an application, such as a machine learning or artificial intelligence application, using distributed compute clusters with heterogeneous compute resources. A discovery subsystem may identify the different compute resources of each compute cluster. The application is divided into a plurality of workloads with each workload associated with resource demands corresponding to the compute resources of one of the compute clusters. Adaptive modeling allows for hyperparameters to be defined for each workload based on the compute resources associated with the compute cluster to which each respective workload is assigned and the associated dataset.
Owner:HEWLETT PACKARD DEVELOPMENT COMPANY LP

GPU cluster resource allocation method and device, equipment and storage medium

The invention discloses a GPU cluster resource allocation method and device, equipment and a storage medium, and relates to the technical field of GPU cluster resource allocation. According to the method, a hardware performance description vector is generated by fusing static hardware parameters and dynamic micro-benchmark test data, the limitation that the GPU performance is represented only by depending on the static parameters is broken through, and the actual operation capacity of different architecture GPUs in a heterogeneous cluster is accurately matched; task feature vectors are generated by extracting task calculation features and resource demands, and quantitative description of reasoning task resource consumption features is achieved; information of hardware, tasks and load dimensions is integrated through a machine learning model, performance degradation characteristics during multi-task parallel are effectively captured, and the accuracy of execution time prediction is improved; scheduling decision operation including candidate node screening, comprehensive cost evaluation and resource reservation backfilling is executed in combination with the predicted execution time, and the task execution efficiency and cluster load balancing are both considered.
Owner:HANGZHOU DIANZI UNIV +2

Recalculation training strategy generation method, electronic equipment and computer program product

The embodiment of the invention is suitable for the technical field of artificial intelligence, and provides a re-calculation training strategy generation method, electronic equipment and a computer program.The method is applied to a heterogeneous cluster and comprises the steps that cluster information of the heterogeneous cluster and model feature information of a to-be-trained model are determined; for each storage and calculation resource in the heterogeneous cluster, according to the cluster information and the model feature information, re-calculation time information and re-calculation performance requirements are determined; determining a re-calculation time information characteristic value under the condition that the re-calculation performance demand is not greater than the performance limit value of the storage and calculation resources; and determining a recalculation training strategy for the to-be-trained model according to the recalculation time information feature value corresponding to each storage and calculation resource. According to the embodiment of the invention, the adaptive re-calculation training strategy can be determined for each storage and calculation resource of the heterogeneous cluster, and the model training efficiency and throughput are improved by adopting the re-calculation training strategy corresponding to each storage and calculation resource to carry out model training.
Owner:HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY

Complex construction site construction equipment control system and method

The invention provides a complex construction site construction equipment control system and method, and the system constructs a basic state semantic model of each type of equipment, and covers all perception and control instruction data; the comprehensive management and control platform dynamically extracts and issues one or more data packet formats from the basic model according to construction scenes, areas and equipment types, and fine adjustment of sensing information and instruction formats is achieved; the edge computing device is deployed at an equipment end or a cluster end and is responsible for dynamic mapping analysis, redundancy removal, priority screening and unified protocol packaging of a multi-heterogeneous protocol; the equipment end reports a null instruction sensing data packet, the comprehensive management and control platform generates an optimal control instruction, compresses recent sensing data and then issues the compressed recent sensing data in the same data packet format; the equipment end decides whether to execute the instruction or not, if not, problem expression is fed back, and a safe closed loop is formed; in a cluster scene, gathering fusion and instruction distribution of multi-level edge computing devices are supported, and efficient collaboration of heterogeneous clusters is achieved.
Owner:SHANGHAI CONSTRUCTION GROUP CO LTD

Method and system for simplifying large model reasoning service deployment in heterogeneous cluster environment

The invention discloses a method and a system for simplifying large model reasoning service deployment in a heterogeneous cluster environment, mainly relates to the technical field of service deployment, and is used for solving the problems of repeatability cost of heterogeneous adaptation, dependence on manual allocation of containers to specific nodes during containerization deployment and version library compatibility in an existing scheme. And crash is caused, and dynamic resource allocation based on load characteristics is lacked. Comprising the steps of performing preset mirror image condition matching on a node hardware feature tag and a model service demand tag of an implementation node and a pre-built reasoning service container mirror image in a mirror image warehouse; when the pre-constructed reasoning service container mirror image with the matching degree of 100% does not exist, dynamically constructing a customized reasoning service mirror image, and storing the customized mirror image in a mirror image warehouse as the pre-constructed reasoning service container mirror image for implementing node matching; and pulling a pre-constructed reasoning service container mirror image for implementing node matching from the mirror image warehouse, mounting a standardized model file mirror image, and starting a reasoning service container.
Owner:南京中孚信息技术有限公司

Resource determination method and apparatus, and program product and storage medium

Disclosed in the embodiments of the present application are a resource determination method and apparatus, and a program product and a storage medium. The method in the embodiments of the present application comprises: receiving an inference request; determining a first computing resource required for when a large model performs inference on the inference request; acquiring first resource information of a heterogeneous cluster; and on the basis of the first computing resource and the first resource information, determining a first inference resource from the heterogeneous cluster, wherein the first inference resource is used by the large model to perform inference on the inference request. In this way, for an inference request received in real time, a computing resource required for performing inference on the inference request is determined in real time, which enables dynamic determination of computing resources corresponding to inference requests of different lengths. Thus, an inference resource corresponding to the inference request can be dynamically determined on the basis of the computing resource, and the corresponding inference resource can be more accurately determined from the heterogeneous cluster, so that inference is performed on the inference request on the basis of the inference resource, such that an inference process of the large model can avoid being affected.
Owner:HUAWEI TECH CO LTD

Heterogeneous cluster parallel multi-stage fuzzy job scheduling method

The invention discloses a heterogeneous cluster parallel multi-stage fuzzy job scheduling method, which comprises the following steps of: receiving a processing request which is submitted by a user and contains multi-stage jobs, and modeling task execution time, data transmission time and request deadline into triangular fuzzy numbers, constructing a scheduling model taking the minimization of the total lease cost and the minimization of the request tardiness as targets; carrying out global exploration by adopting an improved second-generation non-dominated sorting genetic algorithm to generate parent and offspring populations; based on population similarity threshold judgment, dynamically triggering local search guided by a multi-agent near-end strategy optimization algorithm, and adaptively selecting a local search operator for each individual to generate an adjacent population; combining various populations, performing non-dominated sorting and crowding distance calculation, and screening out a new generation of populations; and iterating the process until convergence, and outputting a Pareto optimal solution set. According to the method, the problem of multi-target scheduling with fuzzy time variables in a heterogeneous environment is effectively solved, and the quality of a scheduling scheme and the algorithm search efficiency are improved.
Owner:GUANGDONG UNIV OF TECH

Efficient collective communication method and device for heterogeneous GPU cluster

The invention discloses a heterogeneous GPU cluster-oriented efficient collective communication method and device, and the method comprises the steps: obtaining a target communication operator set and a physical topology corresponding to a heterogeneous GPU cluster, and determining an initial state constraint and a target state constraint; the scheduling time cost of each communication scheme is determined, and the communication scheme corresponding to the minimum scheduling time cost in the multiple scheduling time costs is selected as the efficient collective communication mode of the heterogeneous GPU cluster. According to the method, the communication scheme corresponding to the minimum scheduling time cost of the heterogeneous GPU cluster is determined to serve as the efficient collective communication mode of the heterogeneous GPU cluster based on the actual conditions of topological structure difference, link bandwidth imbalance and the like in the heterogeneous GPU cluster environment; it is ensured that the communication process is not limited by the bandwidth bottleneck, and system resource utilization efficiency and parallel training performance are prevented from being limited.
Owner:NORTHEASTERN UNIV CHINA

Heterogeneous cluster hybrid parallel training method and system combined with freezing mechanism

The invention discloses a heterogeneous cluster hybrid parallel training method and system combined with a freezing mechanism, and belongs to the technical field of neural networks. According to the method, in combination with the characteristics of a target heterogeneous hardware cluster, the calculation time and storage requirements of different hardware of different types of hardware of each layer in a model are predicted; solving the optimal assembly line segmentation and data parallel configuration by adopting dynamic programming to obtain a hybrid parallel training scheme; and when the frozen state changes in the training process, the system reallocates the released video memory and computing power resources, and dynamically adjusts the parallel scheme to realize load balancing and improve the training efficiency. According to the method, resource occupation is effectively reduced, and the large-scale model training performance in a heterogeneous environment is improved.
Owner:SOUTHEAST UNIV

Tail delay optimization job scheduling method and system based on heterogeneous GPU cluster

The invention relates to the technical field of computers, and discloses a tail delay optimization job scheduling method and system based on a heterogeneous GPU cluster. The method comprises the following steps: constructing a cluster physical topological graph containing link delay, bandwidth and hop count; predicting a job communication demand graph through static code analysis and a graph neural network; executing topology matching scheduling, and preferentially deploying a high communication task pair on a high-speed interconnection link in a node or a low-delay link on the same rack; and tail delay is monitored during operation, and cost-aware local rescheduling is triggered. The system comprises a physical topology modeling module, a communication demand prediction module, an affinity scheduling module and a dynamic rescheduling module. Through active prevention and closed-loop optimization, job tail delay is reduced, and heterogeneous cluster service quality and resource utilization efficiency are improved.
Owner:SHENZHEN XINSAIKE SCI&TECH DEV CO LTD

Method and system for memory mode agnostic workload migration in a heterogeneous cluster

A method for managing a workload migration includes: receiving a request from a user that wants to migrate a workload from a source information handling system (IHS) to a target IHS; analyzing the request and a first IHS configuration list to infer the source IHS' configuration and criticality of the request; making, based on the analyzing of the request and first IHS configuration list, a first determination that the request is non-critical; making, based on the first determination, a second determination that the target IHS does not have the source IHS' memory configuration; and waiting, based on the second determination, until the target IHS or a second target IHS has the source IHS' memory configuration; making a third determination that the second target IHS has the source IHS' memory configuration; and migrating, based on the third determination, the workload from the source IHS to second target IHS.
Owner:DELL PROD LP

Heterogeneous cluster leading-subordinate collaborative guidance control method

The invention discloses a heterogeneous cluster leading-subordinate cooperative guidance control method, and the method comprises the steps: a leading aircraft obtains the target position and speed information through a laser seeker, obtains the target acceleration in combination with an observer, generates a guidance instruction, transmits a coordination variable containing the target position and speed information to at least one subordinate aircraft, and carries out the guidance of the at least one subordinate aircraft. On the basis of the coordination variable, the auxiliary aircraft obtains the acceleration of the target in combination with the observer, and then a guidance instruction is generated. In the guidance instruction generation process, a second-order leading-trailing collaborative guidance law is designed in the sight line direction, the remaining flight time of a trailing aircraft can be consistent with the achievement of a leading aircraft in finite time, and collaborative hit is achieved; in the sight line normal direction, a finite time sliding mode guidance law is designed, and the sight line angle of each aircraft can reach an expected value in finite time.
Owner:BEIJING INST OF TECH +1

Electromagnetic scattering calculation method based on heterogeneous GPU cluster

The embodiment of the invention provides an electromagnetic scattering calculation method based on a heterogeneous GPU cluster. The method is applied to the field of computational electromagnetism, and comprises the following steps: constructing a hierarchical bounding volume geometric acceleration structure of a target object and adding an edge index, and meanwhile, completely copying and distributing data of the acceleration structure to a video memory of each computational node in a heterogeneous GPU cluster; generating a random ray path starting from an emission source, modeling an electromagnetic scattering process as a ray path set in a probability space, and generating a ray direction through an importance sampling strategy; determining a static scheduling strategy for the ray batch based on the calculation cost of pre-sampling and geometric enhancement, and carrying out non-uniform distribution on task tiles according to the heterogeneous GPU calculation power; and monitoring the execution state of the heterogeneous GPU cluster in real time, and dynamically adjusting the distribution strategy of the residual ray tasks based on execution feedback. According to the method, the calculation overhead can be remarkably reduced while the calculation precision is ensured, and the efficiency and expandability of complex target electromagnetic scattering calculation are improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Heterogeneous cluster model update task scheduling method based on reinforcement learning

The invention provides a heterogeneous cluster model update task scheduling method based on reinforcement learning, which mainly improves the scheduling efficiency of resources in an unmanned cluster, balances the task processing performance of each resource in the unmanned cluster, and solves the scheduling decision problem of model update tasks in a heterogeneous unmanned cluster. According to the scheme, before an unmanned cluster starts distributed training, scheduling decision making is firstly carried out on a training task of a single model, then task scheduling is executed, and finally distributed training is carried out; the system takes model training completion time and cluster total energy consumption as optimization targets; according to the method, the advantages of Dirichlet distribution and reinforcement learning are combined, task scheduling constraint conditions can be effectively met, meanwhile, the method has high exploration capacity, and therefore an efficient task scheduling strategy is generated; compared with a traditional heuristic algorithm and other reinforcement learning methods, the task scheduling scheme can be directly generated for the cluster without complex mathematical modeling, and the complexity and cost of implementation and maintenance are reduced.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Mimicry transformation system based on service component

The invention provides a mimicry transformation system based on a service component, and the system comprises the steps: determining N distribution requests after determining the installation of a service in a new service request, matching N heterogeneous cluster nodes, creating corresponding M kubernates cluster connections according to the matching of the N heterogeneous cluster nodes and corresponding heterogeneous information, and generating a command to create N service components. The method comprises the steps of storing a cluster node path and service component information into a database table, obtaining relevant configuration and role information of service components, generating serviceRoleUi, determining an installation sequence of each service component according to a dependency relationship existing among the service components, generating a new service command according to the installation sequence, and installing N service components, and different processing modes are designed according to N response results returned by N service component requests, so that the safety characteristic of the service component is improved, and the anti-attack capability of the whole platform is enhanced.
Owner:EAST CHINA INST OF COMPUTING TECH

Heterogeneous UAV cluster air-based task chain closed optimization method

The invention discloses a heterogeneous UAV cluster air-based task chain closed optimization method, and belongs to the technical field of UAV cluster cooperative control. Firstly, task chain nodes and information flow are instantiated based on specific task parameters, and an initial closed task chain is constructed; dimensionality reduction is carried out on task points through AGNES clustering to obtain task groups, and UAV-task group matching is realized in combination with an efficiency function; a task chain path is dynamically adjusted through a closing time model, a final result is generated by optimizing a task chain through an execution precision model, and closing optimization is achieved. According to the method, a double-layer optimization mechanism for decoupling task topology arrangement and task chain attribute closing is designed, the decision space complexity is reduced through a divide-and-conquer method, and rapid response of dynamic task elements is achieved; rapid construction, dynamic adjustment and accurate closing of the task chain are realized, and the task execution efficiency of the heterogeneous UAV cluster is improved.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Heterogeneous cluster-oriented multi-dimensional collaborative division and communication optimization distributed training method

The invention relates to the technical field of deep learning models, and particularly provides a heterogeneous cluster-oriented multi-dimensional collaborative division and communication optimization distributed training method. The method comprises the steps that heterogeneous equipment and a model are analyzed through an analysis module, and input data are provided for optimization solution; a three-dimensional collaborative optimization division algorithm is utilized to bring selection of an activation value re-calculation strategy into a division optimization process, and video memory constraints are converted into soft constraints of calculation cost relaxation; according to the method, topology sensing self-adaptive communication is carried out, communication transmission modes are dynamically switched on the basis of divided link attribute marks, and the throughput and stability of large-model distributed training are improved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1