Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

419 results about "Cluster Node" patented technology

An individual component of a clustering. May contain other nodes.

GPU computing power resource scheduling method and system

The invention relates to the technical field of data analysis, and discloses a GPU computing power resource scheduling method and system, and the method comprises the steps: collecting node hardware parameters and dynamic load indexes of a GPU cluster to construct a multi-dimensional resource feature vector of the GPU cluster, and constructing a resource portrait of the GPU cluster; establishing a node health degree scoring model of the GPU cluster, and generating a health degree score of a cluster node corresponding to the GPU cluster; analyzing a video memory demand of the GPU task request, and calculating an intensive identifier and a communication dependency relationship; determining the SLA weight of the GPU task request, calculating the resource shortage sensitivity of the GPU task request based on the video memory demand, and calculating the target task priority of the GPU task request in combination with the SLA weight; and determining a resource scheduling node group requested by the GPU task in the resource portrait, generating resource scheduling parameters of the resource scheduling node group, and executing scheduling of computing power resources of the GPU cluster based on the resource scheduling parameters. According to the method, the scheduling efficiency of the GPU computing power resources can be improved.
Owner:SHENZHEN DIXI YUNLIAN TECH CO LTD

Unmanned aerial vehicle cluster interaction method and system based on ad hoc network

The invention discloses an unmanned aerial vehicle cluster interaction method and system based on an ad hoc network, and belongs to the technical field of unmanned aerial vehicle communication and cluster control. The method comprises the following steps: acquiring positioning information, energy parameters and task types of an unmanned aerial vehicle group to generate a clustering node set; predicting a cluster moving direction based on the motion acceleration data of the clustering node set, and generating a dynamic path planning instruction; allocating a main communication channel and a standby channel according to the dynamic path planning instruction, and generating a channel allocation result; monitoring a signal quality parameter of the main communication channel according to a channel allocation result, and triggering a link switching instruction when interference is detected; and constructing an interference thermodynamic diagram based on the topological relation of the clustering node set, and generating a power adjustment instruction. According to the invention, through dynamic clustering, mobile prediction, adaptive channel management, interference sensing switching and topology power control, the communication stability, reliability and resource efficiency of the unmanned aerial vehicle cluster in a complex environment are improved.
Owner:SHENZHEN HUIMINGJIE TECH CO LTD

Cluster computing power energy efficiency perception scheduling and green computing system

The invention discloses a cluster computing power energy efficiency perception scheduling and green computing system, which relates to the technical field of computers and comprises a multi-source energy efficiency perception and data acquisition module used for acquiring power consumption, utilization rate, temperature, cooling state, PUE index and environmental data of cluster nodes. According to the invention, through the multi-modal energy efficiency fusion sensing network and the multi-scale convolution and time sequence attention fusion network, multi-source heterogeneous energy efficiency data such as current, voltage, temperature, airflow and the like of a node level can be collected and fused in real time and with high precision, noise is effectively removed, abnormity self-correction is realized, the defect of energy efficiency sensing granularity in the prior art is made up, and the energy efficiency sensing precision is improved. And reliable input is provided for subsequent scheduling decisions. A cross-scale dynamic twinborn collaborative modeling mechanism is adopted, a physical information neural network and a computational fluid mechanics model are coupled, optimization is carried out through a generative adversarial network structure, and accurate prediction of a complex energy consumption evolution curve and a cooling flow field is achieved.
Owner:HEBEI GUOZENG NETWORK TECHNOLOGY CO LTD

Optimized registration method and system based on micro-service architecture

The invention discloses an optimized registration method and system based on a micro-service architecture. The invention relates to the technical field of Nacos systems. The health monitoring center is started and starts to send a health examination request to a downstream service at regular time; the gateway access layer service receives a client request; obtaining a network address of a target service instance through a Nacos service discovery mechanism according to a service name in the request; then forwarding the request to a target service instance for processing; when the configuration information is changed, the Nacos configuration management center pushes an update notification to all subscribed service instances; the service instance pulls the latest configuration information and updates the local cache; according to the method, a secondary hand waving mechanism is implemented in Nacos cluster node, gateway service and downstream service communication, so that the offline condition of the downstream service can be accurately judged, and the problem of data inconsistency caused by node faults or network partition is avoided.
Owner:BEIJING BAIJU YIXING TECH CO LTD

Remote memory exchange system with non-inductive cold and hot perception of user

The invention discloses a user-noninductive cold and hot sensing remote memory exchange system, which belongs to the field of computer storage, and is characterized in that cluster nodes are divided into extensible remote memory service nodes and computing client nodes according to roles; the remote memory node comprises a remote memory exchange server and a memory resource registration module; the computing client node comprises a FrontSwap-based remote memory exchange client, and a remote memory exchange server side is registered as a memory resource which is non-inductive to a user; furthermore, popularity statistics based on a popularity histogram is realized in a kernel mode; the memory exchange behavior is guided through the dynamic decision-making module based on online reinforcement learning, the memory management efficiency and the system performance are remarkably improved, and efficient and non-inductive memory expansion service is provided for user application.
Owner:HUAZHONG UNIV OF SCI & TECH

Chip system and access method

The invention provides a chip system and an access method, relates to the technical field of chips, and improves the data access performance of a high-performance computing system. According to the specific scheme, the chip system comprises a snoop filter, a plurality of processor cluster nodes and a plurality of memories, the processor cluster nodes are coupled through a bus, the snoop filter is coupled with the bus, the processor cluster nodes are in one-to-one correspondence with the memories, and the snoop filter adopts a unified directory format. The snoop filter is used for obtaining an access request initiated by the processor cluster node, and the access request comprises address information. The snoop filter is further used for determining an access mode of the access request based on the address information, wherein the access mode comprises a unified memory access mode and a non-unified memory access mode. The processor cluster node is to access the memory based on the access pattern and the access request. The embodiment of the invention is used for the process of accessing the memory by the processor cluster node.
Owner:HUAWEI TECH CO LTD

Deployment method of reasoning service, electronic equipment and storage medium

The invention discloses an inference service deployment method, electronic equipment and a storage medium, and the method comprises the steps: responding to the creation of an inference service, reading a model information configuration table and an engine information configuration table, and carrying out the matching processing of the model information configuration table and the engine information table according to a pre-configured model name during the creation of the inference service, processing resources, inference engine mirror image identifiers and inference engine starting parameters which can be used for loading inference services on the container scheduling platform are obtained, and workload metadata used for bearing the inference services are created based on the processing resources, the inference engine mirror image identifiers and the inference engine starting parameters; the deployment of the inference service on the container scheduling platform is realized by distributing the workload metadata to the target cluster node. Through the method, the technical problem of relatively low deployment efficiency caused by manually configuring the parameters in the workload metadata depending on manpower in related technologies is solved, and the technical effects of simplifying the deployment process of the model reasoning service and improving the deployment efficiency are achieved.
Owner:JINAN INSPUR DATA TECH CO LTD

Large-scale industrial fault diagnosis system and method based on Bluetooth MESH

The invention relates to the technical field of industrial Internet of Things and wireless sensor networks, in particular to a large-scale industrial fault diagnosis system and method based on Bluetooth MESH, and efficient monitoring and fault early warning of industrial equipment are realized through low-power-consumption sensor nodes, edge computing, a Bluetooth Mesh network and a clustered Mesh tree architecture. The system comprises a client, a server, a gateway, a cluster head and a cluster node, the cluster node comprises a sensor node and a relay node, a lightweight neural network model is built in the sensor node, and edge reasoning and fault detection can be realized locally. Stable communication and real-time cooperation of large-scale nodes are ensured through a decentralized architecture of Bluetooth Mesh, establishment of a dynamic cluster head chain and optimization of an RSSI threshold value. And a low-power-consumption mechanism combining event driving and fixed polling is adopted, so that the energy consumption of the system is remarkably reduced. In-cluster communication loads are reduced through a clustering architecture, and the communication efficiency and the anti-interference capability of the system are improved in combination with channel separation of Mesh and BLE protocols.
Owner:FUDAN UNIVERSITY

MQTT message transmission optimization method and system

The invention discloses an MQTT message transmission optimization method and system. The method comprises the steps of analyzing a theme, extracting a device type, a data feature and a geographic position triple, calculating a hash value, and mapping a device to a specified Broker fragment cluster node to generate a fragment mapping table; processing the equipment data in the edge domain in the fragment mapping table through an edge calculation layer, removing invalid data according to a preset rule, merging the equipment data in the same fragment node, and embedding a fragment node ID for an aggregation message generated after merging; identifying a fragment node ID and routing to a target fragment node, positioning a corresponding shared memory pool, and writing the message into the shared memory pool; and responding to a direct access request of the client, so that the client directly accesses the data from the shared memory pool through the user mode network stack. According to the method, the problem of uneven load is effectively solved, multiple times of state switching in the data transmission process is avoided, and the MQTT message transmission efficiency is improved while the transmission cost is reduced.
Owner:GUANGZHOU SIYUN DATA TECH CO LTD

Underwater dynamic network resource scheduling and time slot conflict optimization method

The invention provides an underwater dynamic network resource scheduling and time slot conflict optimization method, which comprises the following steps of: firstly, acquiring network state characteristics by using underwater sensor nodes to perform cluster head election and cluster member allocation; secondly, dynamically optimizing the sending priority of the nodes in the cluster by utilizing a Q-learning theory, and carrying out time slot allocation; constructing a reward function based on intra-cluster node distribution information, and iteratively and dynamically adjusting the sending priority of the nodes; and finally, a time slot allocation conflict detection mechanism is formulated to avoid resource competition and transmission conflicts among the nodes in the cluster, so that the optimal node allocation scheme is gradually converged. According to the method, the problems of communication conflicts among multiple cluster nodes of the underwater network, low resource allocation efficiency, inaccurate dynamic sending queues and the like are comprehensively considered, the node sending priority and time slot allocation are optimized through Q-learning, and the channel utilization rate is effectively improved. In a complex dynamic underwater network environment, the technology has applicability and effectiveness, and an underwater node sending strategy can be optimized more reasonably.
Owner:HENAN UNIVERSITY

Power load dynamic optimization prediction method based on reinforcement learning

The invention relates to the technical field of power loads, in particular to a power load dynamic optimization prediction method based on reinforcement learning, which comprises the following steps: acquiring real-time voltage frequency power of nodes, tracking frequency difference direction change of adjacent nodes, marking reverse disturbance to generate a distribution map, and identifying a frequent disturbance area according to the distribution map; clustering node groups with consistent fluctuation trends to determine a target area; monitoring power fluctuation inversion to lock energy steering nodes; constructing a prediction input set; comparing prediction and actual trends to extract a deviation interval, adjusting stride and rate of a reinforcement learning model, updating a decision and smoothing a curve, and outputting a power load dynamic prediction curve. According to the method, disturbance space-time dynamic identification is realized through multi-layer correlation analysis, judgment precision is improved through frequency power joint calibration, load path consistency and area extension are reflected through node clustering, power transmission tracking is enhanced through energy node identification, a feedback closed loop is constructed through offset monitoring, and curve continuity and response rate are improved. Stable convergence is predicted to be consistent with the trend under multiple disturbances.
Owner:弘奎(西安)智能科技有限公司

Cluster computing

In some embodiments, a computer cluster system comprises a plurality of nodes and a software package comprising a user interface and a kernel for interpreting program code instructions. In certain embodiments, a cluster node module is configured to communicate with the kernel and other cluster node modules. The cluster node module can accept instructions from the user interface and can interpret at least some of the instructions such that several cluster node modules in communication with one another and with a kernel can act as a computer cluster.
Owner:ADVANCED CLUSTER SYST

Cluster node data processing method and system based on edge side

The invention relates to the technical field of high-availability clusters, and discloses an edge-side-based cluster node data processing method and system, and the method comprises the steps: obtaining edge-side cluster node resources, querying a hardware performance parameter and an idle time period of each dynamic access cluster node, and constructing a node association information group; extracting network condition prediction information of each dynamic access cluster node, and planning a dynamic leader duration in a target cluster data processing period by considering node dynamic access distribution information corresponding to a node association information group of each dynamic access cluster node and a constraint set constructed by the network condition prediction information; an election strategy is generated and is used for executing free interval election when a target operation service is executed, so that the real-time dynamic adjustment of the master control node of the high-availability cluster considering the influence of multiple factors is realized, and the high-availability cluster is helped to perform free release and node switching of the master control node before a fault occurs; and the operation stability and the processing performance of the high-availability cluster are improved.
Owner:GHOSTCLOUD

Mass small file reading optimization method, system, equipment and medium

The invention relates to the technical field of big data, and discloses a massive small file reading optimization method, system and device and a medium, and the method comprises the following steps: scanning small files in a distributed file system by using a Spark calculation engine to obtain metadata information of the small files; based on the metadata information, small files are classified, and the small files with the same type and data relevance are classified into one group; the resource states of cluster nodes are monitored in real time, the resource states comprise the processor utilization rate, the memory occupancy rate and the I / O load, the load score of each node is calculated, and small file task quotas are dynamically allocated; executing small file merging operation in parallel by utilizing a Spark memory calculation engine to generate a merged file containing a plurality of original small files; creating and maintaining index information in an external database, and recording the position and attribute of each original small file in the combined file; and providing a read-write interface through the special service layer, and positioning and accessing specific small file contents in the combined file.
Owner:GUANGDONG POLYTECHNIC NORMAL UNIV

Distributed database built-in monitoring system

The invention discloses a distributed database built-in monitoring system, belongs to the technical field of databases, and aims to solve the technical problem of how to overcome the defects of complex monitoring deployment, high data response delay, coarse information acquisition granularity and low system integration in the prior art. According to the technical scheme, a monitoring platform module is integrated with a database kernel, a front-end and rear-end separation framework is adopted, and visual display and centralized management and control of the database operation state are achieved; the data acquisition module is used for simultaneously starting with a database process through a built-in HTTP (Hyper Text Transport Protocol) server to realize native support of a monitoring function; the distributed state query module is used for discovering cluster node states through a Gossip protocol, supporting node health detection and remote calling state service, and supporting gRPC-based streaming communication and concurrent processing; and the time sequence data management module is used for storing, compressing and querying the collected time sequence monitoring data.
Owner:上海沄熹科技有限公司

Distributed task processing method and device based on spatio-temporal data complexity and storage medium

The invention discloses a distributed task processing method and device based on spatio-temporal data complexity and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: extracting a feature vector of spatio-temporal data of a digital twin model, and inputting the feature vector into a trained random forest model; through the trained random forest model, performing classification prediction on the spatio-temporal data according to the feature vector, and outputting a category label corresponding to the feature vector, the category label representing the complexity of the spatio-temporal data; mapping the category labels of the spatio-temporal data into weight coefficients, and formulating distributed task queues based on the weight coefficients, so that the difference of the sum of weights of the task queues of each processing node is minimum to ensure load balance among the distributed task queues; and distributing the distributed task queue to each cluster node for distributed processing. The technical effect of improving the distributed task processing efficiency is achieved.
Owner:SHENZHEN SMARTCITY TECH DEV GRP CO LTD

Grouping clustering method and application of distributed storage system

The invention discloses a grouping clustering method and application of a distributed storage system. The method comprises the steps that nodes in a target cluster are grouped and belong to M sub-clusters, and M is larger than or equal to 2; respectively distributing corresponding borne services for the M sub-clusters, and limiting service access among the sub-clusters; the metadata service center of the target cluster is deployed in a metadata service sub-cluster, the metadata service sub-cluster is one of M sub-clusters, and the metadata service sub-cluster comprises a main node and a standby node for operating the metadata service center of the target cluster. According to the method disclosed by the invention, the target cluster nodes are grouped, the services are allocated to the sub-clusters, and the access among the sub-clusters is limited, so that the services can be effectively isolated in the target cluster, and the security of important information data of an enterprise is ensured. The metadata service center is deployed in a metadata service sub-cluster and comprises a main node and a standby node, so that the reliability of storage service can be ensured.
Owner:ANCHAO CLOUD SOFTWARE CO LTD

Radio Access Network Architecture and Terminal Apparatus

This application provides a radio access network architecture. The radio access network architecture includes a cluster node and a serving node. The serving node provides task scheduling and executing functions, and the cluster node provides a region-level centralized collaboration function for the serving node and a collaboration function between cross-region cluster nodes.
Owner:HUAWEI TECH CO LTD

Presto task scheduling method and device

The invention discloses a Presto task scheduling method and a Presto task scheduling device. According to the method, after a task request is received, tasks are allocated to matched clusters according to user information and task information in the request, then task types are further analyzed, execution modes of the tasks are determined, whether the tasks belong to task types needing to exclusively share resources or not is judged, and if the tasks belong to resource exclusive tasks, the tasks are allocated to the clusters. Tasks are allocated to reserved single tasks and special resource nodes, and when the tasks exclusively share working node resources, no other types of tasks are executed on the nodes, so that the tasks exclusively share the resources on specified cluster nodes, physical isolation of the working nodes in a cluster is realized, the performance problem caused by resource contention is avoided, and the service life of the cluster is prolonged. Therefore, the method is suitable for a complex query environment shared by multiple users, efficient execution of the important tasks can be ensured even when a large number of tasks are processed, and the overall reliability and stability of the system are improved.
Owner:DUXIAOMAN TECH (BEIJING) CO LTD

Resource allocation system and method based on MLOps platform

The invention discloses a resource allocation system and method based on an MLOps platform, and the system comprises an access controller which obtains a currently executed calculation task, verifies the calculation task, and carries out the access of the calculation task; the scheduler is used for acquiring resource use conditions corresponding to the cluster nodes in multiple dimensions and determining resource load conditions of the cluster nodes according to the resource use conditions; according to the resource load condition, distributing a corresponding cluster node for the calculation task; the control manager is used for managing the execution process of the calculation task; and the command line interface interacts with the user. Through a scheduler, according to a resource load condition, a resource request and limitation of a calculation task (MLOps task) are adjusted, a dynamic resource allocation algorithm is realized, college and optimal utilization of resources are ensured, and full utilization or excessive allocation of resource niches is prevented.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Fine-grained pipeline scheduling method and device for sensing memory difference of cluster nodes

The invention relates to the technical field of deep learning, and discloses a fine-grained pipeline scheduling method and device based on cluster node memory difference perception. The method comprises the following steps: estimating the ratio of the peak memory occupancy to the memory capacity of each GPU node, comparing the ratio with a preset threshold value, and selecting a re-calculation strategy and a back propagation segmentation strategy: by taking minimization of the end-to-end training time of a model as a target, scheduling and modeling micro-batch data as a flow shop problem, analyzing the dependency constraint of each flow line stage, and calculating the flow shop problem; comprising forward-back propagation sequence dependence, stage dependence, operation dependence and memory limitation, and generating a micro-batch operation sequence scheduling scheme meeting the memory capacity limitation; according to the scheduling scheme, an execution sequence queue is created by executing sorting, and the execution time sequence of calculation blocks, communication blocks and re-calculation operation is dynamically coordinated. According to the method, the GPU equipment utilization rate can be remarkably improved, and the end-to-end training completion time of the model is shortened.
Owner:UNIV OF SCI & TECH OF CHINA

Depth learning task node allocation method and system for executing time-aware computing power network heterogeneous GPU (Graphics Processing Unit) cluster

The invention discloses an execution time aware computing power network heterogeneous GPU cluster deep learning task node allocation method and system. The method comprises the following steps: firstly, based on a deep learning task, extracting and preprocessing task features and available node features; secondly, a sampler equally divides new tasks without historical data to available nodes, and each node performs mixed sampling on the tasks until all the tasks estimate execution time data; taking execution time data as a training set, taking the task features and the node features as a test set, and using a regression decision tree model to predict the execution time of the task on each node; performing task allocation on each node by using a cost search algorithm and a short job total JCT priority strategy; and finally, periodically monitoring node resources released in the cluster to obtain an optimal node allocation result. According to the method, task delay and total task JCT are remarkably reduced, cluster node resource changes are monitored in real time, and the resource utilization rate is increased.
Owner:HANGZHOU DIANZI UNIV +1

Distributed system for dynamic perception and adaptation of unmanned aerial vehicle cluster environment

The invention discloses a distributed system for dynamic perception and adaptation of an unmanned aerial vehicle cluster environment, and belongs to the technical field of computers. The distributed system comprises a sensing layer used for collecting state information of each node in an unmanned aerial vehicle cluster in real time, and a data processing layer used for preprocessing and analyzing the collected unmanned aerial vehicle node state information; the decision-making layer is used for generating an adjustment decision for an unmanned aerial vehicle cluster environment based on a preset unmanned aerial vehicle cluster adaptation strategy library and the dynamic feature data; the execution layer is used for performing corresponding adjustment operation on the unmanned aerial vehicle cluster nodes according to the adjustment decision; and the feedback layer is used for collecting new state information after the unmanned aerial vehicle node adjustment operation. Through innovative fusion and hierarchical collaboration of multiple technical modules and dynamic balance of a strategy library, cluster orderliness is maintained, and a self-growing closed loop of'perception-decision-execution-feedback 'is formed.
Owner:XIAN BAOTONG DEFENSE TECHNOLOGY CO LTD

Hierarchical occlusion perception grid infinitesimal streaming and adaptive rendering method

The invention relates to a hierarchical occlusion perception grid infinitesimal streaming transmission and adaptive rendering method, which comprises the following steps of: S1, decomposing a model into grid infinitesimal, recursively constructing a bounding volume hierarchical tree BVH from bottom to top by taking the grid infinitesimal as a leaf node, and setting a semantic error and a geometric error for each sub-node cluster; s2, acquiring camera parameters of a current frame, comparing a total screen space error with an error threshold value, if the error is smaller than the threshold value, stopping traversing, and adding the cluster node into a to-be-rendered list of the current frame; and S3, if the grid infinitesimal corresponding to the cluster node in the to-be-rendered list is not loaded to the GPU memory, sending a loading request to an asynchronous streaming manager, asynchronously loading the corresponding grid infinitesimal from the memory, and rendering a father node of the cluster node. Compared with the prior art, the method has the advantages that large-scale and high-precision model data can be rendered on the lightweight XR all-in-one machine in real time, and the like.
Owner:SHANGHAI JIAOTONG UNIV

Agricultural product data real-time analysis method and system based on big data

The invention relates to an agricultural product data real-time analysis method and system based on big data, and specifically, the method comprises the steps: obtaining multi-source heterogeneous data related to agricultural products, carrying out the preprocessing, obtaining a task DAG according to an analysis task, obtaining all paths from an entrance to an exit in the DAG, taking the sum of the communication time of all nodes on each path as the length of the path, and carrying out the real-time analysis of the agricultural product data. And taking the path with the maximum length as a critical path, clustering nodes on the critical path by adopting a linear clustering mode, calculating DAG granularity based on a linear clustering result, clustering residual nodes, and scheduling execution nodes in the DAG according to a clustering result to obtain an analysis task result.
Owner:JINAN XINDONGLI INFORMATION TECHNOLOGY CO LTD

Container management system control method and device and storage medium

The invention discloses a control method and device of a container management system and a storage medium, and relates to the technical field of container management, and the control method of the container management system comprises the steps: responding to a trigger operation of a dynamic resource demand, and analyzing the operation type of the trigger operation and the operation content of the trigger operation; obtaining an initial resource quota allocated when the triggering operation corresponds to the tenant registration; determining a resource demand variation corresponding to the trigger operation according to the operation type and the operation content; substituting the used resource quantity, the initial resource quota and the resource demand change quantity into a preset resource matching formula, and calculating the real-time resource allocation quantity of the tenant; and adjusting the resource use quota of the tenant according to the real-time resource allocation amount, and synchronously updating the resource scheduling strategy of the cluster node. The technical effect of dynamic resource scheduling of the tenants is achieved.
Owner:SHENZHEN JIETENG TECHNOLOGY CO LTD +1

Routing information processing method and system, electronic equipment and storage medium

The invention provides a routing information processing method and system, an electronic device and a storage medium, and the method comprises the steps: receiving a target service request, extracting request label information corresponding to the target service request, the request label information being used for representing business information associated with the target service request, matching the request label information with each route, and sending the matching result to the target service request; each route comprises at least one first-level route and at least one second-level route, each first-level route is an entrance route of the at least one second-level route, and if a target route matched with the request label information exists, the target service request is forwarded to a service node corresponding to the target route through the target route. By adopting the method, the service request processing capability of different service units under the same function module is improved, and the fine control capability of cluster nodes is also improved.
Owner:BANK OF NINGBO

Clustering network flow scheduling method and equipment based on star flash, and storage medium

The invention relates to the technical field of industrial networks, in particular to a clustering network flow scheduling method based on star flash, which comprises the following steps: S1, collecting service flow information and cluster node state information according to a management G node and a terminal T node; s2, establishing a clustering network flow scheduling model according to the service flow information and the cluster node state information; s3, establishing target optimization and constraint conditions according to the clustering network flow scheduling model; and S4, designing a Markov decision according to target optimization and constraint conditions, training a neural network model according to the Markov decision and a multi-agent near-end strategy, outputting a scheduling decision according to the trained neural network, and obtaining a multi-dimensional resource adaptation result by a solver according to the scheduling decision. And the service flow is transmitted according to the multi-dimensional resource adaptation result and a multi-queue unicast / multicast flow shaping mechanism. According to the invention, the load balancing capability of the star flash clustering network system can be improved, and the end-to-end transmission delay of the service flow can be reduced at the same time.
Owner:SHANGHAI AEROSPACE COMP TECH INST

GPU (Graphics Processing Unit) resource scheduling method based on data awareness and computer equipment

The invention discloses a GPU resource scheduling method based on data awareness and computer equipment, and the method comprises the steps: carrying out the real-time monitoring of a cluster node, obtaining a current cluster load state, and predicting the cluster load state in a future time period; creating services according to an input service definition description file, adding the services into a to-be-scheduled service queue, periodically selecting the services from the to-be-scheduled queue, and allocating computing nodes for the services; distributing a target service for the user request according to the current cluster load state, establishing a session, and analyzing the intention of the user; dynamically adjusting and allocating static resources through deep reinforcement learning based on a hybrid expert model according to a user intention analysis result and a cluster load state; required resources are loaded to a GPU through a multi-level cache architecture, and after calculation is completed, a calculation result is returned to a user through a session. According to the method, the utilization rate and the performance of GPU resources in an interactive GPU acceleration service scene can be improved, and the energy consumption and the cost are reduced.
Owner:BEIJING UNIV OF POSTS & TELECOMM