Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

97 results about "Multi server" patented technology

GPUBox hardware decoupling system based on Retimer card and PCIeSwitch chip

The invention discloses a GPU Box hardware decoupling system based on a Retimer card and a PCIe Switch chip, and belongs to the technical field of computer hardware architecture and high-speed interconnection. According to the system, a Retimer card and a PCIe Switch chip are integrated in an independent GPU Box, and a decoupling link of a CPU server and a GPU acceleration card is constructed; the Retimer card realizes 30-meter long-distance PCIe signal transmission and breaks through physical distance limitation; the PCIe Switch chip pools GPU resources through a dynamic routing and MRIOV technology, supports flexible allocation of computing power by multiple servers, and realizes Peer-to-Peer direct connection communication between GPUs. Aiming at a large model reasoning scene, the system optimizes KV cache bandwidth allocation and video memory and memory cooperative scheduling, so that the 100B parameter model reasoning throughput is greatly improved; and meanwhile, the usability of the system is greatly improved through fault isolation and hot plug design. According to the method, the problems of physical binding of the CPU and the GPU, limited transmission distance, rigid resource allocation and the like in a traditional architecture are solved, and the method is suitable for large-scale AI calculation and distributed GPU cluster deployment.
Owner:HEFEI FENGZHIYI SEMICON CO LTD

Multi-GPU server cluster liquid cooling flow distribution method based on reinforcement learning

The invention relates to the technical field of liquid cooling flow distribution, and discloses a multi-GPU server cluster liquid cooling flow distribution method based on reinforcement learning, and the method comprises the steps: S1, collecting the operation state parameters of each GPU in a multi-GPU server, enabling the operation state parameters to comprise core temperature, voltage, current and load utilization rate, carrying out the thermal behavior modeling through a sliding time window, generating a thermal dynamic feature vector representing the GPU short-time thermal trend; and S2, collecting the flow velocity, the temperature difference of inlet and outlet water and the thermal resistance of the cold plate of each channel of the liquid cooling system in real time, and representing the liquid cooling assembly as a graph structure. According to the method, the technical scheme of joint modeling based on the graph neural network and reinforcement learning is adopted, fusion coding is performed on the GPU thermal dynamic characteristics and the topological state of the liquid cooling system, and an Actor-Critic architecture is introduced to realize self-adaptive regulation and control of the liquid cooling flow, so that the technical effect of improving the heat dissipation efficiency and the energy consumption balance control capability of the multi-GPU server cluster is achieved.
Owner:BEIJING HUAHONG DIGITAL TECH CO LTD

Artificial intelligence content generation task multi-server collaborative optimization method based on deep reinforcement learning

The invention discloses an artificial intelligence content generation task multi-server collaborative optimization method based on deep reinforcement learning, and relates to the crossing field of edge computing and artificial intelligence. According to the method, a user equipment-base station-edge server-core network four-layer architecture is constructed, distributed decision is realized through a deep Q network model, and the method comprises the following steps: defining a task and time delay model; a deep reinforcement learning framework is designed, a decision model is constructed through a state space, an action space and a reward function, and training stability is improved by adopting experience playback and a target network slow update mechanism; and providing an adaptive multi-server selection and load distribution strategy, and solving an optimal distribution proportion based on a time delay equality principle. According to the method, the average unloading time delay of the AIGC tasks is remarkably reduced, the task failure rate is reduced, the dynamic environment adaptability and the extreme scene robustness are improved, and the method is suitable for scenes with strict requirements for time delay and reliability.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

System and method for recommendation and optimization of information technology server resources

A system for recommendation and optimization of information technology (IT) server resources is disclosed. The system includes a server comprising at least one processor configured to access input datasets associated with an IT workload, determine types and counts of server resources based on predefined criteria and a server consolidation configuration, and access a multi-server performance database. The processor utilizes a trained deep learning model to predict infrastructure requirements and generate recommendations for an optimal server configuration by balancing performance, power consumption, and resource utilization. The processor automatically allocates server resources from multiple manufacturers based on the recommendations and generates data for display on a user interface dashboard. The dashboard presents server utilization patterns, recommended configurations, and real-time performance metrics of allocated resources. The system enables intelligent consolidation and efficient server management within a datacenter environment.
Owner:HYBRIDAI PTE LTD

Federal learning method and device based on zero knowledge proof

The invention provides a federal learning method and device based on zero knowledge proof, and the method comprises the steps: obtaining a training task from a task publishing server through a communication address, updating the model parameters of a to-be-trained model according to local privacy data, and obtaining a first target parameter, according to the first target parameter and an updating model parameter of a matching training server, performing aggregation through a task publishing server, and performing repeated iteration until an aggregated model is converged; receiving a random round set, screening model parameters, and performing consistency verification through a task publishing server; meanwhile, local privacy data and initial model parameters are screened, correctness of data traceability, parameter iteration and model training processes is verified through a task publishing server based on zero-knowledge proof, and training is determined to be completed under the condition that verification is passed. According to the method, the problems of relatively low model training efficiency and slow updating caused by multi-server federated learning based on zero knowledge proof in the prior art in order to ensure data security are solved.
Owner:中国邮政储蓄银行股份有限公司

Electric power system monitoring data cloud unified access and resource intelligent prediction system and method based on cloud-side cooperation

The invention provides an electric power system monitoring data cloud unified access and resource intelligent prediction system and method based on cloud edge collaboration. The system is provided with a unified access module, an alarm rule configuration module, a real-time alarm module, a real-time monitoring display module and an intelligent trend prediction and health degree evaluation and maintenance module. According to the system and method, Java language development is adopted, the system and method serve as a central platform to receive and manage monitoring data from a plurality of lower-level power platforms in a unified mode, and standardized data unified access, real-time multi-dimensional monitoring, intelligent trend prediction, dynamic alarm management and system health degree evaluation and maintenance of a CPU, a memory, a hard disk, network resources and middleware services are achieved; the method is very suitable for a power system, can also be widely applied to other industries needing high reliability and real-time performance, and effectively solves the technical problems that monitoring data of multiple servers cannot be managed in a unified mode, the resource prediction precision is low, and system fault response lags behind.
Owner:JIANGSU QIFENG TECHNOLOGY CO LTD

Partition table data export processing method

The invention discloses a processing method for exporting partition table data, which is used for solving the technical problems of low cross-partition query efficiency, inconsistent file size and lack of concurrent processing capability in the prior art. The method comprises the following steps: firstly, reading configuration information from a task detail table to generate an export task, counting the data volume of each partition, and distributing data to an output file with a fixed line number according to a sequential filling strategy; an innovative coding scheme is adopted to record a distribution result, different partition data containing conditions are represented through characters'-'and'. 'and numbers, and compressed coding is carried out on continuous empty partitions; sending a distribution result to a message queue to realize multi-server concurrent processing; the consumer decodes the distribution information and then executes restrictive query according to a partition sequence, and data is exported concurrently by using streaming query in cooperation with multiple threads; and a perfect state monitoring and exception recovery mechanism is established, and an MD5 verification file is generated to ensure data integrity. According to the method, the processing efficiency of partition table data export and the system reliability are effectively improved.
Owner:HAIER CONSUMER FINANCE CO LTD

Computer system and multi-server shared storage system

The application provides a computer system and a multi-server shared storage system, which can be applied to the technical field of server storage. The computer system comprises: a computer system, a non-volatile memory and a plurality of servers which are electrically connected; an arbitration module is used for determining a target server according to a plurality of access signals of the plurality of servers, and sending an authorization signal to the target server and a bus buffer module; the bus buffer module is used for connecting a transmission bus between the target server and the non-volatile memory according to the authorization signal; in the case that a chip select signal in a flash signal is at an effective level, an access address is obtained from the flash signal, the access address is configured by using a configuration partition base address to obtain a target access address, signal conversion is performed on the target access address, a target flash signal is obtained and sent to the non-volatile memory, so that the non-volatile memory manages data stored in a storage partition corresponding to the target access address.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Method of applying heat sink to semiconductor package, method of maintaining processor module in server, and data center system

The disclosure relates to a method of applying a heat sink to a semiconductor package, a method of maintaining a processor module in a server, and a data center system. One or more sensors are integrated into a processor module that includes a semiconductor package and a heat sink. There are multiple processor modules in the server, and data from each processor module is sent to a digital integration monitoring system, which may be part of a data center hardware monitoring system. The server may be one of a number of servers located in the data center. The data may also be sent to a remote data center maintenance center that simultaneously monitors multiple data centers. When compared to different criteria, the data may be used to determine in real time whether a given processor module needs maintenance.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

A multi-server security aggregation system and method based on homomorphic chameleon hashing

A multi-server secure aggregation system and method based on homomorphic chameleon hashing utilizes a bulletin board system maintained by a trusted third party to output, download, and collect data. Public-key encryption and secret sharing ensure data security. Homomorphic chameleon hashing is used to verify aggregation and ensure integrity. This creates a non-interactive secure aggregation model that eliminates the need for secure channels between clients and multiple servers. This system boasts a simple, easy-to-implement architecture, is secure and verifiable, and further improves communication efficiency.
Owner:BEIJING VENUS INFORMATION TECH +1

A method and system for drone identity anonymization and secure data outsourcing

The application discloses a kind of unmanned plane identity anonymization and safe data outsourcing method and system, method is based on integrated supervision platform and unmanned plane, including system initialization stage, integrated supervision platform generates main private key and public parameter;User anonymous and private key generation stage, unmanned plane registers anonymous identity and obtains user private key;Data preprocessing and outsourcing stage, unmanned plane divides data into block and sub-block and encrypts, generates coding vector using erasure code coding, generates verification label for each coding data block, constructs random distribution table according to the number of cloud server, distributes coding vector and label to multiple cloud server storage.In this way, the unmanned plane identity anonymization and safe data outsourcing method described reduces storage redundancy, ensures that data can be reconstructed when multiple servers fail, and guarantees data availability and robust recovery.
Owner:SOUTHWEST PETROLEUM UNIV

SM9-based client-multi-server rapid self-adaptive key negotiation method

The invention provides an SM9-based client-multi-server rapid self-adaptive key negotiation method. The SM9-based client-multi-server rapid self-adaptive key negotiation method mainly comprises the following steps: an initialization stage, a private key extraction stage, a broadcasting stage, an offline preparation stage, an online response stage and a session key generation stage. The method has the following characteristics: 1, the method is suitable for a client-multi-server scene, a client initiates a communication request, a plurality of servers are automatically adapted to communicate with the client according to states such as self load and communication delay, and a key negotiation process only needs three times of communication; 2, aiming at a client-server communication mode, computing load optimization is carried out; different from the traditional P2P key negotiation with the same calculation amount of both parties, the method migrates the calculation overhead to the server side as much as possible; and 3, the client does not need to carry out complex bilinear pairing operation, and particularly, the calculated amount in an online response stage can be almost ignored. And meanwhile, an online / offline optimization technology is adopted, so that the client can carry out pre-calculation before communication, and the online response efficiency is improved.
Owner:SOUTHEAST UNIV +1

Privacy protection method and device for preventing collusion of multiple servers, equipment and medium

The invention provides a privacy protection method and device for preventing collusion of multiple servers, equipment and a medium, and the method comprises the steps: determining a first privacy protection conversion strategy corresponding to a data type according to the data type of disturbance output under a privacy protection disturbance mechanism; obtaining a first noise value corresponding to the first privacy budget according to the original data and conversion information corresponding to the first privacy budget; obtaining a second noise value corresponding to a second privacy budget according to the first noise value and a first privacy protection conversion strategy; determining a second privacy protection conversion strategy corresponding to the data type according to the disturbance output data type; according to the original data, the second noise value and the second privacy protection conversion strategy, a third noise value corresponding to the first privacy budget is obtained, the purpose that noise data corresponding to the claimed privacy budget can be obtained by achieving data collection of any server through data multiplexing is achieved, and therefore extra privacy budget consumption is avoided.
Owner:INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES

Computing power server

The utility model discloses a computing power server, which comprises a mainboard, a display card and a case, the mainboard is horizontally arranged at the bottom of the case, the display card is vertically fixed on the upper surface of the mainboard, the plane of the display card is vertical to the plane of the mainboard, the case is of an L-shaped structure, the mainboard is accommodated at the bottom of the L-shaped structure of the case, and the mainboard is fixed on the upper surface of the case. The vertical part of the L-shaped structure accommodates a display card, and the two cases form an integrated unit when stacked; according to the server structure, the display cards are vertically arranged on the main board and form the L-shaped case in appearance instead of being arranged in a traditional parallel stacking mode, the stacking mode is adopted, the overall width of the server module is remarkably reduced, and compared with a traditional square module layout, the structure allows more server modules to be contained, and the cost of a machine room is reduced.
Owner:SHENZHEN CLOUDSKY TECH CO LTD

Power management method and device in multi-server system, medium and product

The invention discloses a power management method which is executed by a BMC (baseboard management controller) built in a main server in a multi-server system, the multi-server system comprises a plurality of servers and a power management subsystem connected with the servers, the power management subsystem comprises a plurality of power management units connected in parallel, and a selective switch is arranged on a parallel branch where the power management units are located. The multi-server system comprises a master server and a plurality of slave servers, and the method comprises the steps of appointing each acquisition time point for acquiring data after determining that the master server and a BMC (Baseboard Management Controller) in each slave server complete time standard unification; and receiving the real-time power consumption information of each slave server acquired by the BMC in each slave server, performing information grouping on the real-time power consumption information of each slave server according to the acquisition time point, and performing on-off control on the selection switch on the parallel branch where each power management unit is located according to the grouping condition of the real-time power consumption information. And the flexibility, the reliability and the intelligent level of power supply management of the multi-server system are improved.
Owner:HUBEI SILANG WANWEI COMPUTING EQUIPMENT MANUFACTURING CO LTD

Providing executing programs with access to stored block data of others

Techniques are described for managing access of executing programs to non-local block data storage. In some situations, a block data storage service uses multiple server storage systems to reliably store copies of network-accessible block data storage volumes that may be used by programs executing on other physical computing systems, and snapshot copies of some volumes may also be stored (e.g., on remote archival storage systems). A group of multiple server block data storage systems that store block data volumes may in some situations be co-located at a data center, and programs that use volumes stored there may execute on other computing systems at that data center, while the archival storage systems may be located outside the data center. The snapshot copies of volumes may be used in various ways, including to allow users to obtain their own copies of other users' volumes (e.g., for a fee).
Owner:AMAZON TECH INC

Authentication method and system under multi-server architecture

This application discloses an authentication method and system under a multi-server architecture. When a user registers with a registration server, the registration server generates a temporary identity identifier (TID) for the user. U When the user needs the service provided by the application server, enter the TID on the client. U The authentication user identity information γ0 can be calculated, and the client sends the set M0 containing γ0 to the application server. The application server generates relevant authentication information and sends the authentication information and M0 to the registration server. The registration server completes the identity authentication of the user and the application server. The authentication of the user and the application server is completed on the registration server. Each application server does not need to store any information about the user, which reduces the resource burden of the application server. In addition, the user only needs to register with the registration server to obtain a temporary identity to complete the authentication and key exchange. There is no need to register with each application server in the multi-server architecture, which saves time and cost.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

A method and system for automatic operation of operation and maintenance tasks

The application discloses a kind of operation and maintenance task automatic running method and system;Script operation and maintenance environment is dependent on risk high and complex task scheduling inefficient, the present application obtains low code arrangement feature, Playbook structure feature and host load parameter is merged and fused to obtain automatic operation and maintenance feature, and obtains server certificate execution method, utilizes automatic operation and maintenance feature and server certificate execution method to build certificate execution operation and maintenance automation model.The present application considers low code arrangement feature, Playbook structure feature and host load parameter to build multi-parameter correction neural network model, more Playbook structure feature data is generated by reinforcement learning method, so that certificate execution operation and maintenance automation model is more accurate and intelligent to obtain server certificate intelligent execution method, while task runs in container sandbox to isolate risk, combined with the self-recovery mechanism that drag-and-drop flowchart parameter generates executable script, applied to the intelligent batch execution of multiple server certificate update and avoids fault node.
Owner:GUANGZHOU SHANGHANG INFORMATION TECH CO LTD

High-throughput splitting federated learning method based on activated cache

The invention discloses a high-throughput split federated learning method based on an activated cache. The method comprises the following steps: carrying out system initialization and network structure variable definition; performing client parameter synchronization, local feature extraction and independent parameter updating; performing activation feature drift asynchronous calculation and maximum old-age window updating; performing double-branch independent training of the server side based on hybrid retrieval; global federal aggregation based on data volume weighting is carried out; and performing a multi-server integrated reasoning test of the global model. Through an innovative drift sensing activation reuse mechanism and a multi-GPU communication-free integrated training strategy, the throughput performance of the split federated learning system in heterogeneous equipment and a non-independent identically distributed data environment is remarkably improved.
Owner:ZHEJIANG UNIV OF TECH

Server cabinet interface capable of being movably spliced

The utility model relates to the technical field of server cabinets, and discloses a movably spliced server cabinet interface, which comprises a server cabinet body, a vertically distributed strip-shaped fixed frame is arranged on the back side of the server cabinet body, an interface assembly is clamped in the fixed frame in an up-and-down sliding manner, and the interface assembly is connected with the server cabinet body in an up-and-down sliding manner. A position adjusting mechanism used for driving the interface assembly to ascend and descend is installed in the center of the interior of the fixing frame, and a plurality of wiring windows used for being communicated with internal components of the server cabinet body and adjusting the interface assembly to the position are sequentially formed in the cabinet wall of the back side of the server cabinet body from top to bottom. The server cabinet interface capable of being movably spliced not only improves the flexibility and adaptability of server wiring, but also greatly reduces the maintenance cost and improves the operation efficiency through the modular design and the accurate positioning mechanism. The method shows great advantages in wiring operation after splicing of multiple server cabinets, and can adapt to server cabinets with different configurations.
Owner:SHENZHEN QIHUI TECHNOLOGY CO LTD

Monitoring system for data center maintenance

One or more sensors are integrated into a processor module that includes a semiconductor package and a heat sink. A plurality of processor modules is present in a server, and data from each processor module is sent to a digitally integrated monitoring system which can be part of a data center hardware monitoring system. The server may be one of many servers located in a data center. The data may be further sent to a remote data center maintenance center that concurrently monitors several data centers. When compared to different benchmarks, the data can be used to determine in real time whether maintenance is needed for a given processor module.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Cost and reliability-aware profit optimization method for heterogeneous multi-server systems

This invention discloses a cost- and reliability-aware profit optimization method for a heterogeneous multi-server system. The method includes: establishing a heterogeneous multi-server system model to obtain expected service revenue and multi-server system utilization; establishing a soft error reliability model to formulate the application domain investment selection and profit maximization problem as a constrained nonlinear optimization problem; utilizing an iterative method based on a penalty function to obtain the optimal multi-server system configuration, minimum cost, and maximum profit for any single application domain, and converting the profit and configuration optimization problem of the heterogeneous multi-server system into a 0-1 knapsack problem; and utilizing a dynamic programming-based algorithm and a greedy-based algorithm to solve the optimal application domain investment strategy and the corresponding maximum profit. This method can consider budget constraints and the soft error reliability of service requests while configuring the heterogeneous multi-server system configuration, thereby improving cloud service quality and provider profits.
Owner:NANJING UNIV OF SCI & TECH

Method and system for automatic operation of operation and maintenance task

The invention discloses a method and a system for automatically running an operation and maintenance task. In order to solve the problems that a script operation and maintenance environment is high in dependence risk and complex task scheduling is low in efficiency, an automatic operation and maintenance feature is obtained by obtaining a low-code arrangement feature, a Playbook structure feature and host load parameters and combining and fusing the features, and a server certificate execution method is obtained; and constructing a certificate execution operation and maintenance automation model by utilizing the automatic operation and maintenance characteristics and a server certificate execution method. According to the method, a multi-parameter correction neural network model is constructed by considering a low code arrangement feature, a Playbook structure feature and a host load parameter, and more Playbook structure feature data is generated through a reinforcement learning method, so that a certificate execution operation and maintenance automation model obtains a server certificate intelligent execution method more accurately and intelligently; and meanwhile, tasks run in a container sandbox to isolate risks, a self-healing mechanism of an executable script is generated in combination with drag-and-drop flow chart parameters, and the method is applied to intelligent batch execution and fault node avoidance during multi-server certificate updating.
Owner:GUANGZHOU SHANGHANG INFORMATION TECH CO LTD

Method and device for parallel processing and calculation of time delay information of multiple server nodes

The invention relates to the technical field of software development, and discloses a method and device for parallel processing and calculation of time delay information of multiple server nodes, and the method comprises the steps: creating a plurality of communication instances for a target server through a transit server node information set deployed by a rear-end server, the target server being a plurality of servers having communication requirements, the transit server node information set comprises a plurality of pieces of transit server node information, and each communication instance corresponds to one piece of transit server node information; matching the communication examples of all the target servers according to the transit server node information corresponding to all the communication examples of all the target servers to obtain one or more communication example matching groups; the time delay information of the server node corresponding to each communication instance matching group is calculated by obtaining the timestamp information of each communication instance matching group. Therefore, by implementing the method, the calculation of the time delay information of the server node can be accelerated through a multi-communication instance mode.
Owner:深圳鼎匠科技有限公司

Multi-server-oriented anti-quantum two-factor identity authentication method and system

The invention relates to the technical field of Internet, and provides an anti-quantum two-factor identity authentication method and system for multiple servers. According to the method, in a registration stage, biological characteristics of a user and physical characteristics of a client PUF (Physical Unclonable Function) are combined to serve as a two-factor authentication factor, a temporary identity label and temporary password information are generated in combination with an identity label and a password of the user for registration, and a registration center returns login verification information to a client for storage; in the login verification stage, login verification of the user is completed through login verification information stored by the client; and entering a key negotiation stage after the login verification is successful, and negotiating a session key between the client and the server logged in by the user by using a central binomial distribution sampling technology and combining an error coordination mechanism in the stage. According to the method provided by the invention, server leakage attacks and quantum computing attacks can be resisted, so that safe and efficient identity authentication is completed.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Multi-server data synchronization method of access control system, electronic equipment and storage medium

The invention relates to a multi-server data synchronization method of an access control system, a server and a storage medium, and belongs to the technical field of computers, and the method comprises the following steps: step 1, processing data required to be synchronized by the access control system to form a data model; 2, extracting a data unique identifier in a data model of to-be-synchronized data, and putting the data unique identifier into a push queue of a push center; and step 3, the pushing center assembles a complete data model from a local database in real time according to the data unique identifier and pushes the complete data model to a target server. By adopting the method, the queue occupation is smaller, and more to-be-processed tasks can be borne without influencing the performance; rapid broken line recovery can be realized, an increment reissuing mechanism is adopted, full-amount synchronization waste is avoided, and the recovery speed is fast and efficient; the service consistency is higher, and the method is suitable for a multi-service parallel access control system.
Owner:ZKTECO CO LTD

A Multi-User Dependency Task Unloading Method Based on Deep Q-Learning and CH Algorithm

This invention discloses a multi-user dependent task offloading method based on deep Q-learning and the CH algorithm, belonging to the field of task offloading technology. It includes: determining a single BS / SC and multiple servers as the network model and defining communication conditions and computation methods; then using a DAG to establish a dependent task model to complete MEC system modeling; establishing an objective function; proposing a DTO-DQN-CH hybrid optimization algorithm; describing the task offloading optimization using deep Q-learning; introducing a load balancing mechanism based on the CH algorithm; and finally obtaining the complete DTO-DQN-CH process; setting three simulation tasks, using MD-TSDDQN and DTO-GA-CH as control groups, to evaluate the load balancing effect between MECs, the number of UDs, and the impact of the number of MECs on task offloading optimization. This invention can achieve the goal of minimizing energy consumption and latency.
Owner:WUHAN UNIV OF SCI & TECH +2

Server assembly method, apparatus, medium, and program product

The application discloses a kind of server assembly method, equipment, medium and program product, it is related to server technical field, comprising: the multidimensional data of server module is collected and identified, generates module feature dataset;Using multidimensional feature matching algorithm, each dimension data in module feature dataset is weighted matching, generates compatibility score result and filters out module combination, obtains module combination set;According to module combination set, using reinforcement learning algorithm, module assembly path is planned, and the optimal assembly sequence is output;According to the optimal assembly sequence, the assembly process is divided into multiple parallel executable subtasks;According to subtask division result, using multilevel parallel assembly mechanism, assembly operation is executed in parallel in module, between module and multiple server node levels.The automation of assembly process is realized, compatibility error rate is reduced, manual intervention link is reduced, and assembly efficiency and system stability are effectively improved.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

GPU cluster heat flow cooperative control method based on reinforcement learning driving

The invention discloses a GPU cluster heat flow cooperative control method based on reinforcement learning driving, and relates to the technical field of flow distribution. Environmental data is collected and preprocessed; generating a cooling routing path through a reinforcement learning layering strategy; dividing emergency levels for cooling routing paths according to the temperature of each GPU server cluster, and generating a global cooling path through a distributed routing negotiation protocol across the GPU server clusters; the global cooling path is converted into a flow scheduling control instruction through a cooling topology programmer, the temperature change of the GPU server cluster is dynamically monitored, and the temperature change rate is output; and setting a change threshold range, comparing the temperature change rate with the change threshold range, judging whether the flow scheduling control instruction is feasible, and outputting a coping strategy. According to the method, intelligent global optimization of the liquid cooling flow of the multi-GPU server cluster is realized through fusion of a reinforcement learning layering strategy and a distributed collaborative mechanism.
Owner:BEIJING HUAHONG DIGITAL TECH CO LTD

Scale analysis method, computing power scheduling method and device for intelligent computing power cluster

The invention discloses a scale analysis method and device for an intelligent calculation power cluster and a calculation power scheduling method and device. The scale analysis method comprises the following steps: acquiring demand information of a target service scene on a computing power scale and cluster hardware resources, network switching equipment and server configuration dimension features of each cluster; determining a current computing power level and a storage level according to the demand information and cluster hardware resource dimension features; determining the upper limit number of intelligent calculation cards of a parameter surface network and the maximum number of servers of a sample surface network according to network and server characteristics, and obtaining computing power cluster scale matching capability information according to cluster hardware resource dimension characteristics, the upper limit number of the intelligent calculation cards supported by the parameter surface network and the maximum number of the servers supported by the sample surface network; and determining a scale analysis result of each computing power cluster according to the current computing power level and the current storage level of each cluster. According to the method, unified measurement of heterogeneous computing power resources can be realized, and the scale of the computing power cluster can be quantified comprehensively and accurately.
Owner:BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD