Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

620 results about "Performance computing" patented technology

Control method for heat dissipation of high-performance computing cluster platform

The invention discloses a control method for heat dissipation of a high-performance computing cluster platform, which comprises the following steps of: firstly, establishing a three-dimensional temperature distribution model of a high-performance computing cluster, acquiring data such as temperature and power consumption of each node in real time, obtaining deviation data of each node by utilizing a prediction model, and dynamically evaluating a heat dissipation risk level of each node; dividing the whole cluster into different risk areas; computing tasks are intelligently distributed based on risk levels, and high-load tasks are preferentially scheduled to a low-temperature area; and then, starting graded cooling measures for the high-risk area, continuously monitoring the cooling effect, and feeding back the result to an early warning system and a hardware maintenance module to form'prediction-regulation-feedback 'closed-loop control. The heat dissipation efficiency of the high-performance computing cluster platform is remarkably improved, the overheat fault risk is reduced while the computing performance is guaranteed, the service life of key hardware is prolonged through dynamic optimization, and intelligent cooperation of heat dissipation resources and computing tasks is achieved.
Owner:TIANJIN ZHONGDA ZHITENG TECH CO LTD

Integer parallel computing method and device based on distributed storage and computer equipment

The invention belongs to the field of high-performance computing, and relates to an integer parallel computing method and device based on distributed storage and computer equipment, and the method comprises the steps of collecting real-time resource indexes, dynamically identifying fault nodes, triggering task migration, and performing data verification and hard disk fault detection. The weight value of each node is calculated, the nodes are arranged according to the descending order of the weight values, and the nodes with high load capacity are selected to distribute tasks; dynamically distributing a data generation task to a computing node, executing parallel computing, and performing distributed storage on a result; obtaining an operand, converting the operand into a first-order tensor form of a basic operand, serializing tensor data, and sending the serialized tensor data to a parallel computing layer; distributing a search task to a computing node, retrieving storage data in parallel, reading effective data from a storage layer, and combining search results into a partial sum; and summarizing and then outputting. The system has dynamic resource management and fault-tolerant capabilities, and can realize efficient task allocation and load balancing.
Owner:SHENZHEN Y& D ELECTRONICS CO LTD

Parallel task scheduling algorithm for heterogeneous multi-core processor

The invention relates to the technical field of computer architecture and parallel computing, and discloses a parallel task scheduling algorithm for a heterogeneous multi-core processor, which comprises the steps of task modeling, resource mapping, dynamic load balancing, communication optimization, task scheduling decision and execution monitoring. Task allocation is adjusted in real time through dynamic load balancing, cross-core communication delay is reduced in combination with communication optimization, and an efficient task allocation sequence is generated by using an improved genetic algorithm. According to the method, the resource utilization rate and the task execution efficiency of the heterogeneous multi-core processor in a high-performance computing scene can be improved, meanwhile, the robustness and adaptability of an algorithm are enhanced, and the task allocation problem in a complex computing scene is effectively solved.
Owner:SUZHOU DUXUEKEZHENG INTELLIGENT TECH CO LTD

Low-code platform and Wasm high-performance computing integration system and method

The invention discloses a low-code platform and Wasm high-performance computing integration system and method, and relates to the technical field of front-end development. In order to solve the problems that an existing low-code platform is limited in calculation performance and poor in security isolation performance, the scheme adopted by the invention comprises five modules: a description and registration module establishes a metadata structure for a Wasm module; the dynamic loading and asynchronous compiling module realizes loading as required and compiling during operation, and supports caching, version verification and hot replacement; the authority sandbox and resource isolation module constructs a security environment based on authority configuration to prevent illegal operation; the parameter binding and safety bridging module encapsulates a uniform interface to realize standardized interaction; the packaging and life cycle management module packages the module into a visual component, and supports full life cycle management. According to the method, high-performance, high-safety and visual integration of the Wasm module in a low-code platform is realized.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Neural network large model efficient reasoning method based on multiple GPGPUs

The invention belongs to the technical field of artificial intelligence and high-performance computing, and particularly relates to a neural network large model efficient reasoning method based on multiple GPGPUs. The method aims to solve the problems of high communication overhead, non-uniform load, low resource utilization rate, high data transmission delay and the like among multiple processors. Dividing a calculation task into a plurality of sub-graphs through static analysis and mixed granularity partitioning of a model calculation graph; distributing the sub-graphs to the optimal GPGPU based on a weighted cost function in combination with heterogeneous resource perception and a dynamic mapping strategy; a global pipeline scheduling plan is constructed by using communication topology perception, and calculation and communication overlap are maximized; data are loaded in advance through a host side hierarchical caching and asynchronous prefetching mechanism, and transmission delay is hidden; multi-stream concurrent execution and event-based lightweight synchronization are adopted on each GPGPU, so that waiting overhead is reduced. According to the method, the reasoning delay can be remarkably reduced, the throughput and the hardware utilization rate are improved, and the method has good adaptivity and expandability.
Owner:BEIJING TOPMOO TECH

Deployment method and system of large language model

The invention relates to the field of artificial intelligence, in particular to a deployment method and system for a large language model, and the method comprises the steps: (1) a service request end receives at least one piece of service request information, and stores the service request information in a service request message queue; (2) responding to a pre-filling stage calculation unit request and a load condition, and distributing service request information; (3) finishing the processing of the pre-filling stage to obtain at least one operation result; (4) the load balancer determines a task operation calculation unit in the stage according to the load condition of the calculation unit in the decoding stage; and (5) inputting the operation result of the pre-filling stage into the task operation calculation unit of the decoding stage, and outputting the result. The method has the advantages that the pre-filling stage and the decoding stage are deployed on a machine with high-performance computing power and a large memory respectively, load tasks are balanced, maximum hardware utilization is achieved, idle computing power is reduced, overall delay is reduced, throughput is improved, and expansibility and fault tolerance of the system are enhanced.
Owner:HANGZHOU DEEPQUOSUO ARTIFICIAL INTELLIGENCE BASIC TECHNOLOGY RESEARCH CO LTD

Translating Between CXL.mem and CXL.cache Read Transactions

Memory has been playing a major role in the performance, scalability and applicability of General Compute systems, and more recently, in realizing Generative Artificial Intelligence (GenAI) and High-Performance Computing (HPC) systems that scale to thousands of GPUs, CPUs and special-purpose Accelerators. Embodiments herein disclose efficient software-defined protocol terminations and protocol translations utilizing Compute Express Link (CXL), including translations between CXL.mem and CXL.cache protocols. Also disclosed are CXL-based systems, Resource Provisioning Units (RPUs), and Memory Fabric Switches enabling dynamic memory pooling and sharing, host-to-host communication utilizing CXL.mem, CXL.cache and CXL.io, intent-based protocol translations, and optionally seamless interactions between CXL, UALink, NVLink, and / or Ethernet protocols, utilizing a broad range of semantics including IO, Cache, and Memory, optimizing memory access and reducing latency. Some embodiments also enable scalability, flexibility and security in high-performance architectures suited for data centers and next-generation computing environments.
Owner:HYATT GAYA OPAL MS +1

Memory architecture-oriented dual-precision general matrix multiplication optimization method and system

The invention belongs to the related technical field of high-performance computing, and provides a memory architecture-oriented dual-precision general matrix multiplication optimization method and system in order to solve the problems of limited computing power and access efficiency and the like in the prior art. Decomposing the matrix into a plurality of sub-matrix blocks according to the slave core array topology; the slave core receives the sub-matrix blocks issued by the master core, divides the sub-matrix blocks into small sub-matrix blocks based on a uniform blocking rule, loads the small sub-matrix blocks to an independent buffer area of a local data memory based on a DMA double-buffer protocol, divides the small sub-matrix blocks in the buffer area into SIMD vectors according to the SIMD unit characteristics of the slave core, and sends the SIMD vectors to the slave core; vectorization calculation and caching operation are alternately switched according to an iteration period through different independent buffer areas; and after all the slave cores finish calculation, the master core collects results written back to the master memory by the slave cores to obtain a final operation result, and double breakthrough of calculation power and memory access efficiency is realized.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Translating between CXL.mem and CXL.cache read transactions

Memory has been playing a major role in the performance, scalability and applicability of General Compute systems, and more recently, in realizing Generative Artificial Intelligence (GenAI) and High-Performance Computing (HPC) systems that scale to thousands of GPUs, CPUs and special-purpose Accelerators. Embodiments herein disclose efficient software-defined protocol terminations and protocol translations utilizing Compute Express Link (CXL), including translations between CXL.mem and CXL.cache protocols. Also disclosed are CXL-based systems, Resource Provisioning Units (RPUs), and Memory Fabric Switches enabling dynamic memory pooling and sharing, host-to-host communication utilizing CXL.mem, CXL.cache and CXL.io, intent-based protocol translations, and optionally seamless interactions between CXL, UALink, NVLink, and / or Ethernet protocols, utilizing a broad range of semantics including IO, Cache, and Memory, optimizing memory access and reducing latency. Some embodiments also enable scalability, flexibility and security in high-performance architectures suited for data centers and next-generation computing environments.
Owner:HYATT GAYA OPAL MS +1

Methods and systems for tacit knowledge generation using high performance computing in document synthesis

The present disclosure herein addresses the problem of synthesizing a series of documents and extracting or summarizing meaningful information or content embedded as tacit knowledge in the series of documents. The embodiment of the present disclosure provides a system and method for tacit knowledge generation using large language model (LLM) in document synthesis. The method of the present disclosure performs intelligent document generation orchestrating a generative artificial intelligence solution workflow. In the present disclosure, tacit knowledge of subject matter experts in a knowledge base or in a series of documents is extracted. Further a content capturing the tacit knowledge is generated leveraging a large language models (LLMs) framework as the underlying architecture. The system of the present disclosure is artificial intelligence (AI) accelerated, cloud agnostic, latency defined, and security enabled.
Owner:TATA CONSULTANCY SERVICES LTD

DMA (Direct Memory Access) communication device for computing network integration computing architecture and working method of DMA communication device

The invention discloses a direct memory access (DMA) communication device for a computing network convergence computing architecture and a working method of the DMA communication device. The DMA communication device comprises a DMA transaction processing module and a protocol conversion module; the DMA transaction processing module is used for analyzing related DMA read-write requests, realizing processing of the DMA read-write requests and receiving of response data of the DMA read requests and completing processing of interrupt requests at the same time, and the protocol conversion module is used for completing conversion of read-write requests between a DMA interface and an AXI interface and mapping of response states. The invention aims to realize efficient protocol conversion between AXI and DMA interfaces, avoid the problem of DMA read request starvation caused by resource competition and ensure the consistency of DMA data access so as to improve the performance of an accelerator in high-performance calculation and artificial intelligence application, reduce transmission delay and improve data transmission bandwidth and energy efficiency.
Owner:NAT UNIV OF DEFENSE TECH

Cholesky decomposition heterogeneous parallel optimization method and system based on SW architecture

The invention provides a Cholesky decomposition heterogeneous parallel optimization method and system based on an SW architecture, and relates to the technical field of high-performance computing. The method comprises the following steps: performing sub-block division on a symmetric positive definite matrix based on a distributed parallel distribution scheme, and performing iteration to complete matrix decomposition; each sub-block is distributed to different processes through an MPI programming model, data exchange is carried out between the processes through asynchronous communication, and coarse-grained task-level parallel acceleration is carried out; performing two-stage parallel acceleration on four operations in Cholesky decomposition by utilizing the acceleration parallel characteristic of a master core and a slave core of the SW architecture; wherein for GEMM and SYRK operations, column vectors of a matrix are mapped to a slave core array, and columns are divided according to the number of slave cores; the calculation process is optimized through a double-buffering mechanism, vectorization operation and a loop expansion technology, and the parallel efficiency is improved; and for the TRSM operation, the TRSM operation is decomposed into a plurality of TRSV operations, the TRSV operations are allocated to the slave cores for parallel execution, and a circular reading and data broadcasting mode is adopted to reduce data dependence and realize efficient parallel calculation.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Switched Protocol Transformer for High-Performance Computing (HPC) and AI Workloads

Embodiments for communicating using a switch configured to establish multiple types of communication routes. First and second upstream switch ports (USPs) communicate with first and second hosts according to first and second Compute Express Link (CXL) protocols, respectively. A downstream switch port (DSP) communicates with a device according to a third CXL protocol. The switch couples the first USP to the DSP via a first route traversing a single Virtual CXL Switch (VCS), and couples the first USP to the second USP via a second route traversing two VCSs. Optionally, the switch includes a Resource Provisioning Unit (RPU) coupling the two VCSs of the second route, terminating the first and second CXL protocols, and translating between CXL messages conforming to the first and second CXL protocols.
Owner:HYATT GAYA OPAL MS +1

Communication method between graphics processors, product, equipment and medium

The invention discloses a communication method between graphics processors, a product, equipment and a medium, relates to the technical field of high-performance calculation and artificial intelligence acceleration, and is applied to a stand-alone system comprising a plurality of graphics processors and a central photoelectric hybrid switching chip constructed based on interconnection of an electric switching matrix and an optical switching matrix. The graphics processor is connected with the central photoelectric hybrid switching chip through an optical link and an electric link; the method comprises the following steps: performing data classification on to-be-transmitted data of a source graphics processor; if the to-be-transmitted data is a control flow, the source graphics processor is controlled to send the to-be-transmitted data to an electric switching matrix through an electric link, and then the to-be-transmitted data is routed to a target graphics processor in the graphics processors; and if the to-be-transmitted data is a data stream, the source graphics processor is controlled to send the to-be-transmitted data to the optical switching matrix through the optical link, and then the to-be-transmitted data is routed to the target graphics processor. Interconnection between graphics processors is optimized to improve communication efficiency between graphics processors.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Computing resource allocation method for distributed supercomputing center

The invention relates to the technical field of high-performance computing resource management, and discloses a computing resource allocation method for a distributed supercomputing center. The method comprises the following steps: on the basis of obtaining real-time computing task and supercomputing center resource data and uniformly quantifying, integrally predicting resource requirements of future tasks; constructing a mixed integer linear programming model with the minimization of the total operation cost as a single target, wherein the total operation cost is the sum of the energy cost, the carbon emission cost, the data transmission cost and the SLA default penalty cost; solving the model by taking the time-varying electricity price, the green energy ratio, the resource capacity and the network parameters of each center as constraint conditions to generate an optimal resource allocation scheme; and then, by dynamically monitoring the resource state and the task progress, the model is triggered to resolve when the resource utilization rate is detected to be unbalanced or default risks, so that self-adaptive adjustment is realized. According to the invention, global collaborative resource allocation across super computing centers is realized, and operation economy, environmental sustainability and service reliability are considered.
Owner:CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY

High-performance computing memory optimization method and system based on cache classification

The invention discloses a high-performance computing memory optimization method and system based on cache classification, and the method comprises the following steps: S1, dividing a physical memory page into different categories according to the number of cache groups mapped to a target cache, and distributing the divided categories to a target process running in the target cache; and S2, when the target process requests the physical memory page from the operating system, allocating the physical memory page of the same category to the target process based on the category corresponding relationship between the process and the physical memory page. According to the method, the problem of cache conflict among a plurality of processes under the condition of high CPU utilization rate is solved, so that the performance of the whole system is improved.
Owner:KYLIN CORP

Temperature regulation and control method for high-performance computer system

The invention relates to the field of computer heat dissipation control, and discloses a high-performance computer system temperature regulation and control method which comprises the following steps: dividing a computer system into a plurality of hot areas and initializing thermodynamic parameters; collecting temperature, power consumption and process data in real time, identifying a burst process and constructing a dynamic heat capacity model; the heat capacity parameter and the heat dissipation coefficient are updated online through a recursive least square method; the multi-target weight is dynamically adjusted based on the temperature deviation degree and the heat capacity change rate, a distributed optimization model containing heat diffusion coupling constraints is established, and an optimal heat dissipation strategy is solved by adopting an improved ADMM algorithm; model prediction deviation is corrected through a double-loop control mechanism, and data updating circulation is triggered in combination with burst process marks and temperature out-of-limit events. According to the method, through a dynamic modeling, weight self-adaption and closed-loop control fusion mechanism, the temperature control precision and the system safety in a complex thermal coupling scene are remarkably improved, and meanwhile heat dissipation efficiency, noise suppression and energy efficiency balance optimization are achieved.
Owner:QINGDAO HAIKUOTIANGAO INFORMATION TECH CO LTD

Automatic SEM image analysis and process defect detection system based on full database

The invention provides an automatic SEM image analysis and process defect detection system based on a full database, and relates to the technical field of semiconductor manufacturing, the system combines self-adaptive feature extraction and density anomaly detection of a high-performance calculation unit through dynamic feature fusion (Sp1) of a multi-scale SEM image and a multi-channel acquisition device, and realizes automatic analysis of the SEM image. Precise classification of defect types and automatic identification of new defects are achieved, the feature range and the fusion weight are dynamically adjusted, traditional static detection limitation is broken through, adaptive analysis of process context is supported, the detection precision is improved to 98%, the misjudgment and missing judgment rate is reduced by 30%-50%, especially in advanced processes such as 3nm, the new defects can be rapidly identified, rules can be updated, and the method is suitable for large-scale popularization and application. The process research and development efficiency and the product reliability are remarkably improved, and the flexible and intelligent detection capability is provided for semiconductor manufacturing.
Owner:上海芯无双仿真科技有限公司

High-computing-power special processor architecture for big data file and processing method

According to the high-computing-power special processor architecture for the big data file and the processing method, a data preprocessing unit receives an original big data file and formats and cleans the original big data file, and a special acceleration computing chip executes a computing-intensive task and avoids the floating-point number precision problem; the parallel processing unit realizes parallel data processing through a MapReduce programming model, the cache storage architecture adopts a multi-level cache system and an LRU cache elimination algorithm to store an intermediate result, and the instruction set optimization module performs parallel operation on isomorphic continuous storage data. The method can be widely applied to multiple fields of data analysis, machine learning, high-performance calculation and the like, and has great significance in improving the data processing capacity in the big data era.
Owner:HAIER CONSUMER FINANCE CO LTD

Intelligent blasting parameter calculation platform

The invention discloses a blasting parameter intelligent computing platform, which belongs to the field of blasting engineering, and comprises a distributed computing resource scheduling module, which adopts a distributed computing architecture to decompose a computing task into a plurality of sub-tasks, and distributes the sub-tasks to different computing nodes for parallel computing; through an intelligent scheduling algorithm, according to factors such as the load condition and the computing power of each computing node, computing tasks are dynamically allocated, dispersed computing resources in a network are fully utilized, the utilization rate of the computing resources is improved, and dependence on single high-performance computing equipment is reduced; data are processed in real time through the edge calculation preprocessing module, and the data transmission time is shortened; the lightweight calculation model library can quickly complete blasting parameter calculation; and the real-time feedback and dynamic optimization module realizes real-time adjustment and optimization of blasting parameters, so that accurate blasting parameters can be provided in time according to actual conditions in the tunnel construction process, and the construction progress is improved.
Owner:SHANDONG UNIV

Computer architecture with disaggregated memory and high-bandwidth communication interconnects

Conventional high performance computer connections are electron-based systems, which require the memory packages to be as close as mechanically possible to the computation engine. Low power and high bandwidth long distance communication, e.g. photonic or electronic, links can drastically change the architecture of high-performance computers by eliminating the bottlenecks in communication. A computer system comprises: a plurality of memory aggregation devices configured to retrieve data from and store data in a plurality of random access memory modules forming a unified contiguous memory address space disaggregated from a processing unit; one or more computational devices configured for simultaneously launching a plurality of data signals including memory read and / or write requests for the data to the plurality of memory aggregation devices; and a plurality of communication links coupling each of the plurality of memory aggregation devices to each of the one or more computational devices for transferring the data therebetween.
Owner:ADVANCED MICRO DEVICES INC

Container-based parallel computing system

A container-based parallel computing system for executing high-performance computing (HPC) applications. The system leverages container technology to package the applications executed at the nodes in a cluster. To load and execute a job in the parallel computing system, containers are deployed in a cluster that include all the application resources and configuration information that the particular HPC application needs to execute. An event-driven batch scheduler may be used to dynamically allocate resources for executing multi-node jobs in the container-based parallel computing system, handling the coordination of resource allocation for the customer. The scheduler insures that jobs begin executing as fast as possible, and handles failure conditions such as partial scaling. Virtual network interfaces are attached to the containers that allow the containers to connect to and communicate with other containers in the cluster directly through the network interfaces of host machines using IP addresses provided by the virtual network interfaces.
Owner:AMAZON TECH INC

Heterogeneous processor-oriented reciprocal calculation instruction sequence generation method

The invention discloses a reciprocal calculation instruction sequence generation method oriented to a heterogeneous processor, and belongs to the field of compilation optimization and code generation. Aiming at the problems of instruction redundancy, weak precision control, poor hardware adaptation and high manual dependence of an existing method in a heterogeneous environment, characteristics of a reciprocal instruction and an operand are accurately identified by linearly scanning heterogeneous object codes (including vectorization, scalar and complex instruction sequences); in combination with hardware characteristics of RISC / SIMD / VLIW / DSP and the like, a multi-round iteration precision improvement and temporary register optimization allocation strategy is adopted, differential generation logic is formulated, and a high-precision low-redundancy instruction sequence is generated. The method comprises linear code scanning classification, reciprocal instruction and operand identification, cross-architecture generation logic rule formulation, instruction sequence generation and legality verification. Full-process automation is achieved, manual intervention is reduced, the execution efficiency and precision of reciprocal calculation of the heterogeneous processor are improved, and the method is suitable for embedded systems, high-performance calculation and other scenes.
Owner:HUNAN UNIV OF SCI & TECH

Predictive diagnostics in high-performance computing

A development system for predictive diagnostics is provided. During operation, the system can perform a first diagnostic test on a distributed computing system based on a first restriction level indicating resource consumption of a first set of hardware units. The distributed computing system can include a plurality of computing devices with processing and memory resources. The system can generate a first log comprising a first set of parameter values indicating an output of the first diagnostic test at the first restriction level of the distributed computing system. The system can configure a first diagnostic tool with the first set of parameter values to emulate the first diagnostic test. The system can then apply the first diagnostic tool to obtain a second set of parameter values indicating an output of the first diagnostic test at a second restriction level, which can be higher than the first restriction level.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

A lifecycle management system and method for scientific computing programs

This invention discloses a full lifecycle management system and method for scientific computing programs. The system includes a build environment subsystem and a production environment subsystem. The former provides computer resources for the build process of the scientific computing program throughout its lifecycle, while the latter provides computer resources for the testing and deployment processes. This invention, through a full lifecycle management method for scientific computing programs, operates on corresponding computing resources, encompassing a series of steps including querying, triggering scheduling, build execution, result distribution, and test deployment. It automatically generates the executable file of the scientific computing program and configures its runtime dependencies. Simultaneously, it automatically generates a corresponding description file recording the entire lifecycle process. Furthermore, based on version management of these description files and their sets, it achieves full lifecycle traceability and cross-platform migration and deployment of scientific computing programs in a high-performance computing environment.
Owner:HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Software and hardware collaborative RDMA network card SR-IOV parameterization configuration method

The invention discloses a software and hardware collaborative RDMA network card SR-IOV parameterization configuration method, which can reduce redundant calculation steps through the collaboration of FPGA hardware acceleration and driving parameterization configuration, greatly improves the power consumption compared with a traditional pure software design scheme, and enables the VF performance of a network card to be close to the level of a physical network card. According to the method, the QP number, the MAC address and the RDMA enabling mark of the VF are dynamically adjusted through parameterization of the configuration file, and the operation and maintenance management efficiency can be effectively improved without manual item-by-item configuration; forwarding and isolation of CF traffic are achieved through the FPGA hardware layer, traffic sniffing between virtual machines is prevented, and safety is enhanced; the realized VF supports standard Ethernet communication and RDMA at the same time, the resource utilization rate is improved, and the requirements of hybrid scenes such as high performance computing (HPC) and cloud computing are met.
Owner:XIDIAN UNIV

Memory database starting method, system and equipment and medium

The invention provides a memory database starting method, system and device and a medium, and belongs to the technical field of computers. The method comprises the following steps: detecting a current project scale and a performance index of an original server, and determining whether to trigger a memory processing mechanism; if a memory processing mechanism is triggered, selecting a memory server according to high-performance computing and expansibility requirements, accessing an original server by utilizing a hot plug technology, configuring a memory database and initializing a memory computing cluster; configuring a to-be-synchronized database table by utilizing a metadata management function, loading original data to a memory database through a full-amount synchronization and incremental updating mechanism, and starting real-time data synchronization; starting a memory computing cluster, automatically detecting a memory processing state on an original server when a service task is triggered, and routing a large quantity of data tasks meeting conditions to a memory server to execute parallel computing. According to the method, the memory can be dynamically started for data processing, the operation efficiency is improved, and the user experience is improved.
Owner:INSPUR GENERSOFT CO LTD

Fusion computing method and system combining cloud platform high-performance computing power and quantum computing power

The present invention relates to a fusion computing method and system combining cloud platform high-performance computing power and quantum computing power, and belongs to the technical field of quantum computing. The method comprises: submitting a quantum program in a quantum virtual machine by means of a classical computer interface; carrying out initial classical-quantum task splitting by means of a compiler, so as to generate a plurality of computing subtasks; inserting communication primitives before the computing subtasks requiring remote invocation, and by means of the communication primitives, distributing corresponding quantum computing subtasks to a quantum computing backend and distributing corresponding classical computing subtasks to a high-performance computing power cluster; locally executing part of the classical computing subtasks, and remotely invoking, by means of RoCE communication, the classical computing subtasks requiring to be completed by the high-performance computing power cluster and the quantum computing subtasks requiring to be completed by the quantum computing backend; and by means of a resource scheduler, receiving the classical computing subtasks completed by the high-performance computing power cluster and the quantum computing subtasks completed by the quantum computing backend.
Owner:CHINA TELECOM CLOUD TECH CO LTD

Streamlined CXL Memory Fabric with Lightweight Scalable Provisioning

Large-scale compute environments for Artificial Intelligence (AI) and High-Performance Computing (HPC) may utilize memory fabric infrastructures to achieve higher workload performance compared to conventional networking-only infrastructures. However, the role of memory fabrics extends beyond high-end use cases into general compute scenarios in public cloud, private cloud, and hybrid enterprise environments, where Memory-as-a-Service (Memory-aaS) augments and complements the portfolio of services provided to tenants in the datacenter. Embodiments herein disclose a streamlined CXL memory fabric utilizing a Resource Provisioning Unit (RPU), enabling lightweight, scalable provisioning of memory to CPUs, GPUs, and accelerators via CXL.io, CXL.mem, and CXL.cache. These embodiments support translation between asymmetric and symmetric memory transactions and operate in both switch-enabled CXL environments, such as CXL 2.0 and CXL 3.x, and non-switched CXL environments, such as CXL 1.1. The embodiments may further enable multi-tier memory pools and advanced protocol processing in switches, leveraging existing rack-level and pod-level topologies.
Owner:HYATT GAYA OPAL MS +1

Method for automatically generating simulation calculation task on supercomputing platform

The invention relates to the technical field of high-performance computing and artificial intelligence crossing, and discloses a method for automatically generating a simulation computing task on a supercomputing platform, which comprises the following steps of: receiving a task demand of a user; performing semantic understanding and parameter extraction on the input by using a parameter dynamic analysis engine to generate a simulation parameter table; based on the simulation parameter table and the input format specification of the target simulation software, an executable parallel computing task script is automatically generated through a self-adaptive script generator; according to the real-time resource state and the task requirement of the supercomputing platform, a dynamic scheduling algorithm is adopted to distribute the task script to the optimal computing node; in the task execution process, the running state is monitored in real time, and when abnormity is detected, a fault-tolerant and self-repairing mechanism is triggered; by automatically generating the task script, the user parameter configuration time is saved, and the overall scientific research efficiency is improved by more than 30%; natural language input and zero code configuration are supported, and common scientific researchers can quickly master the method.
Owner:HEFEI ADVANCED COMPUTING CENT OPERATION MANAGEMENT CO LTD