Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

478 results about "Performance computing" patented technology

Integer parallel computing method and device based on distributed storage and computer equipment

The invention belongs to the field of high-performance computing, and relates to an integer parallel computing method and device based on distributed storage and computer equipment, and the method comprises the steps of collecting real-time resource indexes, dynamically identifying fault nodes, triggering task migration, and performing data verification and hard disk fault detection. The weight value of each node is calculated, the nodes are arranged according to the descending order of the weight values, and the nodes with high load capacity are selected to distribute tasks; dynamically distributing a data generation task to a computing node, executing parallel computing, and performing distributed storage on a result; obtaining an operand, converting the operand into a first-order tensor form of a basic operand, serializing tensor data, and sending the serialized tensor data to a parallel computing layer; distributing a search task to a computing node, retrieving storage data in parallel, reading effective data from a storage layer, and combining search results into a partial sum; and summarizing and then outputting. The system has dynamic resource management and fault-tolerant capabilities, and can realize efficient task allocation and load balancing.
Owner:SHENZHEN Y& D ELECTRONICS CO LTD

Parallel task scheduling algorithm for heterogeneous multi-core processor

The invention relates to the technical field of computer architecture and parallel computing, and discloses a parallel task scheduling algorithm for a heterogeneous multi-core processor, which comprises the steps of task modeling, resource mapping, dynamic load balancing, communication optimization, task scheduling decision and execution monitoring. Task allocation is adjusted in real time through dynamic load balancing, cross-core communication delay is reduced in combination with communication optimization, and an efficient task allocation sequence is generated by using an improved genetic algorithm. According to the method, the resource utilization rate and the task execution efficiency of the heterogeneous multi-core processor in a high-performance computing scene can be improved, meanwhile, the robustness and adaptability of an algorithm are enhanced, and the task allocation problem in a complex computing scene is effectively solved.
Owner:SUZHOU DUXUEKEZHENG INTELLIGENT TECH CO LTD

Neural network large model efficient reasoning method based on multiple GPGPUs

The invention belongs to the technical field of artificial intelligence and high-performance computing, and particularly relates to a neural network large model efficient reasoning method based on multiple GPGPUs. The method aims to solve the problems of high communication overhead, non-uniform load, low resource utilization rate, high data transmission delay and the like among multiple processors. Dividing a calculation task into a plurality of sub-graphs through static analysis and mixed granularity partitioning of a model calculation graph; distributing the sub-graphs to the optimal GPGPU based on a weighted cost function in combination with heterogeneous resource perception and a dynamic mapping strategy; a global pipeline scheduling plan is constructed by using communication topology perception, and calculation and communication overlap are maximized; data are loaded in advance through a host side hierarchical caching and asynchronous prefetching mechanism, and transmission delay is hidden; multi-stream concurrent execution and event-based lightweight synchronization are adopted on each GPGPU, so that waiting overhead is reduced. According to the method, the reasoning delay can be remarkably reduced, the throughput and the hardware utilization rate are improved, and the method has good adaptivity and expandability.
Owner:BEIJING TOPMOO TECH

Memory architecture-oriented dual-precision general matrix multiplication optimization method and system

The invention belongs to the related technical field of high-performance computing, and provides a memory architecture-oriented dual-precision general matrix multiplication optimization method and system in order to solve the problems of limited computing power and access efficiency and the like in the prior art. Decomposing the matrix into a plurality of sub-matrix blocks according to the slave core array topology; the slave core receives the sub-matrix blocks issued by the master core, divides the sub-matrix blocks into small sub-matrix blocks based on a uniform blocking rule, loads the small sub-matrix blocks to an independent buffer area of a local data memory based on a DMA double-buffer protocol, divides the small sub-matrix blocks in the buffer area into SIMD vectors according to the SIMD unit characteristics of the slave core, and sends the SIMD vectors to the slave core; vectorization calculation and caching operation are alternately switched according to an iteration period through different independent buffer areas; and after all the slave cores finish calculation, the master core collects results written back to the master memory by the slave cores to obtain a final operation result, and double breakthrough of calculation power and memory access efficiency is realized.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Methods and systems for tacit knowledge generation using high performance computing in document synthesis

PendingUS20250299018A1Digital data information retrievalBiological modelsTacit knowledgeSubject-matter expert
The present disclosure herein addresses the problem of synthesizing a series of documents and extracting or summarizing meaningful information or content embedded as tacit knowledge in the series of documents. The embodiment of the present disclosure provides a system and method for tacit knowledge generation using large language model (LLM) in document synthesis. The method of the present disclosure performs intelligent document generation orchestrating a generative artificial intelligence solution workflow. In the present disclosure, tacit knowledge of subject matter experts in a knowledge base or in a series of documents is extracted. Further a content capturing the tacit knowledge is generated leveraging a large language models (LLMs) framework as the underlying architecture. The system of the present disclosure is artificial intelligence (AI) accelerated, cloud agnostic, latency defined, and security enabled.
Owner:TATA CONSULTANCY SERVICES LTD

DMA (Direct Memory Access) communication device for computing network integration computing architecture and working method of DMA communication device

The invention discloses a direct memory access (DMA) communication device for a computing network convergence computing architecture and a working method of the DMA communication device. The DMA communication device comprises a DMA transaction processing module and a protocol conversion module; the DMA transaction processing module is used for analyzing related DMA read-write requests, realizing processing of the DMA read-write requests and receiving of response data of the DMA read requests and completing processing of interrupt requests at the same time, and the protocol conversion module is used for completing conversion of read-write requests between a DMA interface and an AXI interface and mapping of response states. The invention aims to realize efficient protocol conversion between AXI and DMA interfaces, avoid the problem of DMA read request starvation caused by resource competition and ensure the consistency of DMA data access so as to improve the performance of an accelerator in high-performance calculation and artificial intelligence application, reduce transmission delay and improve data transmission bandwidth and energy efficiency.
Owner:NAT UNIV OF DEFENSE TECH

Communication method between graphics processors, product, equipment and medium

The invention discloses a communication method between graphics processors, a product, equipment and a medium, relates to the technical field of high-performance calculation and artificial intelligence acceleration, and is applied to a stand-alone system comprising a plurality of graphics processors and a central photoelectric hybrid switching chip constructed based on interconnection of an electric switching matrix and an optical switching matrix. The graphics processor is connected with the central photoelectric hybrid switching chip through an optical link and an electric link; the method comprises the following steps: performing data classification on to-be-transmitted data of a source graphics processor; if the to-be-transmitted data is a control flow, the source graphics processor is controlled to send the to-be-transmitted data to an electric switching matrix through an electric link, and then the to-be-transmitted data is routed to a target graphics processor in the graphics processors; and if the to-be-transmitted data is a data stream, the source graphics processor is controlled to send the to-be-transmitted data to the optical switching matrix through the optical link, and then the to-be-transmitted data is routed to the target graphics processor. Interconnection between graphics processors is optimized to improve communication efficiency between graphics processors.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Computing resource allocation method for distributed supercomputing center

The invention relates to the technical field of high-performance computing resource management, and discloses a computing resource allocation method for a distributed supercomputing center. The method comprises the following steps: on the basis of obtaining real-time computing task and supercomputing center resource data and uniformly quantifying, integrally predicting resource requirements of future tasks; constructing a mixed integer linear programming model with the minimization of the total operation cost as a single target, wherein the total operation cost is the sum of the energy cost, the carbon emission cost, the data transmission cost and the SLA default penalty cost; solving the model by taking the time-varying electricity price, the green energy ratio, the resource capacity and the network parameters of each center as constraint conditions to generate an optimal resource allocation scheme; and then, by dynamically monitoring the resource state and the task progress, the model is triggered to resolve when the resource utilization rate is detected to be unbalanced or default risks, so that self-adaptive adjustment is realized. According to the invention, global collaborative resource allocation across super computing centers is realized, and operation economy, environmental sustainability and service reliability are considered.
Owner:CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY

Automatic SEM image analysis and process defect detection system based on full database

The invention provides an automatic SEM image analysis and process defect detection system based on a full database, and relates to the technical field of semiconductor manufacturing, the system combines self-adaptive feature extraction and density anomaly detection of a high-performance calculation unit through dynamic feature fusion (Sp1) of a multi-scale SEM image and a multi-channel acquisition device, and realizes automatic analysis of the SEM image. Precise classification of defect types and automatic identification of new defects are achieved, the feature range and the fusion weight are dynamically adjusted, traditional static detection limitation is broken through, adaptive analysis of process context is supported, the detection precision is improved to 98%, the misjudgment and missing judgment rate is reduced by 30%-50%, especially in advanced processes such as 3nm, the new defects can be rapidly identified, rules can be updated, and the method is suitable for large-scale popularization and application. The process research and development efficiency and the product reliability are remarkably improved, and the flexible and intelligent detection capability is provided for semiconductor manufacturing.
Owner:上海芯无双仿真科技有限公司

High-computing-power special processor architecture for big data file and processing method

According to the high-computing-power special processor architecture for the big data file and the processing method, a data preprocessing unit receives an original big data file and formats and cleans the original big data file, and a special acceleration computing chip executes a computing-intensive task and avoids the floating-point number precision problem; the parallel processing unit realizes parallel data processing through a MapReduce programming model, the cache storage architecture adopts a multi-level cache system and an LRU cache elimination algorithm to store an intermediate result, and the instruction set optimization module performs parallel operation on isomorphic continuous storage data. The method can be widely applied to multiple fields of data analysis, machine learning, high-performance calculation and the like, and has great significance in improving the data processing capacity in the big data era.
Owner:HAIER CONSUMER FINANCE CO LTD

Container-based parallel computing system

A container-based parallel computing system for executing high-performance computing (HPC) applications. The system leverages container technology to package the applications executed at the nodes in a cluster. To load and execute a job in the parallel computing system, containers are deployed in a cluster that include all the application resources and configuration information that the particular HPC application needs to execute. An event-driven batch scheduler may be used to dynamically allocate resources for executing multi-node jobs in the container-based parallel computing system, handling the coordination of resource allocation for the customer. The scheduler insures that jobs begin executing as fast as possible, and handles failure conditions such as partial scaling. Virtual network interfaces are attached to the containers that allow the containers to connect to and communicate with other containers in the cluster directly through the network interfaces of host machines using IP addresses provided by the virtual network interfaces.
Owner:AMAZON TECH INC

Heterogeneous processor-oriented reciprocal calculation instruction sequence generation method

The invention discloses a reciprocal calculation instruction sequence generation method oriented to a heterogeneous processor, and belongs to the field of compilation optimization and code generation. Aiming at the problems of instruction redundancy, weak precision control, poor hardware adaptation and high manual dependence of an existing method in a heterogeneous environment, characteristics of a reciprocal instruction and an operand are accurately identified by linearly scanning heterogeneous object codes (including vectorization, scalar and complex instruction sequences); in combination with hardware characteristics of RISC / SIMD / VLIW / DSP and the like, a multi-round iteration precision improvement and temporary register optimization allocation strategy is adopted, differential generation logic is formulated, and a high-precision low-redundancy instruction sequence is generated. The method comprises linear code scanning classification, reciprocal instruction and operand identification, cross-architecture generation logic rule formulation, instruction sequence generation and legality verification. Full-process automation is achieved, manual intervention is reduced, the execution efficiency and precision of reciprocal calculation of the heterogeneous processor are improved, and the method is suitable for embedded systems, high-performance calculation and other scenes.
Owner:HUNAN UNIV OF SCI & TECH

Predictive diagnostics in high-performance computing

A development system for predictive diagnostics is provided. During operation, the system can perform a first diagnostic test on a distributed computing system based on a first restriction level indicating resource consumption of a first set of hardware units. The distributed computing system can include a plurality of computing devices with processing and memory resources. The system can generate a first log comprising a first set of parameter values indicating an output of the first diagnostic test at the first restriction level of the distributed computing system. The system can configure a first diagnostic tool with the first set of parameter values to emulate the first diagnostic test. The system can then apply the first diagnostic tool to obtain a second set of parameter values indicating an output of the first diagnostic test at a second restriction level, which can be higher than the first restriction level.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

A lifecycle management system and method for scientific computing programs

This invention discloses a full lifecycle management system and method for scientific computing programs. The system includes a build environment subsystem and a production environment subsystem. The former provides computer resources for the build process of the scientific computing program throughout its lifecycle, while the latter provides computer resources for the testing and deployment processes. This invention, through a full lifecycle management method for scientific computing programs, operates on corresponding computing resources, encompassing a series of steps including querying, triggering scheduling, build execution, result distribution, and test deployment. It automatically generates the executable file of the scientific computing program and configures its runtime dependencies. Simultaneously, it automatically generates a corresponding description file recording the entire lifecycle process. Furthermore, based on version management of these description files and their sets, it achieves full lifecycle traceability and cross-platform migration and deployment of scientific computing programs in a high-performance computing environment.
Owner:HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Method for automatically generating simulation calculation task on supercomputing platform

The invention relates to the technical field of high-performance computing and artificial intelligence crossing, and discloses a method for automatically generating a simulation computing task on a supercomputing platform, which comprises the following steps of: receiving a task demand of a user; performing semantic understanding and parameter extraction on the input by using a parameter dynamic analysis engine to generate a simulation parameter table; based on the simulation parameter table and the input format specification of the target simulation software, an executable parallel computing task script is automatically generated through a self-adaptive script generator; according to the real-time resource state and the task requirement of the supercomputing platform, a dynamic scheduling algorithm is adopted to distribute the task script to the optimal computing node; in the task execution process, the running state is monitored in real time, and when abnormity is detected, a fault-tolerant and self-repairing mechanism is triggered; by automatically generating the task script, the user parameter configuration time is saved, and the overall scientific research efficiency is improved by more than 30%; natural language input and zero code configuration are supported, and common scientific researchers can quickly master the method.
Owner:HEFEI ADVANCED COMPUTING CENT OPERATION MANAGEMENT CO LTD

Slurm scheduling specification integration method and system

The invention relates to the field of high-performance computing cluster resource scheduling management, and discloses a Slurm scheduling specification integration method and system, and the method comprises the following steps: 1, analyzing a heterogeneous job description file, and extracting a resource demand parameter and a dependency relationship through a regular expression rule base; 2, based on the extracted original parameters, converting the original parameters into SLURM standard parameters through a preset mapping rule; step 3, according to the converted standard parameters; 4, calling a Slurm interface command to submit a script; and step 5, monitoring the execution state of the submitted job, and triggering a re-submission process for the abnormal job with the resource overrun or dependency missing. Through a multi-level analysis architecture and a regular expression rule base, job description files in different formats are effectively compatible, the problem that analysis of a non-standardized input format by a traditional method fails is solved, unified processing of cross-platform job definition is achieved, and the manual adaptation cost is remarkably reduced.
Owner:北京月新时代科技股份有限公司

Switched protocol transformer for high-performance computing (HPC) and AI workloads

Embodiments for communicating using a switch configured to establish multiple types of communication routes. First and second upstream switch ports (USPs) communicate with first and second hosts according to first and second Compute Express Link (CXL) protocols, respectively. A downstream switch port (DSP) communicates with a device according to a third CXL protocol. The switch couples the first USP to the DSP via a first route traversing a single Virtual CXL Switch (VCS), and couples the first USP to the second USP via a second route traversing two VCSs. Optionally, the switch includes a Resource Provisioning Unit (RPU) coupling the two VCSs of the second route, terminating the first and second CXL protocols, and translating between CXL messages conforming to the first and second CXL protocols.
Owner:HYATT GAYA OPAL MS +1

High-performance front-end audio and video processing method and system based on WebAssembly

The invention relates to the technical field of browser-side high-performance computing, in particular to a WebAssembly-based high-performance front-end audio and video processing method and system, and the method comprises the following steps: designing a modularized Wasm runtime container, constructing a zero-overhead memory interaction mechanism, and establishing a hardware adaptive scheduling engine; the method has the beneficial effects that dynamic loading of an audio and video algorithm is realized by designing a modularized Wasm runtime container, serialization overhead of communication between threads is eliminated by utilizing a shared memory mechanism, a hardware self-adaptive dynamic scheduling system is created, and finally three core objectives are achieved: execution efficiency close to native codes is provided in a pure browser environment; establishing a standardized processing pipeline without dependence of a third party; intelligent scheduling and elastic expansion and contraction of terminal heterogeneous computing resources are realized, and Web end infrastructure support is provided for professional audio and video applications.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Rail transit data processing system adopting high-performance computing power node board card

The invention relates to the technical field of rail transit data processing, in particular to a rail transit data processing system adopting a high-performance computing power node board card, which is used for solving the problems that the existing rail transit data processing technology has obvious defects in a complex operation scene, the state prediction depends on a black box model, and the reliability is poor. Explicit modeling of physical constraints such as track curvature, gradient and speed limit is lacked, and real anomaly and model misalignment are difficult to distinguish; structured features and compressed time sequence data are fused through the rail traffic physical credible prediction module, the train operation stage is intelligently recognized, adaptive deep learning, time sequence Bayesian or reinforcement learning sub-models are dynamically called, high-precision state prediction is achieved, meanwhile, physical constraints such as the track curvature, the gradient and the speed limit are combined, and the reliability and reliability of the system are improved. And the causal consistency verification based on the mechanics principle is executed, so that the physical mismatch abnormity is effectively discriminated, and the prediction credibility, the abnormity identification accuracy and the operation and maintenance decision intelligent level are remarkably improved.
Owner:NANJING XINYUANTONG INTELLIGENT TECH CO LTD

Decentralized high-performance multicast method and system based on Gossip and RDMA

The invention discloses a decentralized high-performance multicast method and system based on Gossip and RDMA, and relates to the field of computer network communication. The method has better reliability, realizes zero-loss transmission of data packets through message de-duplication, timeout retransmission, ACK / NACK convergence and a dynamic topology adaptation mechanism, ensures data consistency, and solves the problem of insufficient reliability of traditional UDP multicast; the method has the advantages that the real-time performance is high, the delay is low, the throughput is high, the zero copy and kernel bypass characteristics of the RDMA and the batch processing mechanism of the MPMC are combined, the average delay is lower than 10 microseconds, the throughput in a large data packet scene reaches 374GB / s, and the low-delay requirements of high-frequency transaction, high-performance calculation and the like are met; the method has high expansibility, is based on a decentralized architecture of an improved Gossip protocol, does not need core node maintenance topology, supports dynamic joining or quitting of nodes, can adapt to large-scale cluster expansion, and solves the expansibility bottleneck of centralized multicast.
Owner:CHONGQING UNIV OF TECH

Single namespace for high performance computing system

A data processing architecture controls data processing arbitration in a high performance computing system that includes one or more premises. Individual premises can include one or more server computers executing an instance of a local file system and including one or more temporary data storage devices. Individual instances of the local file system can access files stored in objects of a primary data store. Individual objects of the primary data store can be accessed using a common identifier indicating a storage location of the individual objects in the primary data store.
Owner:GUARDANT HEALTH INC

Inter-thread data sharing method, electronic equipment, storage medium and program product

The invention relates to the technical field of high-performance computing, and provides an inter-thread data sharing method, electronic equipment, a storage medium and a program product.The method comprises the steps that a plurality of threads in a thread block are divided into at least one cooperative thread group, and the threads in each cooperative thread group are from at least two different thread bundle groups; when a write operation of a first thread in any cooperative thread group is received, writing target data into a shared register resource associated with any cooperative thread group; and when a read operation of a second thread in any cooperative thread group for the target data is received, executing a synchronous operation, and reading the target data from the shared register resource after the execution is completed. According to the method, the cooperative thread group spanning different thread beam groups is constructed, and the shared register resources are utilized for data exchange, so that efficient data sharing is directly carried out among the threads spanning the thread beam groups through the register, and the method has the advantages of low delay, high bandwidth and capability of remarkably improving parallel computing performance.
Owner:SHANGHAI BIREN TECH CO LTD

Computer architecture with disaggregated memory and high-bandwidth communication interconnects

Conventional high performance computer connections are electron-based systems, which require the memory packages to be as close as mechanically possible to the computation engine. Low power and high bandwidth communication, e.g. photonic, links can drastically change the architecture of high-performance computers by eliminating the bottlenecks in communication and augment existing memory systems to allow them to be both high capacity and high bandwidth simultaneously. A computer system comprises: a plurality of memory aggregation devices configured to retrieve data from and store data in a plurality of random access memory modules forming a unified contiguous memory address space disaggregated from a processing unit; a plurality of computational devices configured for simultaneously launching a plurality of data signals including memory read and / or write requests for the data to the plurality of memory aggregation devices; and a plurality of communication links coupling each of the plurality of memory aggregation devices to each of the plurality of computational devices for transferring the data therebetween.
Owner:ADVANCED MICRO DEVICES INC

Storage and calculation integrated structure and chip

The invention provides a storage and calculation integrated structure and a chip, and belongs to the technical field of integrated circuits, and the storage and calculation integrated structure comprises at least one near storage core grain which is used for providing large-capacity storage; the in-memory computing unit is used for providing computing processing, and the near-memory core grains provide cache space for data and / or instructions processed by the in-memory computing unit; the at least one logic processing core particle is used for loading the data and / or instructions stored in the near memory core particle into the in-memory computing unit, and updating and storing the data and / or instructions processed by the in-memory computing unit into the near memory core particle; and the logic processing core particle is integrated with the near memory core particle and the in-memory computing unit through hybrid bonding and / or TSV (Through Silicon Via). The storage and calculation integrated structure is applied to high-performance calculation, the high-performance calculation requirement and the large-capacity storage requirement can be met at the same time, the high-performance calculation of the system is maximized, and support is provided for the development of the calculation power technology from the hardware structure.
Owner:XI AN UNIIC SEMICON CO LTD

Consciousness perception quantum gravitational nerve acceleration unified computing platform and method based on formalized verification

The invention discloses an awareness perception quantum gravitational nerve acceleration unified computing platform and method based on formalized verification, and belongs to the field of high-performance computing, medical AI and quantum information processing. A quantum gravitational coupling neural processing unit (QGCNPU); a consciousness information integral calculation core (phi ICC); formalizing a receipt chain verifier (FRCV); and the mirror reflection symmetry correction module (MRSCM) is used for eliminating the continuous projection mistake and the discrete Boolean mistake through topological equivalence verification of a local observer and a global GodCam visual angle. The system adopts an FPGA + ASIC + quantum processor three-layer heterogeneous architecture, and supports real-time pathological modeling (precision gt; gt) of neurodegenerative diseases such as Alzheimer's disease, Parkinson's disease and ALS; 99.7%), drug target discovery (acceleration by 50 times) and consciousness state monitoring (time resolution lt; and meanwhile, the mathematical proving performance and the clinical auditing performance of the calculation process are ensured, and the requirements of the global 152 billion dollar neural science and technology market and the 89 billion dollar quantum calculation industry are met.
Owner:GUANGZHOU KINGPIN IND CO LTD

Heterogeneous computing method, device and equipment and computer readable storage medium

The invention provides a heterogeneous computing method, a heterogeneous computing device, heterogeneous computing equipment and a computer readable storage medium, and is applied to the technical field of data processing. The method comprises the following steps: when a heterogeneous computing power card accesses data, a central processing unit firstly determines whether a target object accessed this time exists in a video memory or not; and under the condition that the target object does not exist in the video memory, the central processing unit reads the target object from the memory and stores the target object in the video memory. Further, the central processing unit can provide the target object loaded into the video memory for the heterogeneous computing power card. In the scheme provided by the embodiment of the invention, the heterogeneous computing data meeting the preset condition is transferred from the video memory to the memory, and the data written into the memory is subjected to a realistic copying strategy, so that the storage space of the heterogeneous computing data can be expanded, and the utilization rate of the video memory space can be improved, thereby realizing video memory oversell, and improving the user experience. And thus, high-performance calculation of the heterogeneous computing power card can be guaranteed.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Sparse matrix multiplication acceleration method, system and equipment based on tensor processor

The invention provides a sparse matrix multiplication acceleration method based on a tensor processor, which relates to the field of high-performance computing, and comprises the following steps: carrying out block processing on a sparse matrix to obtain sparse matrix blocks matched with the register granularity of the tensor processor; performing block processing on the dense matrix to obtain dense matrix blocks; generating a matrix multiplication task set based on the sparse matrix block and the dense matrix block; and distributing the task set to a plurality of threads, controlling the plurality of threads to respectively call the tensor processors to execute matrix multiplication operation, and aggregating calculation results of the threads to obtain an output matrix. Compared with the prior art, the sparse matrix blocks are processed into the sparse matrix blocks matched with the register granularity of the tensor processor, the index coding and decoding overhead, the reconstruction overhead and the communication overhead of the sparse matrix are reduced, and therefore the operation performance of sparse matrix multiplication based on the tensor processor is improved. The system has the same beneficial effects.
Owner:NAT UNIV OF DEFENSE TECH

Power system large model parameter scale compression and high performance computing resource optimization method

The invention relates to the technical field of electric power system intelligence, in particular to an electric power system large model parameter scale compression and high-performance computing resource optimization method, which comprises the steps of obtaining a pre-trained large model, generating a lightweight small model by model distillation, sparsely cutting an optimized structure, performing fine tuning training to improve performance and the like. Through the technical means of model distillation, sparse clipping, semi-supervised fine tuning and the like, the model parameter quantity is compressed to 10%-30% of the original scale, and the reasoning delay is reduced by 40% or above. Meanwhile, the lightweight model shows high applicability and accuracy in scenes such as electric power intelligent questioning and answering, report generation and new energy prediction, and technical support is provided for intelligent transformation of the electric power industry.
Owner:CHINA SOUTHERN POWER GRID COMPANY

Modularization-based high-performance computing device power state analysis method and system

The invention relates to the technical field of electric signal data analysis, in particular to a high-performance computing equipment power state analysis method and system based on modularization. The method comprises the following steps: in a modular assembly of a high-performance computing equipment power supply, collecting contact resistance change data of an equipment power supply connector through a distributed sensor network, and calculating a resistance change rate; carrying out partial discharge monitoring on a cable of an equipment power supply, and recording discharge capacity trend data; taking the resistance change rate and the discharge capacity trend data as equipment electric energy characteristic parameters; according to the invention, the accuracy, comprehensiveness, practicability and reliability of the power supply system are remarkably improved by accurately monitoring electric energy parameters, early warning and positioning mechanical faults, evaluating electromagnetic thermal coupling interference, comprehensively analyzing the power supply state and outputting a report.
Owner:HUIZHOU XINHUIYUAN TECH +1