Computing and archiving research environment using multiple data integration

By combining HPC, scientific cloud, and data archiving services, it solves the problems of high cost and latency in the HPC field, and provides a low-cost, high-efficiency data storage and sharing platform that supports large-scale computing and multiple storage options, suitable for various research and decision support.

CN122070534APending Publication Date: 2026-05-19DEPT OF SCI & TECH ADVANCED SCI & TECH INST
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DEPT OF SCI & TECH ADVANCED SCI & TECH INST
Filing Date
2024-09-25
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

The existing high-performance computing (HPC) field suffers from high prices, large infrastructure investments, and high management costs. Furthermore, existing cloud services have significant latency, making it difficult to achieve the integration, efficient storage, analysis, and sharing of different data.

Method used

By combining high-performance computing (HPC), scientific cloud, and data archiving services, it provides a multi-data integration platform, utilizes SLURM as a batch processing scheduler, uses GlusterFS and LustreFS file systems, integrates OpenStack to implement scientific cloud, and employs data catalog and tape backup for data archiving to meet different storage requirements.

Benefits of technology

It enables low-cost, high-efficiency data storage, analysis, and sharing, supports large-scale computing and multiple storage options, and is suitable for various research and decision support applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122070534A_ABST
    Figure CN122070534A_ABST
Patent Text Reader

Abstract

The present invention relates to a computing and archiving research environment that uses multi-data integration between different information or data with high requirements for data storage and high performance computing. The invention provides a platform for easy storage, analysis and sharing of different scientific data by providing different service combinations such as high performance computing (HPC), scientific clouds and data archiving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] In general, the present invention relates to the realization of different or multiple data integrations between different information or data that have high requirements for data storage and high-performance computing.

[0002] Furthermore, this invention specifically relates to providing a suitable platform for the easy storage, analysis, and sharing of scientific data by offering different combinations of services such as high-performance computing (HPC), scientific cloud, and data archiving. Data generated through these different services can serve as input for a wide range of discoveries and high-impact research. Background Technology

[0003] Continuous innovation in science, technology, and information technology has played a crucial role in the various advancements in technology, materials, algorithms, and design this century. These innovations have driven progress in engineering, robotics, industry, computer science, and medicine by pushing the boundaries of new designs and methods, making them more capable, adaptable, usable, and efficient. These innovations can shape different fields, leading to new applications, improved performance, and enhanced integration across industries and everyday life.

[0004] In computer programming, a specific data type is defined as an attribute associated with a piece of data, which tells the computer system how to correctly interpret a particular value or combination of values. Furthermore, understanding these data types ensures that data is collected in a preferred format and that the value of each attribute matches expectations.

[0005] In software programming, the term "data type" describes the type of values ​​a variable can have, and the type of mathematical, relational, or logical operations that can be performed on it without causing errors. Furthermore, many programming languages ​​use data types such as string, integer, and floating-point to represent different text, integer, and decimal values, respectively. An interpreter or compiler can determine how a programmer plans to use a given dataset by looking up its data type.

[0006] High-performance computing (HPC) is defined as the ability to process diverse data and perform complex calculations at high speeds. It is also defined as the practice of aggregating different computing resources to achieve performance higher than that of a single workstation, server, or computer. It can take the form of a custom-built supercomputer or a group of individual computers (called a cluster).

[0007] HPC can run on-premises, in the cloud, or a hybrid of both. Each specific computer used in a cluster is typically called a node, and each node will be responsible for different tasks. Controller nodes run specific basic services that coordinate the work between nodes, interactive nodes, or login nodes, which can act as host, compute, or worker nodes that users can log in to via a graphical user interface or command line. Algorithms and software run in parallel on each node of the cluster to help perform their given tasks. Finally, HPC typically has three main components: compute, storage, and networking / archiving.

[0008] A prior art related to this invention is EP3146426 entitled "HIGH-PERFORMANCE COMPUTING FRAMEWORKFOR CLOUD COMPUTING ENVIRONMENTS", which relates to various embodiments of high-performance computing architectures in cloud computing environments. A parallel computing application executable by at least one computing device in the cloud computing environment can invoke a message passing interface (MPI) to cause a first VM among a plurality of virtual machines (VMs) in the cloud computing environment to store a message in a queue store of the cloud computing environment, wherein a second VM among the plurality of virtual machines (VMs) is configured to poll the queue store of the cloud computing environment to access the message and perform data processing associated with the message. The parallel computing application can invoke the message passing interface (MPI) to access the result of processing data from the queue store, the result of which is placed in the queue store by the second VM among the plurality of virtual machines (VMs).

[0009] Another prior art related to this invention is KR102231358 entitled "UNIFIED VIRTUALIZATION METHOD AND SYSTEM FOR HIGH-PERFORMANCE CLOUD SERVICE," which relates to a unified virtualization method and system for high-performance cloud services. According to one aspect of the invention, a unified virtualization system for high-performance cloud services includes: multiple compute nodes; and a virtualization server for providing a unified virtualized computing environment by utilizing the resources of the compute nodes through resource isolation technology based on containers in a hypervisor layer. To date, the high-performance computing (HPC) field has had more disadvantages than advantages: exorbitant prices, large infrastructure investments, the need for maintenance experts, and high post-construction management costs. This invention aims to provide updated computing capabilities in the form of container-based cloud services, making them easily accessible to anyone, including developers, through reverse virtualization technology (hyperlink technology), which compensates for latency in existing cloud services (virtualization) by reclaiming old servers in x86-based International Data Corporations (IDCs), computing rooms, or offices. Therefore, the benefits of this invention increase in niche markets within the HPC field, such as the fields of interpretation, various manufacturing design, biology, or new drugs.

[0010] Furthermore, another prior art related to this invention is KR102089450 entitled "DATA MIGRATION DEVICE AND OPERATION METHOD THEREOF," which relates to a data migration device and its operating method. This device is used to process data migration between memories based on monitoring performance changes when an application is running in a high-performance computing (HPC) environment using hybrid memory. The data migration device includes: a monitoring section; a computing section; and a selection section.

[0011] Furthermore, another prior art related to this invention is KR1016250960000 entitled "HIGH PERFORMANCE COMPUTING SYSTEM HAVING CON TO CON CONNECTION STRUCTURE," which relates to a high-performance computing (HPC) system comprising modular arithmetic PCs and PC cartridges for auxiliary keyboard, video, mouse, and audio (KVMA) switches connected to a motherboard in a con-to-con manner. According to this invention, multiple arithmetic PCs responsible for processing data in the high-performance computing system are connected to a motherboard built into the main body in a con-to-con manner, thereby facilitating installation and disassembly. This invention also enables the main body to be conveniently expanded in a compact form, allows for modular configuration of the entire HPC system hardware to minimize cable connections between modules, and facilitates easy control of the arithmetic PCs.

[0012] Finally, another prior art related to this invention is KR1020070011503 entitled "HIGH PERFORMANCE COMPUTING SYSTEM AND METHOD", which relates to a high-performance computing (HPC) node including a motherboard, a switch, and at least two processors. The switch includes eight or more ports integrated on the motherboard, and the processors are operable to perform HPC jobs. Each processor is communicatively coupled to the integrated switch and integrated on the motherboard. Summary of the Invention

[0013] Therefore, the purpose of this invention is to enable the integration of multiple data sets that have high requirements for data storage and high-performance computing.

[0014] The primary objective of this invention is to provide a platform for the easy storage, analysis, and sharing of scientific data by offering a combination of services such as high-performance computing (HPC), scientific cloud, and data archiving. Data generated through these services can serve as input for various discoveries or high-impact research and contribute to science-based strategy and decision-making. Attached Figure Description

[0015] Figure 1 The system architecture of HPC services is shown.

[0016] Figure 2 The system architecture of a scientific cloud service implemented through OpenStack is shown.

[0017] Figure 3 The system architecture of the data archiving service is shown. Detailed Implementation

[0018] I. High-Performance Computing (HPC)

[0019] The HPC service of this invention is used to process large amounts of data that require high-speed and resource-intensive computing as well as powerful computing resources. Compared with ordinary desktop computers, HPC can provide more accurate results. The HPC of this invention comprises different clusters of computing and storage servers that allow for high-speed and resource-intensive computing and processing of large datasets.

[0020] The current capacity of the HPC of this invention is as follows:

[0021] CPU 30 trillion times per second GPU 72 trillion times per second

[0022] The HPC service of this invention uses SLURM as the batch scheduler. The cluster is divided into four different partitions, such as debug, batch, serial, and GPU.

[0023] The following are the specifications for each partition:

[0024] a. Debugging (2 nodes) - 44 cores, 88 threads, 528GB RAM;

[0025] b. Batch processing (14 nodes) – 44 cores, 88 threads, 528GB RAM;

[0026] c. Serial (2 nodes) – 44 cores, 88 threads, 528GB RAM; and

[0027] d. GPU (6 nodes) – 12 cores, 24 threads, 1056GB (1TB)

[0028] RAM NVIDIA Tesla P40

[0029] The home directory ( / home) is a network file system using GlusterFS and is built as the user's home directory. The user's script input data is stored here.

[0030] The temporary directories ( / scratch1 and / scratch2) are COARE parallel file systems using LustreFS. These are built to handle users' I / O-heavy workloads. The output of running jobs (including intermediate files) is stored here.

[0031] II. Science Cloud

[0032] The scientific cloud of this invention provides virtual machines for cloud-based applications and computing. This service is implemented through OpenStack.

[0033] III. Data Archiving

[0034] The data archiving service of this invention provides a redundant repository designed to meet various storage requirements and offer users of this invention multiple storage options and large storage capacity. This service is implemented through a data catalog and tape backup.

[0035] The Data Catalog is a web-based research repository containing a collection of datasets generated from various research findings using the resources of this invention, as well as other data volumes aggregated and contributed by Data Catalog Data Owners. This repository is publicly accessible, and the datasets are indispensable to academia, data analysts, scientists, and other members of the scientific community.

[0036] The data catalog is built with OpenStack Swift as the backend storage and CKAN as the dataset metadata manager. Another option for users to store long-term data is to tape their data for off-site or offline storage.

Claims

1. A computational and archiving research environment using multi-data integration, comprising: High-performance computing (HPC); Science Cloud; as well as Data archiving.

2. The computing and archiving research environment using multiple data integration as described in claim 1, wherein the high-performance computing (HPC) has a current capacity of 30 trillion operations per second for CPUs and 72 trillion operations per second for GPUs.

3. The computational and archiving research environment using multi-data integration as described in claim 1, wherein the system architecture of said high-performance computing (HPC) includes: A batch scheduler using SLURM, which includes debug with 2 nodes, batch with 14 nodes, serial with 2 nodes, and GPU with 6 nodes. Used as the home directory of GlusterFS, a network file system; and Use temporary directories from LustreFS, which is a parallel file system.

4. The computing and archiving research environment using multi-data integration as described in claim 1, wherein the scientific cloud is implemented using OpenStack.

5. The computing and archiving research environment using multi-data integration as described in claim 1, wherein the data archiving is implemented using a data catalog and tape backup, and is built using OpenStack Swift as its backend storage and CKAN as its dataset metadata manager.