Computing and archiving research environment using multiple data integration
The integration of HPC, science cloud, and data archiving services addresses the high costs and complexity of existing HPC systems, facilitating efficient data storage, analysis, and sharing for high-impact research.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2026-03-26
AI Technical Summary
Existing high-performance computing (HPC) systems face high costs, large-scale facility investments, and complex management, while existing data integration solutions lack efficient platforms for easy storage, analysis, and sharing of scientific data.
A platform integrating HPC, science cloud, and data archiving services, utilizing SLURM batch scheduling, GlusterFS, LustreFS, OpenStack, and data catalog with tape backup for efficient data processing, storage, and sharing.
Enables cost-effective, scalable, and manageable data integration for high-impact research by providing easy storage, analysis, and sharing of scientific data across multiple services.
Smart Images

Figure PH2024050014_26032026_PF_FP_ABST
Abstract
Description
[0001] COMPUTING AND ARCHIVING RESEARCH ENVIRONMENT USING MULTIPLE DATA INTEGRATION
[0002] Specification
[0003] Technical Field of the Invention
[0004] The present invention relates, in general, enables different or multiple data integration between different information or data that have a high requirement for data storage and high-performance computation.
[0005] Further, the present invention relates specifically in providing a proper platform for an easy storage, analysis, and sharing of scientific data by providing different combination of services such as high-performance computing (HPC), science cloud, and data archiving. The data generated through these different services can serve as an input to various discoveries and high impact research.
[0006] Background of the Invention
[0007] Continuing innovation in science, technology, information technology and the like has played a very important role in different advancements in technology, materials, algorithms, and designs in our present century. These different innovations have driven advancements in engineering, robotics, industry, computer science and medicine by pushing the boundaries of what new designs and methods can do, making them more capable, adaptable, usable and efficient. These different innovations can shape different fields, leading to new applications, improved performance, and increased integration in various industries and everyday life.
[0008] In computer programming, a specific data type is defined as an attribute associated with a piece of data that tells a computer system how to properly interpret a specific value or combination of values. Moreover, understanding these data types can ensure that a data is collected in a preferred format and that the value of each property is as expected.
[0009] The term “data type” in software programming describes the kind of value that a variable possesses and the kinds of mathematical, relational, or logical operations that can be performed on it without leading to an error. Further, numerous programming languages that utilizes the data types string, integer, and floating point to represent different text, whole numbers, and values with decimal points, respectively. An interpreter or compiler can determine how a programmer plans to use a given set of data by looking up its data type.
[0010] High-performance computing, or also known as HPC, is defined as the ability to process different data and perform different complex calculations at high speeds. It is also defined as the practice of aggregating different computing resources in order to gain performance that is greater than that of a single workstation, server, or a computer. It can take the form of a specific custom-built supercomputers or groups of individual computers known as clusters.
[0011] The HPC can be run on-premises, in the cloud, or as a hybrid of both. Each specific computer that is used in a cluster is often called a node, wherein each node will be responsible for different tasks. Controller nodes run specific essential services that coordinate work between nodes, interactive nodes or login nodes that can act as the hosts that users could log in to, either by graphical user interface or command line, and compute or worker nodes execute the computations. Algorithms and software are run in parallel on each node of the cluster to help perform its given task. Lastly, HPC typically has three main components: compute, storage, and networking / archiving.
[0012] One prior art related to the present invention is EP3146426 entitled “HIGH- PERFORMANCE COMPUTING FRAMEWORK FOR CLOUD COMPUTING ENVIRONMENTS” which relates to various embodiments for a high-performance computing framework for cloud computing environments. A parallel computing application executable by at least one computing device of the cloud computing environment can call a message passing interface (MPI) to cause a first one of a plurality of virtual machines (VMs) of a cloud computing environment to store a message in a queue storage of the cloud computing environment, wherein a second one of the plurality of virtual machines (VMs) is configured to poll the queue storage of the cloud computing environment to access the message and perform a processing of data associated with the message. The parallel computing application can call the message passing interface (MPI) to access a result of the processing of the data from the queue storage, the result of the processing being placed in the queue storage by the second one of the pluralities of virtual machines (VMs). Another prior art related to the present invention is KR102231358 entitled “UNIFIED VIRTUALIZATION METHOD AND SYSTEM FOR HIGH- PERFORMANCE CLOUD SERVICE” which relates to a unified virtualization method and a unified virtualization system for high-performance cloud service. According to one aspect of the present invention, a unified virtualization system for the high-performance cloud service includes: a plurality of computing nodes; and a virtualization server to provide a unified virtualization computing environment by using resources of computing nodes through a resource isolation technology based on a container in a hyper-visor layer. Up to now, a high-performance computing (HPC) field has disadvantages of an astronomical high price, large-scale facility investment, maintenance specialists required, and higher management cost after construction more than advantages. The present invention is to provide renewal computing power in the form of a cloud service based on a container making it easy for anybody including a developer to use through an inverse virtualization technology (hyper-chain technology) of compensating for a latency of existing cloud service (virtualization) by recycling aged servers present inside an x86-based international data corporation (IDC), a computing room, or an office. Accordingly, the interest of the present invention is increased in a niche market in a HPC field such as an interpretation field, various manufacturing design fields, a bio-field, or a new drug field.
[0013] Further, another prior art related to the present invention is KR102089450 entitled “DATA MIGRATION DEVICE AND OPERATION METHOD THEREOF” which relates to a data migration device for processing data migration between memories according to a result of monitoring performance change while an application is running in a high-performance computing (HPC) environment using a hybrid memory, and an operation method thereof. The data migration device comprises: a monitoring part; a calculation part; and a selection part.
[0014] Furthermore, another prior art related to the present invention is KR1016250960000 entitled “HIGH PERFORMANCE COMPUTING SYSTEM HAVING CON TO CON CONNECTION STRUCTURE” which relates to a high- performance computing (HPC) system in which PC cartridges including modularized arithmetic PCs and auxiliary keyboard, video, mouse, and audio (KVMA) switches are coupled to a mainboard in the con-to-con method. According to the present invention, multiple arithmetic PCs in charge of processing data in the high performance computing system are coupled to a mainboard built in a body in the con-to-con method, thereby easily being mounted and separated. The present invention also is able to make the body in compact form, conveniently expand the system, configure hardware of the entire HPC system in a modular form to minimize cable connections between modules, and easily control the arithmetic PCs.
[0015] Lastly, another prior art related to the present invention is KR1020070011503 entitled “HIGH PERFORMANCE COMPUTING SYSTEM AND METHOD” which relates to a High-Performance Computing (HPC) node comprises a motherboard, a switch comprising eight or more ports integrated on the motherboard, and at least two processors operable to execute an HPC job, with each processor communicably coupled to the integrated switch and integrated on the motherboard.
[0016] Summary and Object of the Invention
[0017] It is therefore an object of the present invention is to enable multiple data integration between different data that requires a high requirement for data storage and high-performance computing.
[0018] The primary objective of the present invention is to provide a platform for easy storage, analysis, and sharing of scientific data by providing the combination if services such as high-performance computing (HPC), science cloud, and data archiving The data generated through these services can serve as an input to various discoveries or high-impact research and can contribute to scientific-based policies and decision-making.
[0019] Brief Description of the Drawing
[0020] Figure 1 shows the system architecture for the HPC service.
[0021] Figure 2 shows the system architecture of the science cloud service that is implemented through OpenStack.
[0022] Figure 3 shows the system architecture of the data archiving service
[0023] Detailed Description of the Invention
[0024] I. High Performance Computing (HPC)
[0025] The HPC service of the present invention is utilized for the processing of massive amounts of data that requires high-speed and resource-intensive computations and powerful computing resources. In comparison to an average desktop computer, the HPC can deliver more accurate results. The HPC of the present invention consists of a different cluster of compute and storage servers that allows high-speed and resource-intensive computations and processing of large datasets.
[0026] The current capacity of the HPC of the present invention is as follows:
[0027] The HPC service of the present invention uses SLURM as the batch scheduler. The cluster is divided into 4 different partitions such as debug, batch, serial, and GPU.
[0028] Below are the specifications per partition: a. Debug (2 nodes) - 44 cores, 88 threads, 528 GB RAM; b. Batch (14 nodes) - 44 cores, 88 threads, 528 GB RAM; c. Serial (2 nodes) - 44 cores, 88 threads, 528 GB RAM; and d. GPU (6 nodes) - 12 cores, 24 threads, 1056 GB (1 TB) RAM NVIDIA Tesla P40
[0029] The home directory ( / home) is the network filesystem using GlusterFS and is built to serve as the user’s home directory. Users’ scripts input data are stored here.
[0030] The scratch directories ( / scratch 1 and / scratch2) are the COARE's parallel filesystem using LustreFS. These are built to handle user’s I / O heavy workloads. The output of running jobs including the intermediary files are stored here.
[0031] II. Science Cloud
[0032] The science cloud of the present invention enables the provision of virtual machines for cloud-based applications and computing. This service is implemented through OpenStack.
[0033] III. Data Archiving
[0034] The data archiving service of the present invention provides a repository with redundancy which aims to accommodate various storage requirements and offers multiple store options and large storage capacity for the users of the present invention. This service of the present invention is implemented through the data catalog and tape backup.
[0035] The data catalog is a web-based research repository that contains a collection of datasets that is produced from various research findings using the resources of the present invention, as well as the other bodies of data gathered and contributed by Data Catalog Data Owners. The repository is openly accessible to the public, and the datasets could be integral for academics, data analysts, scientists, and other scholars in the scientific community.
[0036] The data catalog is built with OpenStack Swift as its back- end storage and CKAN as its datasets metadata manager. Another option for users in storing their long-term data is to have their data taped for offsite or off-the-grid storage.
Claims
CLAIMS:1 . A computing and archiving research environment using multiple data integration comprising: a high-performance computing (HPC); a science cloud; and a data archiving.
2. A computing and archiving research environment using multiple data integration according to Claim 1 wherein the high-performance computing (HPC) has a current capacity of 30 Tflops for CPU, and 72 Tflops for GPU.
3. A computing and archiving research environment using multiple data integration according to Claim 1 wherein the system architecture of the high-performance computing (HPC) comprising a batch scheduler that uses SLURM that comprising of a debug with 2 nodes, batch with 14 nodes, serial with 2 nodes, and GPU with 6 nodes; a home directory that uses GlusterFS which is the network filesystem; and a scratch directory that uses LustreFS which is the parallel filesystem.
4. A computing and archiving research environment using multiple data integration according to Claim 1 wherein the science cloud is implemented using OpenStack5. A computing and archiving research environment using multiple data integration according to Claim 1 wherein the data archiving is implemented using data catalog and tape backup and built using OpenStack Swift as its back-end storage and CKAN as its datasets metadata manager.
Citation Information
Patent Citations
High-performance computing method based on cloud computing
CN110109757A
High performance computing system and method
EP3944084A1
Data parallel computing on multiple processors
US10552226B2
High performance computing system
US11625393B2