A bioinformatics analysis system and method based on the supercomputing Internet

By interconnecting multiple supercomputer centers in a unified manner, the supercomputer resource device cluster is solved, and the problem of insufficient single supercomputer resources is achieved is quickly and timely processing and efficient analysis of massive biological information data.

CN114664384BActive Publication Date: 2025-06-03SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210283261.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-22
Publication Date
2025-06-03
Estimated Expiration
2042-03-22

AI Technical Summary

Technical Problem

In the prior art, a single supercomputer resource is insufficient or a single supercomputer resource utilization is insufficient, making it difficult to effectively store and process massive biological information data.

Method used

Through the supercomputing Internet, multiple supercomputing centers are connected uniformly to form a supercomputing resource device group, and the computing power elastic support of the supercomputing Internet can be used to realize the construction of super-large device groups, the coverage of supercomputing Internet and the collaborative sharing of resources.

Benefits of technology

Concentrate more supercomputing resources to ensure the rapid and timely processing of sequencing big data, computing big data and network big data, and improve the efficiency and accuracy of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114664384B_ABST
    Figure CN114664384B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of supercomputing Internet, and particularly relates to a bioinformatics analysis system and method based on the supercomputing Internet. The system includes a bioinformatics analysis platform web portal for receiving job information sent by a user and sending the job information to a data analysis module; the data analysis module for analyzing resource information required by the user job information and sending it to a supercomputing resource scheduling system; the supercomputing resource scheduling system for detecting resource configuration information in a supercomputing cluster and matching the user application resource information to obtain the most suitable supercomputing resource, and sending the job information to the supercomputing cluster; the supercomputing cluster executes the job content, wherein the supercomputing cluster includes a number of supercomputing computing power resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of supercomputing Internet, and particularly relates to a bioinformatics analysis system and method based on the supercomputing Internet. Background Art

[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Supercomputing is an important symbol to measure a country's comprehensive scientific research level and an irreplaceable information technology means to support national security, economic, social development and sustainable development. Supercomputing is widely used in major national public livelihood fields such as basic scientific research, climate and meteorological prediction, and biomedical research and development.

[0004] With the wide application and rapid development of high-throughput sequencing technology in the field of life science, the rapid growth of sequence data generated by biological sequencing, the data volume of biological sequencing in ordinary laboratories can also reach the PT level. The characteristics of large data volume and complex data processing process in sequencing technology also put higher requirements on the high-performance computing service environment, resulting in problems such as insufficient computing power resources of a single supercomputer or insufficient utilization rate of a single supercomputer resource. Effective storage, efficient analysis, and shared utilization of such large-scale data are all difficult problems faced now. Summary of the Invention

[0005] In order to solve the technical problems in the above background art, the present invention provides a bioinformatics analysis system and method based on the supercomputing Internet. The present invention uses the supercomputing Internet to unify and interconnect multiple supercomputer centers to form a set of supercomputing resource device groups.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] The first aspect of the present invention provides a bioinformatics analysis system based on the supercomputing Internet.

[0008] A bioinformatics analysis system based on the supercomputing Internet includes: a bioinformatics analysis platform web portal for receiving job information sent by a user and sending the job information to a data analysis module; the data analysis module for analyzing the resource information required for the user job information and sending it to a supercomputing resource scheduling system; the supercomputing resource scheduling system for detecting the resource configuration information in the supercomputing cluster and matching the user application resource information to obtain the most suitable supercomputing resource and sending the job information to the supercomputing cluster; the supercomputing cluster executes the job content, where the supercomputing cluster includes a number of supercomputing power resources.

[0009] The second aspect of the present invention provides a bioinformatics analysis method based on the supercomputing Internet.

[0010] A bioinformatics analysis method based on the supercomputing Internet, comprising:

[0011] Receiving job information sent by a user;

[0012] Analyzing resource information required for the user's job information;

[0013] Detecting resource configuration information in the supercomputing cluster and matching the user's applied resource information to obtain the most suitable supercomputing resources;

[0014] The supercomputing cluster receives the job information and executes the job content; the supercomputing cluster includes a number of supercomputing computing resources.

[0015] The third aspect of the present invention provides a computer-readable storage medium.

[0016] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the bioinformatics analysis method based on the supercomputing Internet as described in the second aspect above.

[0017] The fourth aspect of the present invention provides a computer device.

[0018] A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in the bioinformatics analysis method based on the supercomputing Internet as described in the second aspect above.

[0019] Compared with the prior art, the beneficial effects of the present invention are:

[0020] The present invention solves the problems of insufficient computing power resources of a single supercomputer or insufficient utilization rate of a single supercomputing resource, etc., and can concentrate more supercomputing computing resources to ensure the rapid and timely processing of sequencing big data, computing big data, and network big data.

[0021] The present invention can better cope with and solve problems such as storage, processing, calculation, and analysis of massive bioinformatics data, make full use of the higher performance of supercomputing computing resources, shorten the data processing time, and give accurate processing results at the same time. Brief Description of the Drawings

[0022] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0023] Figure 1 It is a framework diagram of the bioinformatics analysis system based on the supercomputing Internet shown in the present invention. Detailed Embodiments

[0024] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0025] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0026] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0027] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and systems according to various embodiments of the present disclosure. It should be noted that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Similarly, it should be noted that each block in the flowchart and / or block diagram, and the combination of blocks in the flowchart and / or block diagram, may be implemented using a dedicated hardware-based system for performing the specified functions or operations, or may be implemented using a combination of dedicated hardware and computer instructions.

[0028] Embodiment 1

[0029] This embodiment provides a bioinformatics analysis system based on the supercomputing Internet.

[0030] A bioinformatics analysis system based on the supercomputing Internet includes: a bioinformatics analysis platform web portal for receiving job information sent by a user and sending the job information to a data analysis module; the data analysis module for analyzing the resource information required for the user's job information and sending it to a supercomputing resource scheduling system; the supercomputing resource scheduling system for detecting the resource configuration information in the supercomputing cluster and matching the user's application resource information to obtain the most suitable supercomputing resource and sending the job information to the supercomputing cluster; and the supercomputing cluster executing the job content, where the supercomputing cluster includes a number of supercomputing computing power resources.

[0031] The Supercomputing Internet can unify and interconnect multiple supercomputing centers into a set of supercomputing resource device groups. With the support of the computing power elasticity of the Supercomputing Internet, it can realize the construction of ultra-large device groups, the coverage of the Supercomputing Internet, and the collaborative sharing of resources. Relying on the existing foundation, it can give full play to the platform application of the supercomputing large scientific device group and can quickly and timely analyze and process sequencing big data, computing big data, and network big data.

[0032] The following will be described in conjunction with the attached Figure 1 This embodiment aims to provide a bioinformatics analysis system based on the Supercomputing Internet, including a bioinformatics analysis platform web portal, a data analysis module, a data storage module, a supercomputing resource scheduling system, and multiple supercomputing computing power resources in the Supercomputing Internet.

[0033] The bioinformatics analysis system based on the Supercomputing Internet involved in this embodiment integrates and configures common analysis tools in the bioinformatics analysis process, such as open-source tools like bwa, picard, samtools, FastQC, MultiQC, GATK, SnpEff, etc.

[0034] Step 1: The bioinformatics analysis platform web portal. It is used to accept data submitted by users.

[0035] Step 1.1: The bioinformatics analysis platform web portal accepts user login information and sends it to the data analysis module;

[0036] Step 1.2: The bioinformatics analysis platform web portal accepts job information sent by users and submits it to the data analysis module;

[0037] Step 1.3: The bioinformatics analysis platform web portal accepts viewing information submitted by users and sends it to the data analysis module; accepts the query data returned by the data analysis module and displays it;

[0038] Step 2: The data analysis module. It is used to receive and analyze job information submitted by users.

[0039] Step 2.1: The data analysis module accepts the user requirements submitted in Step 1.1, sends the login information verification to the data storage module; accepts the calculation results submitted in Step 2.2 and sends them to the data storage module;

[0040] Step 2.2: The data analysis module receives the user job information sent by Step 1.2, stores the user job information in the data storage module; extracts the bioinformatics analysis tool name, tool parameters, data set, and the requested allocated CPU and memory, etc. from the job information, converts the user input into a command input for calling the bioinformatics tool for calculation through text parsing, analyzes the resource size required for the user job and sends it to the supercomputer resource scheduling system; receives the calculation result returned by the supercomputer resource scheduling system and sends it to the data analysis module;

[0041] Step 2.3: The data analysis module receives the user query information submitted by Step 1.3, and converts it into a query condition and sends it to the data storage module; receives the query result returned by the data storage module and sends it to the bioinformatics analysis platform web portal;

[0042] Step 3: The data storage module. It is used to store user information, job information submitted by users, and result information after job calculation is completed.

[0043] Step 3.1: The data storage module receives the user data or calculation result data submitted by Step 2.1 and stores it in the database; receives the calculation result data in Step 2.1 and updates the database;

[0044] Step 3.2: The data storage module receives the job record in Step 2.2 and stores it in the database;

[0045] Step 3.3: The data storage module receives the query condition in Step 2.3, retrieves the database and returns the query result to the data analysis module;

[0046] Step 4: The supercomputer resource scheduling system. It is used to detect the resource configuration information in the supercomputer cluster and match the user-requested resource information.

[0047] Step 4.1: The supercomputer resource scheduling system detects the configuration information of all supercomputer computing resources in the supercomputer cluster and marks the attributes;

[0048] Step 4.2: The supercomputer resource scheduling system receives the bioinformatics tool call command and resource application information in Step 2.2, matches the detected configuration information of each supercomputer computing resource with the resource application information, selects the most suitable supercomputer computing resource, and sends the job to the target supercomputer cluster;

[0049] Step 4.3: The supercomputer resource scheduling system receives the calculation result returned by the supercomputer computing resource, and sends the job result to the data analysis module;

[0050] Step 5: The supercomputer computing resource. It is used to execute the job and return the result.

[0051] Step 5.1: The target supercomputer receives the job information in Step 4.2 and executes the job;

[0052] Step 5.2: Wait for the job to complete and send the job result to the supercomputer resource scheduling system.

[0053] Example Two

[0054] This example provides a bioinformatics analysis method based on the supercomputer Internet.

[0055] A bioinformatics analysis method based on the supercomputer Internet includes:

[0056] Receive the job information sent by the user;

[0057] Analyze the resource information required for the user's job information;

[0058] Detect the resource configuration information in the supercomputer cluster and match the user's applied resource information to obtain the most suitable supercomputer resource;

[0059] The supercomputer cluster receives the job information and executes the job content; the supercomputer cluster includes several supercomputer computing power resources.

[0060] Example Three

[0061] This example provides a computer-readable storage medium with a computer program stored thereon, and when the program is executed by a processor, it implements the steps in the bioinformatics analysis method based on the supercomputer Internet described in Example Two above.

[0062] Example Four

[0063] This example provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the bioinformatics analysis method based on the supercomputer Internet described in Example Two above.

[0064] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program code.

[0065] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or the functions specified in multiple blocks.

[0066] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means, and the instruction means implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or the functions specified in multiple blocks.

[0067] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or the functions specified in multiple blocks.

[0068] Those of ordinary skill in the art can understand that to implement all or part of the processes in the above-described method embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-described methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0069] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A bioinformatics analysis system based on the supercomputing Internet, characterized in that, it includes: A bioinformatics analysis platform web portal, a data analysis module, a data storage module, a supercomputing resource scheduling system, and several supercomputing power resources in the supercomputing Internet; The bioinformatics analysis platform web portal is used to receive job information sent by users and send the job information to the data analysis module; The data analysis module is used to extract the bioinformatics analysis tool name, tool parameters, data set, allocated CPU and memory of the job information, convert the user input into a command input for calling bioinformatics tools for calculation through text parsing, analyze the resource size required for the user job information, and send it to the supercomputing resource scheduling system; The data analysis module receives the query information submitted by the user through the bioinformatics analysis platform web portal and converts it into query conditions and sends them to the data storage module; receives the query results returned by the data storage module and sends them to the bioinformatics analysis platform web portal. The bioinformatics analysis platform web portal receives the query data returned by the data analysis module and displays it; The data storage module is used to store user information, job information submitted by users, and job result information; The supercomputing resource scheduling system is used to detect the resource configuration information in the supercomputing cluster and match the user application resource information. The supercomputing resource scheduling system detects the configuration information of all supercomputing power resources in all supercomputing clusters and marks the attributes, and receives the bioinformatics tool call command and resource application information. According to the detected configuration information of all supercomputing power resources in the supercomputing Internet, it matches with the resource application information, selects the most suitable supercomputing power resource, and sends the job to the target supercomputing cluster; the supercomputing cluster executes the job content. When the job is completed, the job result information is sent to the supercomputing resource scheduling system, which is sent to the data analysis module by the supercomputing resource scheduling system and then sent to the bioinformatics analysis platform web portal by the data analysis module. Among them, the supercomputing cluster includes several supercomputing power resources.

2. The bioinformatics analysis system based on the supercomputing Internet according to claim 1, characterized in that, The bioinformatics analysis platform web portal is also used to receive user login information and accept user login information.

3. A bioinformatics analysis method based on the supercomputing Internet, characterized in that, it includes: Receiving job information sent by users; Extracting the bioinformatics analysis tool name, tool parameters, data set, allocated CPU and memory of the job information, converting the user input into a command input for calling bioinformatics tools for calculation through text parsing, and analyzing the resource size required for the user job information; Detecting the resource configuration information in the supercomputing cluster and matching the user application resource information, detecting the configuration information of all supercomputing power resources in all supercomputing clusters and marking the attributes, and receiving the bioinformatics tool call command and resource application information. According to the detected configuration information of all supercomputing power resources in the supercomputing Internet, it matches with the resource application information and selects the most suitable supercomputing power resource; The supercomputer cluster receives the job information, executes the job content, and when the job is completed, sends the job result information to the supercomputer resource scheduling system and finally displays the result to the user; the supercomputer cluster includes a number of supercomputer computing power resources; After the user submits the query information, the corresponding query data is displayed.

4. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, it implements the steps in the bioinformatics analysis method based on the supercomputer Internet as described in claim 3.

5. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the program, it implements the steps in the bioinformatics analysis method based on the supercomputer Internet as described in claim 3.

Citation Information

Patent Citations

  • Computing method and super-computing system for computing task

    CN103279445A

  • Cloud scheduling method of supercomputing resources, cloud scheduling center and system

    CN109951558A

  • Hierarchical storage optimization method for super-large-scale drug data

    CN111210879A

  • Architecture construction method of biological information deep mining analysis system

    CN112151114A