Gene sequencing data analysis system and method, electronic equipment and storage medium

CN120035863APending Publication Date: 2025-05-23BGI HANGZHOU CYCLONESEQ TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280101156.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing gene sequencing data analysis systems that are stand-alone or deployed on a single high-computing server cannot effectively support the massive data analysis needs generated by third-generation sequencing technology, resulting in low data analysis efficiency and slow calculation speed.

Method used

The gene sequencing data analysis system adopts a distributed architecture. Through the collaborative work of data center servers and multiple computing nodes, it uses task queues and virtualization technology to perform data segmentation analysis and realize asynchronous, multi-process process analysis operations.

Benefits of technology

It significantly improves the analysis speed and efficiency of third-generation sequencing data, can meet the analysis needs of massive genetic sequencing data, and improves resource utilization and system availability of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035863A_ABST
    Figure CN120035863A_ABST
Patent Text Reader

Abstract

The invention provides a gene sequencing data analysis system and method, electronic equipment and a storage medium, and relates to the technical field of computers, and the method comprises the following steps: obtaining path information of fragment files generated in a gene sequencing process; creating a fragment file analysis task containing the path information, and filling a task queue with the fragment file analysis task; a plurality of fragmented file analysis results sent by the target computing server are received, the fragmented file analysis results are combined when it is detected that execution of the fragmented file analysis tasks is finished, and the fragmented file analysis results are obtained after the target computing server pulls the fragmented file analysis tasks in the task queue. And calling a plurality of configured computing nodes to check the fragmented file according to path information contained in the fragmented file analysis task, and analyzing a result obtained by the fragmented file by utilizing the gene analysis model. According to the method and the device, the technical problems of low data analysis efficiency and low calculation speed during analysis of the gene sequencing data can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Gene sequencing data analysis system, method, electronic device and storage medium Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a gene sequencing data analysis system, method, electronic device, and storage medium. Background Art

[0002] Gene sequencing is a typical application of high-performance computing. Third-generation gene sequencing technology has gradually become the mainstream sequencing technology. Compared to second-generation sequencing, third-generation sequencing generates unprecedentedly large amounts of data. Its principles dictate that it will generate ever-increasing amounts of data, far exceeding the capabilities of previous sequencing technologies. However, existing data analysis capabilities deployed on single machines or high-performance servers are no longer sufficient to analyze the massive amounts of gene sequencing data at current data growth rates, resulting in low data analysis efficiency and slow computational speeds.

[0003] Summary of the Invention

[0004] The present application provides a gene sequencing data analysis system, method, electronic device and storage medium, the main purpose of which is to solve the current technical problems of low data analysis efficiency and slow calculation speed when analyzing gene sequencing data.

[0005] According to a first aspect of the present application, a gene sequencing data analysis system is provided, comprising: a data center server, a computing server cluster, the computing server cluster comprising a plurality of computing servers, each of the plurality of computing servers being configured with a plurality of computing nodes;

[0006] The data center server is used to obtain the path information of the fragment files generated during the gene sequencing process, create a fragment file analysis task containing the path information, and fill the fragment file analysis task into the task queue;

[0007] The data center server is connected to the computing server cluster, and the data center server is also used to determine at least one target computing server in the computing server cluster to execute the shard file analysis task. The at least one target computing server is used to pull the shard file analysis task from the task queue, and call the configured multiple computing nodes to execute the shard file analysis task, and send the multiple shard file analysis results output by the multiple computing nodes to the data center server.

[0008] Optionally, the computing node is used to retrieve the fragment file based on the path information contained in the fragment file analysis task, wherein the fragment file includes base fragment signals divided according to a specified number or channel source in the upstream gene sequencing step.

[0009] Optionally, a gene analysis model is configured in the computing node, and the computing node is used to analyze the fragment file using the gene analysis model to obtain a fragment file analysis result.

[0010] Optionally, the data center server is used to receive the multiple shard file analysis results sent by the multiple computing nodes, and merge the multiple shard file analysis results when detecting that the execution of the shard file analysis task is completed.

[0011] Optionally, the data center server is further used to generate a task status queue, and the task status queue is used to update and store status information of the shard file analysis task in the task queue.

[0012] Optionally, the data center server is further configured to determine whether the fragment file analysis task has been completed by detecting the status information in the task status queue.

[0013] Optionally, the data center server is also used to clear the redundant intermediate files generated by at least one target computing server and restore the computing resources of the multiple computing nodes configured on the target computing server when detecting the completion of the execution of the shard file analysis task.

[0014] According to a second aspect of the present application, a method for analyzing gene sequencing data is provided, the method being applied to a data center server, the method comprising:

[0015] Obtain the path information of the fragment files generated during gene sequencing;

[0016] Creating a fragment file analysis task containing the path information, and adding the fragment file analysis task to a task queue;

[0017] Receive multiple shard file analysis results sent by the target computing server, and merge the multiple shard file analysis results when detecting the end of the execution of the shard file analysis task, wherein the shard file analysis result is the result obtained by the target computing server pulling the shard file analysis task from the task queue, calling the configured multiple computing nodes to retrieve the shard file according to the path information contained in the shard file analysis task, and using the genetic analysis model to analyze the shard file.

[0018] According to a third aspect of the present application, a method for analyzing gene sequencing data is provided, the method being applied to a computing server, the method comprising:

[0019] Pull the shard file analysis task from the task queue;

[0020] Calling multiple configured computing nodes to execute the shard file analysis task to obtain multiple shard file analysis results, wherein the shard file analysis results are the results obtained by the computing nodes searching for shard files according to the path information contained in the shard file analysis task and analyzing the shard files using the gene analysis model;

[0021] The analysis results of the multiple fragment files are sent to the data center server.

[0022] According to a fourth aspect of the present application, an electronic device is provided, including:

[0023] at least one processor; and

[0024] a memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the second aspect or the third aspect.

[0026] According to a fifth aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the second or third aspect above.

[0027] According to a sixth aspect of the present application, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the method described in the second or third aspect above.

[0028] The present application provides a gene sequencing data analysis system, method, electronic device, and storage medium. The data center server can, after obtaining the path information of the shard file generated during the gene sequencing process, create a shard file analysis task containing the path information and fill the shard file analysis task into the task queue. The server can then receive multiple shard file analysis results sent by the target computing server and, upon detecting the completion of the execution of the shard file analysis task, merge the multiple shard file analysis results. The shard file analysis result is the result obtained by the target computing server pulling the shard file analysis task from the task queue, calling multiple configured computing nodes to retrieve the shard file according to the path information contained in the shard file analysis task, and analyzing the shard file using a gene analysis model. The technical solution disclosed in the present invention can be used by a centralized data center server to allocate data analysis tasks, and by multiple computing nodes to undertake the analysis tasks of the corresponding data shards. The asynchronous and multi-process implementation of the process-based analysis operations within the computing nodes can effectively improve the speed of analyzing the large amount of data generated by third-generation sequencing, thereby improving data analysis efficiency and meeting the analysis needs of massive gene sequencing data under data growth.

[0029] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present application.

[0031] FIG1 is a schematic diagram of the structure of a gene sequencing data analysis system provided in an embodiment of the present application;

[0032] FIG2 is a schematic diagram of a flow chart of a method for analyzing gene sequencing data provided in an embodiment of the present application;

[0033] FIG3 is a schematic diagram of a flow chart of a method for analyzing gene sequencing data provided in an embodiment of the present application;

[0034] FIG4 is a schematic diagram of a principle flow chart of a gene sequencing data analysis provided in an embodiment of the present application;

[0035] FIG5 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0036] In Figure 1:

[0037] 1-Data center server;

[0038] 2-Computing server cluster, 20-Computing servers, 200-Computing nodes. DETAILED DESCRIPTION

[0039] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0040] The following describes the gene sequencing data analysis system, method, electronic device, and storage medium according to the embodiments of the present application with reference to the accompanying drawings.

[0041] Distributed data processing methods that use virtualization technology have been widely adopted in recent years and have developed very rapidly. With the explosive growth of data in the fields of biology, finance, the Internet, the Internet of Things, and artificial intelligence (AI), the computing power of a single physical host has long been bottlenecked in massive data scenarios and cannot support data analysis, storage, AI recognition and other processes. Distributed data processing methods can effectively solve the problems related to massive data, and have the characteristics of resource sharing, good scalability, accelerated computing speed, high reliability, and convenient communication. The advantages are significant, including but not limited to: increasing the system's data capacity and computing power; enhancing system availability, making the system highly available and having certain obstacle avoidance capabilities; modularization makes the system more reusable; development and deployment are more convenient; and it can be expanded in nodes.

[0042] Current domestic methods for analyzing genomic sequencing data, whether second-generation or third-generation, have yet to utilize virtualization technologies like containers. Building deep learning gene recognition clusters allows computing power to be scalable, highly available, and accelerate computing resources. However, existing data analysis capabilities deployed on single machines or high-performance servers are no longer sufficient to analyze the massive amounts of genomic sequencing data at current data growth rates, resulting in low analysis efficiency and computational speed. While other advanced software and hardware solutions exist abroad, these inevitably present technological bottlenecks. With the Third Industrial Revolution and the booming life sciences, the scale and speed of data analysis are crucial. Producing accurate and timely biological data analysis results is crucial for avoiding lagging behind in the global competition in life sciences, and may even allow for a competitive advantage. Therefore, independent high-speed genomic data analysis solutions are currently lacking.

[0043] To solve the above-mentioned technical problems, the present disclosure provides a gene sequencing data analysis system, as shown in FIG1 . The system adopts a distributed architecture of a task center and 1+N computing nodes (one data center server + multiple computing nodes): a centralized data center coordinates data analysis tasks and provides some necessary analysis data resources, and multiple computing nodes undertake the analysis tasks of corresponding data shards. The analysis data resources include a central processing unit (CPU), a graphics processing unit (GPU), a computing accelerator card, etc. Correspondingly, the system includes: a data center server 1, a computing server cluster 2, the computing server cluster 2 includes multiple computing servers 20, and each computing server 20 in the multiple computing servers 20 is respectively configured with multiple computing nodes 200; the data center server 1 is used to obtain the path information of the shard files generated during the gene sequencing process, and create a shard file analysis task containing the path information, and fill the shard file analysis task into the task queue; the data center server 1 is connected to the computing server cluster 2, and the data center server 1 is also used to determine at least one target computing server in the computing server cluster 2 to execute the shard file analysis task, and the at least one target computing server is used to pull the shard file analysis task in the task queue, and call the configured multiple computing nodes to execute the shard file analysis task, and send the multiple shard file analysis results output by the multiple computing nodes to the data center server 1.

[0044] Among them, the determination of the shard file is generated in the upstream gene sequencing step, and it is sharded in a specified number or according to the channel source, such as 5,000 base fragment signals as a shard or all base fragment signals of a channel as a shard. The path of the shard file needs to be filled into the task queue as one of the inputs so that the subsequent process can read the content of the file according to the path for further analysis; the task queue refers to a queue containing task parameters to be submitted to the process pool for execution, which is specifically implemented as a redis queue. After the shard file analysis task is generated, the shard file analysis task can be packaged into elements and filled into the task queue. The purpose of filling into the task queue is to avoid backlog of existing processes while the next process can consume the elements in the queue to perform the following analysis link asynchronously; the computing node is a virtualized computing container node. Virtualization refers to the use of software technology to virtualize a physical host into multiple logical computers. Each logical computer can independently run different operating systems and various applications. Through virtualization technology, each virtual machine (VM) has its own virtual hardware (virtual CPU, network interface card, memory, etc.). The operating system running on the VM assumes it has exclusive access to a single physical host, and the software on the VM runs on a virtual platform, not a real hardware platform. A container is simply a special process running on the host machine, and multiple containers still use the same host operating system kernel. Containers do not rely on the operating system's environment to run applications, but instead use the mechanisms and features provided by the operating system kernel to isolate and restrict application processes.

[0045] Accordingly, when screening target computing servers, the data center server 1 may adopt a random screening method or a screening method according to preset screening rules. As a possible implementation method, the task type of the shard file analysis task may be determined, and the computing server 20 in the computing server cluster 2 that matches the task type may be determined as the target computing server; as a possible implementation method, the working status information of each computing server 20 in the computing server cluster 2 may be obtained, and the computing server 20 whose working status information meets the preset conditions may be determined as the target computing server, wherein the working status information may include working hours, working status, working priority, etc., which are not specifically limited here.

[0046] In specific application scenarios, when processing data, the data center server can play the role of producer, responsible for the creation of shard file analysis tasks, monitoring of task status, and necessary file operations at the end, etc. It also supports receiving and removing tasks. Specifically, when creating a task, it will obtain the list of all shard files to be analyzed and the path and other information to fill the task queue and create status information; when deleting a task, it will send a deletion message to the computing node so that the computing node can restart the service according to the docker_id; input operation: select a data name and a set of analysis parameters on any client's web page, click the submit button to send a request to the data center server, the data center server will parse the parameters to obtain the entire data, and create a task according to the path of the current data shard file and other information to fill in the message queue. Analysis parameters include the name of the deep learning base calling model, whether to include adapter adapters (the base pass signals corresponding to the fixed primer segment preceding the base signal in the read), and base filter length. Parameter parsing involves using the file path in the parameters to determine which data set corresponds to which data folder on the data server. The fragment file path determines which fragment files within that data set should be submitted to the queue for analysis. The list of fragment files to be analyzed contains the paths of the fragment files to be analyzed. Fragment files are generated during the upstream sequencing step and are fragmented by a specified number or by channel source, such as 5000 base fragment signals per fragment or all base fragment signals from a channel. The fragment file path is entered as input into the task queue so that subsequent processes can read the file content according to the path for further analysis. Status information maintains various data structures to indicate which fragment files have completed the corresponding analysis process. Accordingly, the data center server also generates a task status queue, which updates the status information of the fragment file analysis tasks stored in the task queue.

[0047] In specific application scenarios, data analysis can be performed by compute nodes. Compute nodes will pull and consume tasks based on information provided by the data center server. After pulling the corresponding tasks, the compute nodes will each trigger their own data analysis tasks. Each node is relatively independent in terms of resource usage and task execution, and can be considered independent if resources are sufficient. Each compute node generates corresponding sharding results after the data sharding analysis process, which are then fed back to the data center server. The specific data flow is as follows: task queue - retrieve shard files in Hierarchical Data Format Version 5 (HDF5) format via the fiber optic network to the node locally - parse the HDF5 shard files - run the AI ​​model for inference - decode the AI ​​inference results - save the decoded results to the disk in the shard result files - and transmit the results back to the data center server via the fiber optic network. When a compute node retrieves a shard file, it can retrieve it based on the path information contained in the shard file analysis task. The shard file contains base fragment signals divided by a specified number or channel source from the upstream gene sequencing step. Furthermore, each computing node is equipped with a genetic analysis model. After retrieving a shard file, the computing node uses the genetic analysis model to analyze the shard file, obtaining analysis results. These results are then transmitted back to the data center server via the fiber optic network. This means that the data center server can receive analysis results from multiple shard files sent by multiple computing nodes. The genetic analysis models of multiple computing nodes form a genetic analysis cluster, which utilizes artificial intelligence (AI) learning methods and techniques to analyze gene sequencing data. This analysis of gene sequencing data can be transformed into various signals with distinct patterns and characteristics based on the unique physical and chemical properties of its bases. Images and electrical signals are ideal models for AI learning and analysis. Building a genetic analysis cluster distributes the computing resources for gene sequencing data analysis, allowing more computing resource nodes to accelerate and support the AI ​​learning process, thereby improving the speed and accuracy of gene sequencing data analysis.

[0048] Among them, the gene analysis model can be any computing model that can realize the gene sequencing data analysis task, such as a neural network model, a deep learning model, etc. Exemplarily, the gene analysis model in the present disclosure can be selected as a hidden Markov model, a conditional random field and a neural network model, etc., which are not specifically limited here. The gene analysis model configured in each computing node may be the same or different, which is not specifically limited here. In a specific application scenario, before using the gene analysis model to analyze the shard file, it is also possible to perform task training for the gene sequencing data analysis task on the gene analysis model of each computing node based on a certain number of gene sequencing sample data through supervised learning, or unsupervised learning, semi-supervised learning, and reinforcement learning corresponding training methods, until the gene analysis model is determined to have reached a convergence state, and it is determined that the gene analysis model has completed training. The trained gene analysis model can then be directly put into the specific application of the gene sequencing data analysis task corresponding to the computing node, that is, the shard file is input into the trained gene analysis model, and the gene analysis model will directly output the corresponding shard file analysis results.

[0049] In a specific application scenario, after a compute node executes a shard file analysis task and obtains the shard file analysis results, the producer and consumer roles are reversed. The compute node that generated the shard file results provides the corresponding completion information, acting as the producer. The data center server acts as the consumer, monitoring the task. Once the task is completed, it merges the shard file analysis results provided by the compute node to produce the required complete result. The entire analysis process, from start to result merging, forms a closed loop of "total-part-total." Upon completion, intermediate results and other resources are cleared and returned to the overall system process. Output operation: The shard file analysis results receive the HDF5 shard result files sent by the compute node - the shard file analysis results detect that all tasks have been completed - the shard file analysis results read and merge all HDF5 result files to obtain the final analysis result, while removing redundant intermediate files. Accordingly, the data center server determines whether the shard file analysis task has completed by monitoring the status information in the task status queue. Upon detecting the completion of the shard file analysis task, it clears the redundant intermediate files generated by at least one target compute server and restores the computing resources of the multiple compute nodes configured for the target compute server.

[0050] The development environment involved in this disclosure is relatively complex. The specific software and hardware requirements for data center servers are shown in Table 1 below, and the specific software and hardware requirements for compute nodes are shown in Table 2 below. Although this disclosure can run stably on machines with standard CPU configurations, the algorithm module involves reading, writing, and storing large amounts of data, as well as model training and testing. This requires extensive data processing and operations. Running in a high-performance CPU or GPU-based hardware environment can significantly improve efficiency and stability.

[0051] Table 1:

[0052]

[0053] Table 2:

[0054]

[0055] As shown in FIG2 , an embodiment of the present disclosure provides a method for analyzing gene sequencing data. The method is applied to a gene sequencing data analysis system, specifically to a data center server. The method embodiment includes:

[0056] Step 101: Obtain path information of the fragment files generated during the gene sequencing process.

[0057] The fragment files are generated during the upstream gene sequencing step. They are fragmented into a specified number or by channel source, such as 5,000 base fragment signals or all base fragment signals from a channel. Specifically, they can be generated by a third-generation nanopore sequencing device and a supporting PC. The nanopore sequencing device collects the current signals generated when the gene passes through the pore. The sequencing fragment results and other information are then identified and filtered through an algorithm and saved as an h5 file structure on a specially mounted hard drive on the PC. The results are uploaded to the data center server upon completion of the sequencing process.

[0058] Step 102: Create a fragment file analysis task containing path information, and add the fragment file analysis task to the task queue.

[0059] The task queue, specifically implemented as a Redis queue, contains task parameters to be submitted to the process pool for execution. After generating a sharded file analysis task, it can be packaged into elements and added to the task queue. This ensures that the task queue does not backlog existing processes while allowing the next process to consume the elements in the queue and asynchronously perform the next analysis step.

[0060] Step 103: Receive multiple shard file analysis results sent by the target computing server, and merge the multiple shard file analysis results when detecting the end of the shard file analysis task. The shard file analysis results are obtained by calling multiple configured computing nodes to retrieve the shard files according to the path information contained in the shard file analysis task after the target computing server pulls the shard file analysis task from the task queue, and uses the genetic analysis model to analyze the shard files.

[0061] The target computing server is a computing server selected from the computing server cluster to execute the shard file analysis task.

[0062] In a specific application scenario, the target compute server can pull a shard file analysis task from the task queue and invoke multiple configured compute nodes to retrieve the shard files based on the path information contained in the shard file analysis task. After pulling the corresponding shard files, the compute nodes will each trigger their own data analysis tasks, specifically using a genetic analysis model to analyze the shard files and obtain the shard file analysis results. The shard file analysis results can then be sent to the data center server, which can then receive multiple shard file analysis results sent by multiple compute nodes.

[0063] Among them, the gene analysis model can be any computing model that can realize the task of gene sequencing data analysis, such as a neural network model, a deep learning model, etc. Exemplarily, the gene analysis model in the present disclosure can be selected as a hidden Markov model, a conditional random field, and a neural network model, etc., which are not specifically limited here. The gene analysis models configured in each computing node may be the same or different, and are not specifically limited here. It should be noted that there are many types of gene analysis models and they are adapted to different usage scenarios. The algorithm direction in this application is mainly integrated and embedded, which will not be elaborated on, and can support iterative updates of the integrated algorithm content.

[0064] The technical solution disclosed in the present invention is that the data center server can create a shard file analysis task containing the path information after obtaining the path information of the shard file generated in the gene sequencing process, and fill the shard file analysis task into the task queue; then it can receive multiple shard file analysis results sent by the target computing server, and merge the multiple shard file analysis results when detecting the end of the execution of the shard file analysis task, wherein the shard file analysis result is the result obtained by the target computing server pulling the shard file analysis task from the task queue, calling the configured multiple computing nodes to retrieve the shard file according to the path information contained in the shard file analysis task, and analyzing the shard file using the gene analysis model. The technical solution disclosed in the present invention can be used by a centralized data center server to allocate data analysis tasks, and multiple computing nodes to undertake the analysis tasks of the corresponding data shards. The asynchronous and multi-process implementation of the process analysis operations within the computing nodes can effectively improve the speed of analyzing the large amount of data generated by the third-generation sequencing, thereby improving the data analysis efficiency and meeting the analysis needs of the massive gene sequencing data under the data growth rate.

[0065] As shown in FIG3 , an embodiment of the present disclosure provides a method for analyzing gene sequencing data. The method is applied to a gene sequencing data analysis system, specifically to a target computing server. The method embodiment includes:

[0066] Step 201: Pull the shard file analysis task from the task queue.

[0067] For the embodiments of the present disclosure, the target computing server can pull the shard file analysis task to be executed from the task queue. The shard file analysis task can include the path information of the shard file. The multiple computing nodes configured by the target computing server can further read the content in the shard file based on the path information and further analyze it.

[0068] Step 202: Call multiple configured computing nodes to execute the shard file analysis task to obtain multiple shard file analysis results, where the shard file analysis results are the results obtained by the computing nodes searching for the shard files according to the path information contained in the shard file analysis task and analyzing the shard files using the genetic analysis model.

[0069] Compute nodes are virtualized compute container nodes. Virtualization uses software to transform a single physical host into multiple logical computers, each capable of independently running different operating systems and applications. Virtualization technology allows each virtual machine to have its own virtual hardware (virtual CPU, network interface card, memory, etc.), allowing the operating system running on the virtual machine to believe it has exclusive access to a single physical host. The software on the virtual machine runs on a virtual platform, not a real hardware platform. A container is simply a specialized process running on the host machine, and multiple containers share the same host operating system kernel. Containers are independent of the operating system's environment for running applications and can isolate and restrict application processes through the mechanisms and features provided by the operating system kernel.

[0070] For the embodiment of the present disclosure, the computing nodes will pull and consume tasks based on the information provided by the data center server. After retrieving the corresponding shard files, the computing nodes will each trigger their own data analysis tasks. The nodes are relatively independent in terms of resource occupation and task operation. When resources are sufficient, they can be considered independent of each other. Each computing node generates corresponding shard results after the data shard analysis process, and feeds back to the data center server after completion. The specific data flow is: task queue - obtain the shard file in the Hierarchical Data Format Version 5 (HDF5) format to the node locally through the optical fiber network - parse the HDF5 shard file - the data is inferred by the AI ​​model - the AI ​​inference result is decoded - the decoding result is written to the disk into the shard result file - the result is transmitted back to the data center server through the optical fiber network. Among them, when the computing node obtains the shard file, it can retrieve the shard file according to the path information contained in the shard file analysis task, wherein the shard file includes the base fragment signal divided according to the specified number or channel source in the upstream gene sequencing step. Furthermore, each computing node is equipped with a genetic analysis model. After retrieving a shard file, the computing node uses the genetic analysis model to analyze the shard file, obtaining analysis results. These results are then transmitted back to the data center server via the fiber optic network. This means that the data center server can receive analysis results from multiple shard files sent by multiple computing nodes. The genetic analysis models of multiple computing nodes form a genetic analysis cluster, which utilizes artificial intelligence (AI) learning methods and techniques to analyze gene sequencing data. This analysis of gene sequencing data can be transformed into various signals with distinct patterns and characteristics based on the unique physical and chemical properties of its bases. Images and electrical signals are ideal models for AI learning and analysis. Building a genetic analysis cluster distributes the computing resources for gene sequencing data analysis, allowing more computing resource nodes to accelerate and support the AI ​​learning process, thereby improving the speed and accuracy of gene sequencing data analysis.

[0071] Step 203: Send the analysis results of the multiple fragment files to the data center server.

[0072] The technical solution disclosed in the present invention can be used to allocate data analysis tasks by a centralized data center, and multiple computing nodes can be used to undertake the analysis tasks of the corresponding data fragments. The asynchronous and multi-process implementation of process-based analysis operations within the computing nodes can effectively improve the speed of analyzing the large amount of data generated by the third-generation sequencing, thereby improving the efficiency of data analysis and meeting the analysis needs of massive genetic sequencing data under the data growth rate. The present invention creatively integrates the rapidly increasing bioinformatics data in the third-generation sequencing scenario with various advanced technologies such as distributed data analysis and processing, virtualization technology, AI artificial intelligence, GPU parallel computing, asynchronous, multi-process, and message queues in the current computer field. Combining computer principles with specific software code design, it proposes and implements a genetic data analysis solution with easy-to-expand computing resources, iterative algorithms, node obstacle avoidance, and full resource utilization. In addition to the container and message queue technology used in the distributed solution disclosed in the present invention, which use existing computer software products, the rest are native designs and implementations. Using better distributed framework products in the computer field can further improve deployment and management efficiency or self-iterate to form a new set of specialized application framework systems. Moreover, the data analysis algorithm deployed by the present invention can continuously increase the overall analysis speed with iterative optimization. At the same time, with the emergence of new hardware with higher computing power and the resolution of compatibility issues, the project itself can follow up with iterations for further improvement, so that the analysis speed can be increased in line with the hardware improvement.

[0073] In a specific application scenario, as shown in Figure 4, after obtaining the path information of the shard file generated during the gene sequencing process, the data center server (data center / data server) can create a shard file analysis task containing the path information, fill the shard file analysis task into the task queue, and generate a task status queue for updating the status information of the shard file analysis task in the storage task queue; then the data center server can determine at least one target computing server (single / multi-computing card workstation) in the computing server cluster to execute the shard file analysis task, and the target computing server can pull the shard file analysis task in the task queue through data fiber mounting / transmission, and call the configured multiple computing nodes (computing container nodes of single computing card resources) to execute the shard file analysis task. Each computing node is configured with a gene analysis model. When executing a shard file analysis task, the computing node can retrieve the shard file according to the path information contained in the shard file analysis task, and use the gene analysis model to analyze the shard file to obtain the shard file analysis result (shard result), and then transmit the shard file analysis result back to the data center server through the optical fiber network; the data center server can determine whether the shard file analysis task has been completed by detecting the status information in the task status queue. When the execution of the shard file analysis task is completed, the server merges the received multiple shard file analysis results, clears the redundant intermediate files generated by at least one target computing server, and restores the computing resources of multiple computing nodes configured on the target computing server.

[0074] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.

[0075] FIG5 shows a schematic block diagram of an example electronic device 500 that can be used to implement an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0076] As shown in Figure 5, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 into a RAM (Random Access Memory) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0077] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0078] The computing unit 501 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the method for processing FASTQ data. For example, in some embodiments, the method for processing FASTQ data can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method described above can be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the aforementioned communication data processing method in any other appropriate manner (for example, by means of firmware).

[0079] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0080] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0081] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0082] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0083] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0084] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0085] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0086] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved. This is not a limitation herein.

[0087] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.

Claims

1. A gene sequencing data analysis system, characterized in that: include: A data center server and a computing server cluster, wherein the computing server cluster includes multiple computing servers, and each of the multiple computing servers is configured with multiple computing nodes; The data center server is used to obtain the path information of the fragmented files generated during the gene sequencing process, create a fragmented file analysis task containing the path information, and fill the fragmented file analysis task into the task queue; The data center server is connected to the computing server cluster, and the data center server is also used to determine at least one target computing server in the computing server cluster to execute the shard file analysis task. The at least one target computing server is used to pull the shard file analysis task from the task queue, and call the configured multiple computing nodes to execute the shard file analysis task, and send the multiple shard file analysis results output by the multiple computing nodes to the data center server.

2. The gene sequencing data analysis system according to claim 1, characterized in that: The computing node is used to retrieve the fragment file according to the path information contained in the fragment file analysis task, wherein the fragment file includes base fragment signals divided according to a specified number or channel source in the upstream gene sequencing step.

3. The gene sequencing data analysis system according to claim 2, characterized in that: A gene analysis model is configured in the computing node, and the computing node is used to analyze the fragment file using the gene analysis model to obtain a fragment file analysis result.

4. The gene sequencing data analysis system according to claim 1, characterized in that: The data center server is used to receive the multiple shard file analysis results sent by the multiple computing nodes, and merge the multiple shard file analysis results when detecting that the execution of the shard file analysis task is completed.

5. The gene sequencing data analysis system according to claim 4, characterized in that: The data center server is further configured to generate a task status queue, and the task status queue is configured to update and store status information of the shard file analysis tasks in the task queue.

6. The gene sequencing data analysis system according to claim 5, characterized in that: The data center server is further configured to determine whether the fragment file analysis task has been completed by detecting the status information in the task status queue.

7. The gene sequencing data analysis system according to claim 1, characterized in that: The data center server is also used to clear the redundant intermediate files generated by the at least one target computing server and restore the computing resources of the multiple computing nodes configured on the target computing server when detecting the end of the execution of the shard file analysis task.

8. A method for analyzing gene sequencing data, characterized in that: The method is applied to a data center server, and the method includes: Obtain the path information of the fragment files generated during gene sequencing; Creating a fragment file analysis task containing the path information, and adding the fragment file analysis task to a task queue; Receive multiple shard file analysis results sent by the target computing server, and merge the multiple shard file analysis results when detecting the end of the execution of the shard file analysis task, wherein the shard file analysis result is the result obtained by the target computing server pulling the shard file analysis task from the task queue, calling the configured multiple computing nodes to retrieve the shard file according to the path information contained in the shard file analysis task, and using the genetic analysis model to analyze the shard file.

9. A method for analyzing gene sequencing data, characterized in that: The method is applied to a target computing server, and the method includes: Pull the shard file analysis task from the task queue; Calling multiple configured computing nodes to execute the shard file analysis task to obtain multiple shard file analysis results, wherein the shard file analysis results are the results obtained by the computing nodes searching for shard files according to the path information contained in the shard file analysis task and analyzing the shard files using the gene analysis model; The analysis results of the multiple fragment files are sent to the data center server.

10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of claim 8 or 9.

11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to claim 8 or 9.

12. A computer program product, characterized in that A computer program is included which, when executed by a processor, implements the method according to claim 8 or 9.