Multi-sample automated analysis method, system and device based on nanopore sequencing

By automatically configuring and performing analysis tasks, the errors and inefficiency caused by manual parameter setting in multi-sample nanopore sequencing analysis are solved, and a more efficient analysis process is achieved.

CN117423385BActive Publication Date: 2025-05-13ZHEJIANG JIAHE TAIHONG BIOTECHNOLOGY CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311123875.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-01
Publication Date
2025-05-13
Estimated Expiration
2043-09-01

AI Technical Summary

Technical Problem

The prior art requires manual setting of analysis parameters in multi-sample nanopore sequencing analysis, which can easily lead to analysis errors. Since the sequencing time needs to be extended to ensure the amount of data, the analysis time will also increase accordingly, making the efficiency inefficient.

Method used

By configuring the task parameter file, automatically monitor the original sequencing data storage folder, identify the barcode number, and integrate it into analysis tasks, call bioinformatics software and analysis process script files for parallel execution, and automatically store the analysis results.

Benefits of technology

It effectively avoids analysis errors caused by manual setting of analysis parameters, shortens task analysis time, and improves analysis efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117423385B_ABST
    Figure CN117423385B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-sample automated analysis method, system and equipment based on nanopore sequencing, and relates to the field of automated analysis of bioinformatics. First, a task parameter file is configured according to the sample type of multiple samples, including a folder monitoring parameter, a result storage path parameter and a script library parameter; when executing a sequencing task, the task parameter file is read and parsed, and when new sequencing data is monitored, its barcode number is identified and the folder where the sequencing data is located, the file name and the analysis process parameters to be executed are integrated into an analysis task; bioinformatics software and analysis process script files are called, and analysis tasks are executed in parallel, and the analysis results and logs of the analysis tasks are stored; after all analysis tasks are completed, the analysis results are merged in units of barcode numbers. The present invention can effectively avoid analysis errors caused by manually setting analysis parameters, shorten task analysis time, and improve analysis efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automated biological information analysis, and in particular to a multi-sample automated analysis method, system and equipment based on nanopore sequencing. Background Art

[0002] Nanopore sequencing technology is a new generation sequencing method for single-molecule real-time sequencing. A molecular connector is covalently bound to the pore. After the nanopore protein is fixed on the resistor membrane, the nucleic acid is pulled through the nanopore by the motor protein. When the nucleic acid passes through the nanopore, the charge changes, causing the current on the resistor membrane to change. Since the diameter of the nanopore is very small, only a single nucleic acid polymer is allowed to pass through, and the charge properties of the single ATCG base are different, so different bases have different interferences on the current when passing through the protein nanopore. By real-time monitoring and decoding these current signals, the base sequence can be determined, thereby achieving sequencing.

[0003] Barcode is a pre-designed nucleic acid sequence. By adding specific "labels (barcodes)" to different samples, the sequencing data of different samples can be separated by sequence alignment during the sequencing process. Nanopore sequencers and supporting MinKnow software can generate raw sequencing data in real time and identify its barcode number (data format is *.fastq.gz), so as to store these raw sequencing data in folders named after the barcode number.

[0004] The current implementation process of third-generation sequencing analysis can be summarized as follows: the user selects the original data to be sequenced and sets the analysis parameters. The system receives the user input, calls the relevant analysis software or scripts to perform data analysis based on the input, and then outputs the analysis results. In the sequencing analysis task of adding barcodes to multiple samples (hereinafter referred to as multi-sample sequencing tasks), this type of analysis method has the following shortcomings: ① For different types of samples, different analysis parameters need to be set manually and different analysis processes need to be executed, which may lead to the wrong process selection; ② In order to ensure that each sample has a sufficient amount of data, the sequencing time needs to be extended, resulting in an increase in the corresponding analysis time. Summary of the invention

[0005] In response to the problems raised in the above background technology, the present invention provides a multi-sample automated analysis method, system and equipment based on nanopore sequencing to avoid analysis errors caused by manually setting analysis parameters, shorten task analysis time and improve analysis efficiency.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] In one aspect, the present invention provides a multi-sample automated analysis method based on nanopore sequencing, comprising:

[0008] The task parameter file is configured according to the sample types of the multiple samples to be sequenced; the task parameter file includes a folder monitoring parameter, a result storage path parameter and a script library parameter; the folder monitoring parameter is the folder path where the original sequencing data is stored; the result storage path parameter is the folder path where the task analysis results are stored; the script library parameter is the path where the analysis process script file to be executed for each barcode label is located;

[0009] Add different barcode labels to multiple samples and mix them on the machine, and perform sequencing tasks through Nanopore's sequencing software MinKnow;

[0010] After base recognition and barcode splitting, MinKnow software stores the raw sequencing data generated in real time in a folder named after the barcode number;

[0011] Read and parse the task parameter file, and monitor whether new sequencing data is generated in the folder where the original sequencing data is stored according to the folder monitoring parameters;

[0012] When new sequencing data is detected, its barcode number is identified and the folder where the sequencing data is located, the file name, and the analysis process parameters to be executed are integrated into the analysis task as canonical variables;

[0013] Load the analysis task into the analysis pool queue, wait for the analysis process to be idle, take out the analysis task, call the bioinformatics software and the analysis process script file, execute the analysis task in parallel, and store the analysis results of the analysis task and the analysis operation log;

[0014] After all analysis tasks are completed, all analysis results will be merged and displayed in units of barcode numbers.

[0015] Optionally, the reading and parsing of the task parameter file and monitoring whether new sequencing data is generated in the folder where the original sequencing data is stored according to the folder monitoring parameter specifically include:

[0016] Read and parse the task parameter file, and store the parameter information in the task parameter file as system global variables in the form of word-typical variables;

[0017] According to the folder monitoring parameters in the system global variables, monitor whether new sequencing data is generated in the folder where the original sequencing data is stored.

[0018] Optionally, when new sequencing data is detected, its barcode number is identified and the folder where the sequencing data is located, the file name and the analysis process parameters to be executed are integrated into an analysis task as typical variables, specifically including:

[0019] When new sequencing data is detected, its absolute path is called back and the getbarcode function is executed on the absolute path to obtain its barcode number;

[0020] Extract the script library parameters corresponding to the barcode number from the system global variables according to the barcode number, and obtain the analysis process parameters that need to be executed;

[0021] The sequencing data folder, file name, and analysis process parameters to be executed are integrated into analysis tasks as canonical variables.

[0022] Optionally, the analysis pool queue is implemented based on the queue class of python's multiprocessing.

[0023] Optionally, the extracting of the analysis task, calling the bioinformatics software and the analysis process script file, and executing the analysis task in parallel specifically includes:

[0024] The analysis tasks are parsed and extracted using Python and stored as Python word-typical variables. The bioinformatics software and analysis process script files are then called through Python's subprocess module to execute the corresponding analysis process; each analysis task running at the same time is executed in parallel.

[0025] Optionally, the bioinformatics software includes quality control software NanoPlot, NanoFilt, sequence alignment software minimap2, blast, and species identification software Centrifuge, kraken.

[0026] Optionally, storing the analysis results of the analysis task specifically includes:

[0027] The analysis results of the analysis task are stored in txt or csv format; the analysis results include the number of sequences measured under the barcode number, sequencing time, sequence length distribution, sequence quality and sequence alignment results.

[0028] On the other hand, the present invention also provides a multi-sample automated analysis system based on nanopore sequencing, comprising: a data monitoring module, a data analysis module and a result storage module;

[0029] The data monitoring module reads a task parameter file configured according to the sample type of multiple samples to be sequenced; the task parameter file includes a folder monitoring parameter, a result storage path parameter and a script library parameter; the folder monitoring parameter is the folder path where the original sequencing data is stored; the result storage path parameter is the folder path where the task analysis results are stored; the script library parameter is the path where the analysis process script file to be executed for each barcode label is located;

[0030] The data monitoring module parses the task parameter file, monitors whether new sequencing data is generated in the folder where the original sequencing data is stored according to the folder monitoring parameter; when monitoring the generation of new sequencing data, identifies its barcode number and integrates the folder where the sequencing data is located, the file name and the analysis process parameters to be executed into an analysis task as a typical variable; and transmits the analysis task to the data analysis module through the websocket technology;

[0031] The data analysis module loads the analysis task into the analysis pool queue, waits for the analysis process to be idle, takes out the analysis task, calls the bioinformatics software and the analysis process script file, executes the analysis task in parallel, and stores the analysis result of the analysis task and the analysis operation log in the result storage module;

[0032] When the data monitoring module detects that all analysis tasks are completed, it initiates a result merging request to the data analysis module; after receiving the request, the data analysis module merges all analysis results in units of barcode numbers according to the result storage path parameters and stores them in the result storage module.

[0033] In another aspect, the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following method when executing the computer program:

[0034] Obtain a task parameter file configured according to the sample types of multiple samples to be sequenced; the task parameter file includes a folder monitoring parameter, a result storage path parameter, and a script library parameter; the folder monitoring parameter is the folder path where the original sequencing data is stored; the result storage path parameter is the folder path where the task analysis results are stored; and the script library parameter is the path where the analysis process script file to be executed for each barcode tag is located;

[0035] Read and parse the task parameter file, and monitor whether new sequencing data is generated in the folder where the original sequencing data is stored according to the folder monitoring parameters;

[0036] When new sequencing data is detected, its barcode number is identified and the folder where the sequencing data is located, the file name, and the analysis process parameters to be executed are integrated into the analysis task as canonical variables;

[0037] Load the analysis task into the analysis pool queue, wait for the analysis process to be idle, take out the analysis task, call the bioinformatics software and the analysis process script file, execute the analysis task in parallel, and store the analysis results of the analysis task and the analysis operation log;

[0038] After all analysis tasks are completed, all analysis results will be merged and displayed in units of barcode numbers.

[0039] Optionally, the memory is a non-transitory computer-readable storage medium.

[0040] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0041] The present invention provides a multi-sample automated analysis method, system and equipment based on nanopore sequencing. First, according to the sample type of the multi-sample to be sequenced, the task parameter file is configured, and the task parameter file includes a folder monitoring parameter, a result storage path parameter and a script library parameter; when the sequencing software MinKnow of Nanopore executes the sequencing task, the task parameter file is read and parsed, and the folder storing the original sequencing data is monitored according to the folder monitoring parameter to determine whether new sequencing data is generated; when new sequencing data is monitored, its barcode number is identified and the folder where the sequencing data is located, the file name and the analysis process parameters to be executed are integrated into the analysis task as a typical variable; the analysis task is loaded into the analysis pool queue, and when the analysis process is idle, the analysis task is taken out, the bioinformatics software and the analysis process script file are called, the analysis task is executed in parallel, and the analysis result of the analysis task and the analysis operation log are stored; when all analysis tasks are completed, all analysis results are merged and displayed in units of barcode numbers. The present invention effectively avoids the analysis errors caused by manually setting the analysis parameters, greatly shortens the task analysis time, and improves the analysis efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1A flowchart of a multi-sample automated analysis method based on nanopore sequencing according to the present invention;

[0044] Figure 2 A schematic diagram of a task parameter file provided by an embodiment of the present invention;

[0045] Figure 3 A multi-sample automated analysis pipeline operation diagram provided by the present invention;

[0046] Figure 4 This is a schematic diagram of the overall architecture of a multi-sample automated analysis system based on nanopore sequencing according to the present invention;

[0047] Figure 5 A schematic diagram of the getbarcode function provided in an embodiment of the present invention;

[0048] Figure 6 A schematic diagram of the working process of the multi-sample automated analysis system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] The purpose of the present invention is to provide a multi-sample automated analysis method, system and equipment based on nanopore sequencing to avoid analysis errors caused by manually setting analysis parameters, shorten task analysis time and improve analysis efficiency.

[0051] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0052] Figure 1 This is a flow chart of a multi-sample automated analysis method based on nanopore sequencing of the present invention. Figure 1 The present invention provides a multi-sample automated analysis method based on nanopore sequencing, comprising:

[0053] Step 1: Configure a task parameter file according to the sample types of the multiple samples to be sequenced; the task parameter file includes folder monitoring parameters, result storage path parameters, and script library parameters.

[0054] In order to solve the problem that manually setting analysis parameters may easily lead to analysis process errors, the present invention pre-integrates the analysis process into a script file, and pre-integrates the user input required for each sample analysis task into a task parameter file. Figure 2 A schematic diagram of a task parameter file provided by an embodiment of the present invention is shown in FIG. Figure 2 , the task parameter file contains the following information:

[0055] (1) Folder listening parameter (listenDir): the folder path where the original sequencing data of this sequencing task is stored;

[0056] (2) Result storage path parameter (ouputDir): the folder path where the analysis results of this sequencing task are stored;

[0057] (3) Script library parameter (script): the path to the analysis process script file that needs to be executed for each barcode label in this sequencing task;

[0058] (4) Other parameters: such as sample information and sample introduction of each barcode in this sequencing task, etc.

[0059] Step 2: Add different barcode labels to multiple samples and mix them on the machine, and perform the sequencing task through Nanopore's sequencing software MinKnow.

[0060] For multi-sample sequencing tasks, different samples may require different bioinformatics analysis processes. For example, for an unknown sample, it is necessary to prepare a library of all the DNA sequences it contains through random primers, and perform a bioinformatics analysis process for species identification. For a sample of a known pathogen, if you want to know its typing information, you need to perform an analysis process for evolutionary tree analysis. The present invention allows users to execute personalized analysis processes suitable for the sample according to their sample types. Sample types generally include blood samples, tissue samples, fecal samples, etc. For a blind sample, general screening may be required, and the database used in the analysis process is a large library; for a known sample, in order to screen a certain pathogen, the required database can be a bacterial library or a viral library.

[0061] The sequencer used in the sequencing process is officially provided by Nanopore, and the sequencing software MinKnow is Nanopore's official open source software. The analysis software used in the bioinformatics process scripts used in the examples are all open source software in the industry.

[0062] Step 3: After base recognition and barcode splitting, MinKnow software stores the raw sequencing data generated in real time in a folder named after the barcode number.

[0063] Nanopore sequencer and its supporting MinKnow software can generate raw sequencing data in real time and identify its barcode number (data format is *.fastq.gz), thereby storing these raw sequencing data in folders named after the barcode number.

[0064] Step 4: Read and parse the task parameter file, and monitor whether new sequencing data is generated in the folder where the original sequencing data is stored according to the folder monitoring parameters.

[0065] The multi-sample automated analysis system of the present invention is started, the task parameter file is read and parsed, and the parameter information in the task parameter file is stored as a system global variable as a typical variable. Then, according to the folder monitoring parameter in the system global variable, whether new sequencing data is generated in the folder where the original sequencing data is stored can be monitored.

[0066] Step 5: When new sequencing data is detected, its barcode number is identified and the folder where the sequencing data is located, the file name, and the analysis process parameters to be executed are integrated into the analysis task as typical variables.

[0067] When new sequencing data (in files) is detected, its absolute path is called back and the getbarcode function is executed on the absolute path. The getbarcode function obtains its barcode number by matching the file path with a regular expression. According to the barcode number, the script library parameters corresponding to the barcode number are extracted from the system global variables to obtain the analysis process parameters that need to be executed. The folder where the newly generated sequencing data is located, the file name, and the analysis process parameters that need to be executed are integrated into the analysis task as a typical variable.

[0068] Step 6: Load the analysis task into the analysis pool queue, wait for the analysis process to be idle, take out the analysis task, call the bioinformatics software and the analysis process script file, execute the analysis task in parallel, and store the analysis results of the analysis task and the analysis operation log.

[0069] The analysis pool queue is implemented based on Python's multiprocessing queue class, and multi-process parallel analysis is implemented through the multiprocessing pool class, that is, Figure 3 The pipeline operation shown.

[0070] The analysis tasks extracted by python are parsed and stored as python word typical variables, and then the bioinformatics software and the analysis process script file are called through the subprocess module of python to execute the corresponding analysis process. For each analysis task running in the same period of time, they are executed in parallel. The bioinformatics software includes but is not limited to quality control software such as NanoPlot and NanoFilt, sequence alignment software such as minimap2 and blast, and species identification software such as Centrifuge and kraken.

[0071] Step 7: After all analysis tasks are completed, all analysis results are merged and displayed in units of barcode numbers.

[0072] After all analysis tasks are completed, the analysis results of this batch of sequencing tasks are integrated by barcode number. The analysis results are generally saved in txt, csv and other formats that are easy to understand and merge. The analysis results include the number of sequences measured under the barcode number, sequencing time, sequence length distribution, sequence quality, and sequence alignment results. Furthermore, these distributed analysis results can also be combined, counted, and plotted to obtain the overall analysis results.

[0073] Compared with the prior art, the advantages of the method of the present invention are: (1) convenient use of parameters and analysis process scripts. The analysis process script library and parameters are saved on the server in the form of files (mainly in Linux, Python, R and other scripts), so that the same type of analysis tasks do not need to set parameters multiple times, effectively avoiding the analysis error problem caused by manually setting process parameters; (2) after detecting the generation of raw sequencing data, the raw data barcode number will be automatically identified and a personalized analysis process will be executed, greatly improving the analysis efficiency; (3) real-time sequencing and analysis, when the generation of data is detected, the data will be added to the analysis pool queue in time, shortening the time of the analysis phase and alleviating the pressure of big data analysis on the server.

[0074] Figure 4 This is a schematic diagram of the overall architecture of a multi-sample automated analysis system based on nanopore sequencing in the present invention. Figure 4The multi-sample automated analysis system based on nanopore sequencing of the present invention includes: a data monitoring module, a data analysis module and a result storage module. The data monitoring module is used as a websocket client, and the user interface is built based on electron and nodejs technology. The import of task parameter files and the merging of analysis results are all started by it. The data analysis module is used as a websocket server, based on python multiprocessing implementation, and uses the pool and queue classes defined in it to implement Figure 3 The data monitoring module and the data analysis module communicate with each other through websocket technology. The data monitoring module is used to detect whether new sequencing data is generated in the monitored original sequencing data storage folder. Once there is, the folder where the sequencing data is located, the file name and the analysis process parameters to be executed are integrated into an analysis task as a typical variable and passed to the data analysis module. The data analysis module receives the analysis task passed by the data monitoring module and loads it into the analysis pool queue. Without affecting the server load balancing, the analysis tasks are taken out from the analysis pool queue in turn, the bioinformatics software and scripts are called, and the analysis tasks are executed in parallel. The result storage module stores the analysis results of the analysis task, and after the analysis of all the original sequencing files of the multi-sample sequencing task is completed, the result merging function is called to merge the results.

[0075] The specific functions of each module are as follows:

[0076] 1) Data monitoring module: Specifically, the chokidar module of nodejs is used to monitor the events of new files in the folder. Based on the folder monitoring parameters of the task parameter file, once a new file with the suffix gz is generated in the monitored folder and its subfolders, the absolute path of the file will be called back. The absolute path is the path starting from the root directory ( / ), which can uniquely determine the location of the file or directory in the file system. Due to the naming conventions of the sequencing data file and the folder where it is located, the barcode number of the file can be obtained based on JavaScript regular matching. Specifically, the absolute path is executed Figure 5 The getbarcode function shown in the figure obtains the barcode number through regular matching. Based on the system global variables generated by the task parameter file, the path of the corresponding analysis process script file is matched through the barcode parameter, and the required analysis parameters (the folder where the original sequencing data is located, the file name, the barcode type, and the path where the analysis process script file corresponding to the barcode number is located) are integrated into the analysis task as a typical variable, and passed to the data analysis module through the websocket technology.

[0077] 2) Data analysis module: As new sequencing data is continuously generated, analysis tasks are continuously passed and stored in the analysis pool queue (based on the queue class of python's multiprocessing), and multi-process parallel analysis (i.e., the pipeline operation of 3) is implemented through the pool class of multiprocessing. Websocket transmits the analysis task, and the python of the data analysis module parses the analysis task, saves it as a python word-typical variable, and then assembles and executes it through the subprocess module of python. The data analysis module executes the corresponding analysis process according to the script library parameters passed over. After each analysis task is executed, the analysis operation status is transmitted to the websocket client (i.e., the data monitoring module) through websocket technology. The websocket client will record the analysis operation status for the control of the analysis status. The analysis run status includes three situations: completed, pending analysis and failed. Completed means that the analysis task is completed without any warning; pending analysis means that the analysis task is added to the analysis pool queue but not taken out of the analysis pool queue; failed means that the callback file analysis fails. Since the bioinformatics analysis process is step-by-step, the result of the previous step is used as the output of the next step. Therefore, there may be a situation where the previous step does not meet the requirements and cannot be used as the input for the next step, resulting in analysis failure.

[0078] 3) Result storage module: The essence of the analysis task is a typical variable, which contains an outputDir parameter, which points to the analysis result storage path. The result storage module of the present invention is to facilitate data management, and stipulates that all data results are stored in the result storage module under this path. Figure 2 The result storage path " / data / cexufenxi / 20230724" in the task parameter file can be understood as / data / cexufenxi is the result storage module, 20230724 is the name of the multi-sample sequencing task. If there are multiple multi-sample sequencing tasks on the same day, other names can be added, such as 20230724-project2.

[0079] Based on the result storage path parameter in the task parameter file, the analysis results of each analysis task of the data analysis module will be stored in the subfolder numbered with the barcode. At the end of sequencing, a sequencing summary file will be generated in the listening folder directory. When the websocket client listens to the generation of this file and detects that there is no analysis task running on the websocket server, the websocket client will initiate a result merge request to the websocket server. The websocket server receives the request and performs result merging with the barcode number according to the result storage path parameter. Because the original sequencing data files are analyzed one by one in parallel, one file corresponds to one analysis result, and the analysis result of one file cannot represent the overall situation of the entire sample. Since one barcode number represents one sample in the present invention, the overall analysis result of the sample can only be obtained by integrating in units of barcode.

[0080] Figure 6 Schematic diagram of the working process of the multi-sample automated analysis system provided by the embodiment of the present invention. Figure 6 First, manually configure the task parameter file according to the sample type. The task parameter file contains at least the path of the original sequencing file, the analysis result storage path, and the personalized bioinformatics process script path required by the barcode label used for the sample type. The task parameter file type is a json or yaml file. Add different barcode labels to the samples of multiple samples and mix them on the machine, and perform the sequencing task through Nanopore's official sequencing software MinKnow. After base calling and barcode splitting, MinKnow software will assign the sequence to the corresponding folder according to the barcode. When the sequencing software MinKnow configures the sequencing task, a folder will be specified. When the sequencing task is running, a DNA chain will pass through the nanopore of the sequencer. The sequencer will capture the changes in the electrical signal when it passes through the hole, and then perform base calling (base recognition: electrical signal to base ATCG). The number of sequences stored in the file is fixed, and the value of the number is set by the sequencing software. Whenever the number of sequences of this value is measured, these sequences will generate a sequencing original file under the corresponding barcode folder of the original test data generation folder.

[0081] After the above sequencing task is run, the multi-sample automatic analysis system of the present invention is started, the data monitoring module reads and parses the task parameter file, and stores it as a system global variable, and monitors whether new sequencing data is generated in the folder where the original sequencing data is stored according to the folder monitoring parameter; when new sequencing data is monitored, its barcode number is identified and the folder where the sequencing data is located, the file name and the analysis process parameters to be executed are integrated into an analysis task as typical variables; and the analysis task is passed to the data analysis module through the websocket technology.

[0082] The data analysis module loads the analysis task into the analysis pool queue, waits for the analysis process to be idle, takes out the analysis task, calls the bioinformatics software and the analysis process script file, executes the analysis task in parallel, and stores the analysis results of the analysis task and the analysis operation log in the result storage module.

[0083] When the data monitoring module detects that all analysis tasks are completed, it initiates a result merging request to the data analysis module; after receiving the request, the data analysis module merges all analysis results in units of barcode numbers according to the result storage path parameters and stores them in the result storage module.

[0084] Furthermore, the present invention also provides an electronic device, which may include: a processor, a communication interface, a memory and a communication bus. The processor, the communication interface and the memory communicate with each other via the communication bus. The processor may call a computer program in the memory to execute the following method:

[0085] Obtain a task parameter file configured according to the sample types of multiple samples to be sequenced; the task parameter file includes a folder monitoring parameter, a result storage path parameter, and a script library parameter; the folder monitoring parameter is the folder path where the original sequencing data is stored; the result storage path parameter is the folder path where the task analysis results are stored; and the script library parameter is the path where the analysis process script file to be executed for each barcode tag is located;

[0086] Read and parse the task parameter file, and monitor whether new sequencing data is generated in the folder where the original sequencing data is stored according to the folder monitoring parameters;

[0087] When new sequencing data is detected, its barcode number is identified and the folder where the sequencing data is located, the file name, and the analysis process parameters to be executed are integrated into the analysis task as canonical variables;

[0088] Load the analysis task into the analysis pool queue, wait for the analysis process to be idle, take out the analysis task, call the bioinformatics software and the analysis process script file, execute the analysis task in parallel, and store the analysis results of the analysis task and the analysis operation log;

[0089] After all analysis tasks are completed, all analysis results will be merged and displayed in units of barcode numbers.

[0090] In addition, when the computer program in the above-mentioned memory is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-transitory computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks or optical disks.

[0091] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0092] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A multi-sample automated analysis method based on nanopore sequencing, characterized in that: include: Configure the task parameter file according to the sample types of multiple samples to be sequenced; The task parameter file includes a folder monitoring parameter, a result storage path parameter, and a script library parameter; the folder monitoring parameter is the folder path for storing the original sequencing data; The result storage path parameter is the folder path where the task analysis results are stored; The script library parameter is the path where the analysis process script file to be executed for each barcode label is located; Add different barcode labels to multiple samples and mix them on the machine, and perform sequencing tasks through Nanopore's sequencing software MinKnow; After base recognition and barcode splitting, MinKnow software stores the raw sequencing data generated in real time in a folder named after the barcode number; Read and parse the task parameter file, and monitor whether new sequencing data is generated in the folder where the original sequencing data is stored according to the folder monitoring parameters; When new sequencing data is detected, its barcode number is identified and the folder where the sequencing data is located, the file name, and the analysis process parameters to be executed are integrated into the analysis task as canonical variables; Load the analysis task into the analysis pool queue, wait for the analysis process to be idle, take out the analysis task, call the bioinformatics software and the analysis process script file, execute the analysis task in parallel, and store the analysis results of the analysis task and the analysis operation log; After all analysis tasks are completed, all analysis results will be merged and displayed in units of barcode numbers.

2. The multi-sample automated analysis method based on nanopore sequencing according to claim 1, characterized in that: The reading and parsing of the task parameter file and monitoring whether new sequencing data is generated in the folder storing the original sequencing data according to the folder monitoring parameter specifically include: Read and parse the task parameter file, and store the parameter information in the task parameter file as system global variables in the form of word-typical variables; According to the folder monitoring parameters in the system global variables, monitor whether new sequencing data is generated in the folder where the original sequencing data is stored.

3. The multi-sample automated analysis method based on nanopore sequencing according to claim 2, characterized in that: When new sequencing data is detected, its barcode number is identified and the folder where the sequencing data is located, the file name and the analysis process parameters to be executed are integrated into an analysis task as a typical variable, specifically including: When new sequencing data is detected, its absolute path is called back and the getbarcode function is executed on the absolute path to obtain its barcode number; Extract the script library parameters corresponding to the barcode number from the system global variables according to the barcode number, and obtain the analysis process parameters that need to be executed; The sequencing data folder, file name, and analysis process parameters to be executed are integrated into analysis tasks as canonical variables.

4. The multi-sample automated analysis method based on nanopore sequencing according to claim 1, characterized in that: The analysis pool queue is implemented based on the queue class of python's multiprocessing.

5. The multi-sample automated analysis method based on nanopore sequencing according to claim 1, characterized in that: The extracting of the analysis task, calling the bioinformatics software and the analysis process script file, and executing the analysis task in parallel specifically include: The analysis tasks are parsed and extracted using Python and stored as Python word-typical variables. The bioinformatics software and analysis process script files are then called through Python's subprocess module to execute the corresponding analysis process; each analysis task running at the same time is executed in parallel.

6. The multi-sample automated analysis method based on nanopore sequencing according to claim 5, characterized in that: The bioinformatics software includes quality control software NanoPlot and NanoFilt, sequence alignment software minimap2 and blast, and species identification software Centrifuge and kraken.

7. The multi-sample automated analysis method based on nanopore sequencing according to claim 1, characterized in that: The storing of the analysis results of the analysis task specifically includes: The analysis results of the analysis task are stored in txt or csv format; the analysis results include the number of sequences measured under the barcode number, sequencing time, sequence length distribution, sequence quality and sequence alignment results.

8. A multi-sample automated analysis system based on nanopore sequencing, characterized in that: include: Data monitoring module, data analysis module and result storage module; The data monitoring module reads a task parameter file configured according to the sample types of multiple samples to be sequenced; The task parameter file includes a folder monitoring parameter, a result storage path parameter, and a script library parameter; the folder monitoring parameter is the folder path for storing the original sequencing data; The result storage path parameter is the folder path where the task analysis results are stored; The script library parameter is the path where the analysis process script file to be executed for each barcode label is located; The data monitoring module parses the task parameter file and monitors whether new sequencing data is generated in the folder storing the original sequencing data according to the folder monitoring parameter; When new sequencing data is detected, its barcode number is identified and the folder where the sequencing data is located, the file name and the analysis process parameters to be executed are integrated into an analysis task as a typical variable; and the analysis task is transmitted to the data analysis module through the websocket technology; The data analysis module loads the analysis task into the analysis pool queue, waits for the analysis process to be idle, takes out the analysis task, calls the bioinformatics software and the analysis process script file, executes the analysis task in parallel, and stores the analysis result of the analysis task and the analysis operation log in the result storage module; When the data monitoring module detects that all analysis tasks are completed, it initiates a result merging request to the data analysis module; after receiving the request, the data analysis module merges all analysis results in units of barcode numbers according to the result storage path parameters and stores them in the result storage module.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the following method is implemented: Obtaining a task parameter file configured according to the sample types of multiple samples to be sequenced; the task parameter file includes a folder monitoring parameter, a result storage path parameter, and a script library parameter; the folder monitoring parameter is a folder path for storing original sequencing data; The result storage path parameter is the folder path where the task analysis results are stored; The script library parameter is the path where the analysis process script file to be executed for each barcode label is located; Read and parse the task parameter file, and monitor whether new sequencing data is generated in the folder where the original sequencing data is stored according to the folder monitoring parameters; When new sequencing data is detected, its barcode number is identified and the folder where the sequencing data is located, the file name, and the analysis process parameters to be executed are integrated into the analysis task as canonical variables; Load the analysis task into the analysis pool queue, wait for the analysis process to be idle, take out the analysis task, call the bioinformatics software and the analysis process script file, execute the analysis task in parallel, and store the analysis results of the analysis task and the analysis operation log; After all analysis tasks are completed, all analysis results will be merged and displayed in units of barcode numbers.

10. The electronic device according to claim 9, characterized in that: The memory is a non-transitory computer-readable storage medium.

Citation Information

Patent Citations

  • Automatic analysis method for DNA sequencing data

    CN115565609A

  • Analysis sequence control system and analysis sequence control method

    WO2019193810A1