Control method and analysis system
The control method and analysis system address the challenge of analyzing multiple nucleic acid sequence data from the same subject by linking and processing datasets with linking information, ensuring accurate and efficient genetic panel testing.
Patent Information
- Application Number
- JP2021178344
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2041-10-29
AI Technical Summary
Existing systems fail to accurately and efficiently analyze multiple nucleic acid sequence data sets from the same subject obtained by next-generation sequencers, particularly in matched pair testing where tumor and non-tumor samples are analyzed together, due to the lack of consideration for linking and processing data from different samples of the same subject.
A control method and analysis system that receive and link nucleic acid sequence data from multiple library samples of the same subject, using sequencers to generate and transmit datasets with linking information, enabling accurate analysis by a second facility.
Enables rapid and accurate analysis of multiple nucleic acid sequence data from the same subject, ensuring correct pairing and analysis of tumor and non-tumor samples, thereby improving the efficiency and accuracy of genetic panel testing.
Smart Images

Figure 0007794604000001 
Figure 0007794604000002 
Figure 0007794604000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a control method for controlling a computer to analyze, in a second facility, nucleic acid sequence data obtained in a first facility using a sequencer that reads nucleic acid sequences in a genetic panel test. The present invention also relates to an analysis system for analyzing, in a second facility, nucleic acid sequence data obtained in a first facility using a sequencer that reads nucleic acid sequences in a genetic panel test. [Background technology]
[0002] With the advancement of cancer genomic medicine, medical facilities are establishing systems for conducting gene panel testing. Among these, an increasing number of medical facilities are introducing next-generation sequencers (NGS) into their laboratories and accumulating knowledge gained through gene panel testing in order to utilize it in new research.
[0003] On the other hand, genetic panel testing requires data analysis by bioinformaticians. However, there are few bioinformaticians, and it can be difficult to secure such personnel. Therefore, there is a need to outsource the analysis of nucleic acid sequence data obtained at medical facilities for genetic panel testing to external specialized institutions.
[0004] Patent Document 1 describes a system in which a sequencer acquires nucleic acid sequence data, transmits the acquired nucleic acid sequence data to a cloud environment, and analyzes the nucleic acid sequence data in the cloud environment. According to the system described in Patent Document 1, it is possible to analyze the nucleic acid sequence data acquired by the sequencer in the cloud environment. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] U.S. Patent No. 9,444,880 Summary of the Invention [Problem to be solved by the invention]
[0006] NGS typically simultaneously measures a large number (e.g., 16) of test samples (libraries), so a single measurement can obtain multiple nucleic acid sequence data corresponding to multiple libraries collected from multiple subjects. Furthermore, depending on the type of gene panel test, it may be necessary to analyze a set of nucleic acid sequence data corresponding to multiple libraries prepared from the same subject's specimen.
[0007] For example, in matched pair testing, nucleic acid sequence data from tumor samples and non-tumor samples collected from the same subject are analyzed as a set. In such cases, when outsourcing the analysis of nucleic acid sequence data in gene panel testing, the external analysis facility must extract the correct combination of nucleic acid sequence data corresponding to multiple samples from the same subject from the multiple nucleic acid sequence data obtained by NGS, and then analyze the multiple nucleic acid sequence data of the correct combination.
[0008] However, the above Patent Document 1 did not take into consideration the possibility of an external analysis facility extracting multiple nucleic acid sequence data for the same subject from multiple nucleic acid sequence data acquired by the medical facility that requested the analysis, and performing analysis using the multiple nucleic acid sequence data for the subject.
[0009] Therefore, an object of the present invention is to provide a control method and an analysis system that enable a second facility to accurately and quickly perform analysis using multiple nucleic acid sequence data of the same subject based on nucleic acid sequence data acquired by a first facility. [Means for solving the problem]
[0010] The control method of the present invention is a control method for controlling a computer in a second facility to analyze nucleic acid sequence data obtained at a first facility using a sequencer that reads nucleic acid sequences in a genetic panel test, the control method receiving from the first facility via a network a sequence dataset containing multiple pieces of nucleic acid sequence data obtained using the sequencer, the multiple pieces of nucleic acid sequence data corresponding to each of multiple library samples including a first library sample and a second library sample prepared from a specimen of the same subject, and linking information indicating that the first library sample and the second library sample were prepared from the specimen of the same subject; analyzing the first sequence data and the second sequence data corresponding to the first library sample and the second library sample linked by the linking information; and outputting analysis information based on the analysis results of the first sequence data and the analysis results of the second sequence data.
[0011] The analysis system of the present invention is an analysis system for analyzing, in a gene panel test, nucleic acid sequence data obtained at a first facility using a sequencer that reads nucleic acid sequences at a second facility, the analysis system comprising: a first computer that receives, via a network from the first facility, sequence datasets containing nucleic acid sequence data obtained using the sequencer, the sequence datasets corresponding to each of a plurality of library samples, including a first library sample and a second library sample, prepared from a specimen of the same subject, and linking information indicating that the first library sample and the second library sample were prepared from the specimen of the same subject; and transmits the sequence datasets and the linking information obtained from the first facility to a second computer; and the second computer that analyzes the first sequence data and the second sequence data linked by the linking information and outputs analysis information based on the analysis results of the first sequence data and the analysis results of the second sequence data.
[0012] The analysis system of the present invention is an analysis system for analyzing, in a gene panel test, nucleic acid sequence data obtained at a first facility using a sequencer that reads nucleic acid sequences at a second facility, and includes a computer that receives, via a network from the first facility, a sequence dataset containing nucleic acid sequence data obtained using the sequencer, corresponding to each of a plurality of library samples, including a first library sample and a second library sample, prepared from a specimen of the same subject, and linking information indicating that the first library sample and the second library sample were prepared from the specimen of the same subject, analyzes the first sequence data and the second sequence data linked by the linking information, and outputs analysis information based on the analysis results of the first sequence data and the analysis results of the second sequence data. [Effects of the Invention]
[0013] According to the present invention, a control method and an analysis system can be provided that enable accurate and rapid analysis using multiple nucleic acid sequence data of the same subject at a second facility based on nucleic acid sequence data acquired at a first facility. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a schematic configuration diagram of a nucleic acid information transmission and reception system installed in each facility according to a first embodiment. [Figure 2] 10 is a flowchart illustrating a process executed by a sequencer. [Figure 3] FIG. 1 is a diagram showing an example of a sample sheet; [Figure 4] FIG. 10 is a diagram illustrating an example of sequence run data created by a sequencer. [Figure 5] FIG. 10 is a diagram showing a sample sheet of a modified example. [Figure 6] FIG. 10(A) is a diagram showing a sample sheet of another modified example, and FIG. 10(B) is a diagram showing a sample sheet of yet another modified example. [Figure 7]10 is a flowchart illustrating the processes executed by the control units of the data transmitting device, the accepting device, and the nucleic acid sequence analyzing device. [Figure 8] FIG. 10 is a diagram showing an example of a case registration screen displayed on a display unit of the data transmission device. [Figure 9] 10 is a flowchart illustrating details of a consistency verification process executed by a control unit of the accepting device. [Figure 10] 10 is a flowchart showing an example of a processing procedure when a control unit of a nucleic acid sequence analyzer determines a nucleic acid sequence. [Figure 11] FIG. 1 is a schematic diagram showing how a single variant reference sequence is generated. [Figure 12] 10 is a flowchart showing a process performed by a control unit of a nucleic acid sequence analyzer to detect somatic mutations. [Figure 13] 10 is a flowchart showing a process performed by a control unit of a nucleic acid sequence analyzer to detect germ cell mutations. [Figure 14] (A) is a diagram showing an example of a nucleic acid sequence of a somatic mutation, and (B) is a diagram showing an example of a nucleic acid sequence of a germline mutation. [Figure 15] FIG. 10 is a diagram showing an example of a report format for an analysis report. [Figure 16] 10 is a flowchart illustrating processing executed by a sequencer in the second embodiment. [Figure 17] FIG. 10 is a diagram showing an example of a sample sheet according to the second embodiment. [Figure 18] FIG. 11 is a diagram showing an example of a case registration screen displayed on a display unit of a data transmission device in the second embodiment. [Figure 19] 10 is a flowchart illustrating details of a consistency verification process executed by a control unit of the accepting device in the second embodiment. [Figure 20] 11 is a flowchart illustrating processing executed by a sequencer in the third embodiment. [Figure 21] FIG. 10 is a diagram showing an example of a sample sheet according to the third embodiment. [Figure 22] FIG. 11 is a diagram showing an example of a case registration screen displayed on a display unit of a data transmission device in the third embodiment. [Figure 23] 11 is a flowchart illustrating details of a consistency verification process executed by a control unit of the accepting device in the third embodiment. [Figure 24] FIG. 10 is a schematic configuration diagram of a nucleic acid information transmission and reception system installed in each facility according to a fourth embodiment. [Figure 25] 10 is a flowchart illustrating processing executed by each control unit of the data transmitting device and the reception / analysis device. [Figure 26] 13 is a flowchart illustrating processing executed by a sequencer in the fifth embodiment. [Figure 27] 13 is a flowchart illustrating the process performed by the control unit of the reception / analysis device in the fifth embodiment to determine the types of mutations in multiple pieces of nucleic acid sequence data contained in one sequence dataset. [Figure 28] 13 is a flowchart illustrating the processes executed by the control units of the data transmitting device, the accepting device, and the nucleic acid sequence analyzing device in the sixth embodiment. [Figure 29] 13 is a flowchart illustrating processing executed by each control unit of the data transmitting device and the reception / analysis device in the sixth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0015] Hereinafter, examples of embodiments of a control method and an analysis system according to the present invention will be described in detail with reference to the drawings. The embodiments described below are merely examples, and the present invention is not limited to the following embodiments. In addition, in each of the following embodiments, the same components are assigned the same reference numerals in the drawings, and duplicated explanations will be omitted.
[0016] In the following description, tumors may include benign epithelial tumors, benign non-epithelial tumors, malignant epithelial tumors, and malignant non-epithelial tumors. The origin of tumors is not limited. Examples of tumor origins include (1) respiratory system tissues such as the trachea, bronchi, or lungs; (2) digestive tract tissues such as the nasopharynx, esophagus, stomach, duodenum, jejunum, ileum, cecum, appendix, ascending colon, transverse colon, sigmoid colon, rectum, or anus; (3) liver; (4) pancreas; (5) urinary system tissues such as the bladder, ureters, or kidneys; (6) female reproductive system tissues such as the ovaries, fallopian tubes, and uterus; (7) mammary glands; (8) male reproductive system tissues such as the prostate; (9) skin; (10) endocrine system tissues such as the hypothalamus, pituitary gland, thyroid gland, parathyroid gland, and adrenal glands; (11) central nervous system tissues; (12) bone and soft tissues; (13) hematopoietic system tissues such as bone marrow and lymph nodes; and (14) blood vessels.
[0017] In the following description, a sample refers to a sample prepared from a specimen such as tissue, body fluid, or excrement collected from a subject, and contains nucleic acids derived from tumor cells or non-tumor cells. Nucleic acids include deoxyribonucleic acid (hereinafter referred to as DNA) or ribonucleic acid (hereinafter referred to as RNA). Nucleic acids may be present intracellularly or may be present in body fluids after leaking out of cells upon cell destruction or death. Examples of nucleic acids present in body fluids include cell-free DNA (cfDNA) and circulating tumor DNA (ctDNA). Body fluids include blood, bone marrow fluid, ascites, pleural effusion, and cerebrospinal fluid. Excrement includes feces, urine, and sputum. Fluids obtained after washing a part of a patient's body, such as peritoneal lavage fluid or colonic lavage fluid, may also be used as specimens. The amount of nucleic acid contained in a specimen is not limited as long as it is an amount that allows the nucleic acid sequence to be detected. Furthermore, when obtaining nucleic acid sequence data derived from non-tumor cells, a specimen containing nucleic acid derived from non-tumor cells is used. The concentration of non-tumor cells contained in the above-mentioned tissues, body fluids, etc. is not limited as long as the nucleic acid sequence present in the non-tumor cells can be detected. Here, when the tumor cells are derived from a solid tumor, for example, peripheral blood, oral mucosal tissue, skin tissue, etc. can be used as a specimen containing nucleic acid derived from non-tumor cells. When the tumor cells are derived from hematopoietic tissue, for example, oral mucosal tissue, skin tissue, etc. can be used as a specimen containing nucleic acid derived from non-tumor cells.
[0018] Specimens can be collected from fresh tissue, fresh-frozen tissue, paraffin-embedded tissue, etc. Sample collection can be performed according to known methods. In the following explanation, when a sample containing nucleic acid derived from tumor cells and a sample containing nucleic acid derived from non-tumor cells are collected from the same subject, the sample containing nucleic acid derived from non-tumor cells and the sample containing nucleic acid derived from tumor cells may be collected at the same time or at different times.
[0019] Genes to be analyzed for nucleic acid sequence are not limited as long as they are present in the human genome. Preferably, the genes are related to tumor onset, prognosis, and therapeutic effects. In the following description, a genetic mutation may be a disease-related mutation or a genetic sequence polymorphism. Genetic "polymorphisms" include SNVs (Single Nucleotide Variants), VNTRs (Variable Nucleotide Tandem Repeats), STRPs (Short Tandem Repeat Polymorphisms), and microsatellite polymorphisms. Genetic mutations may also be fusion gene mutations.
[0020] In the following description, nucleic acid sequence data is not limited as long as it is data reflecting a nucleic acid sequence. Information about genetic mutations is not limited as long as it is information about genetic mutations possessed by the subject from whom the specimen was collected. For example, information about genetic mutations may include at least a label indicating the name of the gene in which the mutation was detected. Preferably, information about genetic mutations may include a label indicating the name of the gene in which the mutation was detected, information about the detected nucleic acid sequence, and / or information about the amino acid sequence resulting from the mutation. Information about genetic mutations may also include locus information of the gene in which the mutation was detected, reference sequence information, and information about the mutant sequence possessed by the subject. Information about genetic mutations is not limited to information about the presence or absence of a mutation, but may also be, for example, information suggesting the possibility of a genetic mutation (e.g., mosaic mutation).
[0021] (First embodiment) FIG. 1 is a schematic diagram of a nucleic acid information transmission and reception system 1 installed in each facility according to the first embodiment. First, using FIG. 1, the schematic configuration of the nucleic acid information transmission and reception system 1 and an overview of the main information flow in the nucleic acid information transmission and reception system 1 will be described. The nucleic acid information transmission and reception system 1 includes a sequencer 2, a storage (memory device) 3, a data transmission device 5, and an analysis system 4, and the analysis system 4 includes a reception device 6 and a nucleic acid sequence analysis device 7. The data transmission device 5, the reception device 6, and the nucleic acid sequence analysis device 7 are connected to one another via a network 11, which is the Internet. A mutation information database 8 is also connected to the network 11.
[0022] The sequencer 2, storage 3, and data transmission device 5 are installed in an analysis requesting facility 10, such as a hospital (medical facility), testing center, or biomedical research institute. The sequencer 2 is a next-generation sequencer (NGS). Hereinafter, the term "sequencer" refers to a next-generation sequencer. The sequencer 2 is a device that reads nucleic acid base sequence information, and examples of such devices include the MiSeq system (manufactured by Illumina, Inc.), the NextSeq550 system (manufactured by Illumina, Inc.), the Ion GeneStudio S5 system (manufactured by Thermo Fisher Scientific, Inc.), and the Ion Torrent Genexus system (manufactured by Thermo Fisher Scientific, Inc.). The sequencer 2 reads the nucleic acid sequences of multiple library samples (e.g., 16 samples) in a single sequencing run. In a single sequencing run, the sequencer 2 reads nucleic acid sequences from each of multiple library samples, including a first library sample and a second library sample, prepared from specimens collected from the same subject, and generates a sequence dataset containing multiple nucleic acid sequences corresponding to each library sample. The sequencer 2 may generate a sequence dataset corresponding to a single subject, or multiple sequence datasets corresponding to multiple subjects, in a single sequencing run. Linking information indicating that the first library sample and the second library sample were prepared from specimens from the same subject is also input to the sequencer 2. A library sample is a sample prepared for nucleic acid sequence reading and is also called a library. The library sample can be prepared, for example, using the Onco Guide NCC Oncopanel Kit (manufactured by Sysmex Corporation). The linking information indicates that multiple library samples were prepared from specimens from the same subject. The linking information includes sample identification information for identifying the first library sample and the second library sample, and subject identification information for identifying the same subject from whom the specimens corresponding to the first library sample and the second library sample were collected.
[0023] Based on the generated sequence dataset and linking information, the sequencer 2 generates sequencing run data including the sequence dataset and linking information, and stores the data in storage 3. The sequence dataset and sequencing run data will be described in detail below with reference to FIG. 4. Storage 3 is a network-attached storage (NAS). The NAS is configured as a storage device that can be directly connected to a network.
[0024] The data transmission device 5 is a computer. The data transmission device 5 includes an input unit 5a, a display unit 5b, a transmission / reception unit 5c, and a control device 5e, and the control device 5e includes a control unit 5f and a storage unit 5g. The input unit 5a is used to input data and is composed of a keyboard and a mouse. The display unit 5b is composed of a liquid crystal panel and displays images. The display unit 5b may be composed of an organic EL panel. The input unit 5a and the display unit 5b may be composed of a touch panel in which a touch sensor and a display are integrated. The transmission / reception unit 5c is an interface for transmitting and receiving data to and from an external device via a network 11 connected to the data transmission device 5, and is composed of, for example, an interface compatible with Ethernet. The control unit 5f is a CPU, and the storage unit 5g is composed of an SSD and a semiconductor memory.
[0025] Data transmission device 5 reads the sequence run data from storage 3 via transceiver 5c, and transmits the sequence run data to reception device 6 via transceiver 5c and network 11.
[0026] The reception device 6 is installed in a request reception facility 20, for example, a server center. The analysis requesting facility 10 and the request reception facility 20 are different facilities. The reception device 6 may be a computer constituting a cloud system. The server center may be a facility of a cloud service provider or a facility of a company that provides nucleic acid sequence analysis services. The reception device 6 is a computer. The reception device 6 has an input unit 6a, a display unit 6b, a transmission / reception unit 6c, and a control device 6e. The control device 6e includes a control unit 6f and a memory unit 6g. The hardware configurations of the input unit 6a, the display unit 6b, the transmission / reception unit 6c, and the control device 6e are similar to those of the input unit 5a, the display unit 5b, the transmission / reception unit 5c, and the control device 5e, respectively. The reception device 6 transmits the sequencing run data to the nucleic acid sequence analysis device 7 via the transmission / reception unit 6c and the network 11.
[0027] The nucleic acid sequence analysis device 7 is installed at a request receiving facility 30, for example, a data analysis facility. The analysis requesting facility 10 and the request receiving facility 30 are different facilities. The request receiving facility 20 and the request receiving facility 30 are different facilities, but may be the same facility. The nucleic acid sequence analysis device 7 may be a computer constituting a cloud system. The data analysis facility may be a facility of a cloud service provider or a facility of a company that provides nucleic acid sequence analysis services. The nucleic acid sequence analysis device 7 is a computer. The nucleic acid sequence analysis device 7 has an input unit 7a, a display unit 7b, a transmission / reception unit 7c, and a control device 7e. The control device 7e includes a control unit 7f and a memory unit 7g. The hardware configurations of the input unit 7a, the display unit 7b, the transmission / reception unit 7c, and the control device 7e are the same as those of the input unit 5a, the display unit 5b, the transmission / reception unit 5c, and the control device 5e, respectively. The nucleic acid sequence analysis device 7 is capable of accessing a mutation information database 8 via a network 11.
[0028] The mutation information database 8 is composed of, for example, an external public sequence information database or a public known mutation information database. The control device 7e of the nucleic acid sequence analysis device 7 compares each nucleic acid sequence data included in the sequencing run data received from the reception device 6 with the reference nucleic acid sequence data stored in the mutation information database 8, and generates genetic mutation information for each nucleic acid sequence data.
[0029] FIG. 2 is a flowchart illustrating the processing executed by the sequencer 2. The processing executed by the sequencer 2 will be described with reference to FIG. 2. First, in step S1, the sequencer 2 receives a sequencing run ID, a case ID, a sample ID, and an index ID, and generates a sample sheet, which is an electronic file. The sequencing run ID is information identifying sequencing run data. The sample sheet includes a case ID, a sample ID, and an index ID. One sample sheet is generated for each sequencing run, i.e., one cartridge. In one sequencing run, i.e., one cartridge, the nucleic acid sequences of multiple library samples (e.g., 16 samples) are read. The multiple library samples are created by pretreating multiple samples (e.g., 16 samples) prepared from tumor tissues and non-tumor tissues (e.g., blood) of multiple subjects (e.g., 8 subjects) with a reagent, and adding different index sequences to the multiple samples. In the first embodiment, each library sample is a sample prepared from DNA. The case ID is information that identifies the subject from whom each library sample was collected. The sample ID is information that identifies each library sample. The index ID is information that identifies the index sequence added to each library sample.
[0030] FIG. 3 is a diagram showing an example of a sample sheet 35. In the example shown in FIG. 3, sample IDs are associated with case IDs. Library samples with the same case ID are samples prepared from specimens of the same subject. The sample IDs on the sample sheet 35 are an example of sample identification information, and the case IDs are an example of subject identification information as well as an example of linking information indicating that multiple library samples were prepared from specimens of the same subject. Each sample ID is further associated with an index ID and an index array. The index array is information indicating the index array added to the library sample.
[0031] For example, a library sample with a sample ID of 1010 was prepared from a specimen of subject A who has a specific disease, and is a sample with an index ID of 001. The index sequence of this library sample is CGGATTGC. By determining that the nucleic acid sequence read by sequencer 2 contains the partial nucleic acid sequence CGGATTGC, the nucleic acid sequence data can be identified as the nucleic acid sequence data of the library sample with a sample ID of 1010. A library sample with a sample ID of 2019 was prepared from a specimen of subject A who has a specific disease, and is a sample with an index ID of 009. The index sequence of this library sample is ACTATGCA. The library sample with a sample ID of 1010 and the library sample with a sample ID of 2019 have the same case ID (A), which indicates that both library samples were prepared from the specimen of the same subject A. Similarly, because the library sample with sample ID 1013 and the library sample with sample ID 2021 have the same case ID (B), it is indicated that both library samples were prepared from specimens of the same different subject B. Furthermore, in the first embodiment, if the sample ID starts with 1, it indicates that the corresponding library sample is derived from tumor cells, and if the sample ID starts with 2, it indicates that the corresponding library sample is derived from non-tumor cells. Therefore, by referring to the data on the sample sheet, it is possible to identify multiple library samples derived from the same subject and information on whether each library sample is derived from tumor cells or non-tumor cells.
[0032] The notation on the sample sheet may be any notation that includes linking information. For example, as shown in FIG. 5, i.e., a diagram representing a modified sample sheet 35', sample sheet 35 may include an additional column for tumor / non-tumor ID to identify whether each library sample is derived from a tumor or non-tumor specimen. In sample sheet 35', T indicates tumor specimen origin, and N indicates non-tumor specimen origin. In this case, the sample ID does not need to indicate whether the corresponding library sample is derived from tumor cells or non-tumor cells, making it easier to assign sample IDs.
[0033] Furthermore, the sample sheet may omit the case ID column, as shown in FIG. 6(A), i.e., a diagram representing a modified sample sheet 35''. In this case, the sample ID may be composed of a pair of two letters, symbols, or numbers. For example, the first letter, symbol, or number may indicate identification information for identifying the subject, and the second letter, symbol, or number may indicate identification information for identifying the tumor sample or non-tumor sample. For example, in the table of sample sheet 35'', the first capital letter "A", "B", "C", or "D" may indicate identification information for the subject from whom the sample was collected, and the second capital letter "T" or "N" may indicate identification information for the tumor sample and the non-tumor sample. In sample sheet 35'', the sample ID indicates that multiple library samples were prepared from the same subject's sample. For example, the library sample with sample ID AT and the library sample with sample ID AN are both indicated to be samples prepared from subject A's sample. Therefore, the sample ID on the sample sheet 35'' is an example of linking information indicating that multiple library samples have been prepared from specimens of the same subject. In the example described above using Figures 3, 5, and 6(A), the linking information is a sample ID or a case ID. Therefore, the linking information indicates that multiple library samples have been prepared from specimens of the same subject, and also identifies the library sample or the subject. In other words, in the example described above using Figures 3, 5, and 6(A), information that identifies the sample or the subject is also used as linking information, and there is no need to separately input the linking information into the sequencer 2.
[0034] FIG. 6(B) shows another modified example of the sample sheet. As shown in FIG. 6(B), sample sheet 35''' does not include a column of case IDs, but includes a column of paired sample IDs, as compared to sample sheet 35. The paired sample ID is a sample ID that identifies a library sample prepared from a specimen of the same subject as the subject from whom the specimen of the library sample identified by the corresponding sample ID was collected. For example, for a library sample with sample ID 1010, since the paired sample ID is 2019, it can be identified that the library sample with sample ID 2019 was prepared from a specimen of the same subject. Therefore, the paired sample ID is an example of linking information. In this example, in step S1 (see FIG. 2), the paired sample ID is input into the sequencer 2 instead of the case ID. In this way, when information that serves as linking information is input into the sequencer 2 separately from information that identifies the sample or subject, it is not necessary to input information that identifies the sample or subject. In sample sheet 35''', paired sample IDs are entered for both the library sample prepared from a tumor specimen and the library sample prepared from a non-tumor specimen, but either one may be omitted. The notation on the sample sheet may be a notation that adds linking information to the known notation used in sequencer 2. The notation on the sample sheet may be any notation that allows for the identification of the corresponding subject for each library sample and the identification of whether the library sample is derived from tumor cells or non-tumor cells.
[0035] Referring back to FIG. 2, the process subsequently executed by the sequencer 2 will now be described. The user of the sequencer 2 dispenses multiple pre-prepared library samples (e.g., 16 samples) into each well of a cartridge, sets the cartridge in the sequencer 2, and instructs the sequencer 2 to begin sequence reading. When the user instructs the sequencer 2 to begin sequence reading, the sequencer 2 reads the nucleic acid sequence for each of the multiple library samples in step S2. In embodiment 1, the sequencer 2 reads a DNA library sample prepared from a tumor specimen and a DNA library sample prepared from a non-tumor specimen for each of multiple subjects. Next, in step S3, the sequencer 2 generates sequencing run data. Then, in the next step S4, the generated sequencing run data is stored in storage 3, and the process ends.
[0036] FIG. 4 is a diagram illustrating an example of sequence run data 50 created by sequencer 2. Sequence run data 50 is an electronic folder storing electronic files, and the folder name is assigned the sequence run ID 39 received in step S1. The sequence run data 50 stores a sample sheet 35 and nucleic acid sequence data 37 read from each library sample in step S2. Sequencer 2 compares index sequences 38 included in each nucleic acid sequence data 37 with the index sequences included in sample sheet 35, and associates sample IDs having identical index sequences with nucleic acid sequence data 37. Furthermore, multiple nucleic acid sequence data 37 corresponding to multiple sample IDs corresponding to the same case ID constitute a single sequence dataset. In the example of Figure 4, two pieces of nucleic acid sequence data 37 corresponding to the library sample with sample ID 1010 and the library sample with sample ID 2019 corresponding to case A constitute sequence dataset 37-1, and two pieces of nucleic acid sequence data 37 corresponding to the library sample with sample ID 1013 and the library sample with sample ID 2021 corresponding to case B constitute sequence dataset 37-2. The number of sequence datasets included in one sequencing run, i.e., one piece of sequencing run data 50, is not particularly limited, but from the perspective of reading nucleic acid sequences from more subjects in one sequencing run, it is preferably 5 or more.
[0037] FIG. 7 is a flowchart illustrating the processing executed by each control unit of the data transmitting device 5, the receiving device 6, and the nucleic acid sequence analyzing device 7. Referring to FIG. 7, the processing executed by the control unit 5f of the data transmitting device 5 will be described first. When the control unit 5f receives an instruction to analyze the sequencing run data stored in the storage 3 from the user of the data transmitting device 5, it transmits analysis request information to the receiving device 6 in step S20. The analysis request information is information input by the user operating the input unit 5a and includes case information on the subject, information on the type of genetic panel testing, and information on the requesting facility 10. The case information on the subject includes a case ID. The information on the requesting facility 10 includes the name of the requesting facility and identification information for identifying the requesting facility. The analysis request information may include at least one of the case information on the subject, information on the type of genetic panel testing, and information on the requesting facility 10. Furthermore, for example, if there is an agreement between the requesting facility 10 and the request-receiving facility 20 that the transmission of sequencing run data or case information is considered an analysis request, the processing of step S20 may be omitted.
[0038] In step S21, the control unit 5f reads the sequencing run data from the storage 3 and transmits it to the reception device 6. In step S22, the control unit 5f displays a case registration screen on the display unit 5b and receives the registration of case information. The case information includes input information indicating that the library sample derived from the tumor specimen and the library sample derived from the non-tumor specimen were prepared from the same subject. Specifically, the input information includes the sample ID of the library sample derived from the tumor specimen, the sample ID of the library sample derived from the non-tumor specimen, and a case ID corresponding to both sample IDs.
[0039] 8 is a diagram illustrating an example of a case registration screen 40 displayed on the display unit 5b. As shown in FIG. 8, the case registration screen 40 displays a registration section 40a for registering a sequence run ID, a registration section 40b for registering a case ID, a registration section 40c for registering an index ID, a registration section 40d for registering a sample ID, and a registration section 40e for registering an index sequence for a library sample derived from a normal (non-tumor) specimen, a registration section 40f for registering an index ID, a registration section 40g for registering a sample ID, and a registration section 40h for registering an index sequence for a library sample derived from a tumor specimen, and a registration button 40i. The registration sections 40a, 40c, and 40f are configured in a pull-down list format. Expanding the pull-down list displays a list of sequence run IDs and index IDs included in the sequence run data retrieved from the storage 3 by the control unit 5f and not marked with a registered flag (described later). Registration units 40b, 40d, 40e, 40g, and 40h are configured so that a user can operate a keyboard to input numbers, letters, or symbols. The user of data transmission device 5 operates input unit 5a to input information for each subject into registration units 40a to 40h, and when registration for one subject is complete, selects registration button 40i. The user of data transmission device 5 repeats the operation of inputting information for each subject into registration units 40a to 40h and selecting registration button 40i until input for all case IDs included in one sequence run data is completed.
[0040] Registration units 40a-40h may also be configured in a pull-down list format, except for registration units 40a, 40c, and 40f, and registration units 40a, 40c, and 40f may be configured to input numerical values, etc. Furthermore, registration units 40e and 40h may be configured such that when an index ID is input to registration unit 40c or 40f, the corresponding index sequence is read from the sequence run data and displayed on registration unit 40e or 40h. Any screen can be used as the case registration screen as long as it allows for the registration of information that can identify the same subject for each library sample of normal specimens (non-tumor specimens) and tumor specimens.
[0041] Referring again to FIG. 7, the next process executed by the control unit 5f of the data transmission device 5 will be described. When the process of step S22 is completed, i.e., when input for all case IDs included in one sequence run data is completed and the registration button 40i is selected, the control unit 5f transmits the case information input to the registration units 40a to 40h in step S22 to the reception device 6 in step S23. As described above, the case information includes input information indicating that the tumor specimen-derived library sample and the non-tumor specimen-derived library sample were prepared from the same subject. In the first embodiment, the input information includes the normal specimen sample ID input via the registration unit 40d, the tumor specimen sample ID input via the registration unit 40g, and the case ID input via the registration unit 40b. The sample ID is information identifying each library sample, and the case ID is information identifying the subject. In step S24, the control unit 5f adds a flag indicating that the sequence run ID corresponding to the sequence run data transmitted to the reception device 6 in step S21 has been registered.
[0042] Next, the processing executed by the control unit 6f of the reception device 6 will be described. When analysis request information is transmitted from the data transmission device 5, the control unit 6f receives the analysis request information and stores it in the memory unit 6g in step S30. When sequence run data is transmitted from the data transmission device 5, the control unit 6f receives the sequence run data and stores it in the memory unit 6g in step S31. When case information is transmitted from the data transmission device 5, the control unit 6f receives the case information and stores it in the memory unit 6g in step S32. In the subsequent step S33, the control unit 6f verifies consistency and determines whether the linking information included in the sequence run data stored in step S31 is consistent with the input information included in the case information stored in step S32.
[0043] FIG. 9 is a flowchart illustrating the details of the consistency verification process executed by the control unit 6f in step S33. In step S51, the control unit 6f reads out a sequence run ID, a case ID, a normal sample ID, and a tumor sample ID from the stored case information. The sequence run ID is information input via the registration unit 40a of the case registration screen 40 (see FIG. 8), the case ID is information input via the registration unit 40b, the normal sample ID is information input via the registration unit 40d, and the tumor sample ID is information input via the registration unit 40g. In step S52, the control unit 6f reads out information written on a sample sheet of sequence run data to which the same sequence run ID as the read-out sequence run ID is assigned. Then, in step S53, the control unit 6f determines whether the combination of the case ID, normal sample ID, and tumor sample ID read out from the case information in step S51 exists in the sample sheet.
[0044] If the control unit 6f makes a negative judgment in step S53 (if "No"), it sends an error notification to the data transmission device 5 in step S54, indicating that the linking information in the sequence run data and the case information do not match, and terminates the process without executing the processes from step S34 onwards (see FIG. 7). Upon receiving the error notification, the control unit 5f of the data transmission device 5 outputs error information indicating that the linking information in the sequence run data and the case information do not match to the display unit 5b. This output allows the user of the data transmission device 5 to recognize that there is an error in at least one of the information on the sample sheet and the manually entered case information. On the other hand, if the control unit 6f of the receiving device 6 makes a positive judgment in step S53 (if "Yes"), it returns the process to step S34 (see FIG. 7).
[0045] According to the first embodiment, in step S53, it is determined whether the linking information of the sequencing run data and the case information are consistent. If the two pieces of information do not match, the process ends without proceeding to the next step. Therefore, for each subject, the nucleic acid sequence data derived from the tumor sample and the nucleic acid sequence data derived from the non-tumor sample can be accurately linked to that subject. Therefore, even in the case of matched pair testing, in which nucleic acid sequence analysis is performed on multiple nucleic acid sequence data derived from the same subject, it is possible to reliably prevent incorrect analysis due to mismatching of nucleic acid sequence data.
[0046] Referring again to FIG. 7, the process subsequently executed by the control unit 6f of the reception device 6 will be described. If a positive determination is made in step S53 (see FIG. 9), the control unit 6f transmits the sequencing run data stored in step S31 to the nucleic acid sequence analysis device 7 in step S34. Next, the process executed by the control unit 7f of the nucleic acid sequence analysis device 7 will be described. In step S40, the control unit 7f receives the analysis request information transmitted from the reception device 6 and stores it in the memory unit 7g. In step S41, the control unit 7f receives the sequencing run data transmitted from the reception device 6 and stores it in the memory unit 7g. In step S42, the control unit 7f reads out one sequence dataset from the stored sequencing run data. As described above, since the sequence dataset includes multiple nucleic acid sequence data 37 corresponding to the same case ID, the control unit 7f can extract multiple nucleic acid sequence data 37 corresponding to the same case ID as one sequence dataset using the case ID, which is linking information, as a search key.
[0047] In step S43, the control unit 7f analyzes the presence or absence of mutations for each nucleic acid sequence data of the sequence data set extracted in step S42, using the information on the nucleic acid sequences of tumor cells in the mutation information database 8. In step S44, the control unit 7f creates an analysis result report based on the presence or absence of mutations. In step S45, the control unit 7f transmits the analysis result report to the reception device 6. The processing of step S43 will be described in detail below with reference to Figures 10 to 14. The analysis result report will be described in detail below with reference to Figure 15. In step S46, the control unit 7f determines whether all sequence data sets included in the sequencing run data stored in step S41 have been extracted. If all sequence data sets have been extracted (if "Yes"), the control unit 7f terminates the processing. If not all sequence data sets have been extracted (if "No"), the control unit 7f returns the processing to step S42 and executes the processing of steps S42 to S46 again.
[0048] Meanwhile, the control unit 6f of the reception device 6 receives the analysis result report in step S35, transmits the analysis result report to the data transmission device 5 in step S36, and ends the process. The control unit 5f of the data transmission device 5 receives the analysis result report in step S25, stores it in the memory unit 5g, and ends the process. This allows the doctor in charge of the subject to display and view the analysis report stored in the memory unit 5g on the display unit 5b at any time.
[0049] Next, the processing of step S43 by the control unit 7f will be described in detail with reference to Fig. 10. Fig. 10 is a flowchart showing an example of the processing procedure when the control unit 7f of the nucleic acid sequence analysis device 7 determines a nucleic acid sequence. In step S61, the control unit 7f acquires one piece of nucleic acid sequence data 37 (hereinafter referred to as acquired sequence) from the sequence data set extracted in step S42. The control unit 7f also downloads a reference sequence from the mutation information database 8 and stores it in the memory unit 7g.
[0050] A reference sequence is a sequence to which an obtained sequence is mapped in order to determine which region of a gene the obtained sequence corresponds to and which mutation in the gene the obtained sequence corresponds to. For each gene to be analyzed, (1) a wild-type reference sequence, which is a partial or entire sequence of a wild-type exon, can be used as a reference sequence. Furthermore, (2) a single mutant reference sequence, in which a rearranged sequence containing a known polymorphism or mutation is concatenated with a wild-type exon sequence, can be used as a reference sequence. A single mutant reference sequence is a sequence generated for each gene to be analyzed by concatenating two or more rearranged sequences related to the gene to be analyzed. The single mutant reference sequence is used as a mutant reference sequence containing a rearranged sequence when mapping an obtained sequence. Note that, instead of a single mutant reference sequence in which two or more rearranged sequences are concatenated, two or more unconjugated rearranged sequences may be used as the mutant reference sequence.
[0051] FIG. 11 is a conceptual diagram outlining a method for generating a single mutation reference sequence, illustrating an example of a method for generating a mutation reference sequence using publicly known mutation information downloaded from an external mutation information database 8. FIG. 11 illustrates an example in which information about a mutation "C797S" in the gene "EGFR" at chromosome location "xxxx" is newly uploaded to the external mutation information database 8 by research institution P and stored in the mutation information database 8. The information about the mutation "C797S" in the gene named "EGFR" at chromosome location "xxxx" uploaded by research institution P is associated with a mutation ID "yyyy" and an upload date "zz year z month z day," and registered as publicly known mutation information in the external mutation information database 8. The mutation illustrated here as newly uploaded information is a mutation in which the 797th amino acid residue in the protein "EGFR," which is a gene product transcribed and translated from the gene "EGFR," is substituted from cysteine to serine. The external mutation information database 8 is not limited to such mutations, and may also collect and store information on polymorphisms, mutations, methylation, and the like.
[0052] The mutation information database 8 is an external public sequence information database, a public known mutation information database, or the like. Examples of public sequence information databases include NCBI RefSeq (webpage, www.ncbi.nlm.nih.gov / refseq / ), NCBI GenBank (webpage, www.ncbi.nlm.nih.gov / genbank / ), and UCSC Genome Browser. Examples of public known mutation information databases include the COSMIC database (webpage, www.sanger.ac.uk / genetics / CGP / cosmic / ), the ClinVar database (webpage, www.ncbi.nlm.nih.gov / clinvar / ), and dbSNP (webpage, www.ncbi.nlm.nih.gov / SNP / ). The mutation information database 8 may also be a public known mutation information database that includes frequency information for publicly known mutations by race or animal species. Publicly available databases of known mutation information that contain such information include HapMap Genome Browser release #28, Human Genetic Variation Browser (webpage: www.genome.med.kyoto-u.ac.jp / SnpDB / index.html), and 1000 Genomes (webpage: www.1000genomes.org / ).
[0053] Referring again to FIG. 10, the process subsequently executed by the control unit 7f will be described. In step S62, the control unit 7f compares the acquired sequence with the reference sequence to identify positions on the reference sequence where the match rate between the acquired sequence and the reference sequence satisfies a predetermined criterion. The comparison is performed by mapping the acquired sequence to multiple positions on the reference sequence. The match rate is the ratio of the number of bases that match between the acquired sequence and the reference sequence to the number of bases contained in the acquired sequence. The position on the reference sequence is identified by calculating the match rate between the acquired sequence and the reference sequence at each mapped position and identifying the position where the calculated match rate exceeds a predetermined threshold.
[0054] In step S63, the control unit 7f determines whether multiple positions on the reference sequence have been identified, i.e., whether the match rate at multiple positions on the reference sequence meets a predetermined criterion. If the acquired sequence matches a single position on the reference sequence (if "No"), the control unit 7f determines in step S65 whether positions on the reference sequence have been identified for all acquired sequences included in one sequence data set extracted in step S42. If position identification has been completed for all acquired sequences (if "Yes"), the control unit 7f proceeds to step S73 (see FIG. 12). On the other hand, if position identification has not been completed for all acquired sequences (if "No"), the control unit 7f returns the process to step S62 and continues processing.
[0055] If there is a match at multiple positions on the reference sequence in step S63 (if "Yes"), the control unit 7f identifies the position with the highest match rate among the multiple positions as the position on the reference sequence of the acquired sequence in step S64, and proceeds to step S65.
[0056] [Mutation detection] <Detection of somatic mutations> Next, an example of the process by which the control unit 7f detects somatic mutations will be described with reference to Fig. 12. Fig. 12 is a flowchart showing the process by which the control unit 7f of the nucleic acid sequence analyzer 7 detects somatic mutations.
[0057] In step S73, the control unit 7f determines whether or not there is a mismatch between the tumor sequence and the reference sequence at the position on the reference sequence identified in step S62 or S64 for the nucleic acid sequence data of the library sample derived from the tumor specimen (hereinafter referred to as the tumor sequence) among the multiple nucleic acid sequence data included in one sequence dataset acquired in step S61 (see FIG. 10). If there is a mismatch ("Yes"), the control unit 7f proceeds to step S74, and if there is no mismatch ("No"), the control unit 7f proceeds to step S83 (see FIG. 13). In step S74, the control unit 7f determines whether or not there is a mismatch between the normal sequence and the reference sequence at the position on the reference sequence identified in step S62 or S64 for the nucleic acid sequence data of the library sample derived from the non-tumor specimen (hereinafter referred to as the normal sequence) among the nucleic acid sequence data included in the sequence dataset referenced in step S73. If there is no mismatch (if "Yes"), the control unit 7f advances the process to step S75, and if there is a mismatch (if "No"), the control unit 7f advances the process to step S83.
[0058] In step S75, the control unit 7f determines that the mismatched base detected in step S73, i.e., the mutation, is a somatic mutation. In step S76, the control unit 7f searches the mutation information database stored in the mutation information database 8 based on the detected somatic mutation.
[0059] The mutation information stored in the mutation information database of the mutation information database 8 includes a mutation identifier (mutation ID), a gene name, mutation location information (e.g., "CHROM" and "POS"), "REF", "ALT", and "Annotation". The mutation ID is an identifier for identifying a mutation. Among the mutation location information, "CHROM" indicates the chromosome number, and "POS" indicates the position on the chromosome. "REF" indicates a base in the wild type, and "ALT" indicates a base after mutation. "Annotation" indicates information about the mutation. "Annotation" may be information indicating an amino acid mutation, such as "EGFR C2573G" or "EGFR L858R". For example, "EGFR C2573G" indicates a mutation in which the cysteine at residue 2573 of the protein "EGFR" is replaced with glycine.
[0060] In step S77, the control unit 7f assigns mutation information such as gene name, annotation, etc. to the detected somatic mutation based on the search results of step S76. Note that in embodiment 1, the processes of steps S76 and S77 can be omitted.
[0061] <Detection of germline mutations> Next, an example of a process performed by the control unit 7f to detect germ cell mutations will be described with reference to FIG. 13. FIG. 13 is a flowchart showing a process performed by the control unit 7f of the nucleic acid sequence analysis device 7 to detect germ cell mutations. In step S83, the control unit 7f determines whether or not there is a mismatch between the normal sequence and the reference sequence at the position on the reference sequence identified in step S62 or S64 for nucleic acid sequence data (normal sequence) of a library sample derived from a non-tumor specimen, among multiple nucleic acid sequence data included in one sequence data set acquired in step S61 (see FIG. 10). If there is a mismatch ("Yes"), the control unit 7f proceeds to step S84; if there is no mismatch ("No"), the control unit 7f proceeds to step S44 (see FIG. 7).
[0062] In step S84, the control unit 7f determines that the mismatched base detected in step S83, i.e., the mutation, is a germline mutation. In step S85, the control unit 7f searches the mutation information database stored in the mutation information database 8 based on the detected germline mutation. In step S86, the control unit 7f assigns mutation information such as a gene name and annotation to the detected mutation based on the search result of step S85. Note that in embodiment 1, the processes of steps S85 and S86 can be omitted.
[0063] Figure 14(A) shows an example of a nucleic acid sequence having a somatic mutation, and Figure 14(B) shows an example of a nucleic acid sequence having a germline mutation. Referring to Figure 14(A), the sequence data derived from the non-tumor sample (normal sequence) does not have a mismatch with the reference sequence, but the sequence data derived from the tumor sample (tumor sequence) has a base that mismatches with the reference sequence (G in the reference sequence, C in the tumor sequence), i.e., a mutation. In this case, the control unit 7f determines the mutation to be a somatic mutation in step S75 (see Figure 12).
[0064] 14(B), the sequence data (normal sequence) derived from the non-tumor sample contains a base that does not match the reference sequence (A in the reference sequence, while T in the normal sequence), i.e., a mutation. In this case, the control unit 7f determines the mutation to be a germline mutation in step S84 (see FIG. 13).
[0065] [Analysis Result Report] Next, an example of the analysis report created in step S44 (see FIG. 7) will be described. FIG. 15 is a diagram showing an example of the analysis report R1. As shown in FIG. 15, the analysis report R1 includes a summary report area S (hereinafter also referred to as the "summary report area S") for displaying a summary of the analysis results, and a detailed report area D (hereinafter also referred to as the "detailed report area D") for displaying details of the analysis results. The summary report area S includes an area S1 (hereinafter also referred to as the "attribute information area S1") showing attribute information that indicates information about the subject and the test details, and an area S2 (hereinafter also referred to as the "gene mutation list area S2") showing a list of all detected gene mutations. The detailed report area D also includes an area D1 (hereinafter also referred to as the "gene mutation information area D1") showing genes in which somatic mutations have been detected and detailed information about those mutations, and an area D2 (hereinafter also referred to as the "germ cell mutation information area D2") showing genes in which germ cell mutations have been detected and detailed information about those mutations.
[0066] The attribute information region S1 displays information for identifying the patient, such as a patient identifier (patient ID), the name of the attending physician, and the name of the medical institution, as well as information indicating test items such as a gene panel, based on the information stored in the memory unit 7g in step S40. The gene mutation list region S2 displays all detected gene mutations, regardless of whether they are somatic or germline mutations. In the example of the gene mutation list region S2, EGFR, BRAF, and BRCA1 indicate gene names, and L585R, V600E, and K1183R indicate the mutation site in each gene and the amino acid substitution resulting from the mutation. In other words, EGFR_L585R indicates that the 585th codon of the EGFR gene has mutated from a nucleic acid sequence encoding leucine (L) to a nucleic acid sequence encoding arginine (R). The various information displayed in the gene mutation list region S2 uses information acquired by the control unit 7f in steps S77 and S86.
[0067] [Effects of the first embodiment] According to the first embodiment, in a matched pair test in which nucleic acid sequence data of a tumor sample and nucleic acid sequence data of a non-tumor sample from a single subject are analyzed as a set, the reception device 6 receives, via the network 11, from the data transmission device 5, sequence data sets containing multiple nucleic acid sequence data obtained using the sequencer 2 corresponding to multiple library samples, including a first library sample and a second library sample prepared from the same subject's sample, and sequence run data containing linking information indicating that the first library sample and the second library sample were prepared from the same subject's sample, and transmits the sequence run data to the nucleic acid sequence analyzer 7 that analyzes the nucleic acid sequences. Therefore, even if the analysis requesting facility 10 that operates the sequencer 2 is a facility different from the requested facility 30 where the nucleic acid sequence analyzer 7 is installed, the nucleic acid sequence analyzer 7 can accurately and quickly extract from the sequence data sets the correct combination of nucleic acid sequence data corresponding to the multiple library samples from the same subject. This allows the requested facility 30 to analyze the correct combination of multiple nucleic acid sequence data, enabling accurate and fast analysis using multiple nucleic acid sequence data from the same subject. In addition, in the first embodiment, by combining and comprehensively analyzing information on somatic mutations obtained by analyzing nucleic acid sequence data of tumor samples and information on germline mutations obtained by analyzing nucleic acid sequence data of non-tumor samples, analysis can be performed based on more information, making it easier to identify a treatment method appropriate for the subject.
[0068] (Second embodiment) In the first embodiment, a case was described in which a sequence dataset includes a first library sample of DNA derived from tumor cells collected from one subject and a second library sample of DNA derived from non-tumor cells collected from the same subject. In the second embodiment, a sequence dataset includes a first library sample of DNA derived from tumor cells collected from one subject and a second library sample of RNA derived from tumor cells collected from the same subject.
[0069] RNA may be generated due to a DNA fusion gene mutation. Therefore, by identifying the nucleic acid sequence of the RNA, it may be possible to identify the DNA fusion gene mutation. In the second embodiment, information on somatic mutations other than the fusion gene mutation can be obtained from the analysis results of the first sequence data corresponding to the first library sample, and information on the fusion gene mutation can be obtained from the analysis results of the second sequence data corresponding to the second library sample.
[0070] The schematic configuration of the nucleic acid information transmission and reception system 1 of the second embodiment is the same as the configuration shown in FIG. 1. The outline of the processing executed by each control unit of the data transmission device 5, the reception device 6, and the nucleic acid sequence analysis device 7 is the same as the processing shown in FIG. 7. FIG. 16 is a flowchart explaining the processing executed by the sequencer 2 in the second embodiment, and FIG. 17 is a diagram showing an example of a sample sheet 135 that can be used in the second embodiment. FIG. 18 is a diagram showing an example of a case registration screen displayed on the display unit 5b of the data transmission device 5 in the second embodiment. FIG. 19 is a flowchart explaining the details of the processing executed by the control unit 6f in step S33 (see FIG. 7) in the second embodiment.
[0071] The processing executed by the sequencer 2 will be described with reference to FIG. 16. First, in step S1′, the sequencer 2 receives a sequencing run ID, a case ID, a sample ID, and an index ID, and generates a sample sheet, which is an electronic file. As in the first embodiment, the sequencing run ID is information for identifying sequencing run data, and the sample sheet includes the case ID, the sample ID, and the index ID. One sample sheet is generated for each sequencing run, i.e., one cartridge. In one sequencing run, i.e., one cartridge, the nucleic acid sequences of multiple library samples (e.g., 16 samples) are read. The multiple library samples are created by pretreating multiple samples (e.g., 16 samples) prepared from tumor tissue DNA and tumor tissue RNA of multiple subjects (e.g., 8 subjects) with a reagent, and adding different index sequences to the multiple samples.
[0072] FIG. 17 is a diagram showing an example of a sample sheet 135. In the example shown in FIG. 17, sample IDs are associated with case IDs. Library samples with the same case ID are samples prepared from specimens of the same subject, and the case ID on the sample sheet 135 is an example of linking information indicating that multiple library samples were prepared from specimens of the same subject. Each sample ID is further associated with an index ID and an index sequence. The index sequence is information indicating the index sequence added to the library sample.
[0073] For example, a library sample with a sample ID of 1010 was prepared from a specimen of subject A with a specific disease and has an index ID of 001. The index sequence of this library sample is CGGATTGC. A library sample with a sample ID of 3020 was prepared from a specimen of subject A with a specific disease and has an index ID of 009. The index sequence of this library sample is ACTATGCA. The library sample with a sample ID of 1010 and the library sample with a sample ID of 3020 have the same case ID (A), which indicates that both library samples were prepared from the specimen of the same subject A. Similarly, the library sample with a sample ID of 1013 and the library sample with a sample ID of 3024 have the same case ID (B), which indicates that both library samples were prepared from the specimen of the same subject B. Furthermore, in the second embodiment, if the sample ID starts with 1, it indicates that the corresponding library sample is derived from the DNA of tumor cells, and if the sample ID starts with 3, it indicates that the corresponding library sample is derived from the RNA of tumor cells. Therefore, by referring to the data on the sample sheet, it is possible to identify multiple library samples derived from the same subject and information on whether each library sample is derived from DNA or RNA.
[0074] Referring again to FIG. 16, the next process executed by the sequencer 2 will be described. The user of the sequencer 2 dispenses multiple library samples (e.g., 16 samples) prepared in advance into each well of a cartridge, inserts the cartridge into the sequencer 2, and instructs the sequencer 2 to start sequence reading. Upon receiving the user's instruction to start sequence reading, the sequencer 2 reads the nucleic acid sequence of each of the multiple library samples in step S2'. In embodiment 2, the sequencer 2 reads a DNA library sample prepared from a tumor specimen and an RNA library sample prepared from the tumor specimen for each of multiple subjects. Next, in step S3', the sequencer 2 generates sequence run data. The sequence run data is data in which sample sheet 35 of the sequence run data 50 shown in FIG. 4 has been replaced with sample sheet 135. Then, in the next step S4', the generated sequence run data is stored in storage 3, and the process ends.
[0075] The notation on the sample sheet may be any notation that allows for the identification of the corresponding subject for each library sample and the identification of whether the library sample is derived from DNA or RNA.
[0076] As shown in Figure 18, in the second embodiment, the case registration screen displayed on the display unit 5b by the control unit 5f of the data transmission device 5 can be a case registration screen 140 in which, compared to the case registration screen 40 of the first embodiment (see Figure 8), the wording "normal sample" is changed to "DNA sample" and the wording "tumor sample" is changed to "RNA sample."
[0077] 18, the case registration screen 140 displays a registration section 40a for registering a sequence run ID, a registration section 40b for registering a case ID, a registration section 40c for registering an index ID, a registration section 40d for registering a sample ID, and a registration section 40e for registering an index sequence for a library sample derived from DNA of a tumor specimen, a registration section 40f for registering an index ID, a registration section 40g for registering a sample ID, and a registration section 40h for registering an index sequence for a library sample derived from RNA of a tumor specimen, and a registration button 40i. Registration sections 40a, 40c, and 40f are configured in a pull-down list format. Expanding the pull-down list displays a list of sequence run IDs and index IDs included in the sequence run data retrieved from storage 3 by control unit 5f that does not have the aforementioned registered flag attached. Registration units 40b, 40d, 40e, 40g, and 40h are configured so that a user can operate a keyboard to input numbers, letters, or symbols. The user of data transmission device 5 operates input unit 5a to input information for each subject into registration units 40a to 40h, and when registration for one subject is complete, selects registration button 40i. The user of data transmission device 5 repeats the operation of inputting information for each subject into registration units 40a to 40h and selecting registration button 40i until input for all case IDs included in one sequence run data is completed.
[0078] Registration units 40a-40h may be configured in a pull-down list format other than registration units 40a, 40c, and 40f, and registration units 40a, 40c, and 40f may be configured to input numerical values, etc. Furthermore, registration units 40e and 40h may be configured such that when an index ID is input to registration unit 40c or 40f, the corresponding index sequence is read from the sequencing run data and displayed on registration unit 40e or 40h. Any screen can be used as the case registration screen as long as it allows for the registration of information that can identify the same subject for each library sample of tumor specimen DNA and tumor specimen RNA.
[0079] FIG. 19 is a flowchart illustrating the details of the process executed by the control unit 6f in step S33 (see FIG. 7). In step S51′, the control unit 6f reads out a sequence run ID, a case ID, a DNA sample ID, and an RNA sample ID from the stored case information. The sequence run ID is information input via the registration unit 40a of the case registration screen 140 (see FIG. 18), the case ID is information input via the registration unit 40b, the DNA sample ID is information input via the registration unit 40d, and the RNA sample ID is information input via the registration unit 40g. In step S52′, the control unit 6f reads out information written on the sample sheet of the sequence run data to which the same sequence run ID as the read-out sequence run ID is assigned. Then, in step S53′, the control unit 6f determines whether the combination of the case ID, DNA sample ID, and RNA sample ID read out from the case information in step S51′ exists in the sample sheet.
[0080] If the control unit 6f makes a negative judgment ("No") in step S53', in step S54' it sends an error notification to the data transmission device 5 indicating that the linking information of the sequence run data does not match the case information, and terminates the processing without performing the processing from step S34 onwards (see Figure 7).
[0081] Upon receiving the error notification, the control unit 5f of the data transmission device 5 outputs error information to the display unit 5b, indicating that the linking information of the sequence run data and the case information do not match. This output allows the user of the data transmission device 5 to recognize that there is an error in at least one of the sample sheet information and the manually entered case information. On the other hand, if the control unit 6f of the receiving device 6 makes a positive determination ("Yes") in step S53', it returns the process to step S34 (see FIG. 7).
[0082] According to the second embodiment, in step S53', it is determined whether the linking information of the sequence run data and the case information match. If the two pieces of information do not match, the process ends without proceeding to the next step. Therefore, for each subject, the nucleic acid sequence data derived from the DNA sample and the nucleic acid sequence data derived from the RNA sample can be accurately linked to that subject. Therefore, even in the case of matched pair testing, in which nucleic acid sequence analysis is performed on multiple nucleic acid sequence data derived from the same subject, it is possible to reliably prevent incorrect analysis due to mismatched nucleic acid sequence data.
[0083] [Effects of the second embodiment] According to the second embodiment, in a matched pair test in which nucleic acid sequence data of DNA derived from a tumor specimen of a single subject and nucleic acid sequence data of RNA derived from the tumor specimen are analyzed as a set, the reception device 6 receives, via the network 11, from the data transmission device 5, sequence data sets containing multiple nucleic acid sequence data obtained using the sequencer 2 corresponding to multiple library samples, including a first library sample and a second library sample prepared from the same subject's specimen, and sequence run data containing linking information indicating that the first library sample and the second library sample were prepared from the same subject's specimen, and transmits the sequence run data to the nucleic acid sequence analyzer 7 that analyzes the nucleic acid sequences. Therefore, even if the analysis requesting facility 10 that operates the sequencer 2 is a facility different from the requested facility 30 where the nucleic acid sequence analyzer 7 is installed, the nucleic acid sequence analyzer 7 can accurately and quickly extract from the sequence data sets the correct combination of nucleic acid sequence data corresponding to the multiple library samples of the same subject. This allows the requested facility 30 to analyze the correct combination of multiple nucleic acid sequence data, enabling accurate and fast analysis using multiple nucleic acid sequence data of the same subject. In addition, in the second embodiment, by combining and comprehensively analyzing information on somatic mutations other than fusion gene mutations obtained by analyzing nucleic acid sequence data of the DNA of the tumor specimen and information on fusion gene mutations obtained by analyzing nucleic acid sequence data of the RNA of the tumor specimen, analysis can be performed based on more information, making it easier to identify a treatment method appropriate for the subject.
[0084] (Third embodiment) In the first and second embodiments, a case where two library samples derived from the same subject exist in a sequence dataset is described. In the third embodiment, a sequence dataset contains three library samples derived from the same subject. The three library samples are a library sample derived from DNA of a tumor specimen, a library sample derived from RNA of a tumor specimen, and a library sample derived from DNA of a non-tumor specimen, respectively.
[0085] The schematic configuration diagram of the nucleic acid information transmission and reception system 1 of the third embodiment is the same as the diagram shown in FIG. 1. Furthermore, the outline of the processing executed by each control unit of the data transmission device 5, the reception device 6, and the nucleic acid sequence analysis device 7 is the same as the processing shown in FIG. 7. FIG. 20 is a flowchart explaining the processing executed by the sequencer 2 in the third embodiment, and FIG. 21 is a diagram showing an example of a sample sheet 235 that can be employed in the third embodiment. Furthermore, FIG. 22 is a diagram showing an example of a case registration screen displayed on the display unit 5b of the data transmission device 5 in the third embodiment. Furthermore, FIG. 23 is a flowchart explaining the details of the processing executed by the control unit 6f in step S33 (see FIG. 7) in the third embodiment.
[0086] The processing executed by the sequencer 2 will be described with reference to FIG. 20. First, in step S1'', the sequencer 2 receives a sequencing run ID, a case ID, a sample ID, and an index ID, and generates a sample sheet, which is an electronic file. As in the first and second embodiments, the sequencing run ID is information for identifying sequencing run data, and the sample sheet includes the case ID, the sample ID, and the index ID. One sample sheet is generated for each sequencing run, i.e., one cartridge. In one sequencing run, i.e., one cartridge, the nucleic acid sequences of multiple library samples (e.g., 15 samples) are read. The multiple library samples are created by pretreating multiple samples (e.g., 15 samples) prepared from tumor tissue DNA, tumor tissue RNA, and non-tumor tissue DNA of multiple subjects (e.g., five subjects) with a reagent, and adding different index sequences to the multiple samples.
[0087] FIG. 21 is a diagram showing an example of a sample sheet 235. In the example shown in FIG. 21, sample IDs are associated with case IDs. Library samples with the same case ID are samples prepared from specimens of the same subject, and the case ID in sample sheet 235 is an example of linking information indicating that multiple library samples were prepared from specimens of the same subject. Each sample ID is further associated with an index ID and an index sequence. The index sequence is information indicating the index sequence added to the library sample.
[0088] For example, a library sample with a sample ID of 1010 was prepared from a specimen of subject A with a specific disease and has an index ID of 001. The index sequence of this library sample is CGGATTGC. A library sample with a sample ID of 2019 was prepared from a specimen of subject A with a specific disease and has an index ID of 006. The index sequence of this library sample is ACTATGCA. A library sample with a sample ID of 3020 was prepared from a specimen of subject A with a specific disease and has an index ID of 011. The library sample with a sample ID of 1010, the library sample with a sample ID of 2019, and the library sample with a sample ID of 3020 all have the same case ID (A), which indicates that each library sample was prepared from the specimen of the same subject A. Similarly, the library sample with sample ID 1013, the library sample with sample ID 2021, and the library sample with sample ID 3024 have the same case ID (B), which indicates that each library sample was prepared from the specimen of the same subject B. Furthermore, in the third embodiment, if the sample ID starts with 1, it indicates that the corresponding library sample is derived from the DNA of tumor cells, if the sample ID starts with 2, it indicates that the corresponding library sample is derived from the DNA of non-tumor cells, and if the sample ID starts with 3, it indicates that the corresponding library sample is derived from the RNA of tumor cells. Therefore, by referring to the data in the sample sheet, it is possible to identify multiple library samples derived from the same subject and information on whether each library sample is derived from DNA, RNA, or non-tumor.
[0089] Referring again to FIG. 20, the process subsequently executed by the sequencer 2 will be described. The user of the sequencer 2 dispenses multiple library samples (e.g., 15 samples) prepared in advance into each well of a cartridge, inserts the cartridge into the sequencer 2, and instructs the sequencer 2 to start sequence reading. When the user instructs the sequencer 2 to start sequence reading, the sequencer 2 reads the nucleic acid sequence of each of the multiple library samples in step S2''. In embodiment 3, the sequencer 2 reads a DNA library sample prepared from a tumor specimen, an RNA library sample prepared from a tumor specimen, and a DNA library sample prepared from a non-tumor specimen for each of multiple subjects. Next, in step S3'', the sequencer 2 generates sequence run data. The sequence run data is data in which sample sheet 35 of the sequence run data 50 shown in FIG. 4 has been replaced with sample sheet 235. Then, in the next step S4'', the generated sequence run data is stored in storage 3, and the process ends. The notation on the sample sheet may be any notation that allows for the identification of the corresponding subject for each library sample and the identification of whether the library sample is DNA-derived, RNA-derived, or non-tumor-derived.
[0090] As shown in Figure 22, in the third embodiment, the case registration screen displayed on the display unit 5b by the control unit 5f of the data transmission device 5 can be a case registration screen 240 in which, compared to the case registration screen 40 of the first embodiment (see Figure 8), the word "tumor specimen" is changed to "tumor specimen (DNA)" and registration sections 40j to 40l are added for registering information regarding library samples derived from RNA of tumor specimens.
[0091] As shown in Figure 22, this case registration screen 240 displays a registration section 40a for registering a sequence run ID, a registration section 40b for registering a case ID, a registration section 40c for registering an index ID, a registration section 40d for registering a sample ID, and a registration section 40e for registering an index sequence for library samples derived from the DNA of non-tumor specimens, a registration section 40f for registering an index ID, a registration section 40g for registering a sample ID, and a registration section 40h for registering an index sequence for library samples derived from the DNA of tumor specimens, and a registration section 40j for registering an index ID, a registration section 40k for registering a sample ID, and a registration section 40l for registering an index sequence for library samples derived from the RNA of tumor specimens, and a registration button 40i. Registration units 40a, 40c, 40f, and 40j are configured in a pull-down list format. Expanding the pull-down list displays a list of sequence run IDs and index IDs included in sequence run data that the control unit 5f has read from storage 3 and that do not have the aforementioned registered flag attached. Registration units 40b, 40d, 40e, 40g, 40h, 40k, and 40l are configured so that a user can input numbers, characters, or symbols using a keyboard. The user of the data transmission device 5 operates the input unit 5a to input information for each subject into registration units 40a-40l, and when registration for one subject is complete, selects registration button 40i. The user of the data transmission device 5 repeatedly inputs information for each subject into registration units 40a-40l and selects registration button 40i until input for all case IDs included in one sequence run data is completed.
[0092] Registration units 40a-40l may be configured in a pull-down list format, except for registration units 40a, 40c, 40f, and 40j. Registration units 40a, 40c, 40f, and 40j may be configured to accept input of numerical values, etc. Registration units 40e, 40h, and 40l may be configured such that, when an index ID is input to registration unit 40c, 40f, or 40j, the corresponding index sequence is read from the sequencing run data and displayed on registration unit 40e, 40f, or 40l. Any screen may be used as the case registration screen, as long as it allows registration of information that can identify the same subject for each library sample of tumor specimen DNA, tumor specimen RNA, and non-tumor specimen DNA.
[0093] FIG. 23 is a flowchart illustrating the details of the processing executed by the control unit 6f in step S33 (see FIG. 7). In step S51'', the control unit 6f reads out the sequence run ID, case ID, DNA specimen sample ID, RNA specimen sample ID, and non-tumor specimen sample ID from the stored case information. The sequence run ID is information input via the registration unit 40a of the case registration screen 240 (see FIG. 22), the case ID is information input via the registration unit 40b, the non-tumor specimen sample ID is information input via the registration unit 40d, the DNA specimen sample ID is information input via the registration unit 40g, and the RNA specimen sample ID is information input via the registration unit 40k. In step S52'', the control unit 6f reads out information written in the sample sheet of the sequence run data to which the same sequence run ID as the read sequence run ID is assigned. Then, in step S53'', the control unit 6f determines whether the combination of the case ID, sample ID of the DNA sample, sample ID of the RNA sample, and sample ID of the non-tumor sample read from the case information in step S51'' exists in the sample sheet.
[0094] If the control unit 6f makes a negative determination ("No") in step S53", it sends an error notification to the data transmission device 5 in step S54", indicating that the linking information in the sequence run data does not match the case information, and terminates the process without executing the processes from step S34 onwards (see Figure 7). Upon receiving the error notification, the control unit 5f of the data transmission device 5 outputs error information indicating that the linking information in the sequence run data does not match the case information to the display unit 5b. This output enables the user of the data transmission device 5 to recognize that an error exists in at least one of the information on the sample sheet and the manually entered case information. On the other hand, if the control unit 6f of the receiving device 6 makes a positive determination ("Yes") in step S53", it returns the process to step S34 (see Figure 7).
[0095] According to the third embodiment, in step S53'', it is determined whether the linking information of the sequencing run data and the case information are consistent. If the two pieces of information do not match, the process ends without proceeding to the next step. Therefore, for each subject, the nucleic acid sequence data derived from the DNA sample of the tumor specimen, the nucleic acid sequence data derived from the RNA sample of the tumor specimen, and the nucleic acid sequence data derived from the non-tumor specimen can be accurately linked to that subject. Therefore, even in the case of matched pair testing, in which nucleic acid sequence analysis is performed on multiple nucleic acid sequence data derived from the same subject, it is possible to reliably prevent incorrect analysis due to mismatched nucleic acid sequence data.
[0096] [Effects of the third embodiment] According to the third embodiment, in a matched pair test in which nucleic acid sequence data of DNA derived from a tumor specimen, nucleic acid sequence data of RNA derived from the tumor specimen, and nucleic acid sequence data of DNA derived from a non-tumor specimen from a single subject are analyzed as a set, the reception device 6 receives, via the network 11 from the data transmission device 5, sequence data sets containing multiple pieces of nucleic acid sequence data obtained using the sequencer 2 corresponding to multiple library samples, including a first library sample, a second library sample, and a third library sample, prepared from the same subject's specimen, and sequence run data containing linking information indicating that the first library sample, the second library sample, and the third library sample were prepared from the same subject's specimen, and transmits the sequence run data to the nucleic acid sequence analyzer 7 that analyzes the nucleic acid sequences. Therefore, even if the analysis requesting facility 10 that operates the sequencer 2 is a facility different from the requested facility 30 where the nucleic acid sequence analyzer 7 is installed, the nucleic acid sequence analyzer 7 can accurately and quickly extract the correct combination of nucleic acid sequence data corresponding to the multiple library samples from the same subject from the sequence data sets. Therefore, multiple nucleic acid sequence data in the correct combination can be analyzed at the requested facility 30, making it possible to accurately and quickly perform analysis using multiple nucleic acid sequence data from the same subject. Furthermore, in the third embodiment, by combining and comprehensively analyzing three pieces of information: information on somatic mutations other than fusion gene mutations obtained by analyzing nucleic acid sequence data of DNA from a tumor specimen, information on fusion gene mutations obtained by analyzing nucleic acid sequence data of RNA from a tumor specimen, and information on germline mutations obtained by analyzing nucleic acid sequence data of DNA from a non-tumor specimen, analysis can be performed based on more information, making it easier to identify a treatment method appropriate for the subject.
[0097] (Fourth embodiment) In the first embodiment, the case where information is exchanged between the data transmission device 5 and the nucleic acid sequence analysis device 7 is described as being carried out via the reception device 6, but the reception device 6 and the nucleic acid sequence analysis device 7 may also be configured as a single computer.
[0098] 24 is a schematic configuration diagram of a nucleic acid information transmission and reception system 101 installed in each facility according to the fourth embodiment. The nucleic acid information transmission and reception system 101 includes a sequencer 2, a storage (storage device) 3, a data transmission device 5, and a reception and analysis system 104, and the reception and analysis system 104 includes a reception and analysis device 107. The data transmission device 5 and the reception and analysis device 107 are connected to each other via a network 11, which is the Internet. A mutation information database 8 is also connected to the network 11. The hardware configurations of the sequencer 2, the storage (storage device) 3, and the data transmission device 5 are the same as those in the first embodiment. The data transmission device 5 transmits and receives data to and from the reception and analysis device 107 via the network 11.
[0099] The reception / analysis device 107 is installed at a request receiving facility 130, for example, a data analysis facility. The analysis requesting facility 10 and the request receiving facility 130 are different facilities. The reception / analysis device 107 may be a computer constituting a cloud system. The data analysis facility may be a facility of a cloud service provider or a facility of a company providing nucleic acid sequence analysis services. The reception / analysis device 107 is a computer. The reception / analysis device 107 has an input unit 107a, a display unit 107b, a transmission / reception unit 107c, and a control device 107e. The control device 107e includes a control unit 107f and a memory unit 107g. The hardware configurations of the input unit 107a, the display unit 107b, the transmission / reception unit 107c, and the control device 107e are similar to those of the input unit 5a, the display unit 5b, the transmission / reception unit 5c, and the control device 5e of the data transmission device, respectively. The reception / analysis device 107 can access the mutation information database 8 via the network 11.
[0100] 25 is a flowchart illustrating the processing executed by each control unit of the data transmitting device 5 and the reception and analysis device 107. Referring to FIG. 25, the processing executed by the control unit 5f of the data transmitting device 5 will first be described. When the control unit 5f receives an instruction to analyze sequence run data stored in storage 3 from the user of the data transmitting device 5, it transmits analysis request information to the reception device 6 in step S20′. In step S21′, the control unit 5f reads the sequence run data from storage 3 and transmits it to the reception and analysis device 107.
[0101] In step S22', the control unit 5f displays a case registration screen on the display unit 5b and accepts registration of case information. As the case registration screen, the case registration screen 40, 140, or 240 shown in the first to third embodiments can be adopted.
[0102] When the processing of step S22' is completed, in step S23', the control unit 5f transmits the case information entered on the case registration screen in step S22' to the reception and analysis device 107. In step S24', the control unit 5f adds a flag indicating that the sequence run ID corresponding to the sequence run data transmitted to the reception and analysis device 107 in step S21 has been registered.
[0103] Next, the processing executed by the control unit 107f of the reception / analysis device 107 will be described. When analysis request information is transmitted from the data transmission device 5, the control unit 107f receives the analysis request information and stores it in the memory unit 107g in step S30′. When sequence run data is transmitted from the data transmission device 5, the control unit 107f receives the sequence run data and stores it in the memory unit 107g in step S41′. When case information is transmitted from the data transmission device 5, the control unit 107f receives the case information and stores it in the memory unit 107g in step S32′. In the subsequent step S33′, the control unit 107f verifies consistency and determines whether the linking information included in the sequence run data stored in step S41′ is consistent with the information included in the case information stored in step S32′.
[0104] In step S42', the control unit 107f reads out one sequence data set from the stored sequence run data. As described above, since the sequence data set includes multiple nucleic acid sequence data corresponding to the same case ID, the control unit 107f can extract multiple nucleic acid sequence data corresponding to the same case ID as one sequence data set using the case ID, which is linking information, as a search key.
[0105] In step S43′, the control unit 107f analyzes the presence or absence of mutations for each nucleic acid sequence data of the sequence data set extracted in step S42′ using the nucleic acid sequence information of tumor cells in the mutation information database 8. In step S44′, the control unit 107f creates an analysis result report based on the presence or absence of mutations. In step S45′, the control unit 107f transmits the analysis result report to the data transmission device 5. In step S46′, the control unit 107f determines whether all sequence data sets included in the sequencing run data stored in step S41′ have been analyzed. If all sequence data sets have been analyzed (if “Yes”), the control unit 107f terminates the process. If not all sequence data sets have been analyzed (if “No”), the control unit 107f returns the process to step S42′ and executes steps S42′ to S46′ again.
[0106] Meanwhile, in step S25', the control unit 5f of the data transmission device 5 receives the analysis result report, stores it in the storage unit 5g, and ends the process. This allows the doctor in charge of the subject to display and view the analysis report stored in the storage unit 5g on the display unit 5b at any time.
[0107] [Operation and effect of the fourth embodiment] According to the fourth embodiment, the hardware configuration of the reception and analysis system 104 is simplified. Furthermore, since the reception and analysis of sequence run data can be performed on the same computer, the time required to send and receive sequence run data can be reduced, and a decrease in communication speed due to the flow of large amounts of data over the network 11 can be prevented.
[0108] (Fifth embodiment) The fifth embodiment is an embodiment that encompasses embodiments 1 to 4 and their variations. The schematic configuration of the nucleic acid information transmission and reception system 101 may be any of the configurations of embodiments 1 to 3 (see FIG. 1) or the configuration of embodiment 4 (see FIG. 24). FIG. 26 is a flowchart illustrating the processing executed by the sequencer 2 of the fifth embodiment. The processing executed by the sequencer 2 will be described with reference to FIG. 26. First, in step S1'', the sequencer 2 receives a sequence run ID, a case ID, a sample ID, and an index ID, and generates a sample sheet, which is an electronic file.
[0109] Next, the user of the sequencer 2 dispenses multiple library samples prepared in advance into each well of a cartridge, sets the cartridge in the sequencer 2, and instructs the sequencer 2 to start sequence reading. When the user instructs the sequencer 2 to start sequence reading, the sequencer 2 reads the nucleic acid sequence for each of the multiple library samples in step S2''. In the fifth embodiment, the sequencer 2 reads the nucleic acid sequences of multiple library samples collected and prepared from the same subject for each of multiple subjects. Next, in step S3'', the sequencer 2 generates sequence run data. Then, in the next step S4'', the generated sequence run data is stored in storage 3, and the process ends.
[0110] FIG. 27 is a flowchart illustrating the process of determining the type of mutation in multiple nucleic acid sequence data included in one sequence dataset extracted in step S42 (see FIG. 7) or step S42′ (see FIG. 25). The process executed by the control unit 7f or the control unit 107f will be described with reference to FIG. 27. In step S83′, the control unit 7f or the control unit 107f determines whether or not there is a mismatch between an acquired sequence and a reference sequence for one acquired sequence among multiple nucleic acid sequence data included in one sequence dataset acquired in step S61 (see FIG. 10). If there is a mismatch (if "Yes"), the control unit 7f or the control unit 107f proceeds to step S84′; if there is no mismatch (if "No"), the control unit 7f or the control unit 107f proceeds to step S44 (see FIG. 7) or step S44′ (see FIG. 25).
[0111] In step S84', the control unit 7f or the control unit 107f determines the mismatched base detected in step S83', i.e., the type of mutation. In step S99', the control unit 7f or the control unit 107f determines whether all of the multiple nucleic acid sequence data included in the acquired single sequence data set have been compared with the reference sequence. If it is determined that all of the nucleic acid sequence data have been compared (in the case of "Yes"), the control unit 7f or the control unit 107f proceeds to step S85'. If it is determined that all of the nucleic acid sequence data have not been compared (in the case of "No"), the control unit 7f or the control unit 107f returns the process to step S83'.
[0112] In step S85', the control unit 7f or the control unit 107f searches the mutation information database stored in the mutation information database 8 based on each detected mutation. In step S86', the control unit 7f assigns a gene name, annotation, etc. to each detected mutation based on the search result of step S85'. Note that in the fifth embodiment, the processes of steps S85' and S86' can be omitted.
[0113] [Effects of the Fifth Embodiment] According to the fifth embodiment, even if the analysis requesting facility 10 that operates the sequencer 2 is a facility different from the requested facility 30 where the nucleic acid sequence analyzing device 7 is installed or the requested facility 130 where the reception / analysis device 107 is installed, the nucleic acid sequence analyzing device 7 or the reception / analysis device 107 can accurately and quickly extract from the sequence dataset the correct combination of nucleic acid sequence data corresponding to multiple library samples of the same subject. Therefore, the requested facility 30 or 130 can analyze multiple nucleic acid sequence data of the correct combination, making it possible to accurately and quickly perform analysis using multiple nucleic acid sequence data of the same subject.
[0114] (Sixth embodiment) In the fifth embodiment, the analysis request information was transmitted and the integrity verification process was executed, but in the sixth embodiment, the analysis request information is not transmitted and the integrity verification process is not executed. FIG. 28 is a flowchart explaining the process executed by the control units of the data transmitting device 5, the receiving device 6, and the nucleic acid sequence analyzing device 7 in the sixth embodiment. FIG. 29 is a flowchart explaining the process executed by the control units of the data transmitting device 5 and the receiving / analyzing device 107 in the sixth embodiment. As shown in FIG. 28, in the sixth embodiment, the control unit 5f of the data transmitting device 5 executes the processes of steps S21, S24, and S25, but does not execute the processes of steps S20, S22, and S23 (see FIG. 7). The control unit 6f of the receiving device 6 executes the processes of steps S31, S34, S35, and S36, but does not execute the processes of steps S30, S32, and S33 (see FIG. 7). The control unit 7f of the nucleic acid sequence analyzing device 7 executes the processes of steps S41, S42, S43, S44, S45, and S46, but does not execute the process of step S40 (see FIG. 7). As shown in FIG. 29, in the sixth embodiment, the control unit 5f of the data transmitting device 5 executes the processes of steps S21′, S24′, and S25′, but does not execute the processes of steps S20′, S22′, and S23′ (see FIG. 25). The control unit 107f of the receiving / analyzing device 107 executes the processes of steps S41′, S42′, S43′, S44′, S45′, and S46′, but does not execute the processes of steps S30′, S32′, and S33′ (see FIG. 25).
[0115] The present disclosure is not limited to the above-described embodiment and its modifications, and various improvements and modifications are possible within the scope of the claims of the present application and their equivalents.
[0116] For example, the analysis system 4 may be composed of three or more computers. Alternatively, a first library sample may be prepared from a specimen collected from one tumor tissue of one subject, and a second library sample may be prepared from a specimen collected from a tumor tissue different from the one tumor tissue of the same subject. For example, a first library sample may be prepared from a specimen collected from the colon of one subject, and a second library sample may be prepared from a specimen collected from the stomach of the same subject. [Explanation of symbols]
[0117] 1,101 Nucleic acid information transmission and reception system, 2 Sequencer, 3 Storage, 4,104 Analysis system, 5 Data transmission device, 6 Reception device, 7 Nucleic acid sequence analysis device, 8 Mutation information database, 10 Analysis request originating facility, 11 Network, 20 Request reception facility, 30,130 Requested facility, 35,35',135,235,335 Sample sheet, 37 Nucleic acid sequence data, 38 Index, 40,140,240,340 Case registration screen, 50 Sequence run data, 107 Reception and analysis device.
Claims
1. 1. A control method for genetic panel testing, comprising: controlling a computer in a second facility to analyze nucleic acid sequence data obtained in a first facility using a sequencer that reads nucleic acid sequences present in the first facility; receiving, via a network from the first facility at the second facility, a sequence dataset including a plurality of nucleic acid sequence data obtained using the sequencer, the sequence dataset corresponding to each of a plurality of library samples including a first library sample and a second library sample prepared from a specimen of the same subject, and linking information indicating that the first library sample and the second library sample were prepared from the specimen of the same subject; analyzing, at the second facility, first sequence data and second sequence data corresponding to the first library sample and the second library sample linked by the linking information; outputting analysis information based on the analysis results of the first sequence data and the second sequence data; and at the second facility, further receiving input information indicating that the first library sample and the second library sample were prepared from the specimen of the same subject, comparing the linking information with the input information, and determining whether or not there is a match between the linking information and the input information based on the comparison result. Control method.
2. receiving the sequence dataset and the linking information by a first computer, and transmitting the received sequence dataset and the linking information to a second computer; the second computer analyzes the first sequence data and the second sequence data and outputs the analysis information. The control method according to claim 1 .
3. receiving the sequence data set and the linking information by a computer, analyzing the first sequence data and the second sequence data, and outputting the analysis information; The control method according to claim 1 .
4. the first library sample is a sample prepared from a tumor specimen of the subject, and the second library sample is a sample prepared from a non-tumor specimen of the subject; the analysis information includes information on somatic mutations based on the analysis results of the first sequence data and information on germline mutations based on the analysis results of the second sequence data; A control method according to any one of claims 1 to 3.
5. the first library sample is a sample prepared from deoxyribonucleic acid contained in a tumor specimen of the subject, and the second library sample is a sample prepared from ribonucleic acid contained in the tumor specimen; the analysis information includes information on somatic mutations based on the analysis results of the first sequence data and information on fusion gene mutations based on the analysis results of the second sequence data; A control method according to any one of claims 1 to 3.
6. the sequence dataset further comprises third sequence data corresponding to a third library sample prepared from a non-tumor specimen of the same subject; the linking information is information indicating that the third library sample, in addition to the first library sample and the second library sample, has been prepared from a specimen of the same subject; In addition to analyzing the first sequence data and the second sequence data, further analyzing the third sequence data; the analysis information further includes information on germline mutations based on the analysis results of the third sequence data, in addition to information on somatic mutations based on the analysis results of the first sequence data and information on fusion gene mutations based on the analysis results of the second sequence data; The control method according to claim 5.
7. The non-tumor sample is a blood sample collected from the subject. The control method according to claim 4 or 6.
8. 8. The control method according to claim 1, further comprising receiving analysis request information from the first facility via the network, the analysis request information including at least one of case information of the subject, type information of the genetic panel test, and information of the first facility.
9. When the linking information and the input information are consistent, an analysis of the first sequence data and the second sequence data is performed. A control method according to any one of claims 1 to 8.
10. When the linking information and the input information are inconsistent, notifying the first facility of error information based on the inconsistency; A control method according to any one of claims 1 to 9.
11. receiving, together with the sequence dataset, another sequence dataset including a plurality of nucleic acid sequence data obtained using the sequencer, the plurality of sequence dataset corresponding to each of a plurality of library samples, including a fourth library sample and a fifth library sample, each prepared from another specimen of the same subject; A control method according to any one of claims 1 to 10.
12. the first library sample, the second library sample, the fourth library sample, and the fifth library sample are samples whose sequences have been read by the sequencer in the same sequencing run; The control method according to claim 11.
13. receiving the sequence data set and the linking information, analyzing the first sequence data and the second sequence data, and outputting the analysis information are performed by a computer constituting a cloud system; 13. A control method according to any one of claims 1 to 12.
14. The linking information is also used as sample identification information for identifying a library sample, or subject identification information for identifying a subject from whom a specimen of the library sample was collected.
14. A control method according to any one of claims 1 to 13.
15. 1. An analysis system for a gene panel test, in which nucleic acid sequence data obtained at a first facility is analyzed at a second facility using a sequencer that reads nucleic acid sequences present at the first facility, a first computer that receives, via a network from the first facility, a sequence dataset containing nucleic acid sequence data obtained using the sequencer, the sequence dataset corresponding to each of a plurality of library samples, including a first library sample and a second library sample prepared from a specimen of the same subject, and linking information indicating that the first library sample and the second library sample were prepared from the specimen of the same subject, and transmits the sequence dataset and the linking information obtained from the first facility to a second computer; the second computer analyzes the first sequence data and the second sequence data linked by the linking information, and outputs analysis information based on an analysis result of the first sequence data and an analysis result of the second sequence data, the second computer is located at the second facility; the first computer further receives input information indicating that the first library sample and the second library sample were prepared from the specimen of the same subject, compares the linking information with the input information, and determines whether there is a match between the linking information and the input information based on the comparison result. Analysis system.
16. 1. An analysis system for a gene panel test, in which nucleic acid sequence data obtained at a first facility is analyzed at a second facility using a sequencer that reads nucleic acid sequences present at the first facility, a computer that receives, via a network from the first facility, a sequence dataset containing nucleic acid sequence data obtained using the sequencer, the sequence dataset corresponding to each of a plurality of library samples, including a first library sample and a second library sample prepared from a specimen of the same subject, and linking information indicating that the first library sample and the second library sample were prepared from the specimen of the same subject; analyzes the first sequence data and the second sequence data linked by the linking information; and outputs analysis information based on the analysis results of the first sequence data and the analysis results of the second sequence data; the computer is located at the second facility; the computer at the second facility further receives input information indicating that the first library sample and the second library sample were prepared from the same subject's specimen, compares the linking information with the input information, and determines, based on the comparison result, whether there is a match between the linking information and the input information. Analysis system.
Citation Information
Patent Citations
Genetic screening system
JP2004290240A
Gene chromosome inspection management system, inspection management server, client terminal, gene chromosome inspection management method, and program
JP2015170186A
Method for managing requests of examinations by computer, management device, management computer program, and management system
JP2021056891A
Method for supporting expert meetings using computer, support device, computer program for supporting expert meetings, and support system
JP2022011755A
Cloud computing environment for biological data
US9444880B2