Genome information processing method, genome information processing system, and genome information processing program

The genomic information processing method and system use a blockchain to manage the history of genomic data processing steps, addressing personal information protection and secure data sharing in personalized medicine.

WO2025154155A1PCT designated stage expired Publication Date: 2025-07-24NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/000895
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing genomic data management systems fail to adequately protect personal information and ensure secure, transparent processing of genomic data throughout the series of processes required for personalized medicine, lacking comprehensive management of the history of processes from identifying mutations in genomic data.

Method used

A genomic information processing method and system utilizing a blockchain to register and record the history of processes from identifying mutation locations in genomic data, involving multiple corporate servers and a management device to securely manage and share these processes.

Benefits of technology

Ensures secure, transparent management of genomic data processing history, facilitating secure sharing and processing across multiple entities while protecting personal information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024000895_24072025_PF_FP_ABST
    Figure JP2024000895_24072025_PF_FP_ABST
Patent Text Reader

Abstract

This genome information processing system comprises: a company server that performs a series of instances of processing until a location including a mutation is identified from genome data included in a sample from a subject, then issues an instruction to register and record the history of each instance of processing from the series of instances of processing in a blockchain; and a management device that receives the instruction from the company server and registers and records the history of each instance of processing from the series of instances of processing in the blockchain.
Need to check novelty before this filing date? Find Prior Art

Description

Genome information processing method, genome information processing system, and genome information processing program

[0001] The present invention relates to a genome information processing method, a genome information processing system, and a genome information processing program.

[0002] Until now, medical care has been disease-centered, with the main objectives being to identify the causes of disease and develop treatments. It has long been known that the same treatment is not guaranteed to be effective even for the same disease, but individual differences in treatment effectiveness can only be understood by observing the treatment and its effects, making it difficult to create optimal treatment plans for each individual.

[0003] Meanwhile, advances in technologies such as DNA sequencing and identification of single nucleotide polymorphisms (SNPs) within DNA sequences have made it possible to observe how an individual's genes differ from those of others. Therefore, using this information, it has become practical to select the optimal treatment for an individual patient and to select and formulate a therapeutic drug optimized for that individual. For example, Patent Document 1 describes a method for obtaining highly reliable genome mutation information from read information obtained from a next-generation sequencer.

[0004] Japanese Patent Application Laid-Open No. 2017-016665

[0005] The disclosures of the above-mentioned prior art documents are incorporated herein by reference. The following analysis has been carried out by the present inventors.

[0006] Incidentally, genome data, which represents the base sequence of DNA obtained from an individual, is in many cases considered to be a "personal identification code" in the conventional wisdom, and must be strictly managed as personal information. Furthermore, it is expected that there will be many cases in which "genomic information," which is genome data obtained from a patient and to which medical interpretations, such as information on mutations, are added, will be considered to be sensitive personal information and require special care in handling.

[0007] Furthermore, because such genome data is considered to be personal information, it is not enough to simply protect it from leaks, tampering, etc. Specifically, companies that receive genome data must only process it as agreed upon by the data provider, and companies that receive genome data must not only protect it from leaks, tampering, etc., but also strictly manage the processing they perform on the genome data themselves.

[0008] In view of the above-mentioned problems, the object of the present invention is to provide a genome information processing method, a genome information processing system, and a genome information processing program that contribute to managing the history of each step of a series of processes from identifying the location containing a mutation in a subject's sample.

[0009] In a first aspect of the present invention, a genome information processing method is provided, which registers and records the history of each step of a series of processes from identifying the location of a mutation from the genome data contained in a subject's sample in a blockchain.

[0010] In a second aspect of the present invention, a genome information processing system is provided that includes a corporate server that performs a series of processes to identify locations containing mutations from genome data contained in a subject's sample and issues instructions to register and record the history of each of the series of processes in a blockchain, and a management device that receives instructions from the corporate server and registers and records the history of each of the series of processes in a blockchain.

[0011] In a third aspect of the present invention, there is provided a genome information processing program that receives instructions from a corporate server that performs a series of processes from identifying the location of a mutation in genome data contained in a subject's sample to register and record the history of each of the series of processes in a blockchain, and causes a computer to execute a process of registering and recording the history of each of the series of processes in a blockchain. This program can be recorded on a computer-readable storage medium. The storage medium can be a non-transient medium such as a semiconductor memory, a hard disk, a magnetic recording medium, or an optical recording medium. The present invention can also be embodied as a computer program product.

[0012] According to each aspect of the present invention, it is possible to provide a genome information processing method, a genome information processing system, and a genome information processing program that contribute to managing the history of each step of a series of processes from identifying the location containing a mutation in a subject's sample.

[0013] Fig. 1 is a diagram showing an example of the general configuration of a genome information processing system. Fig. 2 is a conceptual diagram showing an example of a series of processes for genome data in personalized medicine. Fig. 3 is a diagram showing a series of processes for acquiring mutation information from genome data. Fig. 4 is a diagram showing an example of the hardware configuration of a company server and a management device used in an embodiment.

[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the present invention is not limited to the embodiments described below. In addition, the same or corresponding elements in each drawing are appropriately designated by the same reference numerals. Furthermore, it should be noted that the drawings are schematic, and the dimensional relationships and ratios of each element may differ from those in reality. There may also be parts in which the dimensional relationships and ratios differ between the drawings.

[0015] 1 is a diagram showing an example of the schematic configuration of a genome information processing system. As shown in FIG. 1, the genome information processing system 10 includes multiple corporate servers 11, 12, and 13, a management device 2, a blockchain 3, and a genome database 4.

[0016] The multiple corporate servers 11, 12, and 13 perform a series of processes up to identifying the location where a mutation is contained from the genome data contained in the subject's sample, and issue instructions to the management device 2 to register and record each history of the series of processes in the blockchain 3. As will be described in detail later, the process up to identifying the location where a mutation is contained from the genome data contained in the subject's sample is composed of multiple consecutive processes. Therefore, although it is conceivable that the series of processes up to identifying the location where a mutation is contained from the genome data contained in the subject's sample be performed by a single corporate server, it is also conceivable that the process be shared and performed by the multiple corporate servers 11, 12, and 13.

[0017] The management device 2 receives instructions from multiple corporate servers 11, 12, and 13 and registers and records the history of each series of processes in the blockchain 3. The genome database 4 records genome data obtained from subject samples and mutation information obtained by performing a series of processes on the genome data. Note that while the genome information processing system 10 shown in FIG. 1 is configured to include multiple corporate servers 11, 12, and 13 and a separate management device 2, the management device 2 may be implemented as part of the functions of the multiple corporate servers 11, 12, and 13, and each of the multiple corporate servers 11, 12, and 13 may directly register and record the history of each series of processes up to identifying the location of a mutation in the genome data contained in the subject sample in the blockchain 3.

[0018] Figure 2 is a conceptual diagram showing an example of a series of processes for genome data in personalized medicine. As shown in Figure 2, personalized medicine is generally carried out by multiple companies, with each company sharing the responsibility.

[0019] Samples such as blood and tissue samples collected from test subjects at medical institutions are sent to sequencing companies. The sequencing companies sequence the samples they receive and obtain genome data (read information) that represents the base sequence of the genome contained in the sample. However, because genome data (read information) is merely the base sequence of the genome, it alone cannot be used to perform personalized medicine. Therefore, the sequencing companies send the genome data (read information) to data analysis companies.

[0020] Data analysis companies analyze genome data (read information) to generate mutation information that identifies the locations containing mutations. Here, mutation information refers to information about mutations in the genome's base sequence, including the location of the base sequence containing the mutation and the type of mutation, such as a single-base mutation, insertion, or deletion in the base sequence. This mutation information is related to the biological characteristics and diseases of individual subjects, and is important information for personalized medicine. However, even if mutation information for an individual subject is obtained, it is not possible to determine a treatment plan for personalized medicine. Therefore, data analysis companies add mutation information to the genome data and send the resulting genome information to ICT (Information and Communication Technology) companies.

[0021] ICT companies use genomic information, which adds mutation information to genomic data, to perform neoantigen analysis and determine treatment strategies. Neoantigens are cancer antigens (markers) that emerge due to genetic mutations in cancer cells, for example, and are important information for determining personalized medicine strategies. Furthermore, since the occurrence of cancer cells does not necessarily mean that these genetic mutations will occur, analysis is required through neoantigen analysis of genomic information, which adds mutation information to genomic data. Meanwhile, once neoantigen information is obtained, neoantigens become targets for personalized medicine.

[0022] ICT companies will send treatment plans derived from neoantigen analysis of genomic information, which adds mutation information to genomic data, to biotechnology companies. The biotechnology companies will manufacture appropriate medicines and other products based on the results of neoantigen analysis. The medicines and other products manufactured by the biotechnology companies will be sent to medical institutions where treatment will be provided as personalized medicine for the subject.

[0023] Note that the above-described processing is an example of a series of processes performed on genome data in personalized medicine, and in actual personalized medicine, processes different from those described above may be performed. For example, neoantigen analysis is an example of a process performed in personalized medicine, and cancer gene panel testing may be performed instead of neoantigen analysis. Cancer gene panel testing is a test that examines mutations in base sequences to obtain information on therapeutic drugs and clinical trials that are expected to be effective. When an ICT company performs a cancer gene panel test, the information obtained from the test on therapeutic drugs and clinical trials that are expected to be effective is sent to a biotechnology company. The biotechnology company manufactures medicines and other products that are appropriate for the treatment plan obtained from the cancer gene panel test, and the medicines and other products manufactured by the biotechnology company are sent to a medical institution where treatment is administered as personalized medicine for the subject.

[0024] As such, personalized medicine does not simply obtain genome data, which is simply the base sequence of the genome, from a specimen; multiple processes are then performed. In particular, the process of generating genome information by adding mutation information to genome data is important for obtaining information useful for personalized medicine from simple genome base sequence information. Moreover, the process of generating genome information by adding mutation information to genome data is not a single process, but a series of processes consisting of multiple processes.

[0025] 3 is a diagram showing a series of processes for obtaining mutation information from genome data. As shown in FIG. 3, the process for obtaining mutation information from genome data (read information) includes a process (S1) of sequencing a subject's sample to generate read information (Fastq file) which is the base sequence of the genome, a process (S2) of generating alignment information (alignment) of the read information by mapping it to a genome reference sequence and evaluating the quality of the alignment information (QC; quality control), and a process (S3) of detecting mutations from the alignment information and generating mutation information (VCF file).

[0026] The process of generating the read information (Fastq file) (S1) involves a sequencer reading the genome sequence from a subject's sample, and the read information (Fastq file) is the output of the sequencer. The read information (Fastq file) records the base sequence and its quality score.

[0027] The process of generating alignment information for read information involves comparing the read information (Fastq file) with a reference sequence and aligning regions with identical sequences. The process of evaluating the quality of the alignment information (quality control; QC) (S2) involves deleting regions with low sequencing accuracy in the alignment information to prevent interference with subsequent analysis.

[0028] In the process of detecting mutation information from the alignment information and generating mutation information (VCF file) (S3), mutations are detected from the alignment information and the mutation information (VCF file) is output (S4). The mutation information (VCF file; Variant Call Format file) records information on mutations in the base sequence.

[0029] The mutation information (VCF file) obtained through the above series of processes is used for further processing such as neoantigen analysis.

[0030] As described above, the process of identifying locations containing mutations in genome data is composed of multiple processes. Therefore, the process of identifying locations containing mutations in genome data itself can also be performed by multiple companies with the roles being shared. In other words, like the genome information processing system 10 shown in Figure 1, it is possible to have multiple company servers 11, 12, and 13, and to perform the process of identifying locations containing mutations in genome data by sharing the roles among the multiple company servers 11, 12, and 13.

[0031] The multiple corporate servers 11, 12, and 13 execute at least a part of the series of processes for identifying locations containing mutations from genome data as described above, and issue instructions to the management device 2 to register and record the history of each of these processes in the blockchain. Here, a hash value or a combination of hash values ​​of a combination of sample ID, subject ID, processing date, processing program name, program version, and processing results (product, evaluation result) is used for each processing history. Note that, although it is preferable to register and record the processing history in the blockchain, it is also possible to include the hash value of the genome data.

[0032] (Hardware Configuration Example) Figure 4 is a diagram showing an example of the hardware configuration of the enterprise server and management device used in the embodiment. That is, the enterprise servers 11, 12, 13 and management device 2 are able to realize the respective functions of the enterprise servers 11, 12, 13 and management device 2 by executing the above-described genome information processing method as a program on an information processing device (computer) 20 employing the hardware configuration shown in Figure 4. However, the hardware configuration example shown in Figure 4 is an example of a hardware configuration that realizes the respective functions of the enterprise servers 11, 12, 13 and management device 2, and is not intended to limit the hardware configuration of the enterprise servers 11, 12, 13 and management device 2. The enterprise servers 11, 12, 13 and management device 2 may include hardware not shown in Figure 4.

[0033] As shown in FIG. 4, the hardware configuration that can be adopted by the enterprise servers 11, 12, 13 and the management device 2 includes a CPU (Central Processing Unit) 21, a main memory device 22, an auxiliary memory device 23, and an IF (Interface) unit 24, which are interconnected by, for example, an internal bus.

[0034] The CPU 21 executes each command included in the genome information processing program executed by the information processing device (computer) 20. The main storage device 22 is, for example, a RAM (Random Access Memory), and temporarily stores various programs, such as the genome information processing program executed by the information processing device (computer) 20, for processing by the CPU 21.

[0035] The auxiliary storage device 23 is, for example, a hard disk drive (HDD), and is capable of storing, for the medium to long term, various programs such as a genome information processing program executed by the information processing device (computer) 20. Various programs such as a genome information processing program can be provided as a program product recorded on a non-transitory computer-readable storage medium.

[0036] The IF unit 24 provides an interface for input and output between the company servers 11 , 12 , and 13 and the management device 2 , for example.

[0037] The information processing device (computer) 20 employing the above-described hardware configuration implements the functions of the company servers 11, 12, and 13 and the management device 2 by executing the genome information processing method described above as a program.

[0038] Some or all of the above embodiments can be described as, but are not limited to, the following supplementary notes. [Supplementary Note 1] A genome information processing method in which a history of each step of a series of processes up to identifying a site containing a mutation from genome data contained in a sample from a subject is registered and recorded in a blockchain. [Supplementary Note 2] The genome information processing method according to Supplementary Note 1, wherein the series of processes include: a process of sequencing the sample from the subject to generate read information, which is a genome base sequence; a process of generating alignment information of the read information by mapping a genome reference sequence; a process of evaluating the quality of the alignment information; and a process of detecting mutation information from the alignment information to generate mutation information. [Supplementary Note 3] The genome information processing method according to Supplementary Note 1 or Supplementary Note 2, wherein the history is a hash value or a combination of hash values ​​of a combination of a sample ID, a subject ID, a processing date, a processing program name, a program version, and processing results (products, evaluation results). [Supplementary Note 4] A genome information processing system comprising: a corporate server that performs a series of processes from identifying locations containing mutations in genome data contained in a sample from a subject, and issues instructions to register and record each history of the series of processes in a blockchain; and a management device that receives instructions from the corporate server and registers and records each history of the series of processes in the blockchain. [Supplementary Note 5] The genome information processing system according to Supplementary Note 4, wherein there are multiple corporate servers, and the series of processes are shared and executed by the multiple corporate servers. [Supplementary Note 6] The genome information processing system according to Supplementary Note 4 or Supplementary Note 5, wherein the management device is implemented as part of the function of the corporate server. [Supplementary Note 7] The genome information processing system according to any one of Supplementary Notes 4 to 6, wherein the series of processes include: a process of sequencing the sample from the subject to generate read information, which is a nucleotide sequence of the genome; a process of generating alignment information of the read information by mapping a reference sequence of the genome; a process of evaluating the quality of the alignment information; and a process of detecting mutation information from the alignment information to generate mutation information.[Appendix 8] The genome information processing system according to any one of Appendices 4 to 7, wherein the history is a hash value or a combination of hash values ​​of a combination of sample ID, subject ID, processing date, processing program name, program version, and processing results (product, evaluation result). [Appendix 9] A genome information processing program that receives an instruction from a company server that performs a series of processes up to identifying locations containing mutations from genome data contained in a subject's sample to register and record the history of each of the series of processes in a blockchain, and causes a computer to execute a process of registering and recording the history of each of the series of processes in a blockchain. [Appendix 10] The genome information processing program according to Appendices 9, wherein the series of processes include a process of sequencing the subject's sample to generate read information, which is a nucleotide sequence of the genome, a process of mapping a genome reference sequence to generate alignment information of the read information, a process of evaluating the quality of the alignment information, and a process of detecting mutation information from the alignment information to generate mutation information.

[0039] The disclosures of the above-cited patent documents and other documents are incorporated herein by reference. Modifications and adjustments of the embodiments and examples are possible within the scope of the entire disclosure of the present invention (including the claims), and further based on the basic technical concepts thereof. Furthermore, various combinations and selections (including partial deletions) of various disclosed elements (including elements of each claim, each element of each embodiment or example, each element of each drawing, etc.) are possible within the scope of the entire disclosure of the present invention. In other words, the present invention naturally embraces various modifications and alterations that would be possible by a person skilled in the art in accordance with the entire disclosure and technical concepts, including the claims. In particular, with regard to the numerical ranges described herein, any numerical value or subrange within the range should be construed as specifically described, even if not otherwise specified. Furthermore, the disclosures of the above-cited documents, when used in part or in whole in combination with the disclosures herein as part of the disclosure of the present invention, in accordance with the spirit of the present invention, are also deemed to be included in the disclosures of this application.

[0040] 2 Management device 3 Blockchain 4 Genome database 10 Genome information processing system 11, 12, 13 Corporate server 20 Information processing device 21 CPU 22 Main memory device 23 Auxiliary memory device 24 IF unit

Claims

1. A genomic information processing method for registering and recording in a blockchain the history of each of a series of processes until a location containing a mutation is identified from genomic data included in a sample of a subject.

2. The series of processes include: a process of sequencing the sample of the subject to generate read information which is a nucleotide sequence of a genome; a process of generating alignment information of the read information by mapping a reference sequence of the genome; a process of evaluating the quality of the alignment information; and a process of detecting mutation information from the alignment information to generate mutation information. The genomic information processing method according to claim 1.

3. The history is a hash value or a combination of hash values of a combination of a sample ID, a subject ID, a processing date, a processing program name, a program version, and a processing result (product, evaluation result). The genomic information processing method according to claim 1 or claim 2.

4. An enterprise server that performs a series of processes until a location containing a mutation is identified from genomic data included in a sample of a subject, and issues an instruction for registering and recording in a blockchain the history of each of the series of processes; and a management device that receives the instruction from the enterprise server and registers and records in a blockchain the history of each of the series of processes. A genomic information processing system comprising the above.

5. There are a plurality of the enterprise servers, and the series of processes are shared and executed by the plurality of enterprise servers. The genomic information processing system according to claim 4.

6. The management device is implemented as a part of the functions of the enterprise server. The genomic information processing system according to claim 4 or claim 5.

7. The series of processes include: a process of sequencing the sample of the subject to generate read information which is a nucleotide sequence of a genome; a process of generating alignment information of the read information by mapping a reference sequence of the genome; a process of evaluating the quality of the alignment information; and a process of detecting mutation information from the alignment information to generate mutation information. The genomic information processing system according to any one of claims 4 to 6.

8. The genomic information processing system according to any one of claims 4 to 7, wherein the history is a hash value of a combination of a sample ID, a subject ID, a processing date, a processing program name, a program version, and a processing result (product, evaluation result), or a combination of hash values.

9. A genomic information processing program that causes a computer to execute a process of registering and recording each history of the series of processes in a blockchain, in response to an instruction for registering and recording each history of the series of processes in the blockchain, from an enterprise server that performs a series of processes until a location containing a mutation is specified from genomic data included in a sample of a subject.

10. The genomic information processing program according to claim 9, wherein the series of processes includes: a process of sequencing the sample of the subject to generate read information that is a nucleotide sequence of a genome; a process of generating alignment information of the read information by mapping a reference sequence of the genome; a process of evaluating the quality of the alignment information; and a process of detecting mutation information from the alignment information to generate mutation information.

Citation Information

Patent Citations

  • Method for selecting variation information from sequence data, system, and computer program

    JP2017016665A

  • Gene big data disease prediction system and auxiliary prediction diagnosis method based on algorithm and block chain

    CN113722765A

  • Program, learning model, information processor, information processing method, and learning model generation method

    JP2020144658A

  • Therapeutic strategy drafting assistance device

    JP2022069002A

  • Medical supply tracking management system

    JP2022083384A