Information processing device, information processing method, and information processing program

JP7911794B2Active Publication Date: 2026-08-27THE UNIV OF TOKYO
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2024509670
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-25
Publication Date
2026-08-27
Estimated Expiration
2042-03-25

AI Technical Summary

Benefits of technology

【0011】 本発明によれば、塩基配列の変異が病気の発生や進行に影響する可能性の程度を、より正確に提示することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007911794000001
    Figure 0007911794000001
  • Figure 0007911794000002
    Figure 0007911794000002
  • Figure 0007911794000003
    Figure 0007911794000003
Patent Text Reader

Abstract

This invention relates to an information processing device relating to genetic information, an information processing method, and an information processing program. Provided is an information processing device that selects a target sequence mutation which is in a specimen and which may be harmful, the information processing device including: a filtering unit 2 that, on the basis of one or more classification criteria, classifies, into categories corresponding to the degree of risk of harm, one or more sequence mutations identified by determining nucleic acid sequences included in the specimen; and a control unit 3 that, on the basis of at least one of the classification criteria, classifies, into the categories corresponding to the degree of risk of harm, base sequences that include a sequence mutation for which the proper affiliation category is known, and compares the classification results to the proper affiliation categories. Also provided are an information processing method and an information processing program.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a base sequence information processing apparatus, an information processing method, and an information processing program.

Background Art

[0002] Conventionally, it is widely known that diseases may occur due to mutations in the base sequences contained in the genetic information of somatic cells. For example, mutations such as single nucleotide polymorphisms (SNPs) and structural polymorphisms (SVs) that occur within genes can cause diseases such as cancer. In recent years, information on what diseases are related to various base sequence mutations in somatic cells has been recorded in databases and is widely used (see Non-Patent Document 1).

[0003] In recent years, with the progress of comprehensive base sequence analysis techniques (e.g., next-generation sequencers (NGS)), it has become possible to analyze the entire genome at the individual level. Therefore, the number of mutations detected in a single mutation analysis has reached an enormous amount of several hundred to several million per sample, and it is not efficient or realistic to interpret the results of each mutation manually. Therefore, there is a need for a device that assists in the human interpretation of analysis results.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] By using the aforementioned database and analyzing the base sequence of a sample, it is possible to determine whether or not a mutation has occurred in the sample's base sequence. However, this information alone is not sufficient to easily determine whether a mutation in the base sequence directly affects a disease (for example, a driver mutation for cancer). This is because many other factors besides the mutation in question must be considered in order to determine whether a mutation in the base sequence directly affects a disease. However, no analysis has been conducted to determine the extent to which a mutation in the base sequence of a sample may affect the development of disease, taking into account such a wide range of factors.

[0006] Therefore, the applicant has filed a patent application for a technology to realize an analytical device that indicates the degree to which mutations in the base sequence are likely to affect the onset and progression of a disease (see specification, International Application No. PCT / JP2020 / 037499).

[0007] The present invention aims to more accurately indicate the degree to which mutations in the base sequence may affect the onset and progression of disease. [Means for solving the problem]

[0008] An information processing device according to one aspect of the present invention, which solves the above problems, is an information processing device for selecting a target sequence mutation in the base sequence of a subject that poses a harmful risk, and comprises: a filtering unit that sequences one or more sequence mutations identified by sequencing the nucleic acid contained in the subject and classifies them into categories according to the degree of harmful risk based on one or more classification criteria; and a control unit that classifies base sequences containing sequence mutations to which the category to which they belong is known into categories according to the degree of harmful risk based on at least one of the classification criteria and compares the classification result with the category to which they belong.

[0009] An information processing method according to one aspect of the present invention is a method for selecting a target sequence mutation in a subject that poses a harmful risk in its base sequence, comprising: a filtering step of classifying one or more sequence mutations identified by sequencing the nucleic acids contained in the subject into categories according to the degree of harmful risk based on one or more classification criteria; and a control step of classifying base sequences containing sequence mutations to which the category to which they belong is known into categories according to the degree of harmful risk based on at least one of the classification criteria, and comparing the results of the classification with the category to which they belong.

[0010] An information processing program according to one aspect of the present invention is configured to cause a computer to function as the above-mentioned information processing device. [Effects of the Invention]

[0011] According to the present invention, it is possible to more accurately indicate the degree to which mutations in the base sequence are likely to affect the onset and progression of a disease. [Brief explanation of the drawing]

[0012] [Figure 1] This is a block diagram showing an example configuration of an information processing device according to one embodiment of the present invention. [Figure 2] This is a functional block diagram showing examples of the functions of an information processing device relating to one embodiment of the present invention. [Figure 3] This is a functional block diagram showing an example of a filtering unit of an information processing device according to one embodiment of the present invention. [Figure 4] This is an explanatory diagram illustrating an example of base sequence information input to an information processing device according to one embodiment of the present invention. [Figure 5] This is a functional block diagram showing an example of a filter processing unit of an information processing device according to one embodiment of the present invention. [Figure 6] This is an explanatory diagram showing an example of output information output by an information processing device according to one embodiment of the present invention. [Figure 7] This is a functional block diagram showing an example of the control unit of an information processing device according to one embodiment of the present invention. [Figure 8] This is a flowchart showing an operation example of a filtering unit of an information processing apparatus according to an embodiment of the present invention. [Figure 9] This is a flowchart showing an operation example of a filter processing unit of an information processing apparatus according to an embodiment of the present invention. [Figure 10] This is a flowchart showing an operation example of a control unit and an adjustment unit of an information processing apparatus according to an embodiment of the present invention. [Figure 11] This is a functional block diagram showing an example of a filter processing unit of an information processing apparatus according to a second embodiment of the present invention. [Figure 12] This is a flowchart showing an operation example of a filter processing unit of an information processing apparatus according to a second embodiment of the present invention.

Mode for Carrying Out the Invention

[0013] Hereinafter, an embodiment of the present invention will be described in detail with reference to the accompanying drawings. However, this embodiment is an example, and the present invention is not limited thereto.

[0014] The first embodiment of the present invention will be described while referring to the drawings.

[0015] The information processing apparatus 1 is an information processing apparatus 1 that selects a target sequence mutation with a harmful risk on a base sequence, and determines the nucleic acid contained in an individual or a specimen (hereinafter also referred to as a subject) to be the target of information processing. It has a filtering unit 2 that classifies one or more sequence mutations identified by sequencing into each category according to the degree of harmful risk based on one or more classification criteria. In addition, the information processing apparatus 1 classifies a base sequence containing a sequence mutation whose category to belong is known into each category according to the degree of harmful risk based on at least one of the classification criteria, and compares the result of the classification with the category to which it should belong. It has a control unit 3. The filtering unit 2 and the control unit 3 of this information processing apparatus 1 will be described in detail later.

[0016] As used herein, "sequence variation" refers to the state of a base sequence variation, including the position and type of the variation. The sequence variation may be, for example, a single nucleotide variation, or may be a structural variation such as a chromosomal translocation spanning multiple genes.

[0017] Information including information representing the sequence variation is referred to as "base sequence information". As information representing the sequence variation, the base sequence information may include information indicating what base or base sequence the original base or base sequence has mutated to at a position where the variation has occurred (such as the position on the chromosome when compared with the reference genomic information (for example, information indicating which base it is from one side of the reference base sequence)). The reference genomic information is, for example, genomic information necessary for NGS analysis, and in humans, examples include GRCh38 (hg38) and GRCh37 (hg19). In addition, the base sequence information may include information extracted by sequence alignment as information representing the sequence variation.

[0018] Also, the base sequence information may be information obtained by sequencing a base sequence using a next-generation sequencer or the like. The base sequence may be a nucleic acid obtained from a subject or may be artificially synthesized. The base sequence information may include, as information obtained by sequencing, for example, files in FASTQ format, SAM (Sequence Alignment Map) format, or BAM format.

[0019] The harmful risk herein means the possibility of developing a disease including cancer. For example, a sequence variation with a harmful risk means that a disease such as cancer may occur due to the variation of the base sequence, and a sequence variation without a harmful risk means a variation of the base sequence without such a possibility. The sequence variation for the purpose of selection by the information processing apparatus 1 is particularly referred to as "target sequence variation".

[0020] Figure 1 is a block diagram showing the schematic configuration of the information processing device 1. As shown in Figure 1, the information processing device 1 comprises a control unit 11, a storage unit 12, a communication unit 13, a display unit 14, an operation reception unit 15, and a drive 16. Each component is connected to the others via a bus 18 so as to be able to communicate with each other.

[0021] The control unit 11 is equipped with a CPU (Central Processing Unit) and, according to the program, controls each component and performs various calculation processes.

[0022] The memory unit 12 includes a ROM (Read Only Memory) for pre-storing various programs and data, a RAM (Random Access Memory) for temporarily storing programs and data as a working area, and a hard disk for storing various programs and data.

[0023] The communication unit 13 communicates with other devices (for example, an information processing device for a terminal that displays analysis results, not shown in the diagram) via a network N including the Internet.

[0024] The display unit 14 consists of a display such as an LCD, a speaker, etc., and outputs various information as images and sounds.

[0025] The operation reception unit 15 is equipped with a touch sensor, a pointing device such as a mouse, a keyboard, etc., and accepts various user operations. The display unit 14 and the operation reception unit 15 may be configured as a touch panel by superimposing a touch sensor (operation reception unit 15) onto the display surface of the display unit 14. The operation reception unit 15 may also have a drive 16.

[0026] A removable media 17, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, may be appropriately mounted in the drive 16. Programs read from the removable media 17 by the drive 16 are installed in the storage unit 12 as needed.

[0027] Furthermore, the removable media 17 can store various types of data stored in the storage unit 12, just as the storage unit 12 does.

[0028] Through the cooperation of various hardware and software components of the information processing device 1 shown in Figure 1, various processes can be executed.

[0029] Figure 2 is a block diagram showing the functional configuration of the control unit 11 of the information processing device 1 according to this embodiment. As shown in Figure 2, the control unit 11 of the information processing device 1 functions as a filtering unit 2, a control unit 3, and an adjustment unit 4 by reading a program and executing processing.

[0030] <Filtering section> The filtering unit 2 classifies one or more sequence mutations identified by sequencing the nucleic acids contained in the subject into categories corresponding to the degree of harmful risk, based on one or more classification criteria. Figure 3 is a block diagram showing an example of the functional configuration for executing various processes related to the filtering unit 2 in the information processing device 1. As shown in Figure 3, the filtering unit 2 includes a first data receiving unit 21, a first setting receiving unit 22, a first filter processing unit 23, a category determination unit 24, and an analysis result output unit 25. (Data Reception Unit 1) The first data receiving unit 21 receives base sequence information that includes one or more sequence mutations identified by sequencing the nucleic acids contained in the subject. Hereinafter, the base sequence information received by the first data receiving unit 21 will also be referred to as the first base sequence information. In addition to information representing sequence mutations, the first base sequence information may also include sample identification information that identifies the individual subject to information processing and the sample obtained from that individual.

[0031] Figure 4 shows an example of the configuration of the first base sequence information received by the first data receiving unit 21 in the information processing device 1 shown in Figure 3.

[0032] As shown in Figure 4, the first nucleotide sequence information includes, for each sequence mutation (each row in the figure), at least the chromosome number (Chr) in which the sequence mutation was found, the start position (Start), the end position (End), the original nucleotide sequence (Ref), the extracted mutated nucleotide sequence (Alt), and the proportion of the mutated nucleotide sequence (allelic frequency: AF). In addition to these, the nucleotide sequence information received by the second data receiving unit 31, described later, includes information about the category to which it should belong, also described later.

[0033] In the first nucleotide sequence information of this example, for each sequence variation (each row in the figure), quality-related indicators such as depth and the number of sequence variations (AltCount) are associated with this information. Note that the length of the nucleotide sequence may be "1" (in this case, the nucleotide sequence information represents one of the bases A, T, C, or G).

[0034] Furthermore, the first base sequence information may include information about the individual's case (such as disease name, treatment history, and tumor percentage).

[0035] Furthermore, the first data receiving unit 21 may accept information (time-series information) on base sequences extracted from the same subject at different time points (there may be multiple time points). In this case, the first data receiving unit 21 may accept input of the time-series of base sequence information to be analyzed.

[0036] (First setting acceptance section) The first setting reception unit 22 receives settings for analyzing the base sequence information received by the first data reception unit 21. These settings include, for example, the setting of the type of filter to be used in the filter processing unit described later, and the setting of classification criteria for each filter.

[0037] (Filtering process) In this embodiment, the filter processing unit performs an assessment of the degree of adverse risk based on various information that influences the interpretation of the analysis results of base sequence mutations. The result of this assessment of the degree of adverse risk is represented by one of the categories MYC1 to MYC4 described later.

[0038] Here, the information that influences the interpretation includes (1) supplementary information about the mutation obtained during the analysis, and (2) information related to the mutation listed in literature and databases. Of these, (1) supplementary information about the mutation obtained during the analysis includes (a) information on detection accuracy and reliability (e.g., the probability that the mutation is not a detection error), (b) allele frequency of the mutation (e.g., an indicator related to the proportion of the total cell population that has the same mutation), and (c) time-series information (e.g., whether the mutation has been repeatedly detected in samples from the same case at other time points).

[0039] Furthermore, (2) information related to the mutation included in the literature and databases may include information indicating whether the mutation is described as a driver mutation for a disease (or how frequently it is described). If the mutation is also registered in an SNP (single nucleotide polymorphism) database, the literature and database may include information indicating the allele frequency of the mutation and how frequently it has been reported as an SNP in the relevant ethnic group. In addition, as a functional prediction, the literature and database may include information indicating whether the mutation affects the three-dimensional structure or function of the encoded protein, for example, whether it has been shown or predicted through experiments to be involved in the pathogenesis of cancer.

[0040] (First filter processing unit) The first filtering unit 23 classifies the sequence variations contained in the base sequence information received by the first data receiving unit 21 into one of the categories MYC1, MYC2, MYC3, and MYC4, which correspond to the degree of harmful risk, based on one or more predetermined classification criteria. A detailed example of the configuration of the first filtering unit 23 will be described later with reference to Figure 5.

[0041] Here, MYC1 and MYC2 are categories with a high risk of adverse effects. For example, MYC1 and MYC2 are likely to have driver mutations in their base sequence. MYC1 has a higher risk of adverse effects than MYC2, indicating a higher probability of it being a true driver mutation.

[0042] MYC3 is a category with a lower risk of harm than MYC1 and MYC2. For example, MYC3 indicates that a mutation in the base sequence is unlikely to be a driver mutation (and therefore will not be treated as a candidate driver mutation). In other words, MYC3 indicates that a sequence mutation has been evaluated as a non-harmful mutation.

[0043] MYC4 is a category with a lower risk of adverse effects than MYC3. For example, MYC4 is a category that indicates a mutation in the base sequence that is considered to have almost zero potential to be a driver mutation, or a mutation in a known SNP or region prone to error.

[0044] Figure 5 is a block diagram showing an example of the detailed functional configuration of the first filter processing unit 23. In Figure 5, the first filter processing unit 23 is equipped with a basic filter 231, a time series filter 232, a database filter 233, a function prediction filter 234, and a quality filter 235.

[0045] <Basic Filters> The basic filter 231 sets a category (e.g., MYC4) to indicate that the sequence mutation being analyzed is benign if it can be determined to be benign. If the basic filter 231 cannot determine that the sequence mutation being analyzed is benign, it sets a category (e.g., MYC3) to indicate that it poses a risk of harm and is not a benign mutation.

[0046] Cases in which a mutation can be judged as benign include those where the overlap between the nucleotide sequence of a known mutation that causes cancer, etc., and the nucleotide sequence corresponding to the sequence mutation is relatively short; where the region where the mutation represented by the sequence mutation is located is an intron region; where the sequence mutation is registered in a database that accumulates mutations without abnormalities, such as an SNP database; or where the sequence mutation can be judged as benign based on the GDI (Gene Damage Index).

[0047] Here, GDI is an indicator that shows how much damage has accumulated in healthy individuals for each gene. It indicates that even if a gene is significantly damaged (shows diversity) from person to person, it may not be considered to pose a harmful risk due to mutation.

[0048] The basic filter 231 accepts from the first setting acceptance unit 22 at least one of the following settings: a threshold for the length of the overlap between the nucleotide sequence of a known mutation that causes cancer, etc., and the mutated nucleotide sequence corresponding to the sequence mutation; information identifying a database for determining whether or not it is an SNP; and a parameter for each database (which is compared with a benign judgment threshold that serves as a criterion for determining whether or not it is benign, or a value registered in the database as the probability of it being an SNP). Based on the accepted settings, the basic filter 231 determines whether or not the sequence mutation to be analyzed is benign.

[0049] For example, basic filter 231 sets a category indicating that a sequence mutation is benign if it is located in a region called a segmental duplication. A segmental duplication is a contiguous region of chromosomes between 10 and 300 kb in which a gene is duplicated in an adjacent region or on a completely separate genome during the evolution of vertebrates. If a sequence mutation is located in a segmental duplication, it is considered a detection error that occurred during the mapping of the sequence result to the reference, and is likely to be a false positive. Therefore, if a sequence mutation is located in a segmental duplication region, it is treated as a benign mutation. Specifically, if a sequence mutation is located in this segmental duplication region and the index for that segmental duplication region exceeds a threshold, it is likely to be an error, and a category indicating a benign mutation is set. Basic filter 231 also sets a category indicating a benign mutation if the region in which the mutation represented by the sequence mutation is located is an intron region.

[0050] Furthermore, even if the above two conditions are not met, the basic filter 231 may set a category indicating a benign mutation based on the results of searching the specified SNP database. For example, if the basic filter 231 finds that the mutation represented by the sequence mutation is registered in the SNP database through a search, and the value registered as the probability of that SNP exceeds a predetermined benignity threshold for that SNP database, it sets a category indicating a benign mutation.

[0051] Furthermore, even if the conditions described above are not met, the basic filter 231 refers to the GDI of the gene in which the sequence mutation exists and sets a category indicating that it is a benign mutation if it is greater than a predetermined GDI threshold.

[0052] This makes it possible for the information processing device 1 to pre-screen out genes that cannot (or have a sufficiently low probability of) become driver mutations for cancer.

[0053] Furthermore, this basic filter 231 may accept a setting from the first setting acceptance unit 22 indicating which of several conditions for determining benignity should be used (or whether to not use any conditions and skip processing by setting the category to MYC3 for all sequence mutations without operating as the basic filter 231).

[0054] In this example, the basic filter 231, when used, will only determine whether or not the specified conditions are met.

[0055] <Time series filter> When the basic filter 231 passes processing (MYC3 is set), the time series filter 232 refers to the sequence mutation information contained in the time series information corresponding to the sequence mutation to be analyzed, and determines whether the same mutation was present in the time series information extracted at different time points.

[0056] The time series filter 232 uses the sequence variant to be analyzed and the corresponding sequence variant included in the time series information. If the same variant exists, it sets a category (for example, by subtracting "1" as a first predetermined amount from the current category) and passes the processing to the quality filter 235. The first predetermined amount is, for example, the minimum value that is subtracted or added to the category related to the sequence variant in a single operation. In this example, the basic filter 231 has passed the processing, so the initial category is MYC3. If the time series filter 232 determines that there is a variant that should be considered problematic, it will subtract "1" as a first predetermined amount from MYC3 and set the category to MYC2.

[0057] On the other hand, the time series filter 232 uses the sequence variant to be analyzed and the corresponding sequence variant included in the time series information. If no identical variants exist, it sets the category as is (in this case, the initial category is MYC3, so it is set to MYC3) and passes the processing to the database filter 233.

[0058] The time series filter 232 may also receive threshold settings from the first setting acceptance unit 22 regarding depth, other sequence quality, mutation allele frequency, etc. For example, if the depth of the corresponding sequence mutation included in the time series information does not exceed the threshold set here (e.g., "20"), the time series filter 232 does not determine whether or not the same sequence mutation was present, but sets the category as is (in this case, the initial category is MYC3, so it is set to MYC3) and passes the processing to the database filter 233.

[0059] Furthermore, in this embodiment, if the first base sequence information received by the first data receiving unit 21 does not contain time-series information, the time-series filter 232 may set the category as is (in this case, since the initial category is MYC3, it will be set to MYC3) without determining whether or not the same sequence mutation exists, and pass the processing to the database filter 233.

[0060] Furthermore, if a setting not to use the time-series filter 232 is input from the first setting acceptance unit 22, the time-series filter 232 does not determine whether or not the same sequence mutation exists, but sets the category as is (in this case, the original category is MYC3, so it is set to MYC3) and passes the processing to the database filter 233.

[0061] <Database Filter> The database filter 233 checks whether the sequence mutation to be analyzed is registered in a database (e.g., the COSMIC Cancer Database) that stores information on predetermined problematic mutations by communicating with the database server. If the sequence mutation is registered in the database, it is categorized as a problematic mutation (having a harmful risk) (for example, by subtracting "1" as the first predetermined amount from the current category) and passed to the quality filter 235. To give an example of the series of processes by each filter, if the basic filter 231 passes the processing of the sequence mutation to be analyzed because it has a harmful risk, and the time-series filter 232 passes the processing with the category unchanged, then if the database filter 233 also determines that it has a harmful risk, the database filter 233 subtracts "1" as the first predetermined amount from MYC3, sets the category to MYC2, and then passes the processing to the quality filter 235.

[0062] Furthermore, if the sequence mutation to be analyzed is not registered in the database containing information on the mutations that should be considered problematic, the database filter 233 leaves the category unchanged and passes the process to the function prediction filter 234. In this example, the category remains MYC3.

[0063] Furthermore, the database filter 233 receives a setting from the first setting acceptance unit 22 regarding which database to use as the database for storing information on the mutations that are the subject of the above-mentioned problem.

[0064] In this configuration, instructions may be given to use multiple databases. In this case, the database filter 233 will categorize the sequence mutation to be analyzed as a problematic mutation if it is registered in any of the databases that store information on the problematic mutations mentioned above.

[0065] <Function Prediction Filter> The functional prediction filter 234 refers to programs (including machine learning programs) that evaluate and predict the harmful risk of mutations, as well as databases that publish evaluation results and predicted values ​​of harmful risk. If the sequence mutation to be analyzed is registered in the program or database as having a harmful risk, it sets a category for the mutation (for example, subtracting "1" as the first predetermined amount from the current category) and passes the processing to the quality filter 235.

[0066] Well-known programs that assess the harmful risks of mutations include SIFT, PolyPhen2, SnpEff, and VEP. Some of these programs and databases also use scoring thresholds or multi-stage evaluations to determine the presence or absence of harmful risks. For example, even if these programs or databases are still in the decision-making stage regarding the presence or absence of harmful risks, this functional prediction filter 234 sets a category for those with harmful risks (for example, by subtracting "1" as the first predetermined amount from the current category) and passes the processing to the quality filter 235.

[0067] Furthermore, the functional prediction filter 234 may, by referring to the programs and databases described above, predict whether deletions or duplications of promoters involved in important gene expression, deletions or insertions that lead to splicing abnormalities in important genes, or deletions or insertions of noncoding RNAs important for the regulation of important gene expression will occur. If these programs are at the stage of determining whether or not there is a risk of harm, the functional prediction filter 234 may set a category as having a risk of harm (for example, by subtracting "1" as the first predetermined amount from the current category) and pass the processing to the quality filter 235.

[0068] To give an example of the series of processes performed by each filter, if the basic filter 231 determines that the sequence mutation to be analyzed has a harmful risk and passes the process, the time series filter 232 passes the process with the category unchanged, and the database filter 233 also passes the process with the category unchanged, and then the functional prediction filter 234 determines that there is a harmful risk, the functional prediction filter 234 will subtract "1" as the first predetermined amount from MYC3 at that time, set the category to MYC2, and then pass the process to the quality filter 235.

[0069] Furthermore, this predictive filter 234 refers to a database that evaluates the harmful risk of mutations. If the mutation related to the sequence mutation to be analyzed is not registered in the database as having a harmful risk (or if it is registered but is unknown, or is registered as benign or presumed to be benign), it sets the category as is and passes the processing to the quality filter 235. In this example, the category remains MYC3.

[0070] Furthermore, this predictive filter 234 also accepts the setting for which database to use from the first setting acceptance unit 22.

[0071] <Quality Filter> The quality filter 235 evaluates the quality of sequencing by using indicators such as the sequencing depth of the sequence mutations to be analyzed, the quality score for each base (e.g., Phred quality score), the mapping quality score to the reference genome, the statistical value of statistical tests (such as Fisher's test) in mutation calling between cancer cells and normal cells, and the degree of bias towards one side of the read sequences supporting mutations in paired-end reads that read the base sequence from both sides. In addition to the depth, there are widely known indicators for this quality, such as the number of sequence mutations, and the quality filter 235 evaluates the quality by combining these (or by accepting such combinations from the first setting acceptance unit 22 and according to the combination of accepted indicators). When multiple indicators are combined, the quality filter 235 will determine that the quality is sufficient only if the condition that the quality is sufficiently high is met by all indicators.

[0072] If the quality filter 235 determines, based on this evaluation, that the sequencing quality of the sequence mutation to be analyzed is sufficient (sufficiently high), it sets a category (for example, by subtracting "1" as a first predetermined amount from the current category) and outputs that category to the category determination unit 24. If the quality filter 235 does not determine that the sequencing quality of the sequence mutation to be analyzed is sufficient (sufficiently high), it leaves the category as is and outputs that category to the category determination unit 24.

[0073] Furthermore, at least one of the classification criteria set for each filter can be changed or selected. Additionally, it is possible to run the filtering unit 2 and the control unit 3 after changing or selecting at least one of the classification criteria. This allows the information processing device 1 to more accurately determine the harmful risks associated with sequence mutations.

[0074] (Category determination section) The category determination unit 24 determines a category value representing the degree of harmful risk for each sequence mutation, according to the category (one of MYC1 to MYC4) for each of the one or more sequence mutations output by the filtering unit. For each of the multiple sequence mutations, the category determination unit 24 generates information (hereinafter referred to as "analysis result information") associating each category value and provides it to the analysis result output unit 25.

[0075] Note that the category value representing the degree of this harmful risk may be a newly calculated value based on MYC1 to MYC4, but for the sake of explanation, MYC1 to MYC4 will be used as is.

[0076] (Analysis result output section) The analysis result output unit 25 outputs the analysis result information by outputting it from the display unit 14 (e.g., a display) in Figure 1, or by transmitting it from the communication unit 13 to other devices not shown.

[0077] Figure 6 shows an example of the structure of the analysis result information output from the information processing device 1. As shown in Figure 6, the analysis result information includes, for each sequence mutation (each row in the figure), at least the chromosome number (Chr) where the base sequence of the sequence mutation is located, the start position (Start), the end position (End), the original base sequence (Ref), the sequence mutation (Alt), and the category value (MYC).

[0078] In addition to the analysis results information in the example in Figure 6, record information R related to the judgment is also associated with each sequence variation (each row in the figure).

[0079] The record information R related to the judgment refers to information that describes how the filters used in the analysis of the target sequence mutations within the filtering processing unit were classified (parameter settings for each filter, judgment content based on classification criteria, etc.).

[0080] As described above, the mutations in the base sequence information received by the first data receiving unit 21 are classified into four levels, MYC1 to MYC4, indicating a risk of harmful effects. This allows users, such as medical specialists, to efficiently identify mutations with a high risk of harmful effects, such as true driver mutations, from among the many existing mutations (e.g., tens of thousands to hundreds of millions). For example, medical specialists can focus on sequence mutations classified as MYC1 or MYC2 to identify true driver mutations.

[0081] <Control Unit> On the other hand, in order to improve the reliability of the classification performed by the information processing device 1, it is necessary to confirm whether the classification process described above is being performed appropriately. Therefore, the information processing device 1 according to this embodiment has a control unit 3 that classifies a base sequence containing a sequence mutation to which the category to which it should belong is known, into each category based on at least one of the classification criteria described above, and compares the result of the classification with the category to which it should belong. If the result of the comparison matches, it can be confirmed that the classification process of the information processing device 1 is being performed appropriately. On the other hand, if the result of the comparison does not match, it can be confirmed that the classification process performed by the information processing device 1 may not be being performed appropriately.

[0082] The control unit 3 according to this embodiment classifies sequence mutations, which include known base sequences, into categories corresponding to the degree of harmful risk based on at least one of the classification criteria, and compares the result of the classification with the category to which it should belong.

[0083] Figure 7 is a block diagram showing an example of a functional configuration for executing various processes related to the control unit 3 in the information processing device 1. As shown in Figure 7, the control unit 3 includes a second data receiving unit 31, a second setting receiving unit 32, a second filter processing unit 33, a comparison unit 34, and a comparison result output unit 35.

[0084] (Second Data Reception Department) The second data receiving unit 31 receives base sequence information (hereinafter also referred to as the second base sequence information) that includes information representing a base sequence containing one or more sequence mutations to which the category to which it should belong is known. Here, a base sequence containing a sequence mutation to which the category to which it should belong is known includes a sequence mutation to which the category to which it should belong is known, and a base sequence to which the category to which it should belong is known, but does not have a mutation. The category to which it should belong is one of MYC1, MYC2, MYC3, and MYC4, which are categories corresponding to the degree of harmful risk described above.

[0085] The structure of the second base sequence information accepted by the second data acceptance unit 31 is equivalent to the structure of the first base sequence information accepted by the first data acceptance unit 21 shown in Figure 4, so no explanation is provided. However, the base sequence information accepted by the second data acceptance unit 31 includes information about the category to which each sequence mutation should belong.

[0086] In one embodiment, the base sequences containing sequence mutations to which the category to which they belong is known may be two or more different types to which they belong. By filtering two or more base sequences to which the category to which they belong in the control unit 3, the comparison results by the comparison unit 34 described later become more detailed, and the accuracy of the filtering process can be understood in more detail.

[0087] Furthermore, in one embodiment, two or more nucleotide sequences belonging to different categories classified by the control unit 3 may include sequence mutations that cause a specific disease (harmful risk) and nucleotide sequences that do not cause a specific disease (no harmful risk). Here, nucleotide sequences that do not cause a specific disease include sequence mutations without harmful risk and nucleotide sequences without mutations. For example, a nucleotide sequence containing a sequence mutation that causes a specific cancer and a nucleotide sequence that does not contain that specific cancer-causing sequence mutation are processed by the control unit 3. This makes it possible to determine whether the judgment function of the second filtering unit 33 is working correctly in both cases where there is a harmful risk and where there is no harmful risk.

[0088] Furthermore, the second base sequence information may include files in formats such as VCF (Variant Call Format), FASTQ, SAM (Sequence Alignment Map), and BAM (Binary Alignment Map), which are output from next-generation sequencers. The VCF format is a file format used when saving base variation data, and when sequencing data is mapped to a reference sequence, it contains information such as the bases on the reference sequence and the bases on the sequencing data mapped to it. FASTQ format files contain the base sequence and the quality of the base call for each base. SAM format files show the result of mapping the FASTQ read sequence to a reference sequence, and BAM format files are a compressed version of SAM format that is easier for computers to process.

[0089] These files may represent base sequences containing arbitrary sequence mutations, and by providing such files to the control unit 3, it becomes possible to classify the arbitrary sequence mutations more accurately. More specifically, for example, if an arbitrary sequence mutation is a hotspot in a gene where mutations are concentrated, by providing the control unit 3 with the above-mentioned file containing information about the hotspot, it becomes possible to classify the mutations in the hotspot more accurately. This also allows the filtering unit 2 to classify the mutations in the hotspot more reliably.

[0090] Furthermore, in one embodiment, if the target sequence mutation in the nucleotide sequence of a subject that poses a risk of adverse effects is a driver mutation for a specific disease, then two or more nucleotide sequences belonging to different categories may include a sequence mutation that is a driver mutation for the specific disease and a nucleotide sequence that is not a driver mutation for the specific disease. For example, if the target sequence mutation in the nucleotide sequence of a sample obtained from a patient is a driver mutation for a certain leukemia, the control unit 3 processes the sequence mutation that is a driver mutation for the leukemia and the nucleotide sequence that does not contain the driver mutation for the leukemia. This makes it possible to determine whether the information processing device 1 has accurately classified the driver mutation for the specific disease.

[0091] (Second setting acceptance section) The second setting reception unit 32 receives settings for analyzing the second base sequence information received by the second data reception unit 31. These settings include, for example, the setting of what classification criteria to use for filtering in the second filtering processing unit 33, which will be described later.

[0092] In the control unit 3, similar to the filtering unit 2, the base sequence information received by the second data reception unit 31 is evaluated by the second filtering unit 33 based on various information that influences the interpretation of the mutation analysis results, including the potential for harmful risks (e.g., the possibility of a driver mutation). This evaluation result, like the evaluation result by the filtering unit 2, is classified into one of categories MYC1 to MYC4. The evaluation (classification) method and the information that influences the interpretation by the second filtering unit 33 are the same as those in the filtering unit 2, so an explanation is omitted.

[0093] (Second filter processing unit) The second filtering unit 33 classifies the base sequences containing sequence mutations to which the category to which they belong is known, as included in the base sequence information received by the second data receiving unit 31, into one of the categories MYC1, MYC2, MYC3, and MYC4, which correspond to the degree of harmful risk, based on at least one classification criterion. MYC1, MYC2, MYC3, and MYC4 are as described in the section on the first filtering unit 23. Furthermore, for the sake of explanation in this specification, the second filtering unit 33 is described separately from the first filtering unit 23, but the classification criteria and filters used in the second filtering unit 33 may be the same as those used in the first filtering unit 23, and the second filtering unit 33 and the first filtering unit 23 may be a common filtering unit.

[0094] (Contrasting section) The comparison unit 34 compares each mutation in the base sequence information received by the second data reception unit 31 with a category (one of MYC1 to MYC4) output by the second filtering unit 33 and a category (one of MYC1 to MYC4) corresponding to the degree of known harmful risk. The comparison unit 34 also provides the comparison results for each mutation to the comparison result output unit 35.

[0095] Note that the value representing this comparison result may be a newly calculated value based on MYC1 to MYC4, but for the sake of explanation, MYC1 to MYC4 will be used as is.

[0096] (Contrast result output section) The comparison result output unit 35 outputs information regarding the comparison results from the comparison unit 34 by outputting it from the display unit 14 (e.g., a display) in Figure 1, or by transmitting it from the communication unit 13 to other devices not shown.

[0097] <Adjustment part> In one embodiment, the information processing device 1 may have an adjustment unit 4 that adjusts the classification criteria in the filtering unit 2 and / or the control unit 3 and / or the classification results in the filtering unit 2 based on the comparison results in the control unit 3. By having the adjustment unit 4, the information processing device 1 can perform calibration of the criteria in the filtering process, and thus more accurately classify the degree of harmful risk of mutations in the base sequence of the subject.

[0098] For example, if the comparison in the comparison unit 34 of the control unit 3 results in a difference between the category output by the filter processing unit and the category corresponding to the known degree of harmful risk for a given sequence mutation, then the degree of harmful risk of the sequence mutation is not accurately classified by the filter processing unit. In such cases, the adjustment unit 4 calibrates the classification criteria of each filter in the filter processing unit based on the comparison results, so that the category output by the filter processing unit matches the known category.

[0099] Furthermore, if the adjustment unit 4, based on the comparison results of the comparison unit 34 in the control unit 3, finds that the category output by the filter processing unit for a given sequence mutation does not match the category corresponding to the degree of known harmful risk, it may choose not to adopt the classification result of the filtering unit 2 and instead perform the classification process by the filtering unit 2 again after the adjustment by the adjustment unit 4 is completed. The adjustment unit 4 may also have a function to display the details of the problem that occurred as an error message based on the comparison results of the comparison unit 34 in the control unit 3. For example, it may display at which stage of the filtering process the problem occurred.

[0100] Furthermore, the sequence mutations classified by the control unit 3 may be those obtained by sequencing a standard nucleic acid composition containing a sequence mutation to which the category to which it should belong is known. That is, a standard nucleic acid composition containing a sequence mutation to which the category to which it should belong is sequenced using a sequencing device such as a next-generation sequencer, and the information of the sequencing result is provided to the control unit 3 for processing. The control unit 3 classifies the information obtained from sequencing the standard composition and compares the results of this classification with the known category to which the standard composition should originally belong, thereby confirming whether the sequencing conditions (for example, sequencing using a sequencing device or pre-processing steps for sequencing) were correct.

[0101] In this case, the sequencing conditions for the standard composition and the sequencing conditions for the nucleic acids contained in the subject may be the same. For example, the sequencing conditions for the standard composition using a next-generation sequencer may be the same as the sequencing conditions for the nucleic acids contained in the subject derived from a patient, etc. As described above, by providing the control unit 3 with the sequencing results of a standard composition whose category to which it should belong is known, it is possible to confirm whether the sequencing conditions were correct or not. Therefore, by making the sequencing conditions for the nucleic acids contained in the subject and the standard composition the same, it is also possible to confirm whether the sequencing conditions for the nucleic acids contained in the subject were correct or not.

[0102] The following will explain the adjustment of each filter in the adjustment unit 4 with specific examples, but the adjustments in the adjustment unit 4 are not limited to these.

[0103] Examples of basic filter adjustments The adjustment unit 4 can adjust the threshold length of the overlapping portion between the nucleotide sequence of a known mutation that causes cancer, etc., and the nucleotide sequence corresponding to the sequence mutation in the basic filter 231. For example, the basic filter 231 sets a category indicating a benign mutation if the sequence mutation is located in this segmental overlapping region and the index of the segmental overlapping region exceeds a threshold, as this indicates a high possibility of error. This threshold can be adjusted. This allows the classification criteria used by the basic filter 231 to set categories to be adjusted.

[0104] Furthermore, the adjustment unit 4 can change the SNP database used in the basic filter 231. The adjustment unit 4 can also configure the basic filter 231 to use multiple SNP databases. In addition, the basic filter 231 sets a category indicating a benign mutation if the mutation represented by the sequence mutation is registered in the SNP database and the value registered in the database as the probability of being an SNP exceeds the benignity threshold in the basic filter 231. The adjustment unit 4 can change the benignity threshold in the basic filter 231. This adjustment also allows the classification criteria for setting the category of a mutation as benign to be adjusted.

[0105] Furthermore, the basic filter 231 refers to the GDI of the gene in which the sequence mutation exists and sets a category indicating a benign mutation if it is greater than a predetermined GDI threshold. The adjustment unit 4 can also adjust this GDI threshold. Through this adjustment, the adjustment unit 4 can also adjust the classification criteria by which the basic filter 231 sets categories.

[0106] Furthermore, the adjustment unit 4 can also change the conditions to be used from among the multiple conditions for determining benignity predetermined by the first setting acceptance unit 22, etc. (or whether to skip processing by setting the category to MYC3 for all sequence mutations without using any conditions and without operating as the basic filter 231).

[0107] Example of adjusting a time series filter The time series filter 232 uses the sequence variant to be analyzed and the corresponding sequence variant included in the time series information, and sets a category for the variant that should be considered problematic when the same variant exists. Here, for example, if multiple time series information is included, the adjustment unit 4 can be adjusted to use a different time series information than the one used by the control unit 3 for the time series filter 232.

[0108] Furthermore, if the time-series filter 232 has pre-configured threshold settings for depth, other sequence quality, mutation allele frequency, etc., the adjustment unit 4 can adjust these settings. For example, it is possible to adjust the categories classified by the time-series filter 232 by changing the depth threshold for corresponding sequence mutations included in the time-series information.

[0109] Examples of adjusting database filters The database filter 233 checks whether the sequence mutation to be analyzed is registered in a database that stores information on mutations by sending information about the sequence mutation to the database server. If it is registered, it sets the category as containing a mutation that should be considered problematic. The adjustment unit 4 can change the database used by the database filter 233. This allows the adjustment unit 4 to adjust the categories set by the database filter 233.

[0110] Examples of adjusting the feature prediction filter. The functional prediction filter 234 refers to a program or database that evaluates the harmful risk of mutations, and if the sequence mutation to be analyzed is registered in the database as having a harmful risk, it sets the category as having a harmful risk mutation. The adjustment unit 4 can be set to refer to a different program or database than the one it referenced, thereby adjusting the category set by the functional prediction filter 234.

[0111] Examples of adjusting the quality filter The quality filter 235 evaluates the quality of sequencing of the sequence mutations to be analyzed using quality indicators such as the sequencing depth of the sequence mutations to be analyzed, the quality score for each base (e.g., Phred quality score), the mapping quality score to the reference genome, statistical tests in mutation calling between cancer cells and normal cells (e.g., Fisher's test), and statistical values ​​of the bias of supporting reads for mutations in paired-end reads. The adjustment unit 4 can adjust the categories set by the quality filter 235 by changing the evaluation criteria for these indicators that represent sequencing quality.

[0112] The adjustment method for each filter by the adjustment unit 4 has been described above. In one embodiment, the information processing device 1 may have a re-execution unit that re-executes the filtering unit 2 and the control unit 3 after the adjustment unit 4 has changed or selected at least one of these classification criteria. This makes it possible to perform classification using calibrated classification criteria and filters, thereby improving the accuracy of the classification performed by the information processing device 1.

[0113] Next, the processing of the information processing device 1 will be explained with reference to the drawings from Figure 8 onward.

[0114] Figure 8 is a flowchart illustrating an example of a series of steps in the filtering unit 2 of the information processing device 1 having the functional configuration shown in Figure 3.

[0115] In step S1, the first setting acceptance unit 22 accepts settings for analyzing base sequence information. Here, the first filter processing unit 23 also accepts settings regarding what classification criteria to use for the filter.

[0116] In step S2, the first data receiving unit 21 determines which predetermined sequence mutations to be processed from the base sequence information extracted by sequence alignment from the genetic information of the subject to be analyzed.

[0117] In step S3, the first filter processing unit 23 applies a filter to the sequence mutations to be processed and outputs the category of the mutations to be processed. Details of the filtering process in the first filter processing unit 23 will be explained separately with reference to Figure 9.

[0118] Next, in step S4, the information processing device 1 determines whether or not it has recorded categories for all sequence mutations.

[0119] If there are sequence mutations for which no category has been recorded, the result is determined as "NO" in step S4, and the process returns to step S2, and the subsequent processing is repeated. In this way, the loop processing of steps S2 to S4 "NO" is repeated, and if the categories of all sequence mutations have been recorded, the result is determined as "YES" in step S4, and the process proceeds to step S5.

[0120] In step S5, the analysis result output unit 25 generates analysis result information and outputs it by displaying it on the display unit 14 (e.g., a display) in Figure 1, or by transmitting it to another device (not shown) via the communication unit 13. This completes the analysis process.

[0121] The details of the filtering process in step S3 are explained below using the flowchart in Figure 9.

[0122] In step S31, the basic filter 231 determines whether or not there is a harmful risk for the sequence mutation to be processed, based on the conditions of the basic filter 231.

[0123] If the sequence mutation to be processed does not pose a harmful risk according to the conditions of the basic filter 231, it is determined to be "NO" in step S31, the category is set to MYC4, and the process proceeds to step S37 or step 35.

[0124] If the process proceeds to step S37, the first filter processing unit 23 outputs a category as the first filter processing unit 23. This completes the filtering process in step S3 of Figure 9, and the process proceeds to step S4. The process if the process proceeds to step S35 will be described later.

[0125] If the sequence mutation to be processed is deemed to pose a harmful risk according to the conditions of the basic filter 231, it is determined to be "YES" in step S31, the category is set to MYC3, and the process proceeds to step S32.

[0126] In step S32, the time series filter 232 determines whether the sequence mutation to be processed poses a harmful risk according to the conditions of the time series filter 232. If the sequence mutation to be processed poses a harmful risk according to the conditions of the time series filter 232, the result is determined as "YES" in step S32, the category is set to MYC2, and the process proceeds to step S35. The processing from step S35 onward will be described later.

[0127] If the sequence mutation to be processed is deemed to pose no harmful risk according to the conditions of the time series filter 232, it is determined to be "NO" in step S32, the category is set to MYC3, and the process proceeds to step S33.

[0128] In step S33, the database filter 233 determines whether or not there is a harmful risk for the sequence mutation to be processed, based on the conditions of the database filter 233.

[0129] If the sequence mutation to be processed is deemed to pose a harmful risk according to the conditions of the database filter 233, it is determined to be "YES" in step S33, the category is set to MYC2, and the process proceeds to step S35. The processing from step S35 onward will be described later.

[0130] If the sequence mutation to be processed is deemed to pose no harmful risk according to the conditions of the time series filter 232, it is determined to be "NO" in step S33, the category is set to MYC3, and the process proceeds to step S34.

[0131] In step S34, the functional prediction filter 234 determines whether or not there is a harmful risk for the sequence mutation to be processed, based on the conditions of the functional prediction filter 234.

[0132] If the sequence mutation to be processed is deemed to pose a harmful risk according to the conditions of the functional prediction filter 234, it is determined to be "YES" in step S34, the category is set to MYC2, and the process proceeds to step S35.

[0133] If the sequence mutation to be processed is deemed to pose no harmful risk according to the conditions of the functional prediction filter 234, it is determined to be "NO" in step S34, the category is set to MYC3, and the process proceeds to step S35.

[0134] In step S35, the quality filter 235 determines whether the quality is sufficient or not.

[0135] If the quality of the results of the processing in steps S31 to S34 (the filtering results of the basic filter 231, time series filter 232, database filter 233, and function prediction filter 234) is sufficient, then in step S35 it is determined to be "YES" and the process proceeds to step S36. In step S36, the quality filter 235 determines that the quality is sufficient and subtracts a first predetermined amount, "1", from the category.

[0136] If the quality of the results of the processing in steps S31 to S34 (the filtering results of the basic filter 231, time series filter 232, database filter 233, and function prediction filter 234) is insufficient, then in step S35 it is determined to be "NO", and the process proceeds to step S37.

[0137] In step S37, the first filter processing unit 23 outputs a category. This completes the filtering process in step S3 of Figure 9, and the process proceeds to step S4.

[0138] Figure 10 is a flowchart illustrating an example of a series of operations in the control unit 3 and adjustment unit 4 of the information processing device 1 having the functional configuration shown in Figure 7.

[0139] In step S1c, the second setting acceptance unit 32 accepts settings for analyzing second nucleotide sequence information relating to a nucleotide sequence containing known sequence mutations to which it should belong. Here, the second filtering processing unit 33 also accepts settings for what classification criteria to use as the basis for the filter.

[0140] In step S2c, the second data receiving unit 31 determines the base sequence to be analyzed. If there are multiple base sequences, it selects and determines the base sequence to be analyzed from among the multiple mutations. In Figure 10, the base sequence to be analyzed by the control unit 3 is shown to be a sequence mutation to which the category to which it belongs is known, but the control unit 3 can also analyze base sequences to which the category to which it should belong is known and which do not contain mutations.

[0141] In step S3c, the second filter processing unit 33 applies a filter to the sequence mutations to be processed and outputs the category of the mutations to be processed. The filtering process in the second filter processing unit 33 is the same as the filtering process in the first filter processing unit 23 explained with reference to Figure 9, so no further explanation is given.

[0142] If the second base sequence information contains multiple sequence mutations, in step S4c, the information processing device 1 determines whether or not a category has been recorded for all sequence mutations. If there are sequence mutations for which a category has not been recorded, the determination in step S4c is "NO", and the process returns to step S2c, and the subsequent processing is repeated.

[0143] In this way, the loop processing of steps S2c to S4c "NO" is repeated, and if all categories of sequence mutations are recorded as a result, step S4c is determined to be "YES", and the process proceeds to step S5c.

[0144] Next, in step S5c, the second data receiving unit 31 compares the sequence variation in the second base sequence information with the category output by the second filtering unit 33 (one of MYC1 to MYC4) and the known category to which it should belong (one of MYC1 to MYC4). If the comparison shows consistency between the category output by the filtering unit and the known category to which it should belong (for example, if they match), the unit outputs a result indicating consistency, and the processing by the control unit 3 ends.

[0145] On the other hand, if the comparison results show that the categories output by the filtering unit for sequence variations in the second base sequence information do not match the known categories to which they should belong (for example, if they do not match), then in step S6c, the adjustment unit 4 adjusts the results of the classification to each of the classification criteria or categories. Details of the adjustment method in the adjustment unit 4 are described in the section on adjustment unit 4.

[0146] After adjustment, the sequence mutations included in the second base sequence information are processed again using steps S2c to S5c. If the categories output by the filtering unit and the categories corresponding to the degree of known harmful risks are consistent, the processing by the control unit 3 is terminated. If the comparison results are not consistent, the processing using steps S2c to S6c may be repeated until consistency is achieved, at which point the processing by the control unit 3 may be terminated.

[0147] Furthermore, the filtering unit 2 may perform processing after the control unit 3 has finished processing.

[0148] Although one embodiment of the present invention has been described above, the present invention is not limited to the above-described embodiment (also referred to as the first embodiment), and any modifications, improvements, etc. that can achieve the objectives of the present invention are considered to be included in the present invention.

[0149] For example, the filter processing unit is not particularly limited to the first filter processing unit 23 and the second filter processing unit 33 shown in Figure 5, and can take various forms having different filter configurations. Below, as a second embodiment of the information processing device 1 according to the present invention, an information processing device 1 employing a third filter processing unit 43 having the configuration shown in the block diagram of Figure 11 will be described. Note that the second embodiment of the information processing device 1 has the same configuration as the first embodiment described above, except for the configuration described below (for example, the third filter processing unit 43 and the adjustment unit 4 that adjusts it), so the description of the configuration similar to the first embodiment will be omitted.

[0150] The third filtering unit 43 in the example shown in Figure 11 is useful in the following type of sequence mutation analysis. First, it is known that the fusion of two genes in a specific combination, due to chromosomal translocation or inversion, can cause the proliferation of cancer cells. For example, the BCR-ABL fusion gene, which is formed by the fusion of the BCR gene and the ABL gene due to chromosomal translocation, is known to cause the proliferation of leukemia cells.

[0151] The third filtering unit 43 includes a basic filter 231, a time-series filter 232, a fusion gene filter 236, a conserved position filter 237, a structure filter 238, and a quality filter 235.

[0152] Furthermore, for each fusion gene, a region in the memory unit 12 stores the base sequences encoding multiple combinations of candidate genes known to cause driver mutations in fusion genes formed by the fusion of two specific combinations of candidate genes. For example, the base sequences encoding the BCR gene and the ABL gene are stored in a region of the memory unit 12.

[0153] In other words, the information processing device 1 can acquire the following information and use it for information processing.

[0154] The information processing device 1 obtains the base sequences of two candidate genes that are candidate driver mutations in a fusion gene (hereinafter referred to as the first fusion gene) formed by the fusion of a specific combination of candidate genes (hereinafter referred to as the first fusion gene), for each first fusion gene. In the example in which the third filter processing unit 43 shown in Figure 11 is employed, the information processing device 1 obtains the base sequences of each of the two candidate genes contained in the multiple first fusion genes stored in the storage unit 12 from the storage unit 12 for each first fusion gene.

[0155] Furthermore, an external server (not shown) may store the base sequences encoding multiple candidate genes for the first fusion gene. The information processing device 1 may obtain the base sequences encoding two candidate genes for the first fusion gene from the external server via the communication unit 13, for each first fusion gene.

[0156] Fusion genes, formed by the fusion of a specific candidate gene with another gene, can sometimes cause the proliferation of cancer cells. For example, fusion genes formed by the fusion of the ALK gene with another gene are known to cause the proliferation of cancer cells. Memory unit 12 stores the base sequences of multiple candidate genes that are candidate driver mutations in fusion genes formed by the fusion of a specific gene with another gene (hereinafter also referred to as the second fusion gene).

[0157] The information processing device 1 obtains the base sequences of candidate genes that are candidate driver mutations in the second fusion gene formed by fusion with other genes. For example, the information processing device 1 obtains the base sequences of multiple candidate genes of the second fusion gene from the storage unit 12. The information processing device 1 may also obtain the base sequences of multiple candidate genes of the second fusion gene from an external server via the communication unit 13.

[0158] The information processing device 1 acquires conserved sequence location information, which indicates the location of conserved sequences, which are base sequences preserved between the genomes of different biological species. For example, the information processing device 1 acquires conserved sequence location information from the storage unit 12. The information processing device 1 may also acquire conserved sequence location information from an external server via the communication unit 13.

[0159] <Basic Filters> The basic filter 231 is the same as the filter processing unit shown in Figure 5, except that it does not perform processing specific to single nucleotide polymorphisms. If the basic filter 231 determines that the sequence mutation to be analyzed is benign, it sets a category (e.g., MYC4) indicating that it is a benign mutation and outputs the result to the filter set as the next filter. If the basic filter 231 does not determine that the sequence mutation to be analyzed is benign, it sets a category (e.g., MYC3) indicating that it is not a benign mutation and passes the processing to the filter set as the next filter.

[0160] The basic filter 231 receives information from the first setting acceptance unit 22 that identifies a threshold length for the overlap between the nucleotide sequence of a known mutation that causes cancer, etc., and the mutated nucleotide sequence corresponding to the sequence mutation, as well as settings for database-specific parameters (which are compared with values ​​registered as benign judgment thresholds, etc., which serve as criteria for determining whether something is benign or not), and determines whether the sequence mutation to be analyzed is benign or not based on these settings.

[0161] Specifically, the basic filter 231 sets a category indicating a benign mutation if the overlap portion between the nucleotide sequence of a known mutation that causes cancer, etc., and the mutated nucleotide sequence corresponding to the sequence mutation is shorter than a predetermined threshold length. In addition, the basic filter 231 also sets a category indicating a benign mutation if the region where the mutation is located, as represented by the sequence mutation, is an intron region.

[0162] Furthermore, even if the above two conditions are not met, the basic filter 231 searches the specified database, and if the variant represented by the sequence variant is registered in the database as a result of the search, and the value registered as the probability of that variant exceeds a predetermined benign judgment threshold for that database, it sets a category indicating that it is a benign variant.

[0163] <Time series filter> The time series filter 232 is the same as the example of the filtering process in Figure 5, except that the value subtracted from the category corresponding to the sequence mutation to be analyzed is different from that of the example of the filtering process in Figure 5, and the output destination of the category after calculation by the time series filter 232 is different from that of the example of the filtering process in Figure 5. The time series filter 232 refers to the sequence mutation information contained in the time series information corresponding to the sequence mutation to be analyzed, and determines whether the same mutation was present in the time series information extracted at different time points.

[0164] The time series filter 232 uses the sequence variant to be analyzed and the corresponding sequence variant included in the time series information to determine the category corresponding to the sequence variant to be analyzed as having a harmful risk (for example, by subtracting "2" as a second predetermined amount from the category) if the same variant exists, and then passes the processing to the structure filter 238. In this example, the basic filter 231 has passed the processing, so the initial category is MYC3. If the time series filter 232 then determines that there is a harmful risk, then "2" will be subtracted from MYC3 as a second predetermined amount to set the category to MYC1. The second predetermined amount is a value greater than the first predetermined amount.

[0165] On the other hand, the time series filter 232 uses the sequence variant to be analyzed and the corresponding sequence variant included in the time series information. If no identical variants exist, it sets the category as is (in this case, the initial category is MYC3, so it is set to MYC3) and passes the processing to the database filter 233.

[0166] The time series filter 232 may also receive threshold settings from the first setting acceptance unit 22 regarding depth, other sequence quality, mutation allele frequency, etc. For example, if the depth of the corresponding sequence mutation included in the time series information does not exceed the threshold set here (e.g., "20"), the time series filter 232 does not determine whether or not the same sequence mutation was present, but sets the category as is (in this case, the initial category is MYC3, so it is set to MYC3) and passes the processing to the database filter 233.

[0167] Furthermore, similar to the example of the filter processing unit in Figure 5, if the first data receiving unit 21 has not received time-series information, the time-series filter 232 may, without determining whether or not the same sequence mutation exists, set the category as is (in this case, since the initial category is MYC3, it is set to MYC3) and pass the processing to the database filter 233.

[0168] Furthermore, if a setting not to use the time-series filter 232 is input from the first setting acceptance unit 22, the time-series filter 232 does not determine whether or not there are the same sequence mutations, but sets the category as is (in this case, the original category is MYC3, so it is set to MYC3) and passes the processing to the fusion gene filter 236.

[0169] <Fusion gene filter> Hereinafter, any base sequence corresponding to any sequence mutation included in the base sequence information will also be referred to as a mutant base sequence. The fusion gene filter 236 determines whether the mutant base sequence contains a fusion gene formed by the fusion of two genes similar to each of the two candidate genes of the first fusion gene acquired by the information processing device 1. More specifically, for each of the multiple first fusion genes acquired by the information processing device 1, the fusion gene filter 236 determines whether the similarity between the two base sequences encoding the two candidate genes of the first fusion gene and at least some of the base sequences included in the mutant base sequence is above a threshold. Similarity is expressed, for example, by the proportion of alignment that matches between the two base sequences. If the proportion of alignment that matches between the two base sequences is above a threshold, the two base sequences are determined to be similar.

[0170] For example, the fusion gene filter 236 calculates the similarity between the base sequence encoded by the BCR gene and the corresponding base sequence in the mutant base sequence in the BCR-ABL first fusion gene, which is obtained by the information processing device 1 through the fusion of the BCR gene and the ABL gene. Next, the fusion gene filter 236 calculates the similarity between the base sequence encoded by the ABL gene and the corresponding base sequence in the mutant base sequence in the BCR-ABL first fusion gene.

[0171] The fusion gene filter 236 determines whether both of the calculated similarities are above a threshold. The threshold is, for example, a value at which the activity of the protein encoded by the first fusion gene is expected to be similar to the activity of the protein represented by the mutated base sequence.

[0172] The fusion gene filter 236 determines that a fusion gene formed by the fusion of two genes similar to each of the two candidate genes of the first fusion gene is included in the mutated base sequence if both of the calculated similarities are above the threshold.

[0173] Meanwhile, the fusion gene filter 236 repeats the same determination for another first fusion gene acquired by the information processing device 1 if at least one of the two similarity values ​​obtained is below a threshold. For all first fusion genes acquired by the information processing device 1, the fusion gene filter 236 determines that for any first fusion gene, the fusion gene formed by the fusion of two genes similar to the two candidate genes of the first fusion gene is not included in the mutated base sequence if at least one of the two similarity values ​​obtained is below a threshold.

[0174] Furthermore, the fusion gene filter 236 may determine that a fusion gene formed by the fusion of two genes similar to the two candidate genes of the first fusion gene is included in the mutant base sequence if the similarity between the base sequences of the two candidate genes of the first fusion gene acquired by the information processing device 1 and the base sequences of the two genes of the fusion gene included in the mutant base sequence is 65% or more and 100% or less. Preferably, the fusion gene filter 236 may determine that a fusion gene formed by the fusion of two genes similar to the two candidate genes of the first fusion gene is included in the mutant base sequence if the similarity between the base sequences of the two candidate genes of the first fusion gene and the base sequences of the two genes of the fusion gene included in the mutant base sequence is 80% or more and 100% or less.

[0175] Furthermore, the fusion gene filter 236 may transmit the mutated base sequence corresponding to the sequence mutation to be analyzed to an external server that stores combinations of candidate genes for multiple first fusion genes. The fusion gene filter 236 checks whether the mutated base sequence contains a fusion gene of two genes similar to each of the two candidate genes for the first fusion gene registered in the database of the external server. If the fusion gene filter 236 receives a notification from the external server indicating that the mutated base sequence contains a fusion gene of two genes similar to each of the two candidate genes for any of the multiple first fusion genes registered in the database of the external server, it may determine that the mutated base sequence contains a fusion gene formed by the fusion of two genes similar to each of the two candidate genes for the first fusion gene.

[0176] The fusion gene filter 236 determines whether the mutated base sequence contains a fusion gene formed by the fusion of a gene with a base sequence similar to the base sequence of a candidate gene for the second fusion gene acquired by the information processing device 1 with another gene. More specifically, for each of the multiple second fusion genes acquired by the information processing device 1, the fusion gene filter 236 calculates the similarity between the base sequence of the candidate gene for the second fusion gene and the base sequence of one of the genes in the fusion gene included in the mutated base sequence. The fusion gene filter 236 then determines whether the calculated similarity is above a threshold. The threshold is a value at which it is assumed that the activity of the protein encoded by the second fusion gene is similar to the activity of the protein indicated by the mutated base sequence.

[0177] The fusion gene filter 236 determines that a fusion gene of a gene similar to a candidate gene for the second fusion gene acquired by the information processing device 1 contains a mutated base sequence if the calculated similarity is above a threshold. The fusion gene filter 236 repeats the same determination for another candidate gene for the second fusion gene acquired by the information processing device 1 if the calculated similarity is below the threshold. For all second fusion genes acquired by the information processing device 1, the fusion gene filter 236 determines that no fusion gene of any gene similar to any candidate gene for the second fusion gene contains a mutated base sequence if the calculated similarity is below the threshold.

[0178] Furthermore, the fusion gene filter 236 may determine that a fusion gene formed by the fusion of a gene with a nucleotide sequence similar to the nucleotide sequence of the candidate gene for the second fusion gene obtained by the information processing device 1 and the nucleotide sequence of one of the genes in the fusion gene included in the mutated nucleotide sequence is included in the mutated nucleotide sequence if the similarity between the nucleotide sequence of the candidate gene for the second fusion gene obtained by the information processing device 1 and the nucleotide sequence of one of the genes in the fusion gene included in the mutated nucleotide sequence is 65% or more and 100% or less. Preferably, the fusion gene filter 236 may determine that a fusion gene formed by the fusion of a gene with a nucleotide sequence similar to the nucleotide sequence of the candidate gene for the second fusion gene and another gene is included in the mutated nucleotide sequence if the similarity between the nucleotide sequence of the candidate gene for the second fusion gene and the nucleotide sequence of one of the genes in the fusion gene included in the mutated nucleotide sequence is 80% or more and 100% or less.

[0179] Furthermore, the fusion gene filter 236 may send the mutated base sequence to an external server that stores multiple second fusion genes. The fusion gene filter 236 checks whether the mutated base sequence contains a fusion gene of a gene similar to one of the multiple candidate second fusion genes registered in the database of the external server. If the fusion gene filter 236 receives a notification from the external server indicating that the mutated base sequence contains a fusion gene of a gene similar to one of the registered candidate second fusion genes, it may determine that the mutated base sequence contains a gene similar to a candidate second fusion gene.

[0180] The fusion gene filter 236 determines a category based on the result of determining whether or not a fusion gene, which is formed by the fusion of two genes similar to each of the two candidate genes of the first fusion gene, is included in the mutated base sequence. For example, if the fusion gene filter 236 determines that a fusion gene, formed by the fusion of two genes similar to each of the two candidate genes of the first fusion gene, is included in the mutated base sequence for any of the multiple first fusion genes acquired by the information processing device 1, it determines a category corresponding to the sequence mutation to be analyzed as having a harmful risk (for example, by subtracting "2" as a second predetermined amount from the category) and passes it to the structure filter 238 for processing.

[0181] In this way, the fusion gene filter 236 can accurately estimate the degree of harmful risk of sequence mutations by categorizing them, by referring to the base sequences of two candidate genes of the first fusion gene that are known to be relatively likely to be driver mutations.

[0182] The fusion gene filter 236 determines a category based on whether or not the mutated base sequence contains a fusion gene in which a gene with a base sequence similar to that of a candidate gene of the second fusion gene has fused with another gene. For example, if the fusion gene filter 236 determines that the mutated base sequence contains a gene similar to one of the candidate genes of the multiple second fusion genes acquired by the information processing device 1, it determines a category corresponding to the sequence mutation to be analyzed as having a harmful risk (for example, by subtracting "1" as the first predetermined amount from the category) and passes the processing to the storage position filter 237.

[0183] If the fusion gene filter 236 determines that the fusion genes of candidate genes similar to the two candidate genes of the first fusion gene acquired by the information processing device 1 do not contain mutated base sequences, or if it determines that the fusion gene of a gene similar to the candidate gene of the second fusion gene does not contain mutated base sequences, it sets the category as is (in this case, the initial category is MYC3, so it is set to MYC3) and passes the processing to the storage position filter 237.

[0184] Even if one of the two candidate gene combinations of a fusion gene is not registered in the memory unit 12, it is known that a driver mutation may still occur in the second fusion gene containing a specific candidate gene. The fusion gene filter 236 can accurately present the degree of harmful risk of sequence mutations by categorizing them by referring to the base sequence of the candidate gene of the second fusion gene.

[0185] <Save Location Filter> Conserved sequences, which are preserved between the genomes of different species, often play important roles in the physiological activity of cells. Therefore, if a mutation occurs at the location of a conserved sequence, the risk of adverse effects from the sequence mutation is relatively high. The conserved location filter 237 determines the category of a sequence based on whether the location of a conserved sequence, which is a base sequence preserved between the genomes of different species, is included in the mutation site of a sequence mutation. Here, the conserved location filter 237 sets a threshold based on a value indicating the degree of conservation (output value of conservation prediction tools such as GERP or phylop PhastCons), and only conserved sequences that exceed this threshold can be used for classification.

[0186] If the conserved location filter 237 determines that the location of a conserved sequence is included in the mutation site, it determines a category corresponding to the sequence mutation to be analyzed as having a harmful risk (for example, by subtracting "1" as the first predetermined amount from the category) and passes the processing to the structure filter 238. On the other hand, if the conserved location filter 237 determines that the location of a conserved sequence is not included in the mutation site, it sets the category as is and passes the processing to the structure filter 238. In this way, the conserved location filter 237 can use information indicating the location of the conserved sequence to present the degree of harmful risk of the sequence mutation corresponding to this mutation site with greater accuracy through categories.

[0187] Furthermore, it is known that the harmful risks associated with structural mutations such as chromosomal translocations, deletions of important genes, and mutations affecting multiple genes are relatively high. The structural filter 238 determines whether the sequence mutation represented by the base sequence information is a structural mutation such as a chromosomal translocation.

[0188] <Structure Filter> The structural filter 238 determines whether the sequence mutation represented by the base sequence information is a chromosomal translocation, and determines a category based on this determination. The structural filter 238 determines whether a chromosomal translocation has occurred by referring to the content and location of the mutation included in the sequence mutation indicated by the base sequence information. Alternatively, the structural filter 238 may determine whether the sequence mutation is a chromosomal translocation by dividing the mutant base sequence corresponding to the sequence mutation into multiple base sequences and identifying the position on the genome for each divided base sequence.

[0189] The structural filter 238 determines whether the sequence mutation represented by the base sequence information is a mutation affecting multiple genes, and determines a category based on this determination. The structural filter 238 determines whether a mutation affecting multiple genes has occurred by referring to the content and location of the mutation included in any of the sequence mutations indicated by the base sequence information. The structural filter 238 may also determine whether the sequence mutation is a mutation affecting multiple genes by dividing the mutant base sequence corresponding to the sequence mutation into multiple base sequences and identifying the position on the genome for each divided base sequence.

[0190] The memory unit 12 has pre-registered information indicating multiple registered genes involved in cell carcinogenesis, etc. The information indicating registered genes includes, for example, identification information for identifying registered genes and information indicating the location of registered genes on chromosomes. The structural filter 238 may determine whether the sequence mutation represented by the base sequence information is a deletion of a registered gene, and may determine a category based on this determination result. The structural filter 238 refers to the content and location of the mutation included in any of the sequence mutations indicated by the base sequence information to determine whether one of the multiple registered genes registered in the memory unit 12 has been deleted.

[0191] The memory unit 12 has pre-registered chromosomal location information of enhancers that control the expression of genes involved in cell carcinogenesis, etc. When the structural filter 238 determines that a translocation, inversion, deletion, etc. has occurred, it may determine whether the sequence mutation represented by the base sequence information of the oncogene registered in the memory unit 12 is a dysregulation abnormality located near an enhancer registered in the memory unit 12, and may determine a category based on this determination result.

[0192] The memory unit 12 has pre-registered information on the orientation of gene regions in the genome (5'→3', 3'→5'). When the structural filter 238 determines that a sequence mutation represented by the base sequence information due to translocation or deletion forms a fusion gene such as a first fusion gene or a second fusion gene, and the two genes forming the fusion gene are designated as the first candidate gene and the second candidate gene, the filter 238 determines whether the orientations of the first candidate gene and the second candidate gene are in the same direction (for example, if the first candidate gene is 5'→3' and the second candidate gene is also 5'→3', or if the first candidate gene is 3'→5' and the second candidate gene is 3'→5'), determines whether a functional fusion gene is formed, and may determine a category based on this determination result.

[0193] The memory unit 12 has pre-registered sequence information related to amino acid translation (codons) of gene regions and RNA splicing. When the structural filter 238 determines that a sequence mutation represented by the base sequence information due to translocation or deletion forms a fusion gene, it may determine whether or not a functional fusion gene is formed based on the information of the above items, and determine a category based on this determination result.

[0194] Furthermore, the structural filter 238 divides the mutated base sequence into multiple base sequences and identifies the genomic location of each divided base sequence. The structural filter 238 may also determine whether or not a deletion of any of the registered genes has occurred by comparing the genomic location of the identified base sequence with the locations of multiple registered genes registered in the memory unit 12.

[0195] The structural filter 238 determines the category corresponding to the sequence variation that will be analyzed as having a harmful risk if it determines that a translocation has occurred. For example, the structural filter 238 subtracts "1" as the first predetermined value from the category corresponding to the sequence variation. On the other hand, if it determines that no translocation has occurred, it leaves the category corresponding to the sequence variation to be analyzed as is.

[0196] If the structural filter 238 determines that a mutation affecting multiple genes has occurred, it determines the category corresponding to the sequence mutation that will be analyzed as having a harmful risk (for example, subtracting "1" as the first predetermined quantity from the category corresponding to the sequence mutation). On the other hand, if the structural filter 238 determines that no structural mutation affecting multiple genes has occurred, it leaves the category corresponding to the sequence mutation unchanged.

[0197] If the structural filter 238 determines that any of the multiple registered genes registered in the memory unit 12 are deleted, it further subtracts a first predetermined amount from the category corresponding to the sequence mutation to be analyzed and passes the data to the structural filter 238 for processing. On the other hand, if the structural filter 238 determines that none of the multiple genes registered in the memory unit 12 are deleted, it leaves the category corresponding to the sequence mutation to be analyzed as is and passes the data to the structural filter 238 for processing. In this way, the structural filter 238 can accurately present the degree of harmful risk of sequence mutations by category by determining whether or not structural mutations such as chromosomal translocations, mutations affecting multiple genes, or deletions of genes involved in cell carcinogenesis have occurred.

[0198] Figure 12 is a flowchart illustrating the details of the filtering process performed by the third filter processing unit 43, which has the functional configuration shown in Figure 11.

[0199] In step S41, the basic filter 231 determines whether the sequence mutation to be processed poses a harmful risk according to the conditions of the basic filter 231. If the sequence mutation to be processed does not pose a harmful risk according to the conditions of the basic filter 231, it is determined to be "NO" in step S41, the category is set to MYC4, and the process proceeds to step S49.

[0200] In step S49, the third filter processing unit 43 outputs a category.

[0201] If the sequence mutation to be processed is deemed to pose a harmful risk according to the conditions of the basic filter 231, it is determined to be "YES" in step S41, the category is set to MYC3, and the process proceeds to step S42.

[0202] In step S42, the time series filter 232 determines whether or not there is a harmful risk for the sequence mutation to be processed, based on the conditions of the time series filter 232.

[0203] If the sequence mutation to be processed is deemed to pose a harmful risk according to the conditions of the time series filter 232, it is determined to be "YES" in step S42, the category is set to MYC2, and the process proceeds to step S47. The processing from step S47 onward will be described later.

[0204] If the sequence mutation to be processed does not pose a harmful risk according to the conditions of the time series filter 232, it is determined to be "NO" in step S42, the category is set to MYC3, and the process proceeds to step S43.

[0205] In step S43, the fusion gene filter 236 determines whether the sequence mutation to be processed contains a fusion gene of genes similar to the two candidate genes of the first fusion gene.

[0206] If the sequence mutation to be processed includes a fusion gene of genes similar to the two candidate genes of the first fusion gene (i.e., there is a risk of adverse effects), the result is determined as "YES" in step S43, the category is set to MYC2, and the process proceeds to step S47. The processing from step S47 onward will be described later.

[0207] If the sequence mutation to be processed does not include a fusion gene of genes similar to the two candidate genes of the first fusion gene (i.e., no harmful risk), the result is determined as "NO" in step S43, the category is set to MYC3, and the process proceeds to step S44.

[0208] In step S44, the fusion gene filter 236 determines whether the sequence mutation to be processed contains a fusion gene of a gene similar to the candidate gene for the second fusion gene.

[0209] In step S45, the conserved location filter 237 determines whether the location of the conserved sequence is included in the mutation site of the sequence mutation to be processed.

[0210] In step S46, the structural filter 238 determines whether the sequence mutation to be processed contains various structural mutations. In each filter from steps S44 to S46, if a harmful risk is determined, the category is set to MYC2. On the other hand, if no harmful risk is determined, the category is set to MYC3.

[0211] In step S47, the quality filter 235 determines whether the quality is sufficient or not.

[0212] If the quality of the results of the processing in steps S41 to S46 (the results of filtering by the basic filter 231, time series filter 232, fusion gene filter 236, conserved position filter 237, and structure filter 238) is sufficient, then in step S47 it is determined to be "YES", and the process proceeds to step S48. In step S47, since the quality was determined to be sufficient, "1" is subtracted from the category.

[0213] If the quality of the results of the processing in steps S41 to S46 (the filtering results of the basic filter 231, time series filter 232, fusion gene filter 236, conserved position filter 237, and structure filter 238) is insufficient, then in step S47, it is determined to be "NO", and the process proceeds to step S49. In this case, since it was determined in step S47 that the quality was insufficient, "1" is not subtracted from the category.

[0214] In step S49, the third filter processing unit 43 outputs a category.

[0215] Below is an example of the adjustment method used by the adjustment unit 4 for each filter of the third filter processing unit 43 in the second embodiment. Note that the adjustment examples for the basic filter 231, time-series filter 232, and quality filter 235 are the same as in the first embodiment and are therefore omitted from this explanation.

[0216] Examples of adjusting fusion gene filters As described above, in one embodiment of the fusion gene filter 236, if the similarity between the two base sequences encoding the two candidate genes of the first fusion gene and at least some of the base sequences included in the mutated base sequence is greater than or equal to a threshold, it is determined that the fusion gene is included in the mutated base sequence. Here, the adjustment unit 4 can adjust the determination result by the fusion gene filter 236 by adjusting the threshold.

[0217] Furthermore, as described above, in one embodiment of the fusion gene filter 236, if the similarity between the base sequences of the two candidate genes of the first fusion gene acquired by the information processing device 1 and the base sequences of the two genes of the fusion gene contained in the mutated base sequence is 65% or more and 100% or less, it can be determined that a fusion gene formed by the fusion of two genes similar to the two candidate genes of the first fusion gene is contained in the mutated base sequence. Here, the adjustment unit 4 can adjust the determination result by the fusion gene filter 236 by adjusting the range of the similarity ratio involved in the determination. For example, it can be determined that a fusion gene is contained in the mutated base sequence if the similarity between the base sequences of the two candidate genes of the first fusion gene and the base sequences of the two genes of the fusion gene contained in the mutated base sequence is 75% or more and 100% or less, or it can be determined that a fusion gene is contained in the mutated base sequence if the similarity is 85% or more and 100% or less.

[0218] Furthermore, as described above, in one embodiment of the fusion gene filter 236, the mutant base sequence corresponding to the sequence mutation to be analyzed is transmitted to an external server that stores combinations of candidate genes of multiple first fusion genes, and it can be determined that the fusion gene is included in the mutant base sequence based on the results of the investigation at the external server. Here, the adjustment unit 4 can adjust the determination result by the fusion gene filter 236 by changing the external server used.

[0219] Furthermore, as described above, in one embodiment of the fusion gene filter 236, for each of the multiple second fusion genes acquired by the information processing device 1, the similarity between the base sequence of the candidate gene of the second fusion gene and the base sequence of one of the fusion genes included in the mutated base sequence is determined for each second fusion gene. The fusion gene filter 236 then determines that the mutated base sequence contains a fusion gene similar to the candidate gene of the second fusion gene acquired by the information processing device 1 if the determined similarity is above a threshold. The adjustment unit 4 can adjust the determination result by the fusion gene filter 236 by adjusting the threshold for this similarity.

[0220] Furthermore, as described above, in one embodiment of the fusion gene filter 236, if the similarity between the base sequence of the candidate gene for the second fusion gene acquired by the information processing device 1 and the base sequence of one of the genes in the fusion gene included in the mutated base sequence is 65% or more and 100% or less, it can be determined that a fusion gene, formed by the fusion of a gene with a base sequence similar to the base sequence of the candidate gene for the second fusion gene and another gene, is included in the mutated base sequence. Here, the adjustment unit 4 can adjust the determination result by the fusion gene filter 236 by adjusting the range of the similarity ratio involved in the determination. For example, it can be determined that a fusion gene is included in the mutated base sequence if the similarity between the base sequence of the candidate gene for the second fusion gene and the base sequence of one of the genes in the fusion gene included in the mutated base sequence is 75% or more and 100% or less, or it can be determined that a fusion gene is included in the mutated base sequence if the similarity is 85% or more and 100% or less.

[0221] Furthermore, as described above, in one embodiment of the fusion gene filter 236, the mutated base sequence may be transmitted to an external server storing multiple second fusion genes, and based on the results of the investigation at the external server, it may be determined that the mutated base sequence contains a gene similar to a candidate gene of the second fusion gene. Here, the adjustment unit 4 can adjust the determination result by the fusion gene filter 236 by changing the external server used.

[0222] 《Example of adjusting the save location filter》 The preserved location filter 237 determines whether the location of the preserved sequence indicated by the preserved sequence location information acquired by the information processing device 1 is included in the mutation site. The classification criteria and determination results of the preserved location filter 237 can be adjusted by changing the threshold value set for determining whether or not it is a preserved sequence.

[0223] Examples of adjusting structural filters The structural filter 238 determines whether or not a chromosomal structural polymorphism (e.g., translocation, deletion, insertion, etc.) has occurred by referring to the content and location of the mutations included in the sequence mutations indicated by the base sequence information. The adjustment unit 4 can adjust the determination result by the structural filter 238 by changing the content and location of the mutations being referenced. Alternatively, the structural filter 238 may determine whether or not the sequence mutation is a chromosomal translocation by dividing the mutant base sequence corresponding to the sequence mutation into multiple base sequences and identifying the position on the genome for each divided base sequence. In this case, the adjustment unit 4 can adjust the determination result by the structural filter 238 by changing the unit of division.

[0224] Furthermore, in one embodiment of the structural filter 238, when it is determined that a translocation, inversion, deletion, etc., has occurred, it may be determined whether the sequence mutation represented by the base sequence information is a disregulated abnormality located near the enhancer of the oncogene, and a category may be determined based on this determination result. Here, the adjustment unit 4 can adjust the determination result by adjusting the criteria that the structural filter 238 uses to determine whether something is a disregulated abnormality.

[0225] Although one embodiment of the present invention has been described above, the present invention is not limited to the embodiments described above, and any modifications, improvements, etc. that can achieve the objectives of the present invention are considered to be included in the present invention.

[0226] Furthermore, the system configuration shown in Figure 1 and the configuration of the control unit 11 of the information processing device 1 shown in Figure 2 are merely illustrative examples for achieving the objectives of the present invention and are not particularly limited.

[0227] Furthermore, the functional block diagrams shown in Figures 2, 3, 5, 7, and 11 are merely illustrative and not particularly limiting. In other words, it is sufficient that the information processing device 1 is equipped with a function that can execute the series of processes described above as a whole, and the functional blocks used to realize this function are not particularly limited to the examples in these figures.

[0228] Furthermore, the location of the functional blocks is not limited to Figures 2, 3, 5, 7, and 11, but can be any location. For example, in the example in Figure 2, the above-mentioned processing is performed on the information processing device 1, but this is not limited to this, and at least part of the processing may be performed on other information processing devices not shown. In other words, the functional blocks necessary for executing the analysis processing are provided on the information processing device 1, but this is merely an example. At least part of the functional blocks located on the information processing device 1 may be provided on other information processing devices not shown.

[0229] The means and methods for performing various processing tasks in the system according to the above embodiment can be implemented by either a dedicated hardware circuit or a programmed computer. The program may be provided, for example, on a computer-readable recording medium such as a flexible disk or CD-ROM, or it may be provided online via a network such as the Internet. In this case, the program recorded on the computer-readable recording medium is usually transferred to and stored in a storage unit 12 such as a hard disk. Furthermore, the program may be provided as a standalone application software, or it may be incorporated into the software of the device as a function of the system.

[0230] In this specification, the step of describing a program to be recorded on a recording medium includes not only processes that are performed chronologically in that order, but also processes that are not necessarily performed chronologically, but are executed in parallel or individually.

[0231] Furthermore, one embodiment of the present invention may include a standard nucleic acid composition used in the information processing device 1 described above, which contains nucleic acids that include sequence mutations to which they belong. It may also include standard nucleic acid data used in the information processing device 1 described above, which includes sequence mutations to which they belong.

[0232] Furthermore, in this specification, the term "system" refers to an overall system composed of multiple devices, means, etc.

[0233] The present invention encompasses the following embodiments and forms.

[0234] [1] An information processing device for selecting a target sequence mutation that poses a harmful risk in a subject, A filtering unit that sorts one or more sequence mutations identified by sequencing the nucleic acids contained in the subject into categories corresponding to the degree of harmful risk, based on one or more classification criteria. An information processing device having a control unit that classifies a base sequence containing a sequence mutation to which the category to which it should belong is known, into each of the categories according to the degree of the harmful risk, based on at least one of the classification criteria, and compares the result of the classification with the category to which it should belong.

[0235] [2] The information processing apparatus according to [1], comprising an adjustment unit that adjusts the classification criteria and / or the classification results in the filtering unit based on the comparison results in the control unit.

[0236] [3] The information processing apparatus according to [1] or [2], wherein the base sequences including the sequence mutation to which the category to which they belong is known are two or more different to which the category to which they belong.

[0237] [4] The information processing apparatus according to [3], wherein the two or more base sequences belonging to different categories include a sequence mutation that causes a specific disease and a base sequence that does not cause the specific disease.

[0238] [5] The target sequence mutation is a driver mutation for a specific disease, The information processing apparatus according to [4], wherein the two or more sequence mutations include a sequence mutation that causes the specific disease and a sequence mutation that does not cause the specific disease.

[0239] [6] The classification criteria described above are changeable or selectable, and the information processing device is as described in any of [1] to [5].

[0240] [7] The information processing apparatus according to [6], which executes the filtering unit and the control unit after changing or selecting the classification criteria.

[0241] [8] The base sequence classified by the control unit is obtained by sequencing a standard composition of nucleic acid containing a known sequence mutation to which it should belong, as described in any of [1] to [7].

[0242] [9] The information processing apparatus according to [8], wherein the conditions for sequencing the standard composition are the same as the conditions for sequencing the nucleic acid contained in the subject.

[0243]

[10] A method for selecting a target sequence mutation that poses a harmful risk in a subject, A filtering step in which one or more sequence mutations identified by sequencing the nucleic acids contained in the subject are classified into categories according to the degree of harmful risk based on one or more classification criteria, An information processing method comprising: a control step of classifying a base sequence containing a sequence mutation to which the category to which it should belong is known, into each of the categories according to the degree of the harmful risk based on at least one of the classification criteria, and comparing the result of the classification with the category to which it should belong.

[0244]

[11] An information processing program that causes a computer to function as an information processing device described in any of [1] to [9].

[0245]

[12] A standard nucleic acid composition containing a nucleic acid with a known sequence mutation to which it should belong, and used in an information processing device described in any of [1] to [9].

[0246]

[13] Standard nucleic acid data used in any of the information processing devices described in [1] to [9], which include sequence variations to which the category to which the data should belong contains known sequence variations. [Industrial applicability]

[0247] The information processing device of the present invention is a device that performs analysis on the possibility that mutations in base sequences may affect the onset and progression of disease, and is capable of presenting more accurate analysis results. Therefore, it is applicable to a wide range of fields, such as the medical field and the life sciences field, and is industrially useful. [Explanation of Symbols]

[0248] 1. Information processing device, 2. Filtering section, 3. Control section, 4...adjustment section, 11. Control Unit, 12...Storage section, 13. Communications Department, 14...display section, 15. Operation reception section, 16...drive, 17. Removable media 18...bus 21. First Data Reception Unit, 22...First setting acceptance section, 23...First filter processing unit, 24...Category determination section, 25...Analysis result output section, 31...Second Data Reception Unit, 32...Second setting acceptance section, 33...Second filter processing unit, 34...Contrast section, 35. Comparison result output section, 43. Third Filter Processing Unit 231...Basic filter, 232...Time series filter, 233...Database filter, 234... Function prediction filter, 235... Quality filter, 236...Fusion gene filter, 237... Save location filter, 238...Structural filter.

Claims

1. An information processing device for selecting a target sequence mutation that poses a harmful risk in a subject, A filtering unit that classifies the base sequence information, including one or more sequence mutations identified by sequencing the nucleic acids contained in the subject, into categories according to the degree of harmful risk for each sequence mutation, based on one or more classification criteria. An information processing device having a control unit that classifies base sequence information containing one or more sequence mutations to which the category to which it should belong is known, into each of the categories corresponding to the degree of the harmful risk, based on the same classification criteria as the one or more classification criteria, for each sequence mutation to which the category to which it should belong is known, and compares the result of the classification with the category to which it should belong.

2. The information processing apparatus according to claim 1, further comprising an adjustment unit that adjusts the classification criteria and / or the classification results in the filtering unit based on the comparison results in the control unit.

3. The information processing apparatus according to claim 1 or 2, wherein the nucleotide sequence information containing one or more sequence mutations to which the category to which it should belong is known comprises two or more types to which the category to which it should belong differs.

4. The information processing apparatus according to claim 3, wherein the two or more types of base sequence information to which the above-mentioned categories belong include a sequence mutation that causes a specific disease and a base sequence that does not cause the above-mentioned specific disease.

5. The aforementioned target sequence mutation is a driver mutation for a specific disease. The information processing apparatus according to claim 3, wherein the two or more types of base sequence information to which the above-mentioned categories belong include a sequence mutation that causes the above-mentioned specific disease and a sequence mutation that does not cause the above-mentioned specific disease.

6. The information processing apparatus according to any one of claims 1 to 5, wherein the classification criteria can be changed or selected.

7. The information processing apparatus according to claim 6, wherein the filtering unit and the control unit are executed after the classification criteria have been changed or selected.

8. The information processing apparatus according to any one of claims 1 to 7, wherein the base sequence information classified by the control unit is obtained by sequencing a standard composition of nucleic acid containing a known sequence mutation to which it should belong.

9. The information processing apparatus according to claim 8, wherein the conditions for sequencing the standard composition are the same as the conditions for sequencing the nucleic acid contained in the subject.

10. A method for selecting a target sequence mutation in a subject that poses a harmful risk, A filtering step of classifying the base sequence information, which includes one or more sequence mutations identified by sequencing the nucleic acids contained in the subject, into categories corresponding to the degree of harmful risk for each sequence mutation, based on one or more classification criteria. An information processing method comprising: a control step of classifying base sequence information containing one or more sequence mutations to which the category to which it should belong is known, into each of the categories corresponding to the degree of the harmful risk, based on the same classification criteria as the one or more classification criteria, and comparing the result of the classification with the category to which it should belong.

11. An information processing program for causing a computer to function as an information processing device according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Genetic mutations for breast cancer risk assessment

    JP2011527565A

  • Methods and systems for identifying causative genomic mutations.

    JP2015501974A

  • Mutation Detection for Cancer Screening and Fetal Analysis

    JP2018512048A

  • Analysis method of nucleic acid sequence of patient samples, presentation method of analysis results, presentation device, presentation program, and analysis system of nucleic acid sequence of patient samples

    JP2021000082A

  • High sensitive genetic variation detection and reporting system based on barcode sequence information

    KR1020200106643A