Information processing device and information processing method

JP2024140387A5Active Publication Date: 2025-05-20HITACHI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023051508
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-28
Publication Date
2025-05-20
Estimated Expiration
2043-03-28

AI Technical Summary

Technical Problem

Current DNA sequencing technologies struggle to accurately align and detect structural variations in genomes due to limitations in read length and the presence of repetitive sequences, leading to errors that can misidentify structural variations.

Method used

An information processing device and method that utilizes k-tuples based on the ratio of label intervals in DNA sequences, incorporating error handling for expansion and contraction, to align label positions accurately and identify structural variations.

Benefits of technology

Enables accurate alignment of DNA sequences, effectively detecting structural variations by handling errors such as expansion, contraction, and label detection failures, improving the reliability of genomic analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To perform labeled positions alignment capable of dealing with apparent expansion and contraction of a target nucleic acid sequence.SOLUTION: An information processing device calculates a plurality of first ratios of intervals between partial sequences in a referenced nucleic acid sequence, constructs an index indicating a combination of the first ratios and information indicating a position of a partial sequence in the nucleic acid sequence corresponding to the combination of the first ratios, calculates a plurality of second ratios of intervals between partial sequences in a target nucleic acid sequence, extracts a combination of the first ratios corresponding to a combination of the second ratios based on a comparison result between the combination of the second ratios and the combination of first ratios indicated by the index, and outputs information indicating a position of a partial sequence corresponding to the extracted combination of the first ratios in the referenced nucleic acid sequence.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing device and an information processing method. [Background technology]

[0002] With the advancement of DNA (DeoxyriboNucleic Acid) sequencing technology, many personal genomes have been revealed. Personal genomes contain many differences from the reference genome. The majority of these are single nucleotide variants (SNVs), in which only one base in the surrounding base sequence differs from the reference genome, but structural variants (SVs), in which a large base sequence of several thousand or more bases changes at once, are also included, although their numbers are smaller than single nucleotide variants.

[0003] SNVs and SVs include not only germline mutations that cause individual differences, but also acquired mutations called somatic mutations, some of which can cause cancer. Accurate detection of such mutations and elucidation of their biological and clinical significance are important issues in cancer treatment and biology research.

[0004] To clarify structural mutations, it is necessary to capture changes in large genomic regions of several thousand bases or more. However, the length of the base sequence that can be read at one time with current DNA sequencing technology is limited. The length of the base sequence that can be read at one time is limited to a maximum of about 1,000 bases with the Sanger method, which was originally used to determine the standard human genome sequence, and to a few hundred bases with the currently mainstream Next Generation Sequencing (NGS).

[0005] NGS can obtain a pair of two base sequences separated by several hundred bases, called paired-end sequencing, but even with paired-end sequencing, only a narrow region of about 1000 bases can be obtained. However, the human genome contains many repetitive sequences such as SINEs (Short Intersparsed Nuclear Elements) and LINEs (Long Intersparsed Nuclear Elements), and there are also repetitive sequences in regions called centromeres and telomeres.

[0006] If only a base sequence of 1,000 bases can be seen at a time, these repetitive sequences cannot be distinguished, and the base sequence of the entire genome cannot be estimated by connecting the obtained base sequences. Long-read sequencing technology, which has become widespread in recent years, can obtain base sequences of tens of thousands of bases at a time, but this is not enough to identify the positions of all repetitive sequences. Therefore, technology to analyze wider regions of the genome is needed.

[0007] A technique that can be used for such purposes is called genome mapping, in which short specific base sequences on the genome are labeled with fluorescent or other labels and their positions on the genome are identified from the pattern of the label intervals. In genome mapping, the DNA that constitutes the genome is amplified and cut to generate many DNA fragments consisting of hundreds of thousands of bases.

[0008] In genome mapping, specific base sequences on each of the many DNA fragments generated are labeled, and the label positions, which indicate the approximate base position from the beginning of each labeled sequence (hereinafter simply referred to as label) that appears in the DNA fragment, are measured. Furthermore, by arranging the label positions of each DNA fragment in ascending order, each DNA fragment can be converted into an ascending numerical sequence. This numerical sequence is hereinafter referred to as measurement data.

[0009] A document disclosing such a genome mapping technique is JP 2009-022274 A (Patent Document 1), which describes "a method for mapping positions on chromosomal DNA, comprising hybridizing a nucleic acid to one type of repeating base sequence in expanded or extended chromosomal DNA, measuring the mutual distances on chromosomal DNA for a set of a plurality of repeating base sequences of the chromosomal DNA by using a label introduced into the hybridized nucleic acid, and then determining the regions or positions on the chromosome of the set and the repeating base sequences contained in the set based on the characteristics of the measured distance" (see Abstract).

[0010] The process of comparing the measurement data obtained by genome mapping with the marker positions obtained from the reference genome sequence or the like to clarify the common and non-common parts is called alignment. If there is no structural mutation in the DNA fragment on which the measurement data is based or no measurement error in the measurement data, each marker position indicated by the measurement data corresponds to one of the marker positions on the reference genome. On the other hand, if a structural mutation occurs, the corresponding positions of the markers on the measurement data and the markers on the reference genome become discontinuous. By capturing such abnormalities in the marker positions, structural mutations can be detected. [Prior art documents] [Patent documents]

[0011] [Patent Document 1] JP 2009-022274 A Summary of the Invention [Problem to be solved by the invention]

[0012] Accurate alignment is necessary to detect structural mutations. If there are many errors in the alignment, abnormalities in the labeling positions caused by the errors may be mistakenly recognized as structural mutations.

[0013] In order to align the measured data, it is necessary to address errors contained in the measured data. One such error is that the overall length of each DNA fragment appears to expand or contract in the measured data. This is caused by the fact that the speed of movement of the molecules during measurement is not uniform.

[0014] Although Patent Document 1 discloses labeling repetitive sequences in a genome and identifying their positions in the genome, it does not disclose a method for comparing the measured labeling interval with the labeling position in a reference genome.

[0015] On the other hand, one aspect of the present invention performs alignment of label positions that can accommodate apparent stretching and contraction of the nucleic acid sequence of interest. [Means for solving the problem]

[0016] In order to solve the above problems, one aspect of the present invention employs the following configuration: An information processing device includes a processor and a memory, the memory holds a first numeric sequence indicating the position of a partial sequence in a reference nucleic acid sequence and a second numeric sequence indicating the measurement position of the partial sequence in a target nucleic acid sequence, the processor calculates a plurality of first ratios of the intervals of the partial sequences in the reference nucleic acid sequence based on the first numeric sequence, constructs an index indicating a combination of the first ratios and information indicating the positions of the partial sequences in the reference nucleic acid sequence corresponding to the combination of the first ratios, calculates a plurality of second ratios of the intervals of the partial sequences in the target nucleic acid sequence based on the second numeric sequence, extracts a combination of first ratios corresponding to the combination of second ratios based on a comparison result between the combination of second ratios and the combination of first ratios indicated by the index, and outputs information indicating the positions of the partial sequences corresponding to the extracted combination of first ratios in the reference nucleic acid sequence. Effect of the Invention

[0017] According to one aspect of the invention, alignment of label positions can be performed that can accommodate apparent stretching of a nucleic acid sequence of interest.

[0018] Problems, configurations and effects other than those described above will become apparent from the following description of the embodiments. [Brief description of the drawings]

[0019] [Figure 1] FIG. 1 is a block diagram showing a configuration example of a genome label position alignment device in a first embodiment. [Diagram 2] FIG. 4 is a diagram illustrating an example of a data configuration of measurement data in the first embodiment. [Diagram 3] 13 is a flowchart illustrating an example of a process of an index construction unit in the first embodiment. [Figure 4] FIG. 2 is an explanatory diagram illustrating an outline example of an index construction process according to the first embodiment; [Diagram 5] FIG. 4 is a diagram illustrating an example of a data configuration of an index in the first embodiment. [Figure 6] 13 is a flowchart illustrating an example of an index search process in the first embodiment. [Figure 7] 13 is a flowchart showing an example of a process of searching for an index while thinning out some of the signs in the first embodiment. [Figure 8] FIG. 2 is an explanatory diagram showing an example of a DNA fragment in which labels have been thinned out in the measurement data in Example 1. [Figure 9] 1 is a flowchart showing an example of an alignment probability calculation process in the first embodiment. [Figure 10] FIG. 2 is a sequence diagram showing an example of the overall processing by the genome label position alignment device. [Figure 11] FIG. 11 is an explanatory diagram illustrating an example of a tree structure constructed by expanding a k-tuple indicated by an index in the second embodiment. [Figure 12] 13 is a flowchart showing an example of alignment processing using a tree structure in the second embodiment. [Figure 13] 13 is a flowchart showing an example of an update process of sets S and T executed in the alignment process using a tree structure in the second embodiment. [Figure 14] FIG. 11 is a diagram illustrating an example of a user interface displayed on an input / output device according to a third embodiment. [Figure 15] FIG. 13 is an explanatory diagram showing an example of a structural mutation detection process in Example 4. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0020] Hereinafter, a genome label position alignment apparatus according to an embodiment of the present invention will be described. Note that the same reference numerals are used to designate the same components in the following drawings, and duplicated explanations will be omitted.

[0021] Errors contained in the measurement data include not only the apparent expansion and contraction of the overall length of the DNA fragment as described above, but also missed detection of a label on the DNA fragment, false detection of a position on the DNA fragment where there is no label, cutting of the DNA fragment during sample preparation, etc. The genome label alignment device of this embodiment realizes alignment processing that can deal with these errors. EXAMPLES

[0022] <Device configuration> 1 is a block diagram showing a configuration example of a genome labeled position alignment device. The genome q labeled position alignment device is configured by a computer 300 including, for example, a CPU (Central Processing Unit) 310, a memory 311, an auxiliary storage device 312, and interfaces 313 to 315. The hardware included in the computer 300 is electrically connected via an internal communication line such as a bus.

[0023] The CPU 310 reads programs and data stored in the memory 311, and executes the programs stored in the memory 311. The CPU 310 includes a processor. The CPU 310 includes, for example, an index construction unit 321, an index search unit 322, and an alignment probability calculation unit 323, which are all functional units. The computer 300 functions as a genome label position alignment device by the CPU 310 executing processes.

[0024] The memory 311 temporarily stores programs executed by the CPU 310 and data used when the programs are executed. The memory 311 includes a ROM (Read Only Memory), which is a non-volatile storage element, and a RAM (Random Access Memory), which is a volatile storage element. The ROM stores immutable programs (e.g., a Basic Input / Output System (BIOS)). The RAM is a high-speed, volatile storage element such as a DRAM (Dynamic Random Access Memory), and temporarily stores the programs executed by the CPU 310 and data used when the programs are executed.

[0025] The memory 311 stores, for example, programs for implementing an index construction unit 321, an index search unit 322, and an alignment probability calculation unit 323, as well as a reference genome label position 330 and measurement data 340.

[0026] For example, CPU 310 functions as index construction unit 321 by operating in accordance with an index construction program loaded into memory 311, functions as index search unit 322 by operating in accordance with an index search program loaded into memory 311, and functions as alignment probability calculation unit 323 by operating in accordance with an alignment probability calculation program loaded into memory 311.

[0027] The auxiliary storage device 312 stores in a non-volatile manner the programs executed by the CPU 310 and data used when the programs are executed. That is, the programs are read from the auxiliary storage device 312, loaded into the memory 311, and executed by the CPU 310.

[0028] The auxiliary storage device 312 is a large-capacity non-volatile storage device such as a hard disk drive (HDD) or a solid state drive (SSD). The auxiliary storage device 312 stores programs that realize the functions of an index construction unit 321, an index search unit 322, and an alignment probability calculation unit 323, as well as a reference genome label position 330 and measurement data 340.

[0029] Each of the interfaces 313 to 315 mediates the transmission and reception of signals and converts protocols, and is connected to an external device. The interface 313 is an I / O interface connected to the input / output device 302 via a wired line or a wireless line. The input / output device 302 includes input devices such as a keyboard and a mouse, and output devices such as a display device and a printer. The interface 313 acquires input information from an operator that is accepted by the input / output device 302. The interface 313 also outputs the execution result of the program to the input / output device 302 in a format that can be visually recognized by the operator.

[0030] The interface 315 is a network interface that is connected to the external storage device 301 via the network 305. The interface 315 controls communications with other devices in accordance with a predetermined protocol.

[0031] The external storage device 301 is a non-transitory storage device that stores data handled by the computer 300. The external storage device 301 includes, for example, a storage device such as an HDD or SSD. The external storage device 301 can store reference genome label positions 330 and measurement data 340.

[0032] Data is transmitted and received between the external storage device 301 and the computer 300 via a network 305. The network 305 includes, for example, a local area network (LAN) and the Internet. The types of the network 305 are not limited to those described above. The network 305 may be configured as a wired network or a wireless network.

[0033] The interface 314 is connected to a drive device that reads and writes data from and to the removable medium 303. The interface 314 includes, for example, a serial interface such as a Universal Serial Bus (USB).

[0034] The removable medium 303 is a non-transitory storage medium that stores data handled by the computer 300. The removable medium 303 includes optical disks such as CDs and DVDs, magnetic disks, and semiconductor memories. The removable medium 303 can store reference genome label positions 330 and measurement data 340.

[0035] In addition, some or all of the programs executed by CPU 310 may be provided to computer 300 from removable media 303, which is a non-transitory storage medium, via interface 314, or from external storage device 301, which is a non-transitory storage device, or an external computer equipped with external storage device 301, via network 305, and stored in non-volatile auxiliary storage device 312, which is a non-transitory storage medium.

[0036] 1, an external storage device 301 and a removable medium 303 are connected to the computer 300 constituting the genome label position alignment device, but these external devices may be omitted if not required. An input / output device 302 equipped with the external storage device 301 may be connected to the computer 300 via a network 305. Instead of being connected to the input / output device 302, the computer 300 may have a built-in device equipped with an input / output function.

[0037] The index construction unit 321 constructs an index based on a k-tuple. The index construction unit 321 constructs an index using a k-tuple based on the ratio of marker intervals as described below, rather than an index using a k-tuple that indicates the interval between marker positions (hereinafter also simply referred to as marker interval). In addition, the index constructed by the index construction unit 321 can also deal with the case where some markers are missing, in order to deal with missed marker detection.

[0038] Note that a k-tuple indicates a combination of k numbers (k is a predefined parameter). Assuming that, as in the conventional technology, each k-tuple is a combination of k labeling intervals on the reference genome, and an index is constructed that is a correspondence table between each k-tuple and the identifier of the label on the reference genome that corresponds to the labeling interval in the k-tuple.

[0039] In this case, simply generating a k-tuple in the measurement data 340 and performing an alignment by comparing the generated k-tuple with the k-tuple indicated by the index cannot address the apparent expansion and contraction of the measurement data 340.

[0040] In addition, even if k-tuples corresponding to these expansions and contractions are generated by expanding and contracting the entire measurement data 340 at various expansion / contractions to deal with the apparent expansion and contraction of the measurement data 340 and the k-tuples corresponding to these expansions and contractions are compared with the k-tuples indicated by the index, in order to improve the accuracy of the alignment, it would be necessary to try out a large number of expansion / contractions that are likely to be correct and adopt the optimal one, which would increase processing time. In addition, even if a huge number of different expansion / contractions were adopted, there would be a risk that all of the adopted expansion / contractions would deviate from the expansion / contraction rate that is optimal for the measurement data 340, and thus accurate alignment might not be possible.

[0041] When measuring DNA fragments, the moving speed of each DNA fragment molecule is not uniform during measurement, so the label interval itself in the measurement data 340 expands and contracts significantly, but the error in the ratio of label intervals in the measurement data 340 is small. Therefore, in this embodiment, the index construction unit 321 constructs an index using k-tuples based on the ratio of label intervals as described above, thereby making it possible to realize highly accurate alignment that does not require expansion or contraction of the measurement data 340.

[0042] The index search unit 322 identifies a position on the reference genome that is likely to correspond to the measurement data, using an index based on the k-tuple constructed by the index construction unit 321. During the search, the index search unit 322 executes a process corresponding to erroneous detection of a marker.

[0043] The alignment probability calculation unit 323 calculates the probability that the generated alignment occurs. When there are multiple candidates for the position where the DNA fragments indicated by the measurement data 340 are aligned, if the probability is high, the alignment is likely to be correct, so the alignment probability calculation unit 323 adopts the alignment with the highest probability. Even if only a portion of the DNA fragments can be aligned, the probability is used to determine whether the alignment is optimal or not.

[0044] Before starting the genome label position alignment process, the reference genome label position 330 and the measurement data 340 are input and stored in the computer 300. The CPU 310 may read the reference genome label position 330 and the measurement data 340 and load them into the memory 311, for example, when the computer 300 is started or when processing is executed.

[0045] The reference genome marker positions 330 and the measurement data 340 may be stored in all of the auxiliary storage device 312, the external storage device 301, and the removable media 303, or may be stored in only some of them. These data may be moved or copied to the external storage device 301 or the removable media 303 when the computer 300 is stopped or when the auxiliary storage device 312 has insufficient free space. Therefore, it is desirable that the reference genome marker positions 330 stored in different storage devices are all the same information. The same applies to the measurement data 340.

[0046] The reference genome label position 330 includes, for each of a plurality of reference genomes, numerical data representing the position of a label present on the reference genome. The measurement data 340 includes, for each of a number of DNA fragments, data obtained by measuring the label position on the DNA fragment.

[0047] 2 is a diagram showing an example of the data configuration of the measurement data 340. The measurement data 340 indicates an ID for identifying each DNA fragment (an example of a target nucleic acid sequence), the molecular length (length of the base sequence) of each DNA fragment, and the position of a label (an example of a partial sequence) on each DNA fragment. Note that blank spaces in the measurement data 340 indicate that no label was observed.

[0048] For example, the molecular length of a DNA fragment with ID "2" is 44951 bases, four markers are measured in the DNA fragment, and the positions of the four markers are 10844th base, 19749th base, 23353rd base, and 35735th base from the beginning of the DNA fragment, respectively. Note that the positions of the markers of the DNA fragment indicated by the measurement data 340 may indicate, for example, the beginning position or the end position of the marker.

[0049] In this embodiment, an example of a short base sequence to be labeled on a DNA fragment is GCTCTTC, which is recognized by an enzyme called Nt.BspQI. As described above, the measurement data 340 indicates the labeling position of GCTCTTC on each DNA fragment in ascending order. In other words, the measurement data 340 indicates information in which the labeling position of each DNA fragment is recorded as a numeric sequence (an example of a second numeric sequence) in ascending order.

[0050] <Index construction> 3 is a flowchart showing an example of the process of the index construction unit 321. The index construction unit 321 constructs a k-tuple for any marker position on the reference genome based on the ratio of the marker interval, and constructs an index indicating the correspondence between the constructed k-tuple and the marker position on the reference genome.

[0051] Step S401: The index construction unit 321 acquires genome sequence data. The genome sequence data indicates the number of each of a plurality of chromosomes that are components of a genome, and the base sequence of each of the plurality of chromosomes (an example of a referenced nucleic acid sequence). The base sequence of each of the plurality of chromosomes is indicated by a character string consisting of characters representing four types of bases, A, T, G, and C, and N representing an unknown base. The genome sequence data is stored in advance in at least one of the memory 311, the auxiliary storage device 312, the removable media 303, and the external storage device 301, for example.

[0052] Step S402: If the index constructing unit 321 determines that all chromosomes have been selected, it completes the index construction process, and if it determines that there are any unselected chromosomes, it transitions to step S403.

[0053] Step S403: The index construction unit 321 selects one unprocessed chromosome. The index construction unit 321 executes construction of a k-tuple and registration in the index for the selected chromosome in the following procedure.

[0054] Step S404: The index construction unit 321 calculates a numerical sequence (an example of a first numerical sequence) indicating the position of the marker (an example of a partial sequence) in the chromosome selected in the most recent step S403, and stores the numerical sequence in association with the chromosome number in the reference genome marker position 330. For example, the above-mentioned GCTCTTC is used as the marker sequence. Note that since genomic DNA is a double helix, a portion where a complementary sequence (a sequence in which A, T, G, and C are replaced with T, A, C, and G, respectively, in reverse order) matches the marker sequence is also labeled in genome mapping. Therefore, when the index construction unit 321 calculates a numerical sequence indicating the marker position, it is necessary to add the position of the marker sequence or its complementary sequence to the numerical sequence without distinction.

[0055] Step S405: The index construction unit 321 calculates the ratio of the marker intervals and the ratio of the marker intervals when some of the markers are thinned out based on the numeric sequence indicating the marker positions calculated in step S404, registers the k-tuples based on the calculated ratio in the index, and returns to step S402.

[0056] A specific example of the process of step S405 will be described with reference to Fig. 4. Fig. 4 is an explanatory diagram showing an outline of an example of the index construction process. In the example of Fig. 4, it is assumed that there are 10 markers, marker 1 to marker 10, in the reference genome, and the marker positions have been obtained. In addition, the distance between marker i and marker j is represented as d(i,j).

[0057] In step S405, first, the index constructing unit 321 calculates the marker interval between adjacent markers and the marker interval between markers that will become adjacent by thinning out some markers according to a predetermined rule, using the numerical sequence indicating the marker positions calculated in step S404. Note that the predetermined rule here is, for example, a rule for thinning out markers when generating k-tuples using k marker interval ratios, and among k+3 consecutive markers, any of the 2nd marker to the k+2nd marker is thinned out. In other words, the targets to be thinned out are markers other than the 1st marker and the k+3rd marker, which are at both ends. Below, an example will be described in which a k-tuple is calculated while sequentially thinning out any one of markers 1 to 10.

[0058] In this case, the index construction unit 321 calculates the marker interval d(i, i+1) for i=1,...,9. Note that the marker interval d(x, y) represents the distance between marker x and marker y, and is expressed in units of the number of bases. Furthermore, the index construction unit 321 calculates the marker intervals of markers that will be adjacent to each other by sequentially thinning out the markers according to the predetermined rule, such as further calculating the marker interval d(1,3) because marker 1 and marker 3 will be adjacent to each other by further thinning out marker 2, for example.

[0059] The index constructor 321 calculates the ratio of adjacent marker intervals and the ratio of marker intervals that will be adjacent when some of the markers are thinned out according to the predetermined rule. That is, the index constructor 321 calculates d(i,i+1) / d(i+1,i+2) for i=1,...,8. Furthermore, the index constructor 321 calculates the ratio of marker intervals that will be adjacent when markers are sequentially thinned out according to the predetermined rule, for example, by further thinning out marker 2, marker interval d(1,3) and marker interval d(3,4) will be adjacent to each other, and so on.

[0060] The index construction unit 321 associates a k-tuple formed by a combination of k ratio values ​​that are consecutive based on the calculated ratio of the marker intervals with a position on the reference genome that corresponds to the k-tuple in the index. Furthermore, the index construction unit 321 associates a k-tuple formed by a combination of k ratio values ​​that become consecutive by thinning out some markers according to the predetermined rule based on the calculated ratio of the marker intervals with a position on the reference genome that corresponds to the k-tuple in the index.

[0061] For example, the number of the chromosome from which the k-tuple was calculated and an integer value representing the order of the k+2 markers from the beginning of the chromosome from which the ratio values ​​contained in the k-tuple were calculated are registered in the index as positions on the reference genome.

[0062] For example, if k = 3, for i = 1,...,6, for each of the numbers of the chromosomes being selected, which are positions on the reference genome, and markers i to i+4, the k-tuples obtained when the markers are not thinned out are d(i,i+1) / d(i+1,i+2), d(i+1,i+2) / d(i+2,i+3), and d(i+2,i+3) / d(i+3,i+4).

[0063] Furthermore, the index construction unit 321 registers in the index the ratio of the intervals between k markers that become adjacent by thinning out the markers sequentially according to the specified rule, for example, as k-tuples for the five markers 1, 3, 4, 5, and 6 that become consecutive by thinning out the marker 2 of the chromosome being selected, the k-tuples consisting of d(1,3) / d(3,4), d(3,4) / d(4,5), and d(4,5) / d(5,6) are registered in the index.

[0064] In addition, when registering the ratio value in the index, the index construction unit 321 may perform an operation to consider similar values ​​as the same value, for example by using a technique such as binning to forcibly consider numbers after the third decimal place as zero, in order to absorb errors in the measurement data 340, which inevitably contains experimental errors.

[0065] Furthermore, when the index constructor 321 calculates the ratio between adjacent first and second marker intervals, if the first marker interval is significantly larger than the second marker interval, the ratio value will be significantly greater than 1.0, whereas if it is smaller, the value will be smaller than 1.0, resulting in asymmetry depending on the magnitude relationship. To prevent this, the index constructor 321 may use the logarithm of the ratio (for example, a logarithm with a predetermined base, such as a common logarithm or a natural logarithm) instead of the ratio value.

[0066] Furthermore, since a k-tuple represents the ratio of k marker intervals, the index constructor 321 may represent the ratio of adjacent marker intervals with an integer value such that the sum of the ratio values ​​matches an integer value given as a parameter, rather than directly calculating the ratio of adjacent marker intervals. For example, if the integer value given as a parameter is 100 and k=3, and the ratios of three adjacent marker intervals are 3000, 3000, and 4000, respectively, then this k-tuple may be represented with integers whose sum is 100, such as 30:30:40.

[0067] 5 is a diagram showing an example of the data configuration of the index constructed in step S405. The index 701 indicates, for example, the correspondence between a k-tuple and a position on the reference genome corresponding to the k-tuple. The index 701 is stored in the same storage device as the reference genome label position 330, for example.

[0068] For example, record 7011 indicates that the k-tuples in "markers 986-990" of "chromosome 9" and the k-tuples in "markers 1532-1536" of "chromosome 10" are all "1.12, 0.98, 0.88". Also, for example, record 7012 indicates that the k-tuples in "markers 312-317" (i.e. markers 312, 313, 314, 316, 317) when "marker 315" of "chromosome 15" is thinned out are "0.54, 0.99, 1.21".

[0069] <Index Search> 6 is a flowchart showing an example of index search processing. The index search unit 322 constructs a k-tuple based on the ratio of label intervals for an arbitrary position of a DNA fragment indicated by the measurement data 340, and identifies a position on a reference genome corresponding to the DNA fragment indicated by the measurement data 340 using an index 701 created by the index construction unit 321.

[0070] Step S 501 : The index search unit 322 receives the measurement data 340 .

[0071] Step S502: If the index search unit 322 determines that all DNA fragments included in the input measurement data 340 have been selected, it ends the index search process, and if it determines that there are unselected DNA fragments, it transitions to step S503.

[0072] Step S503: The index search unit 322 selects one unselected DNA fragment from the measurement data 340. The index search unit 322 may select the DNA fragments in any order, and for example, simply selects the DNA fragments in the order in which the information was input to the measurement data 340.

[0073] Step S504: The index search unit 322 acquires from the measurement data 340 each of the detected labeled positions of the DNA fragment selected in the most recent step S503.

[0074] Step S505: The index search unit 322 calculates each label interval in the DNA fragment based on the label positions acquired in step S504, and calculates the ratio between adjacent label intervals. Then, the index search unit 322 compares each k-tuple, which is the value of the ratio of k consecutive pieces in the DNA fragment, with each k-tuple in the index constructed in step S321, and acquires candidates for the corresponding positions on the reference genome.

[0075] In addition, in step S505, in order to deal with erroneous detection of markers in the measurement data 340, the index search unit 322 further thins out markers from the selected DNA according to the specified rule, compares each of the k-tuples generated when the thinned out markers are not present with each of the k-tuples in the index 701, and obtains candidates for corresponding positions on the reference genome.

[0076] Therefore, in step S505, a correspondence is generated between the label positions indicated by the k-tuples in the DNA fragment being selected and candidate positions on the reference genome.

[0077] Details of the index reference process in step S505 will be described later with reference to FIG. 7. Furthermore, in step S505, the index search unit 322 max is initialized to 0. Since there may be multiple candidates for the positions on the reference genome corresponding to each k-tuple in the DNA fragment, they are processed sequentially using the following procedure.

[0078] Step S506: If the index search unit 322 determines that all of the candidates for the corresponding positions on the reference genome acquired in step S505 have been processed, it transitions to step S511, and if it determines that there are unprocessed candidates for the positions, it transitions to step S507.

[0079] Step S507: The alignment probability calculation unit 323 calculates the probability p that the label positions indicated by the k-tuples of the DNA fragments corresponding to each candidate for the candidate position on the reference genome are aligned. The details of the calculation procedure for the alignment probability p will be described later with reference to FIG. 8.

[0080] Step S508: The alignment probability calculation unit 323 checks whether p>p max If it is determined that p≦p, the process proceeds to step S509. max If so, the process returns to step S506.

[0081] Step S509: The alignment probability calculation unit 323 records the candidate for the selected position and max Then, the alignment probability calculation unit 323 substitutes p into p, and the process returns to step S506. If a position candidate has already been recorded, the alignment probability calculation unit 323 overwrites the position candidate.

[0082] Step S510: The index search unit 322 outputs candidates for the recorded positions, and ends the index search process.

[0083] In the procedure of steps S506 to S508, a method of outputting the position on the reference genome with the highest probability has been described, but a procedure of outputting multiple candidates, rather than just the single highest position, may be added.

[0084] Fig. 7 is a flowchart showing an example of the process of searching for the index 701 while thinning out some of the markers in step S505. In the example of Fig. 7, the predetermined rule for thinning out some of the markers is to sequentially thin out any one of the markers contained in the selected DNA fragment.

[0085] Step S1301: The index search unit 322 sets the variable n to the number of markers observed in the selected DNA fragment of the measurement data 340.

[0086] Step S1302: If n < k + 2, the interval between labels in the DNA is k or less, and the ratio of the interval between labels becomes k - 1 or less. Therefore, since the index search unit 322 cannot refer to the index 701, the process of FIG. 7 ends. If n ≥ k + 2, the index search unit 322 transitions to step S1303.

[0087] Step S1303: The index search unit 322 initializes the variable i to 1.

[0088] Step S1304: The index search unit 322 calculates a total of k ratios from a total of k + 1 label intervals at labels i to i + k + 1 without thinning out the labels from the selected DNA fragment to generate a k - tuple, and refers to the index 701. For example, in the process of referring to the index 701, the index search unit 322 acquires a k - tuple that matches the generated k - tuple from the index 701, and acquires the position on the reference genome corresponding to the acquired k - tuple in the index 701 as the candidate.

[0089] Step S1305: If i ≥ n - k - 1, thinning out the labels would result in the interval between labels being k + 1 or less. Therefore, the index search unit 322 ends the process of FIG. 7. If i < n - k - 1, the index search unit 322 transitions to step S1306.

[0090] Step S1306: The index search unit 322 initializes the variable j to 2. The reason for not initializing the variable j to 1 is that the left - most label does not need to be thinned out (thinning out the left - most label does not generate a new k - tuple).

[0091] Step S1307: If j < k + 2, the index search unit 322 transitions to step S1310. If j ≥ k + 2, the index search unit 322 transitions to step S1308.

[0092] Step S1308: The index search unit 322 calculates the ratio of a total of k marker intervals from a total of k+1 marker intervals obtained when marker i+j-1 is considered not to exist among markers i to i+k+2 of the selected DNA fragment, generates a k-tuple, and refers to the index 701. That is, the index search unit 322 refers to the index 701 with the k-tuple obtained by thinning out marker i+j-1.

[0093] Step S1309: The index search unit 322 adds 1 to the variable j to update the sign to be thinned out to the next sign, and returns to step S1307.

[0094] Step S1310: The index search unit 322 adds 1 to the variable i to update the position of the k-tuple of the selected DNA fragment for which the index 701 is to be referenced, and returns to step S1304.

[0095] The reference genome index registration process in step S405 can also be realized by a process similar to that shown in Fig. 7. Specifically, the "DNA fragment" in the process in Fig. 7 can be read as "chromosome," and the process of referring to index 701 in Fig. 7 can be replaced with a process of registering the generated k-tuple in index 701 in association with a position on the reference genome.

[0096] Fig. 8 is an explanatory diagram showing an example of a DNA fragment in which markers in the measurement data 340 have been thinned out. In the example of Fig. 8, five markers, marker 1 to marker 5, are observed in the DNA fragment of the measurement data 340. If k=2, then in the process of Fig. 7, a k-tuple is generated when no markers are thinned out from the DNA fragment, a k-tuple is generated when marker 2 is thinned out from the DNA fragment, a k-tuple is generated when marker 3 is thinned out from the DNA fragment, and a k-tuple is generated when marker 4 is thinned out from the DNA fragment.

[0097] <Alignment probability calculation> 9 is a flowchart showing an example of the alignment probability calculation process in step S507. In the following, m is the aligned portion of the DNA fragment that is the subject of the probability calculation, and P(m) is the probability that the alignment occurs (p in steps S507 to S509).

[0098] In this embodiment, the alignment probability calculation unit 323 calculates P(m) as P(m)=P scale w1 P pos w2 P ins w3 P del w4 It is calculated using the probability model shown by P scale , P pos , P ins , and P del are the probability of the stretching ratio indicating the ratio between the observed molecular length of the DNA fragment and the molecular length on the reference genome, the probability of the deviation of the observed label interval (from the interval on the genome), the probability of false positive detection of the label, and the probability of missed detection of the label, respectively. In addition, w1, w2, w3, and w4 are weights, and may be w1=w2=w3=w4=1 unless otherwise specified by the user.

[0099] In the above example, P(m)=P scale w1 P pos w2 P ins w3 P del w4- However, P scale w1 , P pos w2 , P ins w3 , P del w4-- may not be taken into consideration (i.e., some of w1, w2, w3, and w4 may be 0).

[0100] The alignment probability calculation unit 323 calculates P scale For example, P scale =f scale(|stretch rate - 1|). For example, f scale (x) has mean 0 and variance σ scale 2 is the probability density function of the normal distribution of σ scale is determined in advance from given real data.

[0101] The alignment probability calculation unit 323 calculates P pos For example, P pos =f pos (|x1-y1|)f pos (|x2-y2|) f pos (|x n -y n For example, f pos (x) is a normal distribution N(0,σ pos 2 ) probability density function, σ pos 2 is determined in advance from real data. Also, x i is the i-th labeling interval in the DNA fragment, y i is the i-th marker interval of the chromosome represented by the corresponding reference genome.

[0102] The alignment probability calculation unit 323 calculates P ins For example, the number of false positives n ins Based on n ins The alignment probability calculation unit 323 calculates P using a function that monotonically decreases as P increases. Specifically, for example, the alignment probability calculation unit 323 calculates P using an exponential function exp(x). ins (n ins )=R ins exp(-n ins ) by P ins For example, R ins is a constant coefficient, which is determined in advance from given actual data.

[0103] Similarly, the alignment probability calculation unit 323 calculates P del For example, the number of missed detections of the label, n del Using P del (n del )=Rdel exp(-n del The calculation process for the occurrence probability of alignment based on the above definition will be described below.

[0104] Step S601: The alignment probability calculation unit 323 obtains the correspondence between the alignment candidate obtained as a result of the reference process in step S505, i.e., the labeled position of the selected DNA fragment and the labeled position on the reference genome. That is, for example, the alignment probability calculation unit 323 specifies the portion where the k-tuples match between the selected DNA fragment and the reference genome indicated by the index 701.

[0105] Step S602: The alignment probability calculation unit 323 calculates the stretching rate of the entire alignment. Specifically, for example, the alignment probability calculation unit 323 can calculate the stretching rate of the entire alignment by calculating the distance between the marker positions at both ends for each of the DNA fragment selected in step S503 and the reference genome (chromosome) indicated by the correspondence obtained in step S601, and obtaining the ratio between the calculated distances between the marker positions at both ends.

[0106] Step S603: In a loop with step S604, the alignment probability calculation unit 323 calculates the deviations |x i -y i In step S603, the alignment probability calculation unit 323 determines whether all landmark intervals have been processed, and if so, proceeds to S605, and if any landmark intervals remain unprocessed, proceeds to step S604.

[0107] Step S604: The alignment probability calculation unit 323 calculates a set of associated marker intervals x i , y i For the label spacing deviation |x i -y i | is calculated and the process returns to step S603.

[0108] Step S605: In a loop with step S606, the alignment probability calculation unit 323 counts the number of erroneous detections of markers on DNA fragments when it is assumed that the marker positions of the DNA fragments indicated by the correspondence obtained in step S601 match the positions on the reference genome. However, the erroneous detection here refers to a marker position on the DNA fragment that does not correspond to a position on the reference genome. If the alignment probability calculation unit 323 has already processed all markers of the measurement data that are not in the reference genome, it proceeds to step S607, and if there are any unprocessed markers of the measurement data that are not in the reference genome, it proceeds to step S606.

[0109] Step S606: The alignment probability calculation unit 323 increments the number of false positives by 1, and returns to step S605. Note that the number of false positives is initialized to 0 in advance.

[0110] Step S607: In a loop with step S608, the alignment probability calculation unit 323 counts the number of missed detections of markers on the DNA fragments when it is assumed that the marker positions of the DNA fragments indicated by the correspondence obtained in step S601 match the positions on the reference genome. Here, the missed detection refers to the marker positions on the reference genome that cannot be associated with the markers on the DNA fragments. If the alignment probability calculation unit 323 has already processed all the markers of the reference genome that are not in the measurement data, it proceeds to step S609, and if there are any unprocessed markers of the reference genome that are not in the measurement data, it proceeds to step S608.

[0111] Step S608: The alignment probability calculation unit 323 increments the number of missed detections by 1, and returns to step S607. Note that the number of missed detections is initialized to 0 in advance.

[0112] Step S609: The alignment probability calculation unit 323 calculates P(m) by substituting the numerical values ​​obtained in the above processes (the stretching ratio obtained in step S602, the deviation in the marker interval obtained in step S604, the number of false detections obtained in step S606, and the number of missed detections obtained in step S608) into the definition equation for P(m), and ends the alignment probability calculation process.

[0113] In addition, it is preferable that the parameters determined from the actual data are not common values ​​for the entire genome, but are set individually for each region of each chromosome. By the above processing, the genome label position alignment device of this embodiment can align the k-tuples based on the ratio of label intervals in each DNA shown in the measurement data 340 with the k-tuples based on the ratio of label intervals in the reference genome.

[0114] As described above, the genome label position alignment device of this embodiment can perform alignment that can deal with the expansion and contraction of DNA fragments during measurement by performing alignment using k-tuples based on the ratio of label intervals.

[0115] Furthermore, the genome label position alignment device registers in index 701 a k-tuple constructed while thinning out some of the labels of the chromosomes of the reference genome, and compares the k-tuple generated while thinning out the positions of the labels of the DNA fragments indicated by the measurement data 340 with the k-tuple indicated by index 701, thereby realizing alignment that can deal with false detection of labels on DNA fragments and missed detection of labels on DNA fragments.

[0116] In the above process, the genome label position alignment device scale , P pos , P ins , and P delSince the corresponding positions on the reference genome are determined based on the alignment probability using each, alignment can be realized that can deal with the expansion and contraction of DNA fragments during measurement, the shift of the label position on the DNA fragment, the false detection of the label on the DNA fragment, and the missed detection of the label on the DNA fragment. In addition, the genome label position alignment device can provide a means for evaluating the reliability of partial alignment caused by the breakage or structural mutation of DNA molecules during sample preparation and performing alignment. EXAMPLES

[0117] In the first embodiment, the genome label position alignment device can align k-tuples based on the ratio of label spacing between the DNA fragments of the measurement data 340 and the reference genome, but in the present embodiment, the genome label position alignment device extends the alignment by sequentially matching labels around the k-tuples registered in the index 701.

[0118] Fig. 10 is a sequence diagram showing an example of the overall processing by the genome label position alignment device. In the processing of Fig. 10, the process in which the index construction unit 321, the index search unit 322, and the alignment probability calculation unit 323 cooperate with each other while referring to the reference genome label position 330 and the measurement data 340 will be described. It is expected that the processing of Fig. 10 makes it possible to align the entire molecule indicated by the measurement data 340 with the reference genome if there is no breakage or structural mutation of the DNA molecule during sample preparation. Note that the processing up to step S1004 is the same as that of the first embodiment.

[0119] Step S1001: The index constructor 321 receives a genome sequence as an input.

[0120] Step S1002: The index constructing unit 321 constructs the index 701 using the k-tuples based on the ratio of the marker intervals through steps S401 to S405.

[0121] Step S1003: The index search unit 322 acquires the measurement data 340, and acquires the labeled positions of the DNA fragments contained in the acquired measurement data 340.

[0122] Step S1004: The index search unit 322 searches for a position on the reference genome that matches the k-tuple by referring to the index 701 while thinning out some of the markers of the DNA fragments in steps S501 to S505.

[0123] Step S1005: The index search unit 322 also compares the ratio values ​​of the marker intervals between the DNA fragment and the reference genome for markers in the vicinity of the marker position where the k-tuple matches in step S1004.

[0124] Step S1006: The alignment probability calculation unit 323 calculates the alignment probability that also reflects the comparison result of the surrounding marker intervals in step S1005.

[0125] Step S 1007 : The alignment probability calculation unit 323 outputs the alignment probability calculated in step S 1006 to the index search unit 322 .

[0126] Step S 1008 : The index search unit 322 determines the optimal alignment based on the alignment probability, and outputs the determined alignment to the input / output device 302 .

[0127] Details of steps S1005 to S1008 will be described later with reference to FIGS.

[0128] 11 is an explanatory diagram showing an example of a tree structure constructed by extending the index 701 in order to compare the ratio of the marker intervals around the k-tuples. In this embodiment, the index 701 is extended, and the index construction unit 321 constructs two tree structures for each k-tuple.

[0129] One tree structure for each k-tuple indicates the ratio of labeling intervals on the upstream side of that k-tuple in the reference genome (the side closer to the beginning of the string representing the chromosome), and the other tree structure for that k-tuple indicates the ratio of labeling intervals on the downstream side (the side closer to the end position of the string representing the chromosome).

[0130] In each tree structure, the node closest to the root represents the ratio of adjacent sign intervals to the k-tuple, and the child nodes represent the ratio of adjacent sign intervals to the parent node. That is, a tree structure showing the ratio of upstream sign intervals represents the ratio of adjacent upstream sign intervals, and a tree structure showing the ratio of downstream sign intervals represents the ratio of adjacent downstream sign intervals.

[0131] The ratio values ​​indicated by the k-tuples and the tree structure are considered to be the same because they are close to each other using methods such as binning. Therefore, the same k-tuple may appear at multiple locations on the reference genome, and since the ratios of adjacent labeling intervals at each location are different, one parent node may have multiple child nodes in the tree structure. This is why the data structure indicating the ratios of labeling intervals around the labeling interval indicated by the k-tuple is a tree structure.

[0132] 11, the index 701 and tree structures 1111 and 1112 constructed by the index construction unit 321 from the marker position 330 of the reference genome will be specifically described. Record 1110 of index 701 indicates a k-tuple in which the combination of the ratios of three adjacent marker intervals is "1.78", "1.34", and "0.97", and the position on the reference genome corresponding to the k-tuple. Tree structure 1111 indicates the ratio of the marker intervals on the upstream side of the k-tuple indicated by record 1110, and tree structure 1112 indicates the ratio of the marker intervals on the downstream side of the k-tuple indicated by record 1110.

[0133] For example, for each position on the reference genome corresponding to the k-tuple indicated by record 1110, tree structure 1111 is constructed by sequentially searching for the ratios of marker intervals adjacent to the upstream side of "1.78", and tree structure 1112 is constructed by sequentially searching for the ratios of marker intervals adjacent to the downstream side of "0.97".

[0134] In the tree structure 1111, the node closest to the root node is "0.87." This indicates that for all positions on the reference genome corresponding to the k-tuple indicated by the record 1110, the ratio of the marker interval adjacent to the upstream side of "1.78" is "0.87."

[0135] Furthermore, in the tree structure 1111, there are a node indicating "1.03" and a node indicating "1.04" as child nodes of "0.87." Therefore, at the position on the reference genome corresponding to the k-tuple indicated by record 1110, the ratio of the two marker intervals upstream of "1.78" (i.e., the ratio of the marker intervals adjacent to the upstream side of "0.87") is "1.03" or "1.04."

[0136] Similarly, by calculating the ratio of the marker spacing upstream of the position on the reference genome, a node indicating "0.61" and a node indicating "0.94" are obtained as child nodes of "1.03", and a node indicating "1.78" is obtained as a child node of "1.04" (it is assumed that a node indicating "0.94" is not obtained as a child node of "1.04").

[0137] Here, when the tree structure 1111 includes a node group that has the same parent node and exhibits similar values ​​(e.g., the difference is within a predetermined value) such as "1.03" and "1.04," for example, an edge 1113 may be added from a node included in the node group to a child node of another node included in the node group, or a search process in Fig. 12 described below may be performed assuming that such an edge exists virtually. Adding edge 1113 increases the likelihood of obtaining a correct alignment by a search that follows edge 1113, even if the calculated marker spacing ratio contains a small error.

[0138] The genome label position alignment device of this embodiment uses the above-described tree structure shown in Fig. 11 to compare the ratio of label intervals of DNA fragments indicated by the measurement data 340 with the ratio of label intervals on the reference genome around the k-tuples aligned by the procedure of embodiment 1, and generates an optimal alignment. The method is described below.

[0139] In order to ensure that the alignment is optimal, the genome label position alignment device sequentially calculates the probability of generating the alignment by the alignment probability calculation unit 323, and can finally obtain the optimal alignment by sequentially examining nodes that can generate the optimal alignment at that time, i.e., the alignment that maximizes the probability p calculated by the alignment probability calculation unit 323. More precisely, the calculation is performed according to the procedures shown in Figs. 12 and 13.

[0140] FIG. 12 is a flowchart showing an example of alignment processing using a tree structure.

[0141] Step S1401: The index search unit 322 initializes the set S with {(root node of the tree, 0)}. The index search unit 322 also initializes the set T with an empty set. The first value (the element on the left side) included in each element of the set S is called a node, and the second value (the element on the right side) is called a score. Hereinafter, the calculation is performed so that the set S is a node in the middle of the search, and T is a set of nodes corresponding to the indicators for which the search has been completed.

[0142] Step S1402: If the set S is an empty set, the index search unit 322 proceeds to step S1410, and if the set S is not an empty set, the index search unit 322 proceeds to step S1403.

[0143] Step S1403: The index search unit 322 extracts one element of the set S with the maximum score, and deletes the extracted element from the set S. The extracted element is defined as (v, p).

[0144] Step S1404: If all child nodes of v have been processed, the index search unit 322 returns to step S1402, and if there are unprocessed child nodes of v, the index search unit 322 proceeds to step S1405.

[0145] Step S1405: The index search unit 322 selects one child node of v, and sets the selected node as u.

[0146] Step S1406: The index search unit 322 selects a marker adjacent to the processed marker (the marker associated with v) in the measurement data 340 as a marker to be associated with u.

[0147] Step S1407: The index search unit 322 and the alignment probability calculation unit 323 execute update processing of the sets S and T for u and the marker selected in step S1406. Details of the update processing of the sets S and T will be described later with reference to FIG.

[0148] Step S1408: The index search unit 322 selects a marker adjacent to the marker corresponding to u in order to deal with the case where the marker to be associated with u is not detected in the DNA fragment being processed.

[0149] Step S1409: The index search unit 322 and the alignment probability calculation unit 323 regard the marker selected in step S1408 as a new u and execute the update process of the sets S and T described later with reference to FIG. 13, and after execution, return to step S1404.

[0150] Step S1410: The index search unit 322 sets the node with the maximum score as u in the set T. The index search unit 322 also outputs the alignment corresponding to the node u, that is, the alignment in which the markers on the genome corresponding to each node selected until finally reaching u by tracing the tree structure correspond to the markers of the DNA fragments, as the optimal alignment.

[0151] FIG. 13 is a flowchart showing an example of an update process for the sets S and T, which is executed in the alignment process using a tree structure.

[0152] Step S1501: The alignment probability calculation unit 323 calculates the alignment probability corresponding to the selected node u by the method shown in FIG. 9, similar to step S507, and sets the calculated probability as q.

[0153] Step S1502: The index search unit 322 judges whether the labeled position corresponding to the node u is the end of the DNA fragment molecule. If the index search unit 322 judges that the labeled position is the end, the process proceeds to step S1503, and if it judges that the labeled position is not the end, the process proceeds to step S1504.

[0154] Step S1503: The index search unit 322 adds (u, q) to the set T.

[0155] Step S1504: The index search unit 322 adds (u, q) to the set S. EXAMPLES

[0156] In the third embodiment, a user interface is provided for setting parameters for adjusting the processing by the genome label position alignment device. This user interface also visualizes the alignment results.

[0157] 14 is a diagram showing an example of a user interface displayed on the input / output device 302. The user interface 200 includes, for example, an input data setting area 210, a parameter setting area 220, and an alignment result display area 230.

[0158] The input data setting area 210 is an area for setting the source of the reference genome label position 330 and the measurement data 340. In the example of Fig. 14, the source is specified by a file name, but may be specified by a URL (Uniform Resource Locator) or the like as necessary.

[0159] The parameter setting area 220 is an area for setting parameters used in processing by the genome label position alignment device. In the parameter setting area 220, for example, the number (k) of ratio values ​​in the k-tuple can be set.

[0160] Also, in the parameter setting area 220, it is possible to set, for example, the number of signs to be thinned out by the index construction unit 321 and the index search unit 322 in order to deal with missed detections and erroneous detections. In the above example, the number of signs to be thinned out is 1, but the number of signs to be thinned out can also be set to 2 or more. Note that as the number of signs to be thinned out increases, the data size of the index 701 increases, but error resistance increases.

[0161] In the parameter setting area 220, for example, it is possible to set parameters w1, w2, w3, and w4 for weighting errors used when the alignment probability calculation unit 323 calculates the probability P(m). By being able to set these weights in the parameter setting area 220, it becomes possible to tune the importance of various errors.

[0162] Information for visualizing the alignment obtained as a result of the calculation is displayed in the alignment result display area 230. In the alignment result display area 230, the positions on the reference genome corresponding to the DNA fragments of the measurement data 340 are displayed, and the corresponding markers between the two can be confirmed.

[0163] In addition, in the alignment result display area 230, it is possible to distinguish between markers corresponding to each other on the DNA fragment and the reference genome, which are associated by k-tuples (an example of a non-extended portion, the k-tuple region associated by a solid line in FIG. 14) and surrounding regions associated by using a tree structure (an example of an extended portion, the k-tuple region associated by a dotted line in FIG. 14). In addition, the alignment result display area 230 can assist the user in verifying the calculation results by displaying the ratio value of the marker interval and the expansion / contraction rate of the entire molecule of the DNA fragment. EXAMPLES

[0164] In this embodiment, a genome-labeled position alignment system detects structural variations present in the genome of a subject. The genome-labeled position alignment system includes a genome mapping device and the genome-labeled position alignment device described in the first to third embodiments.

[0165] 15 is an explanatory diagram showing an example of a structural mutation detection process. The genome mapping device 1200 acquires genomic DNA collected from a subject. The genome mapping device 1200 performs genome amplification and fragmentation on the acquired genomic DNA of the subject. Furthermore, the genome mapping device 1200 acquires measurement data 340 by measuring the position of the marker in each DNA fragment.

[0166] The genome label position alignment device can identify the position on the reference genome corresponding to each marker indicated by the measurement data 340 by aligning the DNA fragment indicated by the measurement data 340 with the reference genome label position 330 through k-tuple-based alignment processing using the ratio of marker intervals described in Examples 1 and 2.

[0167] The genome label position alignment device determines whether there is a structural mutation in the subject's genome based on the comparison result between the label position in the DNA fragment indicated by the measurement data 340 and the label position in the reference genome that corresponds to the label position of the DNA fragment.

[0168] Specifically, for example, when the genome label position alignment device determines that the label positions on the subject's genome are discontinuous (specifically, when the labels of the reference genome corresponding to the labels that are consecutive on the DNA fragment are not consecutive), or when the label interval is abnormally large or small (for example, the difference between the label interval in the DNA fragment indicated by the measurement data 340 and the label interval in the reference genome corresponding to the label interval is equal to or larger than a predetermined value or smaller than a predetermined value), it can determine that there is a structural mutation in the genome of the subject. The genome label position alignment device outputs a list of the structural mutations detected in this way to, for example, the input / output device 302, so that the user of the genome label position alignment device can comprehensively grasp the structural mutations present in the genome of the subject.

[0169] The present invention is not limited to the above-mentioned embodiment, and various modifications are included. For example, the above-mentioned embodiment has been described in detail to clearly explain the present invention, and is not necessarily limited to those having all the configurations described. It is also possible to replace a part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. It is also possible to add, delete, or replace a part of the configuration of each embodiment with another configuration.

[0170] In addition, the above-mentioned configurations, functions, processing units, processing means, etc. may be realized in part or in whole by hardware, for example, by designing them as integrated circuits. In addition, the above-mentioned configurations, functions, etc. may be realized in software by a processor interpreting and executing a program that realizes each function. Information such as the program, table, file, etc. that realizes each function can be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD.

[0171] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and not all control lines and information lines in the product are necessarily shown. In reality, it can be considered that almost all components are connected to each other. [Explanation of symbols]

[0172] 300 computer, 302 input / output device, 310 CPU, 311 memory, 312 auxiliary storage device, 313 interface, 321 index construction unit, 322 index search unit, 323 alignment probability calculation unit, 330 reference genome marker position, 340 measurement data, 701 index

Claims

1. An information processing device, A processor and a memory, the memory holds a first numeric sequence indicating a position of a partial sequence in a reference nucleic acid sequence, and a second numeric sequence indicating a measurement position of the partial sequence in a target nucleic acid sequence; The processor, calculating a plurality of first ratios of intervals of the subsequences in the reference nucleic acid sequence based on the first numerical sequence; constructing an index indicating the first ratio combination and information indicating the position of a subsequence in the reference nucleic acid sequence corresponding to the first ratio combination; calculating a plurality of second ratios of the intervals of the partial sequences in the target nucleic acid sequence based on the second numerical sequence; extracting a first ratio combination corresponding to the second ratio combination based on a comparison result between the second ratio combination and the first ratio combination indicated by the index; An information processing device that outputs information indicating the position of the partial sequence corresponding to the extracted first ratio combination in the referenced nucleic acid sequence.

2. 2. The information processing device according to claim 1, The information processing device, wherein the processor calculates, based on the first numerical sequence, the ratio of the spacing between adjacent partial sequences in the referenced nucleic acid sequence and the ratio of the spacing between partial sequences that would become adjacent when some of the partial sequences are thinned out from the referenced nucleic acid sequence based on a predetermined rule, as the multiple first ratios.

3. 2. The information processing device according to claim 1, The information processing device, wherein the processor calculates, based on the second numerical sequence, the ratio of the spacing between adjacent partial sequences in the target nucleic acid sequence and the ratio of the spacing between adjacent partial sequences that will become adjacent when some of the partial sequences are thinned out from the target nucleic acid sequence based on a predetermined rule, as the multiple second ratios.

4. 2. The information processing device according to claim 1, The processor, Identifying a first ratio combination from the index that matches the second ratio combination; calculating, for each of the identified combinations of first ratios, a probability that the identified combination of first ratios corresponds to the combination of second ratios based on a predetermined probability model; An information processing device that extracts a combination of first ratios corresponding to the combination of second ratios based on the calculated probability.

5. 5. The information processing device according to claim 4, The information processing device, wherein the specified probability model is a model reflecting at least one of the following: the probability of an elongation ratio indicating the ratio between the observed molecular length and the correct molecular length of the target nucleic acid sequence, the probability of deviation between the measured position of the partial sequence in the target nucleic acid sequence and the correct position, the probability of false detection of the partial sequence in the target nucleic acid sequence, and the probability of missed detection of the partial sequence in the target nucleic acid sequence.

6. 2. The information processing device according to claim 1, The processor, For the first ratio combinations indicated by the indexes, expand the first ratio combinations based on the ratio of intervals between subsequences adjacent to the subsequences in the referenced nucleic acid sequence corresponding to the indexes; Expanding the second ratio combination based on the ratio of intervals of subsequences adjacent to the subsequences in the nucleic acid sequence of interest corresponding to the second ratio combination; An information processing device that extracts a first ratio combination corresponding to the second ratio combination based on a comparison result between the expanded second ratio combination and the expanded first ratio combination.

7. 7. The information processing device according to claim 6, The processor, Identifying the expanded first ratio combinations that match the expanded second ratio combinations; information indicating the position of a subsequence of the reference nucleic acid sequence indicated by the first ratio that matches the second ratio in the extended portion of the extended first ratio combination that matches the extended second ratio combination; and information indicating the position of a partial sequence of the referenced nucleic acid sequence indicated by the first ratio that matches the second ratio in the non-extended portion of the extended first ratio combination that matches the extended second ratio combination.

8. 2. The information processing device according to claim 1, connected to an input device, The processor, an information processing device that accepts, via the input device, an input of the number of first ratios included in each of the first ratio combinations and the number of second ratios included in each of the second ratio combinations.

9. 2. The information processing device according to claim 1, The processor, determining whether or not there is a structural mutation in the target nucleic acid sequence based on a comparison result between a measured position of a partial sequence in the target nucleic acid sequence indicated by the combination of the second ratios and a position of a partial sequence in a reference nucleic acid sequence indicated by the extracted combination of the first ratios; When it is determined that a structural mutation is present in the target nucleic acid sequence, the information processing device outputs information indicating the structural mutation.

10. An information processing method by an information processing device, The information processing device includes a processor and a memory. the memory holds a first numeric sequence indicating a position of a partial sequence in a reference nucleic acid sequence, and a second numeric sequence indicating a measurement position of the partial sequence in a target nucleic acid sequence; The information processing method includes: The processor calculates a plurality of first ratios of intervals of the subsequences in the reference nucleic acid sequence based on the first numerical sequence; The processor constructs an index indicating the first ratio combinations and information indicating the positions of subsequences in the reference nucleic acid sequence corresponding to the first ratio combinations; the processor calculates a plurality of second ratios of the intervals of the subsequences in the target nucleic acid sequence based on the second numerical sequence; The processor extracts a first ratio combination corresponding to the second ratio combination based on a comparison result between the second ratio combination and the first ratio combination indicated by the index; An information processing method in which the processor outputs information indicating the position of the partial sequence corresponding to the extracted first ratio combination in the reference nucleic acid sequence.