Information processing device, information processing method, and information processing program
Patent Information
- Application Number
- JP2025561601
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-12
AI Technical Summary
Current DNA sequencing technologies struggle to detect large-scale structural variants in genomes due to limitations in reading longer base sequences, and existing alignment methods fail when label intervals are short, leading to measurement errors that disrupt alignment accuracy.
An information processing apparatus that uses the ratio of intervals between two or more labels to identify the position of a subsequence in a reference nucleic acid sequence, where at least one of the intervals used in the ratio is set to intervals between three or more labels, allowing for alignment that accounts for apparent elongation and contraction of nucleic acid molecules.
This approach enables accurate alignment even when label intervals are short, effectively addressing measurement errors and improving the detection of structural variations in genomes.
Abstract
Description
Information processing device, information processing method, and information processing program
[0001] The present invention relates to an information processing device.
[0002] Individual genomes, especially those of cancerous cells, contain numerous differences, or mutations, from the reference genome. Among these mutations, structural variants (SVs), which involve large changes in the base sequence of several thousand or more bases at once, are fewer in number than small-scale mutations but play an important role. However, current major DNA sequencing technologies are limited to reading a single sequence of several hundred bases, making it difficult to capture such large-scale sequence changes. Therefore, a separate, low-cost technology is needed to analyze even larger regions of the genome.
[0003] A technique called genome mapping can be used for such purposes. This involves labeling specific short base sequences (hereinafter referred to as labels) of about seven bases on the genome with fluorescence or other means, and identifying the position on the genome from the distance on the genome of the labeled locations, i.e., the pattern of label spacing. In genome mapping, the DNA that constitutes the genome is amplified and cut to generate a large number of DNA fragments consisting of hundreds of thousands of bases. On each of these DNA fragments, the approximate base position from the beginning of the DNA at which the label appears, i.e., the label position, is measured. The label position observed for each DNA fragment can be expressed as a numerical sequence. This numerical sequence will also be referred to as measurement data hereinafter.
[0004] A document disclosing a technique similar to this is Patent Document 1, which describes "a method for mapping locations on chromosomal DNA, comprising hybridizing a nucleic acid to one type of repeat base sequence in unfolded or extended chromosomal DNA, measuring the mutual distances on the chromosomal DNA for sets of multiple repeat base sequences in the chromosomal DNA by using a label introduced into the hybridized nucleic acid, and then determining the regions or positions on the chromosome of the sets and the repeat base sequences contained in the sets based on the characteristics of the measured distances" (Claim 1).
[0005] Meanwhile, the process of comparing each measurement data obtained by genome mapping with the marker positions obtained from the reference genome sequence to identify common and different parts is called alignment. If there are no genome mutations or measurement errors, each marker position indicated by the measurement data corresponds to one of the marker positions on the reference genome. On the other hand, if a structural mutation occurs, the corresponding positions between the markers on the measurement data and the markers on the reference genome become discontinuous. Structural mutations can be detected by detecting such abnormalities in the marker positions.
[0006] As a technique related to the present application, the inventor of the present invention provided, in Japanese Patent Application No. 2023-051508 (hereinafter referred to as the prior application), an alignment method using the ratio of the distance between adjacent labels, which is a quantity that is invariant to the apparent expansion and contraction of the molecule.
[0007] JP 2009-022274 A
[0008] To detect structural mutations, accurate alignment is required, which requires addressing errors contained in the measurement data. One such error is that the overall length of each DNA fragment appears to expand or contract in the measurement data. This is caused by the fact that the migration speed of molecules during measurement is not uniform. Patent Document 1 discloses labeling repeat sequences on a genome and identifying their positions in the genome, but does not disclose a method for comparing the labeling positions between the measured labeling interval and a reference genome.
[0009] Japanese Patent Application No. 2023-051508 attempts to perform alignment appropriately even with the above-mentioned errors by using the ratio of label spacing. However, since the label positions measured by genome mapping have an error of several hundred bases, if the label spacing is short, such as several hundred bases, the error can significantly change the ratio value, making alignment impossible.
[0010] The present invention has been made in consideration of the above-mentioned problems, and aims to perform alignment that can deal with apparent elongation and contraction of the target nucleic acid sequence even when the labeling interval is short.
[0011] The information processing device of the present invention identifies the position of a subsequence in a reference nucleic acid sequence of a target nucleic acid sequence, for example, by using a ratio of the spacing between two or more labels, and sets at least one of the spacing used as the denominator in the ratio value or the spacing used as the numerator in the ratio value to the spacing between three or more labels possessed by the target nucleic acid sequence or the reference nucleic acid sequence.
[0012] According to the information processing device of the present invention, even when a portion where the labeling interval is short is included, it is possible to perform alignment that can deal with the apparent elongation and contraction of the nucleic acid molecule. Problems, configurations, and effects other than those described above will become clear from the description of the following embodiments.
[0013] FIG. 1 is a block diagram showing an example of the configuration of a genome marker position alignment device according to embodiment 1. FIG. 2 is a diagram showing an example of the data configuration of measurement data 140. FIG. 3 is a flowchart showing an example of processing by the index construction unit 121. FIG. 4 shows an example of a method for mitigating the problem of measurement errors. FIG. 5 shows another example of a method for mitigating the problem of measurement errors. FIG. 6 shows an example of a command used to specify a combination of marker intervals for the genome marker position alignment device. FIG. 7 is a schematic diagram showing processing when an apparatus according to embodiment 2 calculates a ratio of marker intervals. FIG. 8 is a schematic diagram showing processing when an apparatus according to embodiment 3 calculates a ratio of marker intervals. FIG. 9 shows an example of a user interface 200 displayed by the input / output device 102.
[0014] Hereinafter, a genome label position alignment device according to an embodiment of the present invention will be described. In the following drawings, the same components are designated by the same reference numerals, and duplicated explanations will be omitted.
[0015] As described above, when calculating the ratio of label spacing to correspond to the apparent expansion or contraction of the overall length of a DNA fragment, if the label spacing is short, the magnitude of the ratio may be significantly affected by measurement errors in the label positions. The genome label alignment device of this embodiment realizes alignment processing that can address these problems.
[0016] <First Embodiment> Fig. 1 is a block diagram showing an example of the configuration of a genome-labeled position alignment device according to a first embodiment of the present invention. Fig. 1 is the same as Fig. 1 of the prior application except for the reference numerals, and will be described again in this specification. The genome-labeled position alignment device (information processing device) is configured by a computer 100. The computer 100 includes, for example, a CPU (Central Processing Unit) 110, a memory 111, an auxiliary storage device 112, and interfaces 113 to 115. The hardware included in the computer 100 is electrically connected, for example, via an internal communication line such as a bus.
[0017] The CPU 110 reads programs and data stored in the memory 111 and executes the programs stored in the memory 111. The CPU 110 includes a processor. The CPU 110 includes, for example, an index construction unit 121, an index search unit 122, and an alignment probability calculation unit 123, which are all functional units. The computer 100 functions as a genome marker position alignment device by the CPU 110 executing processing.
[0018] The memory 111 temporarily stores programs executed by the CPU 110 and data used during program execution. The memory 111 includes a non-volatile storage element, read-only memory (ROM), and a volatile storage element, random access memory (RAM). The ROM stores unchanging programs (e.g., a basic input / output system (BIOS)). The RAM is a high-speed, volatile storage element, such as dynamic random access memory (DRAM), and temporarily stores programs executed by the CPU 110 and data used during program execution.
[0019] The memory 111 stores, for example, programs that realize an index construction unit 121, an index search unit 122, and an alignment probability calculation unit 123. The memory 111 further stores a reference genome label position 130 and measurement data 140.
[0020] For example, CPU 110 functions as index construction unit 121 by operating in accordance with an index construction program loaded into memory 111, functions as index search unit 122 by operating in accordance with an index search program loaded into memory 111, and functions as alignment probability calculation unit 123 by operating in accordance with an alignment probability calculation program loaded into memory 111.
[0021] The auxiliary storage device 112 stores in a nonvolatile manner the programs executed by the CPU 110 and data used when the programs are executed. That is, the programs are read from the auxiliary storage device 112, loaded into the memory 111, and executed by the CPU 110.
[0022] The auxiliary storage device 112 is a large-capacity, non-volatile storage device such as a hard disk drive (HDD) or a solid state drive (SSD). The auxiliary storage device 112 stores programs that realize the functions of the index construction unit 121, the index search unit 122, and the alignment probability calculation unit 123. The auxiliary storage device 112 also stores a reference genome label position 130 and measurement data 140. Here, the label position means, for example, a position where a predefined short partial sequence of 10 bases or less appears in the sequence.
[0023] Each of the interfaces 113 to 115 is a device that mediates the transmission and reception of signals and converts protocols, and is connected to an external device. The interface 113 is an I / O interface that is connected to the input / output device 102 via a wired or wireless line. The input / output device 102 includes input devices such as a keyboard, a mouse, a touch panel, a numeric keypad, a scanner, a microphone, and a sensor, as well as output devices such as a display device, a printer, and a speaker. The interface 113 acquires input information from an operator that is accepted by the input / output device 102. The interface 113 also outputs the results of program execution to the input / output device 102 in a format that can be visually recognized by the operator.
[0024] The interface 115 is a network interface that is connected to the external storage device 101 via the network 105. The interface 115 controls communication with other devices in accordance with a predetermined protocol.
[0025] The external storage device 101 is a non-transitory storage device that stores data handled by the computer 100. The external storage device 101 includes, for example, a storage device such as an HDD or SSD. The external storage device 101 can store reference genome marker positions 130 and measurement data 140.
[0026] Data is transmitted and received between the external storage device 101 and the computer 100 via a network 105. The network 105 includes, for example, a local area network (LAN) and the Internet. Note that the types of the network 105 are not limited to those described above. The network 105 may be configured as a wired or wireless network.
[0027] The interface 114 is connected to a drive device that reads and writes data from and to the removable medium 103. The interface 114 includes, for example, a serial interface such as a USB (Universal Serial Bus).
[0028] The removable medium 103 is a non-transitory storage medium that stores data handled by the computer 100. The removable medium 103 includes optical disks such as CDs and DVDs, magnetic disks, and semiconductor memories. The removable medium 103 can store reference genome marker positions 130 and measurement data 140.
[0029] Some or all of the programs executed by CPU 110 may be provided to computer 100 from removable media 103, which is a non-transitory storage medium, via interface 114, or from external storage device 101, which is a non-transitory storage device, or from an external computer equipped with external storage device 101, via network 105, and stored in non-volatile auxiliary storage device 112, which is a non-transitory storage medium.
[0030] 1, an external storage device 101 and removable media 103 are connected to a computer 100 constituting the genome label position alignment device, but these external devices can be omitted if not required. An input / output device 102 equipped with the external storage device 101 may also be connected to the computer 100 via a network 105. Instead of being connected to the input / output device 102, the computer 100 may incorporate a device equipped with input / output functions.
[0031] The index construction unit 121 constructs an index based on k-tuples. The index construction unit 121 does not construct an index using k-tuples that indicate the interval between marker positions (hereinafter simply referred to as marker interval), but constructs an index using k-tuples based on the ratio of marker intervals, as will be described later. Furthermore, the index constructed by the index construction unit 121 is also capable of dealing with the case where some markers are missing, as in the prior application, in order to deal with the possibility of marker detection failure.
[0032] A k-tuple is a combination of k numbers (k is a predefined parameter). Assuming that, as in the prior art, each k-tuple is a combination of k labeling intervals on the reference genome, and an index is constructed as a correspondence table between each k-tuple and the identifier of the label on the reference genome that corresponds to the labeling interval in the k-tuple.
[0033] In this case, simply generating a k-tuple in the measurement data 140 and performing an alignment by comparing the generated k-tuple with the k-tuple indicated by the index cannot address the apparent expansion and contraction of the measurement data 140.
[0034] Furthermore, even if k-tuples are generated by expanding or contracting the entire measurement data 140 at various expansion / contraction rates to deal with the apparent expansion / contraction of the measurement data 140 and the k-tuples corresponding to these expansion / contraction rates are compared with the k-tuples indicated by the index, it would be necessary to try a huge number of expansion / contraction rates to improve the accuracy of the alignment. This would increase processing time. Moreover, even if a huge number of expansion / contraction rates were used, there is a risk that all of the expansion / contraction rates tried would deviate from the optimal expansion / contraction rate for the measurement data 140, which could result in inaccurate alignment.
[0035] When measuring DNA fragments, the migration speed of each DNA fragment molecule differs from molecule to molecule, so the individual label intervals in the measurement data 140 themselves expand and contract significantly, but the expansion rate is constant for one measurement data 140. Therefore, in this embodiment, as described above, the index construction unit 321 constructs an index using k-tuples based on the ratio of label intervals, thereby making it possible to achieve highly accurate alignment without having to try multiple expansion rates for the measurement data 140.
[0036] The index search unit 122 identifies positions on the reference genome that are likely to correspond to each piece of measurement data, using the index based on the k-tuple constructed by the index construction unit 121. During the search, the index search unit 122 executes processing to deal with erroneous detection of markers.
[0037] The alignment probability calculation unit 123 calculates the probability that the generated alignment will occur. When there are multiple candidates for the position where the DNA fragments indicated by the measurement data 140 are aligned, if the probability is high, the alignment is likely to be correct, so the alignment probability calculation unit 123 adopts the alignment with the highest probability. Even if only part of the DNA fragments can be aligned, the probability is used to determine whether the alignment is significant or a coincidence.
[0038] Before the start of the genome label position alignment process, the reference genome label position 130 and the measurement data 140 are input and stored in the computer 100. The CPU 110 may read the reference genome label position 130 and the measurement data 140 and load them into the memory 111, for example, when the computer 100 is started or when processing is executed.
[0039] The reference genome marker positions 130 and the measurement data 140 may be stored in all of the auxiliary storage device 112, the external storage device 101, and the removable media 103, or may be stored in only some of them. These data may be moved or copied to the external storage device 101 or the removable media 103 when the computer 100 is stopped or when the auxiliary storage device 112 runs out of free space. Therefore, it is desirable that the reference genome marker positions 130 stored in different storage devices all contain the same information. The same applies to the measurement data 140.
[0040] As will be described later, the present application extends the k-tuple creation method compared to the prior application, and improves alignment accuracy by mitigating ratio changes due to measurement errors even when the marker interval is short.
[0041] The reference genome label position 130 includes, for each of a plurality of reference genomes, numerical data representing the position of a label present on the reference genome. The measurement data 140 includes, for each of a large number of DNA fragments, data obtained by measuring the label position on the DNA fragment.
[0042] 2 is a diagram showing an example of the data configuration of the measurement data 140. The measurement data 140 indicates an ID for identifying each DNA fragment (an example of a target nucleic acid sequence), the molecular length (length of the base sequence) of each DNA fragment, and the position of a label (an example of a partial sequence) on each DNA fragment. Note that blank spaces in the measurement data 140 indicate that no label was observed.
[0043] For example, the molecular length of a DNA fragment with ID "2" is 44,951 bases, four labels are measured in the DNA fragment, and the positions of the four labels are 10,844 bases, 19,749 bases, 23,353 bases, and 35,735 bases from the beginning of the DNA fragment, respectively. Note that the positions of the labels of the DNA fragment indicated by the measurement data 140 may indicate, for example, the position of the beginning or the position of the end of the label.
[0044] In this embodiment, an example of a short base sequence to be labeled on a DNA fragment is GCTCTTC, which is recognized by an enzyme called Nt.BspQI. As described above, the measurement data 140 shows the labeled positions of GCTCTTC on each DNA fragment in ascending order. In other words, the measurement data 140 represents information in which the labeled positions of each DNA fragment are converted into a numeric sequence (an example of a second numeric sequence) in ascending order.
[0045] 3 is a flowchart showing an example of processing by the index construction unit 121. The index construction unit 121 constructs a k-tuple based on the ratio of the marker interval for any marker position on the reference genome, and constructs an index indicating the correspondence between the constructed k-tuple and the marker position on the reference genome, by the following steps.
[0046] Step S301: The index construction unit 121 acquires genome sequence data. The genome sequence data indicates the number of each of a plurality of chromosomes and the base sequence of each of the plurality of chromosomes (an example of a referenced nucleic acid sequence). The base sequence of each of the plurality of chromosomes is indicated by a character string consisting of letters representing four types of bases, A, T, G, and C, and N representing an unknown base. The genome sequence data is stored in advance in at least one of the memory 111, the auxiliary storage device 112, the removable media 103, and the external storage device 101, for example.
[0047] Step S302: If the index construction unit 121 determines that all chromosomes have been selected, it completes the index construction process, and if it determines that there are unselected chromosomes, it proceeds to step S303.
[0048] Step S303: The index construction unit 121 selects one unprocessed chromosome. The index construction unit 121 constructs a k-tuple for the selected chromosome and registers it in the index in the following procedure.
[0049] Step S304: The index construction unit 121 calculates a numerical sequence (an example of a first numerical sequence in the claims) indicating the position of the marker (an example of a partial sequence in the claims) in the chromosome selected in the most recent step S303, and stores the numerical sequence in association with the chromosome number in the reference genome marker position 130. For example, the aforementioned GCTCTTC is used as the marker sequence. It should be noted that, since genomic DNA is a double helix, a portion of the complementary sequence (a sequence in which A, T, G, and C are replaced with T, A, C, and G, respectively, in reverse order) that matches the marker sequence is also labeled in genome mapping. Therefore, when calculating the numerical sequence representing the marker position, the index construction unit 321 must add the position of the marker sequence or its complementary sequence to the numerical sequence without distinction.
[0050] Step S305: The index constructor 121 registers the k-tuple calculated as described below in the index based on the numeric sequence indicating the marker position calculated in step S304, and returns to step S302.
[0051] 4 and 5 show specific methods for calculating the k-tuple in step S305. These calculations can be performed by the index construction unit 121. When calculating the ratio value of the jth (j=1, 2, ...) k-tuple starting from marker number i, where d(i1, i2) is the marker interval between markers i1 and i2, the prior application used the marker interval ratio r = d(i+j-1, i+j) / d(i+j, i+j+1). However, when d(i+j-1, i+j) or d(i+j, i+j+1) is short, on the order of several hundred bases, the measurement error is also several hundred bases, and the ratio value r may change significantly due to the measurement error. Therefore, in this embodiment, the problem of measurement error is mitigated by taking the sum of multiple marker intervals.
[0052] FIG. 4 shows an example of a method for mitigating the problem of measurement error. In FIG. 4, instead of d(i+j, i+j+1) used as the denominator in the prior application, the sum of all marker intervals in the range in which k-tuples are constructed is used. That is, d(i, i+k) is uniformly used as the denominator value for k-tuples starting from marker i. Therefore, for example, the first ratio value (ratio value 1 in FIG. 4) is r = d(1+1-1, 1+1) / d(1,1+3) = d(1,2) / d(1,4) = marker interval 1 / (marker interval 1 + marker interval 2 + marker interval 3).
[0053] FIG. 5 shows another example of a method for mitigating the problem of measurement errors. In FIG. 5, the numerator and denominator are each the sum of multiple marker intervals. For example, if the numerator and denominator are each the sum of two marker intervals (intervals between three markers), then r = d(i + j - 1, i + j + 1) / d(i + j, i + j + 2). Therefore, for example, the first ratio value (ratio value 1 in FIG. 5) is r = d(1 + 1 - 1, 1 + 1 + 1) / d(1 + 1, 1 + 1 + 2) = d(1, 3) / d(2, 4) = (marker interval 1 + marker interval 2) / (marker interval 2 + marker interval 3).
[0054] FIG. 6 shows an example of a command used to specify a combination of labeling intervals for the genome label position alignment system. In FIG. 6, the combination of labeling intervals is given by four vectors A, B, C, and D. Vector ABCD is configured so that any combination in the prior application, FIG. 4, or FIG. 5 can be specified. Vector ABCD is a vector of length k+1 consisting of 0 or 1. When the xth element (x=1, 2, ...) in each vector is 1, and the starting label number is y, d(x+y-1, x+y) is summed. Vectors A and B represent the labeling intervals to be summed in the numerator, and vectors C and D represent the labeling intervals to be summed in the denominator. An example is shown here where k=3. Note, however, that this example includes cases where the number of labeling intervals to be summed in the numerator or denominator is greater than those described in FIGS. 4 and 5. The length of the vector is set to match the maximum number of labeling intervals that can be used.
[0055] As with the denominator in Fig. 4, when a common value is used without changing the number of the marker interval to be used even when the number of the ratio value changes, the value of the denominator is specified using vector C. Similarly, when a common value is used for the numerator, the value of the numerator is specified using vector A. As with the numerator in Fig. 4 or the numerator in Fig. 5, when the number of the marker interval to be used changes when the number of the ratio value changes, the value of the numerator is specified using vector B. Similarly, as with the denominator, when the number of the marker interval to be used changes when the number of the ratio value changes, the value of the denominator is specified using vector D.
[0056] The example on the left side of Figure 6 is an example of a combination of label intervals in the prior application. As shown in the middle of Figure 4 or the middle of Figure 5, for both the numerator and denominator, as the ratio number increases, the label interval used also increases accordingly, so vectors A and C are not used, but vectors B and D are used. Therefore, A = C = 0000. For the numerator, only one label interval number that is the same as the ratio number is added together, so B = 1000. For the denominator, only one label interval number that is one step ahead of the ratio number is added together, so D = 0100.
[0057] The example in the center of Figure 6 is an example of a combination of indicator intervals in Figure 4. As shown in the bottom of Figure 4, for the numerator, as the ratio number increases, the indicator interval used also increases, so vector A is not used and A = 0000, while only one indicator interval with the same number as the ratio number is added together, resulting in B = 1000. For the denominator, the ratio number is common within the same tuple even if it increases, so vector D is not used and D = 0000, while three indicator intervals are used starting from the indicator interval with the same number as the first ratio number, resulting in C = 1110.
[0058] The example on the right side of Figure 6 is an example of a combination of the indicator intervals in Figure 5. As shown in the bottom of Figure 5, for both the numerator and denominator, as the ratio number increases, the indicator interval used also increases accordingly, so vectors A and C are not used, and vectors B and D are used. Therefore, A = C = 0000. For the numerator, an interval of three indicators is used starting from the indicator interval that is the same as the ratio number, so B = 1110. For the denominator, an interval of three indicators is used starting from the indicator interval that is one after the ratio number, so D = 0111.
[0059] The device of the present invention stores the k-tuple calculated as above in an index and uses it, as in the prior application. The alignment process after referencing the k-tuple index is the same as in the prior application.
[0060] As described above, the genome label position alignment device in embodiment 1 can alleviate the problem of short label spacings changing significantly due to measurement errors by adding up multiple label spacings when calculating ratio values in alignments using k-tuples.
[0061] 7 is a schematic diagram showing the process of calculating the ratio of labeling intervals by an apparatus according to a second embodiment of the present invention. In the second embodiment, in order to prevent errors caused by short labeling intervals, a k-tuple is constructed for data from which short labeling intervals have been removed in the reference genome.
[0062] Specifically, the parameter W 1 For d(x, x+1) < W 1 When x is an integer such that i≦x<j, the labels i to j are removed and a k-tuple is added when they are replaced with a virtual new label i'. The position of label i' is the midpoint between the positions of labels i and j. When referring to the k-tuple index constructed after removing short labeling intervals based on FIG. 7, the k-tuple constructed for the data from which short labeling intervals have been removed in the measurement molecule is used together with the normal k-tuple. However, the alignment process after referring to the k-tuple index constructed based on FIG. 7 is the same as in the prior application.
[0063] The device of the present invention extends the alignment in the same way as the prior application by sequentially associating labels around the k-tuples registered in the index constructed based on FIG.
[0064] <Third embodiment> Fig. 8 is a schematic diagram showing the process when an apparatus according to a third embodiment of the present invention calculates the ratio of labeling intervals. In this embodiment, in order to prevent errors caused by short labeling intervals, a corrected value d'(x, x+1) = f(d(x, x+1)) is used instead of the short labeling interval d(x, x+1) in the reference genome. The function f(t) is a function used for correction, and the parameter W 2 When t is large, it is close to t, and as t approaches 0, W 2 An example of such a function is the exponential function f(x) = x + W 2 ×exp(-Cx) can be used. exp(·) is the exponential function, C is a constant coefficient, and f(x)≧W for any x. 2 is selected to satisfy.
[0065] As mentioned in the problem to be solved by the invention, when the marker interval is short, a large error may occur when calculating the ratio of the marker interval. Therefore, as shown in FIG. 8, when the marker interval is short, the marker interval is set to the lower limit W 2 By making the correction so that the ratio becomes as shown above, it is possible to avoid such large errors. As this process forcibly changes the sign spacing, it is generally best not to perform it. However, when the sign spacing is small, the disadvantage of the ratio becoming too large outweighs the benefits, so we have decided to allow the correction shown in Figure 8 to that extent.
[0066] <Fourth Embodiment> Fig. 9 shows an example of a user interface 200 displayed by the input / output device 102. The user interface 200 includes, for example, an input data setting area 210, an alignment result display area 220, and an alignment calculation result 230 in the index construction unit 121. The input data setting area 210 specifies a data file to be input to the computer 100. The alignment result display area 220 shows the result of alignment performed by the computer 100. The alignment calculation result 230 shows the calculation result of the ratio value described in Figs. 4 and 5.
[0067] <Modifications of the Present Invention> The present invention is not limited to the above-described embodiment, and various modifications are included. For example, the above-described embodiment has been described in detail to clearly explain the present invention, and is not necessarily limited to an embodiment including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0068] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the program, table, and file that implements each function can be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD.
[0069] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected.
[0070] 100 Computer, 102 Input / Output Device, 110 CPU, 111 Memory, 112 Auxiliary Storage Device, 113 Interface, 121 Index Construction Unit, 122 Index Search Unit, 123 Alignment Probability Calculation Unit, 130 Reference Genome Label Position, 140 Measurement Data, 701 Index
Claims
1. An information processing apparatus for specifying the position of a subsequence of a reference nucleic acid sequence corresponding to a target nucleic acid sequence, comprising: - a storage device that stores data describing the labeled positions of the target nucleic acid sequence; - a processor that specifies the position of the subsequence using the data, - wherein the processor specifies the interval between labels of the target nucleic acid sequence and the reference nucleic acid sequence, - and the processor specifies the position of the subsequence of the target nucleic acid sequence in the reference nucleic acid sequence using the ratio of two or more of the intervals, - and the processor uses, as at least one of the intervals used as the denominator in the value of the ratio or the interval used as the numerator in the value of the ratio, an interval between three or more labels of the target nucleic acid sequence or the reference nucleic acid sequence.
2. The processor: - specifies a first label interval between a first label of the target nucleic acid sequence and a second label of the target nucleic acid sequence, and a second label interval between a third label of the target nucleic acid sequence and the second label, respectively, - and the processor uses the first label interval as the numerator and obtains the denominator by summing at least the first label interval and the second label interval.
3. The processor calculates a first ratio and a second ratio as the ratio from a plurality of the intervals, - and the processor uses an interval between three or more labels included in the plurality of the intervals as a common denominator between the first ratio and the second ratio, - and the processor uses, as the numerator of the first ratio, a first label interval between a first label of the target nucleic acid sequence and a second label of the target nucleic acid sequence, - and the processor uses, as the numerator of the second ratio, a second label interval between a third label of the target nucleic acid sequence and the second label.
4. The processor identifies the first label interval between the first label and the second label of the target nucleic acid sequence, the second label interval between the third label and the second label of the target nucleic acid sequence, and the third label interval between the fourth label and the third label of the target nucleic acid sequence, respectively. The processor obtains the numerator by summing at least the first label interval and the second label interval, and the processor obtains the denominator by summing at least the second label interval and the third label interval. The information processing apparatus according to claim 1, characterized in that.
5. The processor identifies the first label interval between the first label and the second label of the target nucleic acid sequence, the second label interval between the third label and the second label of the target nucleic acid sequence, the third label interval between the fourth label and the third label of the target nucleic acid sequence, and the fourth label interval between the fifth label and the fourth label of the target nucleic acid sequence, respectively. The processor calculates a first ratio and a second ratio as the ratio from a plurality of the intervals. The processor obtains the numerator of the first ratio by summing at least the first label interval and the second label interval. The processor obtains the denominator of the first ratio by summing at least the second label interval and the third label interval. The processor obtains the numerator of the second ratio by summing at least the second label interval and the third label interval. The processor obtains the denominator of the second ratio by summing at least the third label interval and the fourth label interval. The information processing apparatus according to claim 1, characterized in that.
6. The processor is configured to receive an instruction to specify the interval used as the numerator and the interval used as the denominator using the number of the label, and calculate the ratio of the number corresponding to the number of the label according to the instruction. The instruction includes: when the numerator is common across the ratios of a plurality of numbers and the numerator is calculated by summing the intervals between one or more consecutive labels, the number of labels constituting the numerator and the number of the label at which the summing starts; when the denominator is common across the ratios of a plurality of numbers and the denominator is calculated by summing the intervals between one or more consecutive labels, the number of labels constituting the denominator and the number of the label at which the summing starts; when the numerator changes as the number of the ratio changes and the numerator is calculated by summing the intervals between one or more consecutive labels, the number of labels constituting the numerator and the number of the label at which the summing starts; when the denominator changes as the number of the ratio changes and the denominator is calculated by summing the intervals between one or more consecutive labels, the number of labels constituting the denominator and the number of the label at which the summing starts. The information processing apparatus according to claim 1, characterized in that it is constituted by the above.
7. When the interval between the first label of the target nucleic acid sequence and the second label of the target nucleic acid sequence is less than or equal to a threshold value, the processor replaces the first label and the second label with one label. The processor obtains the ratio for the target nucleic acid sequence after the replacement. The information processing apparatus according to claim 1, characterized in that it is so configured.
8. The processor is configured to identify a partial sequence of the target nucleic acid sequence by comparing the interval between the labels of the reference nucleic acid sequence with the ratio. When the interval of the reference nucleic acid sequence or the target nucleic acid sequence is less than or equal to a lower limit value, the processor corrects the interval of the reference nucleic acid sequence to be the lower limit value or a value greater than the lower limit value, and then compares it with the ratio. The information processing apparatus according to claim 1, characterized in that it is so configured.
9. The data describes the labeled positions of the reference nucleic acid sequence, the processor specifies the interval between the labels of the reference nucleic acid sequence as a reference interval, the processor calculates the ratio of two of the reference intervals as a reference ratio, and the processor compares the ratio calculated from the target nucleic acid sequence with the reference ratio to identify the correspondence between the subsequence in the reference nucleic acid sequence and the subsequence in the target nucleic acid sequence. The information processing apparatus according to claim 1, characterized in that.
10. The information processing apparatus further includes an input / output device that provides a user interface, and the user interface has an input data setting unit that designates data input to the information processing apparatus. The information processing apparatus according to claim 1, characterized in that.
11. The information processing apparatus further includes an input / output device that provides a user interface, and the user interface includes an alignment result display unit that presents the result of the processor identifying the position of the subsequence, and an alignment calculation result display unit that presents the value of the ratio calculated by the processor. The information processing apparatus according to claim 1, characterized in that.
12. An information processing method for identifying the position of a subsequence of a reference nucleic acid sequence corresponding to a target nucleic acid sequence, the method including the step of reading the data from a storage device that stores data describing the labeled positions of the target nucleic acid sequence and using the data to identify the position of the subsequence. In the identifying step, the interval between the labels of the target nucleic acid sequence is identified. In the identifying step, the position of the subsequence of the reference nucleic acid sequence corresponding to the target nucleic acid sequence is identified using the ratio of two of the intervals. In the identifying step, at least one of the intervals used as the denominator or the intervals used as the numerator in the value of the ratio is the interval between three or more labels of the target nucleic acid sequence or the reference nucleic acid sequence. An information processing method, characterized by that.
13. An information processing program for causing a computer to execute a process of specifying the position of a partial sequence of a reference nucleic acid sequence corresponding to a target nucleic acid sequence, the program causing the computer to: read data storing the labeled positions of the target nucleic acid sequence from a storage device and use the data to execute a step of specifying the position of the partial sequence; in the specifying step, cause the computer to execute a step of specifying the interval between labels of the target nucleic acid sequence; in the specifying step, cause the computer to execute a step of specifying the position of the partial sequence of the reference nucleic acid sequence corresponding to the target nucleic acid sequence using the ratio of two of the intervals; in the specifying step, cause the computer to execute a step of setting at least one of the interval used as the denominator in the value of the ratio or the interval used as the numerator in the value of the ratio to be the interval between three or more labels of the target nucleic acid sequence or the reference nucleic acid sequence. An information processing program characterized by the above.