A human mtDNA rapid comparison method and device with adaptive learning control capability
By dividing mtDNA into multi-level sub-libraries and adaptively sorting and aligning them, the problem of low mtDNA alignment efficiency under large library capacity is solved, and fast and accurate mtDNA alignment is achieved.
Patent Information
- Application Number
- CN202310708134.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-14
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2043-06-14
AI Technical Summary
Existing mtDNA comparison methods are inefficient in large library volumes and cannot effectively eliminate meaningless comparisons. This is especially true in multi-ethnic countries and when maternal kinship is close, making it difficult to perform rapid and accurate comparisons.
The mtDNA is divided into multiple primary sub-libraries, and further divided into secondary sub-libraries. Through adaptive sorting and difference order calculation, locus values are prioritized to eliminate meaningless alignments and improve alignment efficiency.
It significantly improves the speed and accuracy of comparison under large storage capacity conditions, is suitable for groups with close maternal kinship, reduces the number of comparisons, and reduces meaningless calculations.
Smart Images

Figure CN116844639B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of DNA search, and in particular to a human mtDNA rapid comparison method and device with adaptive learning control capability. Background Art
[0002] Mitochondrial DNA (mtDNA) is the genetic material found in mitochondria, which generate energy (ATP) for cells. It is a specialized form of deoxyribonucleic acid found within mitochondria, the organelles that provide cells with energy (ATP). A single mitochondrion typically contains multiple DNA molecules, each carrying its own DNA—mtDNA. Since mitochondria are primarily transmitted through egg cells, they are generally inherited from the mother.
[0003] Human mtDNA is a double-stranded, closed circular DNA molecule approximately 16,569 base pairs long, encompassing 37 genes. According to the Mitochondrial DNA Testing Specification for Forensic Evidence Identification (SF / Z JD0105008-2018), the presence of two or more base differences in samples can rule out the possibility that the two samples came from the same individual or maternal line. Due to the multi-copy nature of mitochondria within cells and the relatively short length and indestructibility of their DNA, mtDNA is widely used in scenarios such as maternal kinship identification, identification of corpses from major disasters, and identification of highly degradable and difficult specimens.
[0004] There are currently two main conventional methods for mtDNA comparison. The first method is to directly compare two complete mtDNA sequences to count their differences, but this is only suitable for situations where the two sequences are exactly the same length. Although the mtDNA of most people (about 70%) is 16569 in length, there are still many insertions or deletions between individuals, that is, there are various mutations such as 16567- or 16571+. The second method can be called the rCRS comparison method. This method first calculates the difference sequence between each sample and the standard sequence (the international standard uses the Anderson reference sequence or the Cambridge reference sequence, abbreviated as rCRS), and then compares them based on the difference sequence. Although the second method has an extra pre-step, its compatibility is better than the first method, and the number of sequences to be compared subsequently will be significantly reduced. It is the current mainstream mtDNA sequence comparison method. However, this method still has great limitations. Its comparison efficiency is low, and it cannot be optimized according to a specific human population. The specific reasons are as follows:
[0005] On the one hand, the mutation rate of mtDNA is higher (compared to nuclear DNA), and it will show irregular changes in the actual massive population (such as a country with a vast territory and multi-ethnic groups like China). Therefore, the differential sequence of rCRS will also show irregular changes. The maximum difference may be as many as hundreds of bases, and the minimum difference may be only dozens of bases, and the distribution position is relatively random.
[0006] On the other hand, the mtDNA of people in the same region or with similar maternal clans are relatively close, and the length and sequence of rCRS are also very similar. It is often impossible to screen out the final result without comparing the last few bases.
[0007] In summary, in the actual comparison process, rCRS can basically only compare all database sequences and target sequences one by one, and then eliminate and sort them according to the number of differences, although the speed has been greatly improved compared to the full sequence. However, as the library capacity increases, especially considering that mtDNA also has heterogeneity problems (that is, the mtDNA of different sources, different organs or different cell types of the same person is not completely consistent, resulting in a person may have multiple samples), especially because databases in most places are often regional (that is, rCRS sequences are relatively similar), the actual number of base comparisons for a sample will usually easily rise to 100 million times (taking a library capacity of 1 million, 50 rCRS differences, and 2 types of samples for one person in the library as an example), and the library capacity in actual combat situations will continue to expand, and the number of differences will also be greater. This may not be a big problem for individual cases, but it is very tricky for daily business volumes, especially for large-scale screening. Therefore, the above problems need to be solved urgently. Summary of the Invention
[0008] In order to overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a method and device for rapid human mtDNA comparison with adaptive learning and control capabilities, which can eliminate a large number of meaningless comparisons, adaptively control the comparison order, and quickly perform human mtDNA searches in a large library capacity.
[0009] To achieve the above objectives, the present invention proposes a rapid human mtDNA comparison method with adaptive learning control capabilities, comprising the following steps:
[0010] Divide the mtDNA into multiple primary sub-libraries, and further divide the primary sub-libraries into secondary sub-libraries. Adaptively sort each sample into a database to form a database.
[0011] After the database is established, the total length of the mtDNA to be compared is calculated, and the primary sub-library corresponding to the mtDNA to be compared is selected according to the total length;
[0012] Calculating the difference sequence between the mtDNA to be compared and the standard sequence, and simultaneously confirming the locus value and the total number of differences of the mtDNA to be compared, and selecting a secondary sub-library according to the total number of differences;
[0013] According to the priority determined by the adaptive sorting, the locus value of the mtDNA to be compared is compared with the locus value of the mtDNA in the database.
[0014] Preferably, further dividing the primary sub-library into secondary sub-libraries comprises the following steps:
[0015] In each of the primary sub-libraries, the entire mtDNA library sample is divided into thirty-three regions according to a specified allocation rule, forming thirty-three position intervals, wherein the allocation rule is formed by allocating all base positions according to gene functional regions;
[0016] The difference sequences between the mtDNA library sample and the standard sequence are placed on the loci according to the position interval, the locus values are recorded, and the difference numbers are counted to form a difference number sum value, and multiple secondary sub-libraries are formed according to the difference number sum value, wherein the standard sequence is the Anderson reference sequence.
[0017] Preferably, dividing the mtDNA into multiple primary sub-libraries comprises the following steps:
[0018] The mtDNA library samples are divided into multiple primary sub-libraries according to their lengths, so as to exclude mtDNA library samples that necessarily have two or more base differences with the mtDNA to be compared.
[0019] Preferably, the adaptive sorting is performed for each sample stored in the database, including the following steps:
[0020] For each sample stored, the distribution frequencies of the thirty-three loci values were counted, and the comparisons of the loci were sorted in descending order according to the difference frequencies.
[0021] Preferably, calculating the difference sequence between the mtDNA to be compared and the standard sequence, simultaneously confirming the locus value and the sum of the difference numbers of the mtDNA to be compared, and selecting the secondary sub-library according to the sum of the difference numbers comprises the following steps:
[0022] Calculating the difference sequence between the mtDNA to be compared and the standard sequence, and simultaneously confirming the thirty-three loci values and the total number of differences of the mtDNA to be compared;
[0023] Based on the sum of the number of differences, a secondary sub-library that does not have a two-base difference standard with the mtDNA to be compared is selected.
[0024] Preferably, comparing the locus value of the mtDNA to be compared with the locus value of the mtDNA in the database according to the priority determined by the adaptive sorting comprises the following steps:
[0025] If the difference between the locus value of any mtDNA library sample and the locus value of the mtDNA to be compared is greater than or equal to 2, the mtDNA library sample currently being compared is excluded;
[0026] If the cumulative difference between the locus value of the mtDNA library sample and the locus value of the mtDNA to be compared is less than 2, or the locus values are exactly the same, base alignment of the mtDNA library sample and the mtDNA to be compared is performed;
[0027] When base alignment is performed between the mtDNA library sample and the mtDNA to be compared, the alignment is performed in the order of the number of loci with the highest frequency to the lowest frequency difference.
[0028] Preferably, the method further comprises:
[0029] In response to the completion of the sequential alignment, resulting sequences that meet the difference within two points are obtained, and the mtDNA library samples belonging to the same sample individual are summarized;
[0030] Sorting the mtDNA library samples according to their closeness to the mtDNA to be compared, wherein the closeness is ranked from high to low in terms of complete match and heterogeneity;
[0031] If there is one base difference between the mtDNA library sample and the mtDNA to be compared, the samples are sorted in order according to the differences caused by insertions and deletions, the differences caused by transitions, and the differences caused by transversions.
[0032] To achieve the above-mentioned object, the present invention further provides a human mtDNA rapid comparison device with adaptive learning control capability, comprising:
[0033] A database creation module is used to divide mtDNA into multiple primary sub-libraries, and further divide the primary sub-libraries into secondary sub-libraries. Each time a sample is stored in the library, it is adaptively sorted to form a database;
[0034] A primary sub-library selection module is used to calculate the total length of the mtDNA to be compared after the database is established, and select the primary sub-library corresponding to the mtDNA to be compared according to the total length;
[0035] A secondary sub-library selection module is used to calculate the difference sequence between the mtDNA to be compared and the library sample, obtain the locus value and the sum of the difference numbers of the mtDNA to be compared, and confirm the secondary sub-library according to the sum of the difference numbers;
[0036] The sequential base comparison module is used to compare the gene locus value of the mtDNA to be compared with the gene locus value of the mtDNA in the database according to the priority determined by the adaptive sorting.
[0037] Preferably, the module database creation module is specifically used to:
[0038] Dividing the mtDNA library samples into multiple primary sub-libraries according to their lengths, so as to exclude mtDNA library samples that necessarily have two or more base differences with the mtDNA to be compared;
[0039] In each of the primary sub-libraries, the entire mtDNA library sample is divided into thirty-three regions according to a specified allocation rule, forming thirty-three position intervals, wherein the allocation rule is formed by allocating all base positions according to gene functional regions;
[0040] The difference sequences between the mtDNA library sample and the standard sequence are placed on the loci according to the position interval, the locus values are recorded, and the number of differences is counted to form a total number of differences value, and multiple secondary sub-libraries are formed according to the total number of differences value, wherein the standard sequence is the Anderson reference sequence;
[0041] For each sample stored, the distribution frequencies of the thirty-three loci values were counted, and the comparisons of the loci were sorted in descending order according to the difference frequencies.
[0042] Preferably, the secondary sub-library selection module is specifically used to:
[0043] Calculating the difference sequence between the mtDNA to be compared and the standard sequence, and simultaneously confirming the thirty-three loci values and the total number of differences of the mtDNA to be compared;
[0044] Based on the sum of the number of differences, a secondary sub-library that does not have a two-base difference standard with the mtDNA to be compared is selected.
[0045] Compared with the prior art, an embodiment disclosed in the present invention comprehensively considers the length polymorphism of mtDNA and 37 functional loci of mtDNA, and adaptively learns to control the priority of the comparison position by continuously counting the frequency of the difference positions, thereby continuously reducing the number of final comparisons, greatly improving the comparison efficiency. It not only improves the comparison speed when the maternal mtDNA differences are large and the capacity library is large, but also can be widely applied when the maternal relationship is closer. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Flow chart of the steps of the human mtDNA rapid comparison method disclosed in the present invention;
[0047] Figure 2 A flowchart of a specific embodiment of the present invention;
[0048] Figure 3 This is a structural framework diagram of the human mtDNA rapid comparison device disclosed in the present invention. DETAILED DESCRIPTION
[0049] The following describes the embodiments of the present invention using specific examples and accompanying drawings. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. The present invention may also be implemented or applied through other different specific examples, and the details in this specification may be modified and altered based on different viewpoints and applications without departing from the spirit of the present invention.
[0050] Figure 1 Flow chart of the steps of the human mtDNA rapid comparison method disclosed in the present invention. Figure 1 As shown, the present invention provides a human mtDNA rapid comparison method with adaptive learning control capability, comprising the following steps:
[0051] Step S10: Divide the mtDNA into multiple primary sub-libraries, and further divide the primary sub-libraries into secondary sub-libraries. Adaptively sort each sample into a database to form a database.
[0052] Step S11, after the database is established, the total length of the mtDNA to be compared is calculated, and the primary sub-library corresponding to the mtDNA to be compared is selected according to the total length;
[0053] Step S12, calculating the difference sequence between the mtDNA to be compared and the standard sequence, and confirming the locus value and the total difference value of the mtDNA to be compared, and selecting a secondary sub-library according to the total difference value;
[0054] Step S13, comparing the locus value of the mtDNA to be compared with the locus value of the mtDNA in the database according to the priority determined by the adaptive sorting.
[0055] Specifically, this method pre-establishes an mtDNA database and further divides the database into multiple first-level sub-libraries. Under the first-level sub-libraries, multiple second-level sub-libraries are established, and the library samples are adaptively sorted to form a database that can quickly compare human mtDNA. Subsequently, the first-level sub-libraries, second-level sub-libraries and specific loci are compared.
[0056] Preferably, in step S10, dividing the mtDNA into a plurality of primary sub-libraries comprises the following steps:
[0057] The mtDNA library samples are divided into multiple primary sub-libraries according to their lengths, so as to exclude mtDNA library samples that necessarily have two or more base differences with the mtDNA to be compared.
[0058] Specifically, mtDNA is pre-sorted by length. During the subsequent comparison process, library samples with different lengths from the mtDNA to be compared can be identified as necessarily differing by two or more bases from the mtDNA to be compared, meaning that the mtDNA to be compared and these samples do not originate from the same individual or maternal lineage. In one embodiment of the present invention, data of various lengths, such as 16569, 16568, and 16570, are first sorted into libraries. If the mtDNA to be compared is 16569 in length, sequences in sub-libraries of 16567 or less and 16571 or more necessarily differ by two or more bases from it. Therefore, direct comparison is unnecessary, and only sub-libraries of 16568, 16569, and 16570 need to be considered. The same applies to other cases.
[0059] Preferably, in step S10, further dividing the primary sub-library into secondary sub-libraries comprises the following steps:
[0060] In each of the primary sub-libraries, the entire mtDNA library sample is divided into thirty-three regions according to a specified allocation rule, forming thirty-three position intervals, wherein the allocation rule is formed by allocating all base positions according to gene functional regions;
[0061] The difference sequences between the mtDNA library sample and the standard sequence are placed on the loci according to the position interval, the locus values are recorded, and the difference numbers are counted to form a difference number sum value, and multiple secondary sub-libraries are formed according to the difference number sum value, wherein the standard sequence is the Anderson reference sequence.
[0062] Specifically, the difference sequence between the mtDNA library sample and the standard sequence is the difference sequence between the original sequence of a person and the Anderson reference sequence (rCRS standard sequence). It should be noted that the rCRS standard sequence (referred to as rCRS) is the original sequence of the human mitochondrial DNA sequence that Anderson completed the measurement for the first time in 1981 at Sanger Laboratory in Cambridge, England. It is also known as the Anderson sequence or the Cambridge sequence. Because it is widely adopted internationally, it has formed a de facto international standard sequence and is used as a ruler for mitochondrial sequences. In addition, since mtDNA has a total of 37 loci, but some are distributed in the light chain, i.e., the inner ring, and some are distributed in the heavy chain, i.e., the outer ring, and there are 4 loci that overlap, so it is divided into 33 position intervals. In addition, in a specific embodiment of the present invention, the 33 position intervals of the allocation rule are specifically enumerated.
[0063] Preferably, in step S10, each time a sample is stored, adaptive sorting is performed, including the following steps:
[0064] For each sample stored, the distribution frequencies of the thirty-three loci values were counted, and the comparisons of the loci were sorted in descending order according to the difference frequencies.
[0065] Specifically, the ones with higher difference frequencies (that is, more diverse values) are placed at the front, and the ones with more unique values are placed at the back, which can reflect the type and characteristics of the database and improve the comparison efficiency.
[0066] Preferably, in step S12, the calculation of the difference sequence between the mtDNA to be compared and the standard sequence, the confirmation of the locus value and the sum of the difference number of the mtDNA to be compared, and the selection of the secondary sub-library according to the sum of the difference number include the following steps:
[0067] Calculating the difference sequence between the mtDNA to be compared and the standard sequence, and simultaneously confirming the thirty-three loci values and the total number of differences of the mtDNA to be compared;
[0068] Based on the sum of the number of differences, a secondary sub-library that does not have a two-base difference standard with the mtDNA to be compared is selected.
[0069] Specifically, comparing the difference sequence with the standard sequence (rCRS) is a traditional technical method, that is, taking rCRS as the basic library, and taking the human mitochondrial coding in hg19 as the reference library, performing mutation annotation, then all those different from rCRS are counted as mutations, and can be directly calculated and converted using annovar software, and what is obtained is the difference sequence, which will not be repeated here. It should be noted that, in a specific embodiment of the present invention, an example of the rCRS difference sequence obtained by existing technical means is provided, that is, the data of an individual, and the data can be used for the above-mentioned library construction link, and can also be used for the comparison link here. In addition, in a specific embodiment disclosed by the present invention, the calculation of the locus value and the difference quantity sum value is specifically as follows: for example, the difference value of "73G", the value of "73" on rCRS is "A", and the position of "73" belongs to the MT-RNR2 (1-556) locus interval, then the locus value of MT-RNR2 + 1 can be obtained, if there are other difference values on MT-RNR2, if there is one, a "+ 1" operation is performed, and the sum of the difference values is the difference quantity sum value.
[0070] Preferably, in step S13, comparing the locus value of the mtDNA to be compared with the locus value of the mtDNA in the database according to the priority determined by the adaptive sorting comprises the following steps:
[0071] If the difference between the locus value of any mtDNA library sample and the locus value of the mtDNA to be compared is greater than or equal to 2, the mtDNA library sample currently being compared is excluded;
[0072] If the cumulative difference between the locus value of the mtDNA library sample and the locus value of the mtDNA to be compared is less than 2, or the locus values are exactly the same, base alignment of the mtDNA library sample and the mtDNA to be compared is performed;
[0073] When base alignment is performed between the mtDNA library sample and the mtDNA to be compared, the alignment is performed in the order of the number of loci with the highest frequency to the lowest frequency difference.
[0074] Specifically, if the difference in the locus values is directly greater than or equal to 2, no base alignment is required and the bases can be directly excluded. In one embodiment of the present invention, for example, if MT-RNR2 is the first mtDNA locus to be compared, and the library sample has a value of 1 while the comparison sample has a value of 4, then the difference can be directly determined to be 3. If the loci values are identical or the cumulative difference is less than 2, the final base alignment is performed, and the positions and values must be exactly the same. Taking MT-RNR2 as an example, if the library sample is 73G, the comparison sample must also be 73G to be considered identical; insertions or deletions of 73 are not considered to be consistent. Because the alignment is performed first on loci with high-frequency differences, inconsistent data is often eliminated at the first few loci, eliminating the need for a full alignment. This method is particularly effective for databases with strong regional distribution or close maternal lineages, avoiding a large number of meaningless alignment calculations. To summarize the above, this method retains the idea of differential sequence in the rCRS method, but is completely different from the other party's rule. It first considers the length polymorphism of mtDNA, while also considering the 37 functional loci of mtDNA (the mutation itself is still biologically reasonable). Then, by continuously counting the frequency of differential positions, it adaptively learns to control the priority of the comparison position to continuously reduce the number of final comparisons and thus improve the comparison efficiency.
[0075] Preferably, the method further comprises the steps of:
[0076] In response to the completion of the sequential alignment, resulting in result sequences that differ within two points, the mtDNA library samples belonging to the same individual are summarized. Specifically, this step is intended to address or reduce the problem of heterogeneity. This means that each person's blood, internal organs, skin, hair, and other different types of organs and tissues may have subtle differences in mtDNA due to the long-term influence of the external environment and living environment. If a simple comparison is taken between a person's blood and their hair, there may be differences. Therefore, it is best to compare samples of the same type.
[0077] The mtDNA library samples are ranked according to their proximity to the mtDNA to be compared, with the degree of closeness, from highest to lowest, indicating complete match and heteroplasmy. Specifically, this step is a mechanism for evaluating results. A complete match suggests a high probability of origin or identity. Similarly, a discrepancy can be interpreted in two ways: if blood is compared to blood, the discrepancy is likely a true difference between two individuals; if blood is compared to hair, the discrepancy is likely due to heteroplasmy within the same individual, rather than a true individual difference.
[0078] If the mtDNA library sample differs by a single base from the mtDNA to be compared, the samples are sorted by differences due to insertions and deletions, transitions, and transversions. Specifically, this step also serves as a result evaluation mechanism. This sorting method is based on the probability of different mutations, with insertions and deletions having the highest probability, followed by A→G and T→C, and finally A→C and T→G having the lowest probability. The higher the probability, the more likely it is the same maternal lineage; the lower the probability, the more likely it is another maternal lineage.
[0079] Figure 2 This is a flowchart of a specific embodiment of the present invention. In this embodiment, the following steps are included:
[0080] Step S20: Establish a multi-level database, compare the library samples with the standard sequence, continuously perform statistical analysis on the comparison results, and place the loci that are more likely to produce differences at a higher priority level so that the difference value can be quickly obtained in the next comparison.
[0081] Specifically, first, a multi-level database is established; second, the library samples are compared with the standard sequence. It should be pointed out that the traditional method actually compares the bases one by one (characters such as ATCG) to see if they are the same. The comparison essence of this method converts the base content into corresponding numbers, and the range is more accurate and more suitable for computer calculations; third, the comparison results are continuously statistically analyzed, and the loci that are more likely to produce differences are placed at a higher priority in the comparison; finally, it is easier to compare the difference value in advance in the next comparison, thereby continuously shortening the comparison time.
[0082] Furthermore, step S20 includes the following steps:
[0083] Step S201: divide the mtDNA into different primary sub-libraries according to their lengths, so as to separate data of different lengths into sub-libraries in advance.
[0084] For example, data of various lengths such as 16569, 16568, and 16570 are first divided into sub-libraries. The purpose of the sub-library division is to make full use of the standard for two base differences in the "Specifications for Mitochondrial DNA Testing of Forensic Evidence Identification" (SF / Z JD0105008-2018). Specifically, if there are two or more base differences (excluding length heterogeneity) between the library sample and the sequence of the sample to be compared, it can be ruled out that the two samples come from the same individual or the same maternal line. In other words, if the length of the mtDNA to be compared is 16569, then the sequences in the sub-libraries of 16567 or less and 16571 or more must have two or more base differences with it, so there is no need to compare them directly. Only the three sub-libraries of 16568, 16569, and 16570 need to be considered, and the same applies to other cases.
[0085] Step S202 : In each primary sub-library, each mtDNA library sample is divided into 33 regions according to the allocation rule, forming 33 position intervals.
[0086] Specifically, the primary sub-library is divided according to the base length of each individual, for example, the majority is 16569, and the others are dispersed in a certain proportion between 1656X and 1657X, or even more primary sub-libraries; however, for the majority of 16569, this partitioning efficiency is far from sufficient, so it is necessary to further count and quantify the range of 33 loci in this sub-library, so that the population can be further divided, and so on. It should be noted that mtDNA has a total of 37 loci, but some are distributed in the light chain, i.e., the inner loop, and some are distributed in the heavy chain, i.e., the outer loop. There are four loci that overlap, so the entire mtDNA is divided into 33 position intervals.
[0087] The specific allocation rules are shown in the following table:
[0088]
[0089]
[0090] It should be noted that in order to cover all base positions and avoid duplication, the allocation rules presented in the above table have made slight adjustments to the base positions, but can cover the basic locus functional region, including the intermediate connecting band.
[0091] Step S203, the difference sequence between the mtDNA library sample and the standard sequence is placed on the locus according to the position interval, and the 33 locus values and the total difference value of the mtDNA library sample are obtained.
[0092] Specifically, the present invention provides an existing rCRS difference sequence example (length 1-16569), specifically:
[0093] 73G 150T 263G 315.1C 489C 523a 524c 750G 752T 1107C 2492R 2706G3107del 4742C 4769G 4883T 5178A 5301G 7028T 8701G 8860G 9180G 9540C 10397G10398G 10400T 10531Y 10873C 11719A 11944C 12026G 12705T 13356C 14766T 14783C15043A 15301A 15326G 15883A 16172C 16182c 16183c 16189C 16193.1c 16223T16362C 16519C
[0094] Now we will explain it in detail based on the above-mentioned allocation rules and existing difference sequence examples. For example, at position 73 of a certain sample, the standard sequence is base A, and the sample is base G. Then, according to the position interval, the MT-RNR2 of this sample (representing a region, such as the range of MT-RNR2 here is 1-556, which is called a locus) is added by 1, that is, the difference quantity value is added by 1, and its true value is recorded at the same time, that is, 73G (allele value, representing a single position, such as "73", its value is G, called an allele value); if the entire sample has no other difference sequences in the interval of 1-556, then the final value of MT-RNR2 of this sample is recorded as 1 (difference quantity value); and so on, and finally the 33 locus values of the sample and the sum of the difference quantities with the standard sequence are obtained. It should be noted that the standard sequence here uses the rCRS standard sequence, which is common knowledge in the technical field. Please refer to the end of the article for details.
[0095] Step S204: performing secondary database division based on the sum of the difference quantities, so as to further divide the data in the database.
[0096] Step S205: For each mtDNA library sample stored, the distribution frequencies of the 33 loci of all mtDNA are counted, and the comparison order of all loci is sorted according to the difference frequency.
[0097] Specifically, each time a sample is added to the database, the distribution frequency of the values of the 33 loci is counted and the comparisons of these loci are ranked in order, with loci with higher difference frequencies (i.e., more diverse values) placed higher, and loci with more uniform values placed lower. The more data is added to the database, the more this order reflects the type and characteristics of the database, and the higher the efficiency ratio.
[0098] It should be noted that the above steps S201 to S205 are not shown in the drawings.
[0099] Step 21: compare the mtDNA to be compared with the mtDNA library samples in the database.
[0100] Furthermore, step S21 includes the following steps:
[0101] Step 211 , calculating the total length of the mtDNA to be compared, and selecting a primary sub-library that does not have two base differences with the mtDNA to be compared for comparison.
[0102] In step S212, based on the differential sequences between the mtDNA to be compared and the standard sequence, the 33 loci values and the sum of the differences of the mtDNA to be compared are determined. Based on the sum of the differences, the corresponding secondary sub-library for the mtDNA to be compared is determined, and the secondary sub-library that does not meet the two-base difference standard with the mtDNA to be compared is selected for comparison. In other words, the selection of the secondary sub-library for the mtDNA to be compared also follows the two-base difference standard. Specifically, the method for calculating the differential sequences, loci values, and sum of the differences is the same as described in step S203 above and will not be repeated here.
[0103] Step S213: If the difference between the locus value of any mtDNA library sample and the locus value of the mtDNA to be compared is greater than or equal to 2, the mtDNA library sample currently being compared is excluded.
[0104] Specifically, if the value difference of the locus itself is directly greater than or equal to 2, there is no need to perform base matching and it can be directly excluded. For example, if MT-RNR2 is the first locus of the mtDNA to be compared, the value of the library sample is 1, and the value of the sample to be compared is 4, then it can be directly determined that its difference is equal to 3.
[0105] Step S214: If the cumulative difference between the locus value of the mtDNA library sample and the locus value of the mtDNA to be compared is less than 2, or the locus values are exactly the same, base alignment of the mtDNA library sample and the mtDNA to be compared is performed.
[0106] That is to say, if the locus values are the same or the cumulative difference is <2, the final base matching will be carried out. The position and value must be exactly the same. Taking MT-RNR2 as an example, if the library sample is 73G, the sample to be compared must also be 73G to be considered the same. Insertions or deletions of 73 do not match.
[0107] In step S215 , when performing base comparison between the mtDNA library sample and the mtDNA to be compared, the comparison is performed in the order of the number of loci with the highest frequency to the lowest frequency difference.
[0108] Specifically, because the comparison starts with high-frequency differential loci, inconsistent data is often eliminated at the first few loci, eliminating the need to compare all the loci. This approach is particularly effective for databases with strong regional distribution or similar maternal lineages, avoiding a large number of meaningless comparisons and calculations.
[0109] It should be noted that the above steps S211 to S215 are not shown in the accompanying drawings.
[0110] In step S22, upon completion of the comparison, the resulting sequences that differ by less than two points are sorted and, if necessary, stored in a database. Specifically, whether the mtDNA to be compared needs to be stored depends on its actual source and business needs. For example, samples from the field are often stored as evidence for future investigations. Samples from personnel who were not matched are often not stored because they are not relevant to the case.
[0111] Furthermore, step S22 includes the following steps:
[0112] In step S221, mtDNA library samples belonging to the same sample individual are summarized to reduce heterogeneity. It should be noted that the problem caused by heterogeneity is that one individual may have multiple samples.
[0113] Specifically, this step is to solve or reduce the problem of heterogeneity. That is, each person's blood, internal organs, skin, hair and other different types of organs and tissues may have subtle differences in mtDNA due to the long-term influence of the external and living environment. If you simply compare a person's blood and his hair, there may be differences, so it is best to compare samples of the same type.
[0114] Step S222: Sort the mtDNA library samples by their degree of similarity to the mtDNA to be compared, with the degree of similarity being ranked from highest to lowest, based on the degree of complete match and the degree of heterogeneity. That is, completely matched samples are ranked first, while samples with heterogeneity are ranked last.
[0115] Specifically, this step is a result evaluation mechanism. If they are completely consistent, then it can be considered that the probability that they come from the same maternal clan or are the same person is extremely high. For the same difference, there may be two situations. The first is blood compared with blood, then this difference may be a real difference between two people; if it is blood compared with hair, then this difference may be caused by the heterogeneity of the same person, rather than real individual differences.
[0116] In step S223, if there is a single base difference between the mtDNA library sample and the mtDNA to be compared, the samples are sorted by differences caused by insertions and deletions, differences caused by transitions, and differences caused by transversions. Specifically, differences caused by transitions refer to DNA bases that maintain the same number of loops, such as A→G and T→C; differences caused by transversions refer to DNA bases that change the number of loops, such as A→C and T→G.
[0117] Specifically, this step also serves as a result evaluation mechanism. This ranking is based on the probability of different mutations, with insertions and deletions having the highest probability, followed by A→G and T→C, and finally A→C and T→G having the lowest probability. The higher the probability, the more likely it is the same maternal lineage; the lower the probability, the more likely it is another maternal lineage.
[0118] It should be noted that the above steps S221 to S223 are not shown in the accompanying drawings.
[0119] Figure 3 This is a structural framework diagram of the human mtDNA rapid comparison device disclosed in the present invention. Figure 3 As shown, the present invention provides a human mtDNA rapid comparison device with adaptive learning control capability, comprising:
[0120] The database creation module 30 is used to divide the mtDNA into multiple primary sub-libraries, and further divide the primary sub-libraries into secondary sub-libraries. Each time a sample is stored in the library, it is adaptively sorted to form a database;
[0121] The primary sub-library selection module 31 is used to calculate the total length of the mtDNA to be compared after the database is established, and select the primary sub-library corresponding to the mtDNA to be compared according to the total length;
[0122] The secondary sub-library selection module 32 is used to calculate the difference sequence between the mtDNA to be compared and the library sample, obtain the locus value and the sum of the difference numbers of the mtDNA to be compared, and confirm the secondary sub-library according to the sum of the difference numbers;
[0123] The sequential base matching module 33 is used to compare the gene locus value of the mtDNA to be compared with the gene locus value of the mtDNA in the database according to the priority determined by the adaptive sorting.
[0124] Preferably, the module database creation module 30 is specifically used to:
[0125] Dividing the mtDNA library samples into multiple primary sub-libraries according to their lengths, so as to exclude mtDNA library samples that necessarily have two or more base differences with the mtDNA to be compared;
[0126] In each of the primary sub-libraries, the entire mtDNA library sample is divided into thirty-three regions according to a specified allocation rule, forming thirty-three position intervals, wherein the allocation rule is formed by allocating all base positions according to gene functional regions;
[0127] The difference sequences between the mtDNA library sample and the standard sequence are placed on the loci according to the position interval, the locus values are recorded, and the number of differences is counted to form a total number of differences value, and multiple secondary sub-libraries are formed according to the total number of differences value, wherein the standard sequence is the Anderson reference sequence;
[0128] For each sample stored, the distribution frequencies of the thirty-three loci values were counted, and the comparisons of the loci were sorted in descending order according to the difference frequencies.
[0129] Preferably, the secondary sub-library selection module 32 is specifically used to:
[0130] Calculating the difference sequence between the mtDNA to be compared and the standard sequence, and simultaneously confirming the thirty-three loci values and the total number of differences of the mtDNA to be compared;
[0131] Based on the sum of the number of differences, a secondary sub-library that does not have a two-base difference standard with the mtDNA to be compared is selected.
[0132] Preferably, the sequential base comparison module 33 is specifically used to:
[0133] If the difference between the locus value of any mtDNA library sample and the locus value of the mtDNA to be compared is greater than or equal to 2, the mtDNA library sample currently being compared is excluded;
[0134] If the cumulative difference between the locus value of the mtDNA library sample and the locus value of the mtDNA to be compared is less than 2, or the locus values are exactly the same, base alignment of the mtDNA library sample and the mtDNA to be compared is performed;
[0135] When base alignment is performed between the mtDNA library sample and the mtDNA to be compared, the alignment is performed in the order of the number of loci with the highest frequency to the lowest frequency difference.
[0136] Preferably, the apparatus further comprises an alignment result collating module for summarizing mtDNA library samples belonging to the same sample individual in response to completion of the sequential alignment and obtaining result sequences that meet the difference within two points;
[0137] Sorting the mtDNA library samples according to their closeness to the mtDNA to be compared, wherein the closeness is ranked from high to low in terms of complete match and heterogeneity;
[0138] If there is one base difference between the mtDNA library sample and the mtDNA to be compared, the samples are sorted in order according to the differences caused by insertions and deletions, the differences caused by transitions, and the differences caused by transversions.
[0139] It should be noted that this device corresponds to the above-mentioned human mtDNA rapid comparison method with adaptive learning control capabilities. For other undescribed parts, please refer to the contents of the method and will not be repeated here.
[0140] In addition, to ensure the completeness of the description, the present invention provides the rCRS standard sequence described in the above method and device, specifically:
[0141]
[0142]
[0143]
[0144]
[0145]
[0146]
[0147]
[0148]
[0149]
[0150]
[0151]
[0152] In summary, the present invention provides a rapid human mtDNA comparison method and device with adaptive learning control capabilities. Meaningless data is eliminated using length polymorphism as a primary sub-library for library classification, and data is secondary diverted using the sum of the number of differences as a secondary sub-library for library classification. The entire mtDNA is divided into 33 regions according to functional areas. During the library construction process, the comparisons are continuously sorted according to the number of difference sequences counted in the entire database, thereby achieving the purpose of adaptive learning control and greatly improving the human mtDNA comparison speed.
[0153] One aspect of the present invention can achieve the following beneficial effects:
[0154] (1) Compared with the current mainstream rCRS comparison technology, it has a huge speed improvement, especially in the case of large differences in maternal mtDNA, such as a database randomly sampled nationwide. Taking a database with a capacity of 1 million, 50 rCRS differences, and 2 types of samples per person in the database as an example, if the mainstream rCRS method is used to exclude samples based on 2 site differences, and the probability that the random difference population conforms to the normal distribution is that each sample must be compared with 25 difference points before it can be determined, then each sample to be compared must be compared 1 million x 25 x 2 = 50 million times. According to this method, the first length polymorphism sub-library can exclude at least 3-4 sub-libraries. Similarly, according to the normal distribution probability, the middle 3 standard deviations of 68.26% can also exclude more than 30% of the sub-libraries from comparison. Then, the second sub-library is divided according to the number of differences. Similarly, according to the normal distribution probability, the middle 3 standard deviations of 68.26% are taken, and more than 30% of the sub-libraries are excluded from comparison. Finally, because the comparison order of the 33 locations is sorted by the type of difference value, all possible variants can be eliminated in the first 3-5 locations. Assuming that the 50 rCRS are roughly evenly distributed across the 33 locations, the first 5 locations can be eliminated with a maximum of 8 comparisons. Therefore, the number of comparisons required for a single sample in this method is approximately 1 million x 8 x 2 x 0.6826 x 0.6826 = 7.455 million, which is over 6.7 times faster.
[0155] (2) The comparison speed of the present invention is greatly improved compared with the existing technology, which is also reflected in the case where the maternal relationship is closer, such as when a local database (such as a provincial or municipal database) is established. At this time, more similarities will be shown in the rCRS difference sequence. The existing rCRS standard sequence is essentially a sequence of a specific foreigner, so local databases are often more likely to produce a larger number (hundreds) of differences, and many of these differences will be repetitive. In this case, the number of comparisons of the mainstream method will increase rapidly, and the comparison order of the 33 locations in this method is sorted according to the type of difference value, which will produce stronger adaptability, so that highly similar locations are sorted more backward, thereby avoiding a large number of meaningless comparisons, making the comparison speed faster and the comparison efficiency higher.
[0156] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any skilled artisan may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be as set forth in the appended claims.
Claims
1. A rapid human mtDNA alignment method with adaptive learning control capability, comprising the following steps: Divide the mtDNA into multiple primary sub-libraries, and further divide the primary sub-libraries into secondary sub-libraries. Adaptively sort each sample into a database to form a database. After the database is established, the total length of the mtDNA to be compared is calculated, and the primary sub-library corresponding to the mtDNA to be compared is selected according to the total length; Calculating the difference sequence between the mtDNA to be compared and the standard sequence, and simultaneously confirming the locus value and the total number of differences of the mtDNA to be compared, and selecting a secondary sub-library according to the total number of differences; According to the priority determined by the adaptive sorting, the locus value of the mtDNA to be compared is compared with the locus value of the mtDNA in the database; the first-level sub-library is further divided into a second-level sub-library, comprising the following steps: In each primary sub-library, the entire mtDNA library sample is divided into 33 regions according to the specified allocation rules, forming 33 position intervals, among which, The allocation rule is formed by allocating all base positions according to the gene functional region; The difference sequences between the mtDNA library sample and the standard sequence are placed on the loci according to the position interval, the locus values are recorded, and the difference numbers are counted to form a difference number sum value, and multiple secondary sub-libraries are formed according to the difference number sum value, wherein the standard sequence is the Anderson reference sequence.
2. A human mtDNA rapid comparison method with adaptive learning control capability according to claim 1, characterized in that: The method of dividing the mtDNA into a plurality of primary sub-libraries comprises the following steps: The mtDNA library samples are divided into multiple primary sub-libraries according to their lengths, so as to exclude mtDNA library samples that necessarily have two or more base differences with the mtDNA to be compared.
3. A human mtDNA rapid comparison method with adaptive learning control capability according to claim 2, characterized in that: The adaptive sorting is performed for each sample stored in the database, including the following steps: For each sample stored, the distribution frequencies of the thirty-three loci values were counted, and the comparisons of the loci were sorted in descending order according to the difference frequencies.
4. A human mtDNA rapid comparison method with adaptive learning control capability according to claim 1, characterized in that: The step of calculating the difference sequence between the mtDNA to be compared and the standard sequence, simultaneously confirming the locus value and the total difference value of the mtDNA to be compared, and selecting a secondary sub-library according to the total difference value comprises the following steps: Calculating the difference sequence between the mtDNA to be compared and the standard sequence, and simultaneously confirming the thirty-three loci values and the total number of differences of the mtDNA to be compared; Based on the sum of the number of differences, a secondary sub-library that does not have a two-base difference standard with the mtDNA to be compared is selected.
5. The human mtDNA rapid comparison method with adaptive learning control capability according to claim 1, characterized in that: The method comprises comparing the locus value of the mtDNA to be compared with the locus value of the mtDNA in the database according to the priority determined by the adaptive sorting, comprising the following steps: If the difference between the locus value of any mtDNA library sample and the locus value of the mtDNA to be compared is greater than or equal to 2, the mtDNA library sample currently being compared is excluded; If the cumulative difference between the locus value of the mtDNA library sample and the locus value of the mtDNA to be compared is less than 2, or the locus values are exactly the same, base alignment of the mtDNA library sample and the mtDNA to be compared is performed; When performing base comparison between the mtDNA library sample and the mtDNA to be compared, the comparison is performed in the order of the number of loci with the highest frequency to the lowest frequency difference.
6. A human mtDNA rapid comparison method with adaptive learning control capability according to claim 1, characterized in that: The method further comprises: In response to the completion of the sequential alignment, resulting sequences that meet the difference within two points are obtained, and the mtDNA library samples belonging to the same sample individual are summarized; Sorting the mtDNA library samples according to their closeness to the mtDNA to be compared, wherein the closeness is ranked from high to low in terms of complete match and heterogeneity; If there is one base difference between the mtDNA library sample and the mtDNA to be compared, the samples are sorted in order according to the differences caused by insertions and deletions, the differences caused by transitions, and the differences caused by transversions.
7. A human mtDNA rapid comparison device with adaptive learning control capability, comprising: A database creation module is used to divide mtDNA into multiple primary sub-libraries, and further divide the primary sub-libraries into secondary sub-libraries. Each time a sample is stored in the library, it is adaptively sorted to form a database; A primary sub-library selection module is used to calculate the total length of the mtDNA to be compared after the database is established, and select the primary sub-library corresponding to the mtDNA to be compared according to the total length; A secondary sub-library selection module is used to calculate the difference sequence between the mtDNA to be compared and the library sample, obtain the locus value and the sum of the difference numbers of the mtDNA to be compared, and confirm the secondary sub-library according to the sum of the difference numbers; The sequential base matching module is used to compare the locus value of the mtDNA to be compared with the locus value of the mtDNA in the database according to the priority determined by the adaptive sorting; the database creation module is specifically used to: Dividing the mtDNA library samples into multiple primary sub-libraries according to their lengths, so as to exclude mtDNA library samples that necessarily have two or more base differences with the mtDNA to be compared; In each primary sub-library, the entire mtDNA library sample is divided into 33 regions according to a specified allocation rule, forming 33 position intervals, wherein the allocation rule is formed by allocating all base positions according to gene functional regions; The difference sequences between the mtDNA library sample and the standard sequence are placed on the loci according to the position interval, the locus values are recorded, and the number of differences is counted to form a total number of differences value, and multiple secondary sub-libraries are formed according to the total number of differences value, wherein the standard sequence is the Anderson reference sequence; For each sample stored, the distribution frequencies of the thirty-three loci values were counted, and the comparisons of the loci were sorted in descending order according to the difference frequencies.
8. The human mtDNA rapid comparison device with adaptive learning control capability according to claim 7, characterized in that: The secondary sub-library selection module is specifically used for: Calculating the difference sequence between the mtDNA to be compared and the standard sequence, and simultaneously confirming the thirty-three loci values and the total number of differences of the mtDNA to be compared; Based on the sum of the number of differences, a secondary sub-library that does not have a two-base difference standard with the mtDNA to be compared is selected.
Citation Information
Patent Citations
Reverse index structure based STR data storage and paternity test sorting comparison method
CN105260395A
STR rapid comparison method and system based on maximum frequency virtual individual
CN110060737A