A whole-genome-based high-throughput gene sequencing method

Through the method of cleavage of genes with the whole genome recognition library and fluorescent markers, the problem of slice contamination and merger time in high-throughput gene sequencing is solved, and efficient and accurate gene sequencing results are achieved.

CN118581202BActive Publication Date: 2025-05-27SUZHOU GEEK GENE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410815424.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2025-05-27
Estimated Expiration
2044-06-24

AI Technical Summary

Technical Problem

Existing high-throughput gene sequencing methods are susceptible to contamination during the slice process, resulting in deviations in the detection results, and the reorganization of large numbers of slices requires a lot of time.

Method used

The whole genome recognition library is used to replicate and cleave the gene sequence to be measured through fluorescently labeled cleaved genes to form an ordered collection of cleaved genes and merge them using a fluorescently labeled mechanism to ensure the accuracy and efficiency of the slice sequencing results.

Benefits of technology

It effectively avoids sequencing errors caused by contamination of single slices, shortens detection time, and improves the accuracy and merging efficiency of sequencing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118581202B_ABST
    Figure CN118581202B_ABST
Patent Text Reader

Abstract

The present invention discloses a high-throughput gene sequencing method based on the whole genome, which relates to the technical field of gene detection and includes: establishing a whole genome recognition library; forming a set of test gene replication sequences from at least one gene replication sequence to be tested; forming at least one set of test cleavage genes; obtaining one of the test gene replication sequences in the set of test gene replication sequences; forming a set of test gene replication fragments from at least one test gene replication fragment; obtaining at least one set of test gene replication fragments; obtaining the categories of test gene replication fragments; obtaining the sequencing results of the test gene replication fragments; and obtaining the sequencing results of the test gene replication sequences. By using fluorescence-labeled cleavage genes to perform split detection on the test gene replication sequences, it avoids sequencing errors caused by contamination of a single slice, can greatly reduce the complexity of merging, and thus effectively shortens the detection time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gene detection, and specifically relates to a high-throughput gene sequencing method based on the whole genome. Background Art

[0002] DNA sequencing is the basis for measuring the main characteristics of various life forms. Since the discovery of the DNA double helix structure in the 1950s, scientists around the world have been working on determining the original sequences of genomes of different species. This task is called genome sequencing, aiming to reveal the genomic composition and gene arrangement order of different organisms. High-throughput sequencing technology is an important analysis tool in the field of biology, which can quickly and accurately determine DNA sequences or RNA sequences. The emergence of high-throughput sequencing technology has greatly promoted the development of genomics, transcriptomics and bioinformatics.

[0003] However, when the existing gene sequencing methods are used for sequencing, in the high-throughput sequencing mode, genes need to be sliced. If the slices are contaminated, the detection results are likely to deviate. In addition, a large number of slices are in a disordered state, and it takes a lot of time to re-combine the detection results of a large number of slices. Summary of the Invention

[0004] To solve the above technical problems, a high-throughput gene sequencing method based on the whole genome is provided. This technical solution solves the problems in the above background art that when the existing gene sequencing methods are used for sequencing, in the high-throughput sequencing mode, genes need to be sliced. If the slices are contaminated, the detection results are likely to deviate. In addition, a large number of slices are in a disordered state, and it takes a lot of time to re-combine the detection results of a large number of slices.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A high-throughput gene sequencing method based on the whole genome, comprising:

[0007] Establish a whole genome recognition library, which is composed of at least one existing gene fragment;

[0008] Obtain at least one gene sequence to be tested, replicate the gene sequence to be tested to obtain at least one replicated gene sequence to be tested, and at least one replicated gene sequence to be tested forms a set of replicated gene sequences to be tested, and pair the set of replicated gene sequences to be tested with the gene sequence to be tested;

[0009] Form at least one set of test cleavage genes, and the cleavage genes in at least one set of test cleavage genes are different from each other. The cleavage genes are fluorescently labeled. Among them, the set of test cleavage genes is an ordered set, and the cleavage genes are genes that are symmetric about the left and right;

[0010] Pair the gene sequence to be tested with the test cleavage gene set, and pair the test cleavage gene set paired with the gene sequence to be tested with the set of replicated gene sequences of the gene sequence to be tested formed by replicating the gene sequence to be tested;

[0011] Obtain one of the replicated gene sequences of the gene sequence to be tested in the set of replicated gene sequences of the gene sequence to be tested as the characteristic replicated gene sequence to be tested;

[0012] Use the cleavage gene in the test cleavage gene set to cleave the characteristic replicated gene sequence to be tested to obtain at least one replicated gene fragment to be tested. The at least one replicated gene fragment to be tested forms a set of replicated gene fragments to be tested. When the cleavage gene cleaves, it is divided into two parts with the symmetry center as the break point, namely the first break point gene and the second break point gene;

[0013] The characteristic replicated gene sequence to be tested synchronously traverses the replicated gene sequences to be tested in the set of replicated gene sequences to be tested to obtain at least one set of replicated gene fragments to be tested;

[0014] Classify the replicated gene fragments to be tested in the at least one set of replicated gene fragments to be tested to obtain the category of replicated gene fragments to be tested;

[0015] Perform gene sequencing on the replicated gene fragments to be tested in the category of replicated gene fragments to be tested to obtain the sequencing results of the replicated gene fragments to be tested;

[0016] Based on the cleavage gene, merge the sequencing results of the replicated gene fragments to be tested to obtain the sequencing results of the replicated gene sequences to be tested, and assign the sequencing results of the replicated gene sequences to be tested to the gene sequence to be tested corresponding to the set of replicated gene sequences to be tested where it is located;

[0017] Among them, at least one gene sequence to be tested synchronously obtains the sequencing results.

[0018] Preferably, the establishment of the whole genome recognition library includes the following steps:

[0019] Based on big data, obtain the existing gene fragments and the corresponding sequencing results of the gene fragments to form a preliminary whole genome recognition library;

[0020] Verify the preliminary whole genome recognition library, and based on big data, obtain at least one sample gene verification set;

[0021] Compare the sample gene sequences in the sample gene verification set with the gene fragments in the preliminary whole genome recognition library. When the sample gene sequence appears in the preliminary whole genome recognition library, no treatment is performed. Otherwise, the sample gene sequence is supplemented into the preliminary whole genome recognition library;

[0022] After verification, use the preliminary whole genome recognition library as the whole genome recognition library.

[0023] Preferably, the step of replicating the gene sequence to be measured to obtain at least one replicated gene sequence to be measured includes the following steps:

[0024] Using polymerase chain reaction, amplify the DNA fragment of the gene sequence to be measured to a sufficient quantity;

[0025] Select a vector to insert the gene sequence to be measured into a recipient cell, and the vector includes plasmids and viruses;

[0026] Use restriction enzymes and ligases to ligate the DNA fragment of the gene sequence to be measured with the vector;

[0027] Introduce the recombinant DNA fragment into the recipient cell;

[0028] Use antibiotics to screen out the recipient cells that have successfully received the gene sequence to be measured;

[0029] Cultivate and amplify the screened recipient cells to obtain at least one replicated gene sequence to be measured.

[0030] Preferably, the step of forming at least one set of test cleavage genes includes the following steps:

[0031] Obtain at least one palindromic cleavage gene in the whole genome recognition library;

[0032] Divide at least one cleavage gene equally to form at least one set of test cleavage genes;

[0033] Form breakpoints at the symmetry centers of the cleavage genes, and the cleavage genes are fluorescently labeled;

[0034] Number the cleavage genes in each set of test cleavage genes.

[0035] Preferably, the step of using the cleavage genes in the set of test cleavage genes to cleave the characteristic replicated gene sequence to be measured to obtain at least one replicated gene fragment to be measured includes the following steps:

[0036] According to the numbers of the cleavage genes in the set of test cleavage genes, insert the cleavage genes into the characteristic replicated gene sequence to be measured;

[0037] The cleavage genes break at the breakpoints to form a first breakpoint gene and a second breakpoint gene, the characteristic replicated gene sequence to be measured is divided into at least one replicated gene fragment to be measured, and the first breakpoint gene and the second breakpoint gene are respectively connected to the ends of the replicated gene fragment to be measured that breaks at the breakpoint.

[0038] Preferably, the step of synchronously traversing the replicated gene sequences in the set of replicated gene sequences to be measured by the characteristic replicated gene sequence to be measured to obtain at least one set of replicated gene fragments to be measured includes the following steps:

[0039] When the characteristic gene replication sequence to be measured traverses the set of gene replication sequences to be measured synchronously, the characteristic gene replication sequence to be measured is cleaved by the cleavage genes in the corresponding test cleavage gene set, and the positions where the characteristic gene replication sequence to be measured is cleaved are all kept consistent, forming at least one set of gene replication fragments to be measured.

[0040] Preferably, the steps of classifying the gene replication fragments to be measured in at least one set of gene replication fragments to be measured to obtain the categories of gene replication fragments to be measured include the following:

[0041] Identify the fluorescently labeled parts at the ends of the gene replication fragments to be measured, and summarize the gene replication fragments to be measured with the same fluorescently labeled parts at the ends into the same category to obtain the categories of gene replication fragments to be measured.

[0042] Preferably, the steps of performing gene sequencing on the gene replication fragments to be measured in the categories of gene replication fragments to be measured to obtain the sequencing results of the gene replication fragments to be measured include the following:

[0043] Identify the fluorescently labeled parts at the ends of the gene replication fragments to be measured, and sequence the fragments of the gene replication fragments to be measured other than the fluorescently labeled parts to obtain preliminary sequencing results;

[0044] The gene replication fragments to be measured traverse the categories of gene replication fragments to be measured to obtain at least one preliminary sequencing result;

[0045] Obtain the gene fragment with the highest degree of coincidence with the preliminary sequencing result in the whole-genome recognition library as the preliminary gene fragment;

[0046] Obtain the preliminary gene fragment with the highest occurrence frequency as the target gene fragment;

[0047] Take the sequencing result of the target gene fragment as the sequencing result of the gene replication fragment to be measured.

[0048] Preferably, the steps of merging the sequencing results of the gene replication fragments to be measured to obtain the sequencing result of the gene replication sequence to be measured include the following:

[0049] Use the fluorescence merging mechanism to merge the gene replication fragments to be measured. When using the fluorescence merging mechanism, merge two of the gene replication fragments to be measured, and satisfy that the fluorescently labeled parts at the ends of the two gene replication fragments to be measured at the merging position are in a mirror-image consistent relationship;

[0050] According to the merging order of the gene replication fragments to be measured, merge the sequencing results of the gene replication fragments to be measured to obtain the sequencing result of the gene replication sequence to be measured.

[0051] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0052] By replicating at least one gene sequence to be measured to obtain at least one replicated gene sequence to be measured, and using a fluorescence-labeled cleavage gene to perform cleavage detection on the replicated gene sequence to be measured, a large number of slices can be measured simultaneously. By using multiple identical replicated gene sequences to be measured for slice sequencing, the sequencing results of the slices can be determined based on the slice results at the same position, thereby avoiding sequencing errors caused by contamination of a single slice. At the same time, since the gene slices are labeled with different fluorescent markers, the sequencing results of the gene slices can be merged according to the fluorescent markers. Due to the setting mechanism of the fluorescent markers, the complexity of the merging can be greatly reduced, and the detection time can be effectively shortened. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic flow diagram of a high-throughput gene sequencing method based on the whole genome of the present invention;

[0054] Figure 2 It is a schematic flow diagram of establishing a whole genome recognition library of the present invention;

[0055] Figure 3 It is a schematic flow diagram of replicating a gene sequence to be measured to obtain at least one replicated gene sequence to be measured according to the present invention;

[0056] Figure 4 It is a schematic flow diagram of forming at least one set of test cleavage genes according to the present invention;

[0057] Figure 5 It is a schematic flow diagram of using the cleavage genes in the set of test cleavage genes to cleave a characteristic replicated gene sequence to be measured to obtain at least one replicated gene fragment to be measured according to the present invention;

[0058] Figure 6 It is a schematic flow diagram of performing gene sequencing on the replicated gene fragments to be measured in the category of replicated gene fragments to be measured to obtain the sequencing results of the replicated gene fragments to be measured according to the present invention;

[0059] Figure 7 It is a schematic flow diagram of merging the sequencing results of the replicated gene fragments to be measured to obtain the sequencing results of the replicated gene sequence to be measured according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations.

[0061] Referring to Figure 1 as shown, a high-throughput gene sequencing method based on the whole genome includes:

[0062] Construct a whole-genome recognition library, which consists of at least one existing gene fragment;

[0063] Obtain at least one gene sequence to be tested, replicate the gene sequence to be tested to obtain at least one replicated gene sequence to be tested, and the at least one replicated gene sequence to be tested forms a set of replicated gene sequences to be tested. Pair the set of replicated gene sequences to be tested with the gene sequence to be tested;

[0064] Form at least one set of test cleavage genes. The cleavage genes in the at least one set of test cleavage genes are different from each other, and the cleavage genes are fluorescently labeled. Among them, the set of test cleavage genes is an ordered set, and the cleavage genes are genes that are symmetric about the center;

[0065] Pair the gene sequence to be tested with the set of test cleavage genes, and pair the set of test cleavage genes paired with the gene sequence to be tested with the set of replicated gene sequences formed by replicating the gene sequence to be tested;

[0066] Obtain one of the replicated gene sequences to be tested in the set of replicated gene sequences to be tested as the characteristic replicated gene sequence to be tested;

[0067] Use the cleavage genes in the set of test cleavage genes to cleave the characteristic replicated gene sequence to be tested to obtain at least one replicated gene fragment to be tested. The at least one replicated gene fragment to be tested forms a set of replicated gene fragments to be tested. When the cleavage gene cleaves, it is divided into two parts with the center of symmetry as the break point, namely the first break-point gene and the second break-point gene;

[0068] The characteristic replicated gene sequence to be tested synchronously traverses the replicated gene sequences to be tested in the set of replicated gene sequences to be tested to obtain at least one set of replicated gene fragments to be tested;

[0069] Classify the replicated gene fragments to be tested in the at least one set of replicated gene fragments to be tested to obtain the categories of replicated gene fragments to be tested;

[0070] Perform gene sequencing on the replicated gene fragments to be tested in the categories of replicated gene fragments to be tested to obtain the sequencing results of the replicated gene fragments to be tested;

[0071] Based on the cleavage genes, merge the sequencing results of the replicated gene fragments to be tested to obtain the sequencing results of the replicated gene sequences to be tested, and assign the sequencing results of the replicated gene sequences to be tested to the gene sequence to be tested corresponding to the set of replicated gene sequences to be tested where it is located;

[0072] Among them, at least one gene sequence to be tested synchronously obtains the sequencing results.

[0073] In this solution, high-throughput gene sequencing is used to perform multiple gene sequencing simultaneously at one time, and at least one gene sequence to be tested is sequenced. During sequencing, replication and cutting are both carried out, and then, based on the comprehensive results of replication, the sequencing result of the gene sequence to be tested is obtained;

[0074] Among them, when the gene sequence to be tested is sequenced, a set of replicated gene sequences to be tested is generated, and the replicated gene sequences in the set of replicated gene sequences to be tested are all synchronously sequenced according to the sequencing method of the characteristic replicated gene sequence.

[0075] Refer to Figure 2 As shown, establishing a whole-genome recognition library includes the following steps:

[0076] Based on big data, existing gene fragments and the corresponding sequencing results of the gene fragments are obtained to form a preliminary whole-genome recognition library;

[0077] The preliminary whole-genome recognition library is verified, and based on big data, at least one sample gene verification set is obtained;

[0078] The sample gene sequences in the sample gene verification set are compared with the gene fragments in the preliminary whole-genome recognition library. When the sample gene sequence appears in the preliminary whole-genome recognition library, no processing is done. Otherwise, the sample gene sequence is supplemented into the preliminary whole-genome recognition library;

[0079] After verification, the preliminary whole-genome recognition library is used as the whole-genome recognition library.

[0080] The purpose of establishing the whole-genome recognition library is to summarize all the genes obtained in existing research, so that during sequencing, the sequencing result of the gene can be judged by comparison. Because the gene may be contaminated, therefore, it is impossible to directly judge based on its sequencing result, which may lead to errors.

[0081] Refer to Figure 3 As shown, replicating the gene sequence to be tested to obtain at least one replicated gene sequence to be tested includes the following steps:

[0082] Using polymerase chain reaction, the DNA fragment of the gene sequence to be tested is amplified to a sufficient quantity;

[0083] A vector is selected to insert the gene sequence to be tested into the recipient cell, and the vector includes plasmids and viruses;

[0084] Using restriction enzymes and ligases to ligate the DNA fragment of the gene sequence to be tested with the vector;

[0085] The recombinant DNA fragment is introduced into the recipient cell;

[0086] Use antibiotics to screen out recipient cells that have successfully received the gene sequence to be tested;

[0087] Culture and amplify the screened recipient cells to obtain at least one replicated sequence of the gene to be tested.

[0088] The reason for replicating the gene sequence to be tested is that there may be contamination in the environment when the gene sequence to be tested is segmented. Therefore, direct slicing and sequencing will result in errors. Thus, the gene sequence to be tested is replicated to obtain at least one replicated sequence of the gene to be tested. The segmentation methods for at least one replicated sequence of the gene to be tested are the same. Therefore, the replicated gene fragments at the same positions in at least one replicated sequence of the gene to be tested are the same. By comparing the sequencing results of at least one identical replicated gene fragment with the whole-genome recognition library, the comprehensive comparison results can eliminate the influence of minor contamination on the sequencing results, thereby improving the sequencing accuracy.

[0089] Refer to Figure 4 As shown, forming at least one set of test cleavage genes includes the following steps:

[0090] Obtain at least one pair of symmetric cleavage genes from the whole-genome recognition library;

[0091] Divide at least one cleavage gene equally to form at least one set of test cleavage genes;

[0092] Form breakpoints at the symmetric centers of the cleavage genes, and the cleavage genes are fluorescently labeled;

[0093] Number the cleavage genes in each set of test cleavage genes.

[0094] All the cleavage genes used in the sets of test cleavage genes are different from each other, and the cleavage genes are symmetric. Therefore, when they are broken at the symmetric points as breakpoints, a first breakpoint gene and a second breakpoint gene will be formed. The first breakpoint gene and the second breakpoint gene are mirror-image consistent, and the first breakpoint genes and second breakpoint genes generated by different cleavage genes are all different. Therefore, they can be used as the basis for merging after cleavage sequencing. Moreover, since both the first breakpoint gene and the second breakpoint gene are fluorescently labeled, it is only necessary to detect whether the fluorescent parts of the replicated gene fragments to be tested are mirror-image consistent. If they are consistent, it indicates that they can be merged. Thus, there will be no confusion during merging, and it is easy to identify the merging.

[0095] Refer to Figure 5 As shown, using the cleavage genes in the set of test cleavage genes to cleave the characteristic replicated sequence of the gene to be tested to obtain at least one replicated gene fragment to be tested includes the following steps:

[0096] Insert the cleavage gene into the feature gene to be tested replication sequence according to the number of the cleavage gene in the test cleavage gene set;

[0097] The cleavage gene breaks at the breakpoint to form a first breakpoint gene and a second breakpoint gene. The feature gene to be tested replication sequence is divided into at least one gene to be tested replication fragment. The first breakpoint gene and the second breakpoint gene are respectively connected to the ends of the gene to be tested replication fragment broken at the breakpoint.

[0098] Here, the test cleavage gene set used is in a paired relationship with the gene to be tested replication sequence set where the feature gene to be tested replication sequence is located, that is, the division of the gene to be tested replication sequences in the gene to be tested replication sequence set is all carried out in the same way using the same test cleavage gene set.

[0099] The feature gene to be tested replication sequence synchronously traverses the gene to be tested replication sequences in the gene to be tested replication sequence set. Obtaining at least one gene to be tested replication fragment set includes the following steps:

[0100] When the feature gene to be tested replication sequence synchronously traverses the gene to be tested replication sequence set, the feature gene to be tested replication sequence is cut by the cleavage genes in the corresponding test cleavage gene set, and the positions where the feature gene to be tested replication sequence is cut are all kept consistent, forming at least one gene to be tested replication fragment set.

[0101] Classify the gene to be tested replication fragments in at least one gene to be tested replication fragment set. Obtaining the gene to be tested replication fragment categories includes the following steps:

[0102] Identify the fluorescently labeled parts at the ends of the gene to be tested replication fragments, and summarize the gene to be tested replication fragments with the same fluorescently labeled parts at the ends into the same category to obtain the gene to be tested replication fragment categories.

[0103] Taking two gene to be tested replication fragments as an example, the gene to be tested replication fragments with the same fluorescently labeled parts at the ends refer to: the fluorescently labeled parts at the ends of the two gene to be tested replication fragments correspond to each other identically;

[0104] When the fluorescently labeled parts at the ends of the gene to be tested replication fragments are the same, it means that they were in the same position in at least one gene to be tested replication sequence before division, and the gene to be tested replication fragments are generated from the same gene to be tested sequence. Therefore, they actually correspond to the same fragment.

[0105] Refer to Figure 6 As shown, perform gene sequencing on the gene to be tested replication fragments in the gene to be tested replication fragment categories. Obtaining the sequencing results of the gene to be tested replication fragments includes the following steps:

[0106] Identify the fluorescently labeled part at the end of the replicated fragment of the gene to be tested, sequence the fragment of the replicated fragment of the gene to be tested except for the fluorescently labeled part, and obtain a preliminary sequencing result;

[0107] Traverse the categories of replicated fragments of the gene to be tested to obtain at least one preliminary sequencing result;

[0108] Retrieve the gene fragment with the highest degree of match with the preliminary sequencing result from the whole-genome identification library as the preliminary gene fragment;

[0109] Obtain the preliminary gene fragment with the highest occurrence frequency as the target gene fragment;

[0110] Use the sequencing result of the target gene fragment as the sequencing result of the replicated fragment of the gene to be tested.

[0111] Due to the influence of contamination, in order to avoid contamination, it is necessary to use the result with the highest alignment match as the sequencing result. Thus, the interference of contamination can be excluded.

[0112] Refer to Figure 7 As shown, the steps for merging the sequencing results of the replicated fragments of the gene to be tested to obtain the sequencing result of the replicated sequence of the gene to be tested are as follows:

[0113] Use a fluorescence merging mechanism to merge the replicated fragments of the gene to be tested. When using the fluorescence merging mechanism, merge two of the replicated fragments of the gene to be tested, and ensure that the fluorescently labeled parts at the ends of the merged parts of the two replicated fragments of the gene to be tested are in a mirror-image consistent relationship;

[0114] Merge the sequencing results of the replicated fragments of the gene to be tested in the order of merging of the replicated fragments of the gene to be tested to obtain the sequencing result of the replicated sequence of the gene to be tested.

[0115] The basis for the merging is that the fluorescently labeled parts at the ends are in a mirror-image consistent relationship, which indicates that they were merged into one before being split at this point. Therefore, merging can be carried out based on this, and further, the sequencing result of the replicated sequence of the gene to be tested can be obtained.

[0116] Furthermore, this solution also proposes a storage medium on which a computer-readable program is stored. When the computer-readable program is called, it executes the above-mentioned high-throughput gene sequencing method based on the whole genome.

[0117] It can be understood that the storage medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; an optical medium, such as a DVD; or a semiconductor medium, such as a solid-state drive (SSD).

[0118] In summary, the advantages of the present invention are as follows: by replicating at least one test gene replication sequence for the gene sequence to be tested and using a cleavage gene labeled with fluorescence to perform split detection on the test gene replication sequence, a large number of slices can be measured simultaneously. By using multiple identical test gene replication sequences for slice sequencing, the sequencing results of the slices can be determined based on the slice results at the same position, thereby avoiding sequencing errors caused by contamination of a single slice. At the same time, since the gene slices are labeled with different fluorescence, the sequencing results of the gene slices can be combined according to the fluorescence labeling. Due to the setting mechanism of the fluorescence labeling, the complexity of the combination can be greatly reduced, and the detection time can be effectively shortened.

[0119] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the principles described in the specification are only the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A high-throughput gene sequencing method based on the whole genome, characterized in that: include: Establishing a whole genome identification library, wherein the whole genome identification library is composed of at least one existing gene fragment; Acquire at least one gene sequence to be tested, replicate the gene sequence to be tested to obtain at least one gene replication sequence to be tested, the at least one gene replication sequence to be tested forms a gene replication sequence set to be tested, and pair the gene replication sequence set to be tested with the gene sequence to be tested; Forming at least one test cleavage gene set, wherein the cleavage genes in the at least one test cleavage gene set are different from each other, and the cleavage genes are fluorescently labeled, wherein the test cleavage gene set is an ordered set, and the cleavage genes are bilaterally symmetric genes; Pairing the gene sequence to be tested with the test cutting gene set, and pairing the test cutting gene set paired with the gene sequence to be tested with the gene copy sequence set to be tested formed by copying the gene sequence to be tested; Obtaining one of the gene replication sequences to be tested in the set of gene replication sequences to be tested as a characteristic gene replication sequence to be tested; Using the cutting gene in the test cutting gene set to cut the characteristic test gene replication sequence to obtain at least one test gene replication fragment, the at least one test gene replication fragment forms a test gene replication fragment set, and when cutting, the cutting gene is divided into two parts with the symmetry center as the breakpoint, namely the first breakpoint gene and the second breakpoint gene; The characteristic gene replication sequence to be tested synchronously traverses the gene replication sequence to be tested in the gene replication sequence set to be tested to obtain at least one gene replication fragment set to be tested; Classifying the gene replication fragments to be tested in at least one set of gene replication fragments to be tested to obtain categories of the gene replication fragments to be tested; Performing gene sequencing on the gene replication fragments to be tested in the gene replication fragment category to obtain sequencing results of the gene replication fragments to be tested; Based on the cut gene, the sequencing results of the gene replication fragments to be tested are merged to obtain the sequencing results of the gene replication sequence to be tested, and the sequencing results of the gene replication sequence to be tested are assigned to the gene sequence to be tested corresponding to the gene replication sequence set to be tested; Wherein, at least one gene sequence to be tested obtains sequencing results simultaneously; The forming of at least one test cutting gene set comprises the following steps: Obtain at least one bilaterally symmetrical cutting gene in the whole genome identification library; equally dividing the at least one cleavage gene to form at least one test cleavage gene set; A breakpoint is formed at the symmetric center of the cut gene, which is fluorescently labeled; Numbering each cleavage gene in the test cleavage gene set; The method of using the cutting gene in the test cutting gene set to cut the characteristic test gene replication sequence to obtain at least one test gene replication fragment comprises the following steps: According to the number of the cutting gene in the test cutting gene set, the cutting gene is inserted into the characteristic gene copy sequence to be tested; The cutting gene is broken at the breakpoint to form a first breakpoint gene and a second breakpoint gene, the characteristic test gene replication sequence is divided into at least one test gene replication fragment, and the first breakpoint gene and the second breakpoint gene are respectively connected to the ends of the test gene replication fragment broken at the breakpoint; The characteristic gene replication sequence to be tested synchronously traverses the gene replication sequence to be tested in the gene replication sequence set to be tested, and obtains at least one gene replication fragment set to be tested, which includes the following steps: When the characteristic gene copy sequence to be tested synchronously traverses the set of gene copy sequences to be tested, the characteristic gene copy sequence to be tested is cut by the cutting gene in the corresponding test cutting gene set, and the positions where the characteristic gene copy sequence to be tested is cut are kept consistent, forming at least one set of gene copy fragments to be tested; The step of classifying the gene replication fragments to be tested in at least one set of gene replication fragments to be tested to obtain the categories of the gene replication fragments to be tested comprises the following steps: Identify the fluorescently labeled portion at the end of the gene replication fragment to be tested, and group the gene replication fragments to be tested with the same fluorescently labeled portion at the end into the same category to obtain the category of the gene replication fragment to be tested; The method of performing gene sequencing on the gene replication fragments to be tested in the gene replication fragment category to obtain the sequencing results of the gene replication fragments to be tested comprises the following steps: Identify the fluorescently labeled portion at the end of the gene replication fragment to be tested, and sequence the fragment of the gene replication fragment to be tested except the fluorescently labeled portion to obtain a preliminary sequencing result; The gene duplication fragment to be tested traverses the categories of the gene duplication fragment to be tested, and obtains at least one preliminary sequencing result; Obtain the gene fragment with the highest degree of consistency with the preliminary sequencing result in the whole genome identification library as the preliminary gene fragment; Obtaining the most frequently occurring prepared gene fragment as the target gene fragment; The sequencing result of the target gene fragment is used as the sequencing result of the gene replication fragment to be tested; The step of combining the sequencing results of the gene replication fragments to obtain the sequencing results of the gene replication sequence to be tested comprises the following steps: Using a fluorescence merging mechanism to merge the gene replication fragments to be tested, when using the fluorescence merging mechanism, the two gene replication fragments to be tested are merged, and the fluorescent labeled parts at the ends of the two gene replication fragments to be tested where they are merged are in a mirror-image consistent relationship; The sequencing results of the gene replication fragments to be tested are merged according to the merging order of the gene replication fragments to be tested to obtain the sequencing results of the gene replication sequence to be tested.

2. A whole genome based high-throughput gene sequencing method according to claim 1, characterized in that: The establishment of the whole genome identification library comprises the following steps: Based on big data, existing gene fragments and the sequencing results corresponding to the gene fragments are obtained to form a preliminary whole genome identification library; Verify the prepared whole genome identification library, and obtain at least one sample gene verification set based on big data; Compare the sample gene sequence in the sample gene verification set with the gene fragments in the prepared whole genome identification library. If the sample gene sequence appears in the prepared whole genome identification library, no processing is performed. Otherwise, the sample gene sequence is added to the prepared whole genome identification library. After verification, the prepared whole genome identification library will be used as the whole genome identification library.

3. A whole genome based high-throughput gene sequencing method according to claim 2, characterized in that: The method of duplicating the gene sequence to be tested to obtain at least one duplicated gene sequence to be tested comprises the following steps: Using polymerase chain reaction, amplify the DNA fragment of the gene sequence to be tested to a sufficient amount; Select a vector to insert the gene sequence to be tested into the recipient cell, and the vector includes plasmids and viruses; Use enzyme digestion and ligase to connect the DNA fragment of the gene sequence to be tested to the vector; introducing the recombinant DNA fragment into recipient cells; Use antibiotics to select recipient cells that have successfully accepted the gene sequence to be tested; The screened recipient cells are cultured and amplified to obtain at least one gene replication sequence to be tested.

Citation Information

Patent Citations

  • Method for testing gene sequences

    CN102766688A

  • High-throughput gene sequencing equipment for cell gene mutation research

    CN117070331A