Multi-gene discovery network construction method and device, equipment and storage medium
By constructing a multi-gene discovery network, generating DNA sequences, and performing sequencing and decoding, the error problems in gene encoding, decoding, sequencing, and biochemical evolution processes have been solved, enabling rapid and efficient gene discovery and improving the construction speed and efficiency of multi-gene discovery networks.
Patent Information
- Application Number
- CN202210938867.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-05
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2042-08-05
AI Technical Summary
In the existing gene encoding, decoding, sequencing, and biochemical evolution processes, errors such as substitution errors, insertion errors, and deletion errors exist, which lead to changes in the positional distance between genes and low efficiency in discovering other genes.
By constructing a multi-gene discovery network, genomic information is obtained, DNA sequences are generated and sequenced, gene location information is decoded, the existence of encoded genes is determined, and DNA sequences are generated using the discovery network and stored information, followed by sequencing and decoding, to quickly and efficiently find the corresponding genes.
It can quickly and efficiently find the corresponding gene during decoding, saving gene search time, and can quickly detect gene reporting errors, avoiding various variants and sequencing result errors, thus improving the speed and efficiency of constructing multi-gene discovery networks.
Smart Images

Figure CN115458050B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics technology, and in particular to a method, apparatus, device, and storage medium for constructing a multi-gene discovery network. Background Technology
[0002] Bioinformatics is an interdisciplinary field that uses methods from applied mathematics, informatics, statistics, and computer science to study biological problems. As early as the 1860s, the academic community proposed the concept of DNA-based data storage. After nearly sixty years of development, research related to DNA storage has gradually become an important branch of bioinformatics.
[0003] In the study of encoding and decoding a genome composed of multiple genes, during encoding, each gene is encoded to obtain a genome that does not change the biological activity of the genome. During decoding, the sequence obtained from sequencing is decoded. The first step is to find all the genes in the genome. During sequencing and biochemical evolution, there is a certain probability of errors occurring, such as incorrect substitution, incorrect insertion, and incorrect deletion, which may change the positional distance between genes. Summary of the Invention
[0004] The main objective of this invention is to provide a method, apparatus, device, and storage medium for constructing a multi-gene discovery network, aiming to solve the technical problem in the prior art where errors such as substitution errors, insertion errors, and deletion errors exist in gene encoding, decoding, sequencing, and biochemical evolution processes, which can change the positional distance between genes and lead to low efficiency in discovering other genes.
[0005] In a first aspect, the present invention provides a method for constructing a multi-gene discovery network, the method comprising the following steps:
[0006] Obtain the genomic information input by the user, construct a discovery network among all genes to be encoded based on the genomic information, and determine the storage information of each gene to be encoded;
[0007] A DNA sequence is generated based on the discovery network and the stored information, and the DNA sequence is sequenced to obtain sequencing results;
[0008] The sequencing results are decoded to obtain gene location information, and the presence of the encoded gene is determined based on the gene location information.
[0009] Optionally, the step of obtaining the genomic information input by the user, constructing a discovery network among all genes to be encoded based on the genomic information, and determining the storage information of each gene to be encoded includes:
[0010] Obtain the genomic information input by the user, and construct a discovery network based on the genomic information;
[0011] Obtain the spacer sequences between each gene to be encoded, and form multiple circular genomes based on the spacer sequences;
[0012] The number of genes and genome decoding information of each circular genome are obtained. Based on the number of genes and genome decoding information, the associated genes of each gene to be encoded are obtained. Based on the associated genes, the storage information to be embedded for each gene to be encoded is determined.
[0013] Optionally, obtaining the genomic information input by the user and constructing a discovery network based on the genomic information includes:
[0014] Obtain the genomic information input by the user and abstract each gene to be encoded into a node;
[0015] Construct a directed unweighted connected graph where each node has two in-degrees and two out-degrees, and generate a discovery network based on the directed unweighted connected graph.
[0016] Optionally, obtaining the spacer sequences between each gene to be encoded, and forming multiple circular genomes based on the spacer sequences, includes:
[0017] Based on the genomic information, all genes to be encoded are identified in a clockwise direction to obtain the spacer sequences between each gene to be encoded, and each gene to be encoded is divided into multiple circular genomes based on the spacer sequences.
[0018] Optionally, the step of obtaining the number of genes and genome decoding information for each circular genome, obtaining the associated genes for each gene to be encoded based on the number of genes and the genome decoding information, and determining the storage information to be embedded for each gene to be encoded based on the associated genes includes:
[0019] The number of genes contained in each circular genome is obtained, and the corresponding genome decoding information is obtained by solving each circular genome based on the number of genes.
[0020] Based on the number of genes and the gene decoding information, the associated genes of each gene to be encoded are determined;
[0021] Obtain the target natural number of the associated gene to be embedded, and use the target natural number as the storage information for each gene to be encoded.
[0022] Optionally, the step of generating a DNA sequence based on the discovery network and the stored information, and sequencing the DNA sequence to obtain sequencing results, includes:
[0023] The location information of each gene to be encoded is obtained from the discovery network, and each gene to be encoded is embedded according to the location information and the storage information to generate a DNA sequence.
[0024] The DNA sequence is subjected to a biochemical process, and the processed DNA sequence is sequenced to obtain the sequencing results.
[0025] Optionally, decoding the sequencing results to obtain gene location information, and determining whether a encoded gene exists based on the gene location information, includes:
[0026] The sequencing results are iterated through based on the start codons to generate a start table, and the sequencing results are iterated through based on the stop codons to generate a stop table. The start table is a list of the positions of all start codons in the genome, and the stop table is a list of the positions of all stop codons in the genome.
[0027] The sequencing results are decoded according to the start table and the end table to find the target gene located in the preset indirect sequence length range in the sequencing results;
[0028] Obtain the gene location information of the target gene, and determine whether the encoded gene exists in other genes besides the target gene based on the gene location information.
[0029] Secondly, to achieve the above objectives, the present invention also proposes a multi-gene discovery network construction device, the multi-gene discovery network construction device comprising:
[0030] The information acquisition module is used to acquire the genomic information input by the user, construct a discovery network among all genes to be encoded based on the genomic information, and determine the storage information of each gene to be encoded.
[0031] A sequencing module is used to generate a DNA sequence based on the discovery network and the stored information, and to sequence the DNA sequence to obtain sequencing results;
[0032] The decoding module is used to decode the sequencing results, obtain gene location information, and determine whether the encoded gene exists based on the gene location information.
[0033] Thirdly, to achieve the above objectives, the present invention also proposes a multi-gene discovery network construction device, the multi-gene discovery network construction device comprising: a memory, a processor, and a multi-gene discovery network construction program stored in the memory and executable on the processor, the multi-gene discovery network construction program being configured to implement the steps of the multi-gene discovery network construction method as described above.
[0034] Fourthly, to achieve the above objectives, the present invention also proposes a storage medium storing a multi-gene discovery network construction program, wherein when the multi-gene discovery network construction program is executed by a processor, it implements the steps of the multi-gene discovery network construction method described above.
[0035] The multi-gene discovery network construction method proposed in this invention acquires user-inputted genomic information, constructs a discovery network among all genes to be encoded based on the genomic information, and determines the storage information of each gene to be encoded; generates a DNA sequence based on the discovery network and the storage information, and sequences the DNA sequence to obtain sequencing results; decodes the sequencing results to obtain gene location information, and determines whether a encoded gene exists based on the gene location information. This method can quickly and efficiently find the corresponding gene during decoding, saving gene search time, and can quickly detect gene reporting errors, avoiding sequencing result errors caused by various variations and sequencing itself, thus improving the speed and efficiency of multi-gene discovery network construction. Attached Figure Description
[0036] Figure 1 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of the present invention;
[0037] Figure 2 This is a flowchart illustrating the first embodiment of the multi-gene discovery network construction method of the present invention;
[0038] Figure 3 This is a flowchart illustrating the second embodiment of the multi-gene discovery network construction method of the present invention;
[0039] Figure 4 This is a flowchart illustrating the third embodiment of the multi-gene discovery network construction method of the present invention;
[0040] Figure 5 This is a flowchart illustrating the fourth embodiment of the multi-gene discovery network construction method of the present invention;
[0041] Figure 6 This is a first schematic diagram of the circular genome in the multi-gene discovery network construction method of the present invention;
[0042] Figure 7 This is a flowchart illustrating the fifth embodiment of the multi-gene discovery network construction method of the present invention;
[0043] Figure 8 This is a first structural diagram of the discovery network process in the multi-gene discovery network construction method of the present invention;
[0044] Figure 9 This is a second schematic diagram of the network discovery process in the multi-gene discovery network construction method of the present invention;
[0045] Figure 10 This is a third schematic diagram of the network discovery process in the multi-gene discovery network construction method of the present invention;
[0046] Figure 11This is a schematic diagram of the fourth construction step in the multi-gene discovery network construction method of the present invention.
[0047] Figure 12 This is the fifth construction diagram of the discovery network process in the multi-gene discovery network construction method of the present invention;
[0048] Figure 13 This is a flowchart illustrating the sixth embodiment of the multi-gene discovery network construction method of the present invention;
[0049] Figure 14 This is a flowchart illustrating the seventh embodiment of the multi-gene discovery network construction method of the present invention;
[0050] Figure 15 This is a functional block diagram of the first embodiment of the multi-gene discovery network construction device of the present invention.
[0051] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0052] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0053] The solution of this invention mainly involves: acquiring user-inputted genomic information, constructing a discovery network among all genes to be encoded based on the genomic information, and determining the storage information of each gene to be encoded; generating a DNA sequence based on the discovery network and the storage information, and sequencing the DNA sequence to obtain sequencing results; decoding the sequencing results to obtain gene location information, and determining whether a encoded gene exists based on the gene location information. This method can quickly and efficiently find the corresponding gene during decoding, saving gene search time, and can quickly detect gene reporting errors. It avoids sequencing result errors caused by various variations and sequencing itself, and solves the technical problem in the prior art where errors such as substitution errors, insertion errors, and deletion errors exist in gene encoding, decoding, sequencing, and biochemical evolution, which can change the positional distance between genes and lead to low efficiency in discovering other genes.
[0054] Reference Figure 1 , Figure 1 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of the present invention.
[0055] like Figure 1As shown, the device may include: a processor 1001, such as a CPU; a communication bus 1002; a user interface 1003; a network interface 1004; and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0056] Those skilled in the art will understand that Figure 1 The device structure shown does not constitute a limitation on the device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0057] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating device, a network communication module, a user interface module, and a multi-gene discovery network construction program.
[0058] The device of the present invention calls the multi-gene discovery network construction program stored in the memory 1005 through the processor 1001 and performs the following operations:
[0059] Obtain the genomic information input by the user, construct a discovery network among all genes to be encoded based on the genomic information, and determine the storage information of each gene to be encoded;
[0060] A DNA sequence is generated based on the discovery network and the stored information, and the DNA sequence is sequenced to obtain sequencing results;
[0061] The sequencing results are decoded to obtain gene location information, and the presence of the encoded gene is determined based on the gene location information.
[0062] The device of the present invention, through processor 1001 calling the multi-gene discovery network construction program stored in memory 1005, also performs the following operations:
[0063] Obtain the genomic information input by the user, and construct a discovery network based on the genomic information;
[0064] Obtain the spacer sequences between each gene to be encoded, and form multiple circular genomes based on the spacer sequences;
[0065] The number of genes and genome decoding information of each circular genome are obtained. Based on the number of genes and genome decoding information, the associated genes of each gene to be encoded are obtained. Based on the associated genes, the storage information to be embedded for each gene to be encoded is determined.
[0066] The device of the present invention, through processor 1001 calling the multi-gene discovery network construction program stored in memory 1005, also performs the following operations:
[0067] Obtain the genomic information input by the user and abstract each gene to be encoded into a node;
[0068] Construct a directed unweighted connected graph where each node has two in-degrees and two out-degrees, and generate a discovery network based on the directed unweighted connected graph.
[0069] The device of the present invention, through processor 1001 calling the multi-gene discovery network construction program stored in memory 1005, also performs the following operations:
[0070] Based on the genomic information, all genes to be encoded are identified in a clockwise direction to obtain the spacer sequences between each gene to be encoded, and each gene to be encoded is divided into multiple circular genomes based on the spacer sequences.
[0071] The device of the present invention, through processor 1001 calling the multi-gene discovery network construction program stored in memory 1005, also performs the following operations:
[0072] The number of genes contained in each circular genome is obtained, and the corresponding genome decoding information is obtained by solving each circular genome based on the number of genes.
[0073] Based on the number of genes and the gene decoding information, the associated genes of each gene to be encoded are determined;
[0074] Obtain the target natural number of the associated gene to be embedded, and use the target natural number as the storage information for each gene to be encoded.
[0075] The device of the present invention, through processor 1001 calling the multi-gene discovery network construction program stored in memory 1005, also performs the following operations:
[0076] The location information of each gene to be encoded is obtained from the discovery network, and each gene to be encoded is embedded according to the location information and the storage information to generate a DNA sequence.
[0077] The DNA sequence is subjected to a biochemical process, and the processed DNA sequence is sequenced to obtain the sequencing results.
[0078] The device of the present invention, through processor 1001 calling the multi-gene discovery network construction program stored in memory 1005, also performs the following operations:
[0079] The sequencing results are iterated through based on the start codons to generate a start table, and the sequencing results are iterated through based on the stop codons to generate a stop table. The start table is a list of the positions of all start codons in the genome, and the stop table is a list of the positions of all stop codons in the genome.
[0080] The sequencing results are decoded according to the start table and the end table to find the target gene located in the preset indirect sequence length range in the sequencing results;
[0081] Obtain the gene location information of the target gene, and determine whether the encoded gene exists in other genes besides the target gene based on the gene location information.
[0082] This embodiment, through the above-described scheme, acquires the genomic information input by the user, constructs a discovery network among all genes to be encoded based on the genomic information, and determines the storage information of each gene to be encoded; generates a DNA sequence based on the discovery network and the storage information, and sequences the DNA sequence to obtain sequencing results; decodes the sequencing results to obtain gene location information, and determines whether the encoded gene exists based on the gene location information. This allows for rapid and efficient locating of the corresponding gene during decoding, saving gene search time, quickly detecting gene reporting errors, and avoiding sequencing result errors caused by various variations and the sequencing process itself.
[0083] Based on the above hardware structure, an embodiment of the multi-gene discovery network construction method of the present invention is proposed.
[0084] Reference Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the multi-gene discovery network construction method of the present invention.
[0085] In a first embodiment, the method for constructing a multi-gene discovery network includes the following steps:
[0086] Step S10: Obtain the genomic information input by the user, construct a discovery network among all genes to be encoded based on the genomic information, and determine the storage information of each gene to be encoded.
[0087] It should be noted that the genomic information refers to the gene information corresponding to the genome input by the user. Based on the genomic information, a discovery network can be constructed among all genes to be encoded, thereby determining the information that each gene should store.
[0088] Step S20: Generate a DNA sequence based on the discovery network and the stored information, and sequence the DNA sequence to obtain sequencing results.
[0089] It is understandable that at the encoding end, a DNA sequence can be generated by combining the stored information with the discovery network algorithm corresponding to the discovery network, and the DNA sequence can be sequenced to obtain the corresponding sequencing results.
[0090] Step S30: Decode the sequencing results to obtain gene location information, and determine whether the encoded gene exists based on the gene location information.
[0091] It should be understood that decoding the sequencing results can obtain the gene location information corresponding to the gene in the sequencing results, thereby determining whether the encoded gene exists based on the gene location information.
[0092] This embodiment, through the above-described scheme, acquires the genomic information input by the user, constructs a discovery network among all genes to be encoded based on the genomic information, and determines the storage information of each gene to be encoded; generates a DNA sequence based on the discovery network and the storage information, and sequences the DNA sequence to obtain sequencing results; decodes the sequencing results to obtain gene location information, and determines whether the encoded gene exists based on the gene location information. This allows for rapid and efficient locating of the corresponding gene during decoding, saving gene search time, quickly detecting gene reporting errors, and avoiding sequencing result errors caused by various variations and the sequencing process itself.
[0093] Furthermore, Figure 3 This is a flowchart illustrating the second embodiment of the multi-gene discovery network construction method of the present invention, as shown below. Figure 3 As shown, a second embodiment of the multi-gene discovery network construction method of the present invention is proposed based on the first embodiment. In this embodiment, step S10 specifically includes the following steps:
[0094] Step S11: Obtain the genomic information input by the user, and construct a discovery network based on the genomic information.
[0095] It should be noted that after obtaining the genomic information input by the user, a gene discovery network can be constructed based on the genomic information. This discovery network enables all genes to be quickly and efficiently identified during gene decoding of the sequencing sequence.
[0096] Step S12: Obtain the spacer sequences between each gene to be encoded, and form multiple circular genomes based on the spacer sequences.
[0097] It is understood that the spacer sequence is the base sequence between each gene that needs to be encoded, and multiple circular genomes are formed by the spacer sequence. After the circular genomes are formed, they can serve as the basis for the genome to be encoded.
[0098] Step S13: Obtain the number of genes and genome decoding information of each circular genome; obtain the associated genes of each gene to be encoded based on the number of genes and the genome decoding information; and determine the storage information to be embedded for each gene to be encoded based on the associated genes.
[0099] It should be understood that each circular genome corresponds to the number of genes and genome decoding information for each gene. The number of genes is the number of genes contained in each circular genome, and the genome decoding information is the gene information after decoding the circular genome. Through the number of genes and the genome decoding information, the associated genes of each gene to be encoded can be obtained, that is, the genes associated with each gene. Through the associated genes, the storage information to be embedded for each gene to be encoded can be determined by solving the calculation.
[0100] This embodiment, through the above scheme, obtains the genomic information input by the user, constructs a discovery network based on the genomic information; obtains the spacer sequences between each gene to be encoded, and forms multiple circular genomes based on the spacer sequences; obtains the number of genes and genome decoding information of each circular genome, obtains the associated genes of each gene to be encoded based on the number of genes and the genome decoding information, and determines the storage information to be embedded in each gene to be encoded based on the associated genes. This enables the determination of the storage information to be embedded in each gene to be encoded, and the corresponding gene to be found quickly and efficiently during decoding.
[0101] Furthermore, Figure 4 This is a flowchart illustrating the third embodiment of the multi-gene discovery network construction method of the present invention, as shown below. Figure 4 As shown, based on the second embodiment, a third embodiment of the multi-gene discovery network construction method of the present invention is proposed. In this embodiment, step S11 specifically includes the following steps:
[0102] Step S111: Obtain the genomic information input by the user and abstract each gene to be encoded into a node.
[0103] It should be noted that obtaining the genomic information input by the user allows each gene to be encoded to be abstracted into corresponding nodes.
[0104] Step S112: Construct a directed unweighted connected graph in which each node has two in-degrees and two out-degrees, and generate a discovery network based on the directed unweighted connected graph.
[0105] It is understandable that each gene is abstracted as a node in graph theory, and a directed unweighted connected graph with two in-degrees and two out-degrees is constructed for each node. The discovery network is then generated at the encoding end through the directed unweighted connected graph.
[0106] This embodiment uses the above-described scheme to obtain the genomic information input by the user, abstract each gene to be encoded into a node, construct a directed unweighted connected graph with two in-degrees and two out-degrees for each node, and generate a discovery network based on the directed unweighted connected graph. This enables the rapid generation of the discovery network and improves the speed and efficiency of gene discovery.
[0107] Furthermore, Figure 5 This is a flowchart illustrating the fourth embodiment of the multi-gene discovery network construction method of the present invention, as shown below. Figure 5 As shown, based on the second embodiment, a fourth embodiment of the multi-gene discovery network construction method of the present invention is proposed. In this embodiment, step S12 specifically includes the following steps:
[0108] Step S121: Identify all genes to be encoded in a clockwise direction according to the genomic information to obtain the spacer sequence between each gene to be encoded, and divide each gene to be encoded into multiple circular genomes according to the spacer sequence.
[0109] It should be noted that, based on the genome learning, all genes to be encoded are identified in a clockwise direction to obtain the soda ash sequence between genes, i.e., the spacer sequence, and then each gene to be encoded is divided into multiple circular genomes based on the spacer sequence.
[0110] In the specific implementation, see Figure 6 , Figure 6 This is a first schematic diagram of the circular genome in the multi-gene discovery network construction method of the present invention, as shown below. Figure 6 As shown, the circular genome Genome consists of 5 genes, each labeled in clockwise order as: Dr.gene0, Dr.gene1, Dr.gene2, Dr.gene3, and Dr.gene4. The spacer sequences between genes are pure base sequences.
[0111] This embodiment uses the above-described scheme to identify all genes to be encoded in a clockwise direction based on the genomic information, thereby obtaining the spacer sequences between each gene to be encoded. Based on the spacer sequences, each gene to be encoded is divided into multiple circular genomes, which can quickly form circular genomes and find the corresponding genes quickly and efficiently during subsequent decoding.
[0112] Furthermore, Figure 7 This is a flowchart illustrating the fifth embodiment of the multi-gene discovery network construction method of the present invention, as shown below. Figure 7As shown, based on the second embodiment, a fifth embodiment of the multi-gene discovery network construction method of the present invention is proposed. In this embodiment, step S13 specifically includes the following steps:
[0113] Step S131: Obtain the number of genes contained in each circular genome, and solve each circular genome according to the number of genes to obtain the corresponding genome decoding information.
[0114] It should be noted that by obtaining the number of genes contained in each circular genome, the circular genome can be solved based on the number of genes. The sequencing results of the DNA sequences in each circular genome can be decoded. The decoding process provides fault tolerance for base sequence errors and fragment loss between genes. The goal is to quickly find all genes that have embedded information during coding from the sequenced results.
[0115] In the specific implementation, refer to Figure 6 Let n be the number of genes contained in the circular genome, and solve for p based on n.
[0116] When n is odd:
[0117]
[0118] Solve for p using this system of equations, and find solutions to this equation that are less than or equal to n-1. If p has no solution, then...
[0119] When n is even:
[0120]
[0121] Solve for p using this system of equations, and find solutions to this equation that are less than or equal to n-1. If p has no solution, then...
[0122] In this example, n = 5, then solve the following equation
[0123]
[0124] Solving for p, we get p = 3.
[0125] Step S132: Determine the associated genes of each gene to be encoded based on the number of genes and the gene decoding information.
[0126] It is understood that the number of genes and the gene decoding information can be used to determine the other related genes associated with each gene.
[0127] It should be understood that in a circular gene, starting from a certain gene, the genes are numbered sequentially in a clockwise direction, as follows: Dr.gene0, Dr.gene1, ..., Dr.gene0. n-1 For each gene Dr.gene i (0≤i≤n-1); the other two genes associated with it are jointly determined by n and p, which is solved in the solution.
[0128] In the specific implementation, n and p are used to solve for the other two genes pointer1 and pointer2 associated with each gene.
[0129] In a circular gene sequence, starting from a given gene, the genes are numbered sequentially in a clockwise direction: Dr.gene0, Dr.gene1, ..., Dr.gene... n-1 For each gene Dr.gene i (0≤i≤n-1), the other two genes associated with it are jointly determined by n and p solved in step 1.1. Let Dr.gene i The other two related genes are Pointer1 and Pointer2, then:
[0130] like or but If the solution to the equation listed in step 1 is i, then when i = 0: When i = (p-1) -1 When (mod n): otherwise,
[0131] like and If it is not a solution to the equation listed in step 1, then
[0132] In this example, but To find the solution to the equations listed in step 1, the steps for solving Pointer1 and Pointer2 for each gene are as follows:
[0133] 1) Encode Dr.gene0, where i = 0:
[0134]
[0135] Therefore, it was found that two side bars should be added to the network graph.<Dr.gene0,Dr.gene1> ,<Dr.gene0,Dr.gene4> ,like Figure 8 As shown, Figure 8This is a first structural diagram of the discovery network process in the multi-gene discovery network construction method of the present invention, wherein the arrow from Dr.gene0 to Dr.gene1 represents Pointer1, and the arrow from Dr.gene0 to Dr.gene4 represents Pointer2. Dr.gene0 should store the position information of Dr.gene1 and Dr.gene4.
[0136] 2) Encode Dr.gene1, where i = 1:
[0137]
[0138] Therefore, it was found that two edges should be added to the network graph, namely...<Dr.gene1,Dr.gene2> ,<Dr.gene1,Dr.gene3> ,like Figure 9 As shown, Figure 9 This is a second schematic diagram of the discovery network process in the multi-gene discovery network construction method of the present invention. The arrows from Dr.gene0 to Dr.gene1 and from Dr.gene1 to Dr.gene2 represent Pointer1, and the arrows from Dr.gene0 to Dr.gene4 and from Dr.gene1 to Dr.gene3 represent Pointer2. Dr.gene1 should store the location information of Dr.gene2 and Dr.gene3.
[0139] 3) Encode Dr.gene2, where i = 2:
[0140]
[0141] Therefore, it was found that two side bars should be added to the network graph.<Dr.gene2,Dr.gene3> ,<Dr.gene2,Dr.gene1> ,like Figure 10 As shown, Figure 10 This is a third schematic diagram of the discovery network process in the multi-gene discovery network construction method of the present invention. The arrows from Dr.gene0 to Dr.gene1, from Dr.gene1 to Dr.gene2, and from Dr.gene2 to Dr.gene3 represent Pointer1. The arrows from Dr.gene0 to Dr.gene4, from Dr.gene1 to Dr.gene3, and from Dr.gene2 to Dr.gene1 represent Pointer2. Dr.gene2 should store the location information of Dr.gene3 and Dr.gene1.
[0142] 4) Encode Dr.gene3, where i = 3, which is (p-1). -1(mod n):
[0143]
[0144] Therefore, it was found that two side bars should be added to the network graph.<Dr.gene3,Dr.gene4> ,<Dr.gene3,Dr.gene0> ,like Figure 11 As shown, Figure 11 This is a fourth schematic diagram of the discovery network process in the multi-gene discovery network construction method of the present invention. The arrows from Dr.gene0 to Dr.gene1, from Dr.gene1 to Dr.gene2, from Dr.gene2 to Dr.gene3, and from Dr.gene3 to Dr.gene4 represent Pointer1. The arrows from Dr.gene0 to Dr.gene4, from Dr.gene1 to Dr.gene3, from Dr.gene2 to Dr.gene1, and from Dr.gene3 to Dr.gene0 represent Pointer2. Dr.gene3 should store the position information of Dr.gene4 and Dr.gene0.
[0145] 5) Encode Dr.gene4, where i = 4:
[0146]
[0147] Therefore, it was found that two side bars should be added to the network graph.<Dr.gene4,Dr.gene0> ,<Dr.gene4,Dr.gene2> ,like Figure 12 As shown, Figure 12 This is the fifth schematic diagram of the discovery network process in the multi-gene discovery network construction method of the present invention. The arrows from Dr.gene0 to Dr.gene1, from Dr.gene1 to Dr.gene2, from Dr.gene2 to Dr.gene3, from Dr.gene3 to Dr.gene4, and from Dr.gene4 to Dr.gene0 represent Pointer1. The arrows from Dr.gene0 to Dr.gene4, from Dr.gene1 to Dr.gene3, from Dr.gene2 to Dr.gene1, from Dr.gene3 to Dr.gene0, and from Dr.gene4 to Dr.gene2 represent Pointer2. Dr.gene4 should store the position information of Dr.gene0 and Dr.gene2.
[0148] Step S133: Obtain the target natural number of the associated gene to be embedded, and use the target natural number as the storage information for each gene to be encoded to be embedded.
[0149] It should be understood that the information to be embedded in the associated genes, namely the target natural number, is obtained so that the target natural number can be used as the storage information to be embedded into each gene to be encoded.
[0150] In the specific implementation, let Dr.gene i The other two associated genes are Pointer1 and Pointer2. Pointer1 and Pointer2 can be used to solve for the information to be embedded in each gene: tagleft1, tagright1, tagleft2, and tagright2. For each gene, Dr.gene... i (0≤i≤n-1), based on the calculated Pointer1 and Pointer2, we obtain the four natural numbers to be embedded in the gene: tagleft1, tagright1, tagleft2, tagright2; tagleft1 is Dr.gene. i The distance between the tail of the gene and the head of Pointer1 is calculated clockwise, and tagright1 is the distance between the tail of the gene and the tail of Pointer1 calculated clockwise; tagleft2 is the distance between the head of the gene and the head of Pointer2 calculated counterclockwise, and tagright2 is the distance between the head of the gene and the tail of Pointer2 calculated counterclockwise.
[0151] This embodiment, through the above scheme, obtains the number of genes contained in each circular genome, solves each circular genome based on the number of genes, and obtains the corresponding genome decoding information; solves the associated genes related to each gene to be encoded based on the number of genes and the gene decoding information; obtains the target natural number of the associated gene to be embedded, and uses the target natural number as the storage information to be embedded for each gene to be encoded, thus determining the storage information to be embedded for each gene to be encoded, and finding the corresponding gene quickly and efficiently during decoding.
[0152] Furthermore, Figure 13 This is a flowchart illustrating the sixth embodiment of the multi-gene discovery network construction method of the present invention, as shown below. Figure 13 As shown, based on the first embodiment, a sixth embodiment of the multi-gene discovery network construction method of the present invention is proposed. In this embodiment, step S20 specifically includes the following steps:
[0153] Step S21: Obtain the location information of each gene to be encoded from the discovery network, and perform embedded encoding on each gene to be encoded according to the location information and the storage information to generate a DNA sequence.
[0154] It should be noted that each gene in each gene to be encoded has different location information. By embedding each gene to be encoded according to the location information and the storage information, a DNA sequence can be generated.
[0155] Understandably, each gene in the genome is encoded, and a specific information embedding algorithm is used to embed the location information of the gene that each node in the graph points to into that node. Finally, a DNA sequence containing information related to the discovery network is output.
[0156] Step S22: Perform biochemical processing on the DNA sequence, and sequence the processed DNA sequence to obtain sequencing results.
[0157] It should be understood that after the DNA sequence is processed by a biochemical process, the corresponding sequencing results can be obtained; the encoded DNA sequence is output according to the discovery network algorithm for multiple genes, and the synthesized DNA sequence is sequenced after a biochemical process to obtain the sequencing results of the DNA sequence. This result may be the positive strand or the reverse complementary strand of the original DNA sequence and may be cut from any position of the circular gene.
[0158] In practice, the DNA sequence is subjected to a series of biological processes, including but not limited to culturing, amplification, and storage. The resulting DNA sequence is then sequenced to obtain the sequencing results. Biochemical processing may introduce substitution errors, i.e., a base at a certain position is replaced with another base; insertion errors, i.e., a base is added at a certain position; and deletion errors, i.e., a base is deleted at a certain position.
[0159] Understandably, the synthesized DNA sequence is generated by the output of a discovery network algorithm for multiple genes. After undergoing a biochemical process, the DNA sequence is sequenced to obtain the sequencing result of the DNA sequence. This result may be the positive strand or the reverse complementary strand of the original DNA sequence and may be cut from any position of the circular gene.
[0160] It is important to note that the fault tolerance of the algorithm in this embodiment is limited to errors occurring in the base sequences of non-gene sequences within the genome; and this algorithm does not provide the function of restoring the original sequence, but only focuses on whether all genes can still be found during decoding even with errors; when substitution errors occur in non-gene sequences, no matter how many substitution errors there are, it will not affect the algorithm's ability to find all genes during decoding; when deletion and insertion errors occur in non-gene sequences in a mixed manner without causing a change in sequence length, it will also not affect the algorithm's ability to find all genes during decoding; when addition or deletion errors occur in non-gene sequences and cause a change in the base sequence length, it depends on the parameters input by the user during decoding, which represent the time cost the user is willing to pay for decoding.
[0161] This embodiment, through the above-described scheme, obtains the location information of each gene to be encoded from the discovery network, performs embedded encoding on each gene to be encoded based on the location information and the storage information, and generates a DNA sequence; the DNA sequence is then processed through a biochemical process, and the processed DNA sequence is sequenced to obtain sequencing results; the corresponding gene can be found quickly and efficiently during decoding, saving gene search time.
[0162] Furthermore, Figure 14 This is a flowchart illustrating the seventh embodiment of the multi-gene discovery network construction method of the present invention, as shown below. Figure 14 As shown, a seventh embodiment of the multi-gene discovery network construction method of the present invention is proposed based on the first embodiment. In this embodiment, step S30 specifically includes the following steps:
[0163] Step S31: Traverse the sequencing results according to the start codons to generate a start table, and traverse the sequencing results according to the stop codons to generate a stop table. The start table is a list of the positions of all start codons in the genome, and the stop table is a list of the positions of all stop codons in the genome.
[0164] It should be noted that traversing the sequencing results based on the start codon generates a start table, and traversing the sequencing results based on the stop codon generates a stop table. Since genes all begin with a start codon and end with a stop codon, in order to find genes, the sequencing results are traversed to generate two lists: a start table and a stop table. The start table lists the positions of all start codons in the genome, and the stop table lists the positions of all stop codons in the genome.
[0165] Step S32: Decode the sequencing results according to the start table and the end table, and find the target gene located in the preset indirect sequence length range in the sequencing results.
[0166] Understandably, the decoding algorithm needs to first exhaustively search for start and stop codons and then pair them up in pairs to find the first gene; based on the start table and the stop table, the sequencing results can be decoded to find the target gene located in the preset indirect sequence length range in the sequencing results.
[0167] In the specific implementation, the first gene Dr.gene′0 is found through exhaustive search; the DNA sequence sequencing result S2 is obtained, with a maximum allowable error cost of bits max. The main output is the location information of all genes. Here, S2 may be the sequence obtained by cutting the genome sequence S1 in the discovery network at any position or the sequence obtained by cutting the reverse complementary strand of S1 at any position; correspondingly, two lists start′ and stop′ are also generated for the reverse complementary strand S′2 of S2.
[0168] 1) Set the initial value min, representing the minimum length of the gene. Set the step size to switch, and let Min_len = min, Max_len = min + switch.
[0169] 2) Pair the start in the start table with the stop in the stop table to filter out sequences whose base sequence length from start to stop is in the interval [Min, Min+switch] and use a specific decoding algorithm to determine whether the gene encoded in step 1 can be found.
[0170] If any gene can be found, integrate S into a sequence starting with Dr.gene′0_start, record the starting position of this gene as Dr.gene′0_start and Dr.gene′0_stop, and exit the algorithm.
[0171] If the start and stop tables are not found after traversing them, proceed to step 3).
[0172] 3) Pair up the start in the start' table and the stop in the stop' table to filter out sequences whose base sequence length from start to stop is in the interval [Min, Min+switch] and use a specific decoding algorithm to determine whether the gene encoded in step 1 can be found.
[0173] If any gene can be found, integrate S′ into a sequence starting with Dr.gene′0_start, record the starting position of this gene as Dr.gene′0_start and Dr.gene′0_stop, and exit the algorithm.
[0174] If the start' and stop' tables are not found after traversing them, proceed to step 4).
[0175] 4) Let Min_len = Max_len, Min_len = Max_len, Max_len = Max_len + switch, and go to step 2).
[0176] Step S33: Obtain the gene location information of the target gene, and determine whether there is a encoded gene in other genes besides the target gene based on the gene location information.
[0177] It should be understood that obtaining the gene location information of the target gene allows for determining whether the encoded gene exists in other genes besides the target gene. By exhaustively searching and finding the first gene, Dr.gene′0, other genes can be found. Let the currently found gene be Dr.gene′, and the starting positions of Dr.gene′ in the genome be Dr.gene′_start and Dr.gene′_stop, respectively. Initially, let Dr.gene′ = Dr.gene′0.
[0178] 1) Extract the four stored values from Dr.gene: tagleft1, tagright1, tagleft2, tagright2;
[0179] 2) Find all start values located in [Dr.gene′_start-tagleft1-max,Dr.gene′_start-tagleft1+max] and add them to stack Stack0; find all stop values located in [Dr.gene′0_stop-tagright1-max,Dr.gene′0_stop-tagright2+max] and add them to stack Stack1;
[0180] The start element located at [Dr.gene′_stop+tagleft2-max,Dr.gene′_stop+tagleft2+max] is added to stack Stack0′, and all stop elements located at [Dr.gene′0_stop+tagright2-max,Dr.gene′0_stop+tagright2+max] are added to stack Stack1′.
[0181] 3) Pop from Stack0 to get start, pop from Stack1 to get stop, or pop from Stack0' to get start, pop from Stack1' to get stop; for sequences with position interval [start, stop], use a specific decoding algorithm to determine whether the encoded gene can be found.
[0182] If the encoded gene is found, then assign Dr.gene′ to this gene, record the location information of this gene, and which gene was used to find this gene. Then determine whether a discovery network graph with each gene as a node can be constructed with an in-degree and out-degree of 2 for each node. If yes, it means that all genes have been found, exit this algorithm, and output the sequence S′ and the location information of the found genes. If not, it means that there are still genes that have not been found, and go to step 1.
[0183] If the encoded gene cannot be found, repeat step 3) until Stack0 or Stack1 is empty. At this point, exit the algorithm and output sequence S′ and the location information of the found gene.
[0184] It should be understood that the multi-gene discovery network algorithm decodes the sequencing results of the DNA sequence during decoding. The decoding process provides fault tolerance for base sequence errors and fragment loss that occur between genes. The goal is to quickly find all genes that embed information during coding from the sequence of the sequencing results.
[0185] Furthermore, the genome to be encoded must be a circular genome, and it contains genes capable of encoding and storing information, as well as the base sequence between each gene.
[0186] Furthermore, the input to the multi-gene discovery network algorithm encoding algorithm includes the genome sequence to be encoded and the location information of the gene sequence that can be encoded to store discovery network-related information.
[0187] Furthermore, in the encoding algorithm of the multi-gene discovery network algorithm, the genes that can be used to construct the discovery network must begin with a specific start codon and end with a specific stop codon.
[0188] Furthermore, in the encoding algorithm of the multi-gene discovery network algorithm, the genes that can encode and store information must be able to store at least four 64-bit natural numbers.
[0189] The input to the decoding algorithm of the multi-gene discovery network algorithm includes at least the sequencing results of the DNA sequence.
[0190] Furthermore, the fault tolerance provided by the multi-gene discovery network algorithm is based on the base sequence between each gene, and the supported error types include any number of substitution errors, insertion errors, and deletion errors, tolerating the above errors with the base as the smallest unit.
[0191] Furthermore, the decoding algorithm of the multi-gene discovery network algorithm needs to first exhaustively search for start and stop codons and then pair them up in pairs to find the first gene.
[0192] Furthermore, the decoding algorithm of the multi-gene discovery network algorithm needs to find other genes based on the first found gene, similar to starting from a node in a graph and traversing to all other nodes in the graph. That is, it uses the fast divergence of the discovery network algorithm to efficiently find all genes.
[0193] Before and after biological processes that may introduce insertion, deletion, substitution errors, and fragment loss, a discovery network is constructed during the encoding of the genome to be encoded, and genes are searched during genome decoding. In the encoding process, each gene in the circular genome can be abstracted as a node in graph theory. A discovery network graph with two out-degrees and two in-degrees is constructed by embedding the positional information of two other genes into each gene using a coding algorithm that incorporates this information. This discovery network ensures that even if a gene is erroneous and cannot be found, or if fragment loss occurs in the genome, it can minimize the impact on finding other genes. In the decoding process, for the sequencing sequence, an exhaustive search of start and stop codons is first used to find the first gene. After finding the first gene, the location of all genes is efficiently found based on the discovery network graph, and a certain degree of error tolerance is also provided. When the genome contains serious errors, or when the sequencing sequence is wild-type and uncoded, the decoding algorithm can detect and report the errors.
[0194] This invention constructs a discovery network with genes as the smallest computational unit during encoding, supporting a flexible number of genes contained in the genome of the discovery network, with no limit on the number of supported genes.
[0195] During decoding, when the location information of a gene is known, this invention can use that gene to quickly find all other genes at a diffusion rate of 2^n under ideal conditions.
[0196] During decoding, this invention can determine whether all genes have been identified by assessing whether the graph constructed during encoding can be reconstructed based on the pointing relationships between the identified genes.
[0197] This embodiment, through the above-described scheme, generates a start table by traversing the sequencing results based on start codons and a stop table by traversing the sequencing results based on stop codons. The start table is a list of the positions of all start codons in the genome, and the stop table is a list of the positions of all stop codons in the genome. The sequencing results are decoded based on the start and stop tables to find the target gene located within a preset indirect sequence length range. The gene location information of the target gene is obtained, and the presence of encoded genes in other genes besides the target gene is determined based on the gene location information. This provides a certain degree of error tolerance; when the error in the genome is severe, or when the sequencing sequence is wild-type and uncoded, the decoding algorithm can detect and report the error, and can quickly and efficiently find the corresponding gene during decoding, saving gene search time.
[0198] Accordingly, the present invention further provides a multi-gene discovery network construction device.
[0199] Reference Figure 15 , Figure 15 This is a functional block diagram of the first embodiment of the multi-gene discovery network construction device of the present invention.
[0200] In a first embodiment of the multi-gene discovery network construction device of the present invention, the multi-gene discovery network construction device includes:
[0201] The information acquisition module 10 is used to acquire the genomic information input by the user, construct a discovery network among all genes to be encoded based on the genomic information, and determine the storage information of each gene to be encoded.
[0202] The sequencing module 20 is used to generate a DNA sequence based on the discovery network and the stored information, and to sequence the DNA sequence to obtain sequencing results.
[0203] The decoding module 30 is used to decode the sequencing results, obtain gene location information, and determine whether the encoded gene exists based on the gene location information.
[0204] The information acquisition module 10 is further configured to acquire genomic information input by the user, construct a discovery network based on the genomic information; obtain the spacer sequence between each gene to be encoded, form multiple circular genomes based on the spacer sequence; acquire the number of genes and genome decoding information of each circular genome, acquire the associated genes of each gene to be encoded based on the number of genes and the genome decoding information, and determine the storage information to be embedded in each gene to be encoded based on the associated genes.
[0205] The information acquisition module 10 is also used to acquire the genomic information input by the user, abstract each gene to be encoded into a node, construct a directed unweighted connected graph in which each node has two in-degrees and two out-degrees, and generate a discovery network based on the directed unweighted connected graph.
[0206] The information acquisition module 10 is further configured to identify all genes to be encoded in a clockwise direction according to the genomic information, obtain the spacer sequence between each gene to be encoded, and divide each gene to be encoded into multiple circular genomes according to the spacer sequence.
[0207] The information acquisition module 10 is further configured to acquire the number of genes contained in each circular genome, solve each circular genome according to the number of genes, and obtain the corresponding genome decoding information; solve the associated genes associated with each gene to be encoded according to the number of genes and the gene decoding information; acquire the target natural number of the associated genes to be embedded, and use the target natural number as the storage information to be embedded for each gene to be encoded.
[0208] The sequencing module 20 is further configured to obtain the location information of each gene to be encoded from the discovery network, perform embedded encoding on each gene to be encoded according to the location information and the storage information to generate a DNA sequence, perform biochemical processing on the DNA sequence, and sequence the processed DNA sequence to obtain sequencing results.
[0209] The decoding module 30 is further configured to traverse the sequencing results according to the start codons to generate a start table, and traverse the sequencing results according to the stop codons to generate a stop table. The start table is a list of the positions of all start codons in the genome, and the stop table is a list of the positions of all stop codons in the genome. The module decodes the sequencing results according to the start table and the stop table to find the target gene located in the sequencing results within a preset indirect sequence length range. The module obtains the gene location information of the target gene and determines whether there is a encoded gene in other genes besides the target gene based on the gene location information.
[0210] The steps for implementing each functional module of the multi-gene discovery network construction device can be referred to in the various embodiments of the multi-gene discovery network construction method of the present invention, and will not be repeated here.
[0211] Furthermore, embodiments of the present invention also propose a storage medium storing a multi-gene discovery network construction program, which, when executed by a processor, performs the following operations:
[0212] Obtain the genomic information input by the user, construct a discovery network among all genes to be encoded based on the genomic information, and determine the storage information of each gene to be encoded;
[0213] A DNA sequence is generated based on the discovery network and the stored information, and the DNA sequence is sequenced to obtain sequencing results;
[0214] The sequencing results are decoded to obtain gene location information, and the presence of the encoded gene is determined based on the gene location information.
[0215] Furthermore, when the multi-gene discovery network construction program is executed by the processor, it also performs the following operations:
[0216] Obtain the genomic information input by the user, and construct a discovery network based on the genomic information;
[0217] Obtain the spacer sequences between each gene to be encoded, and form multiple circular genomes based on the spacer sequences;
[0218] The number of genes and genome decoding information of each circular genome are obtained. Based on the number of genes and genome decoding information, the associated genes of each gene to be encoded are obtained. Based on the associated genes, the storage information to be embedded for each gene to be encoded is determined.
[0219] Furthermore, when the multi-gene discovery network construction program is executed by the processor, it also performs the following operations:
[0220] Obtain the genomic information input by the user and abstract each gene to be encoded into a node;
[0221] Construct a directed unweighted connected graph where each node has two in-degrees and two out-degrees, and generate a discovery network based on the directed unweighted connected graph.
[0222] Furthermore, when the multi-gene discovery network construction program is executed by the processor, it also performs the following operations:
[0223] Based on the genomic information, all genes to be encoded are identified in a clockwise direction to obtain the spacer sequences between each gene to be encoded, and each gene to be encoded is divided into multiple circular genomes based on the spacer sequences.
[0224] Furthermore, when the multi-gene discovery network construction program is executed by the processor, it also performs the following operations:
[0225] The number of genes contained in each circular genome is obtained, and the corresponding genome decoding information is obtained by solving each circular genome based on the number of genes.
[0226] Based on the number of genes and the gene decoding information, the associated genes of each gene to be encoded are determined;
[0227] Obtain the target natural number of the associated gene to be embedded, and use the target natural number as the storage information for each gene to be encoded.
[0228] Furthermore, when the multi-gene discovery network construction program is executed by the processor, it also performs the following operations:
[0229] The location information of each gene to be encoded is obtained from the discovery network, and each gene to be encoded is embedded according to the location information and the storage information to generate a DNA sequence.
[0230] The DNA sequence is subjected to a biochemical process, and the processed DNA sequence is sequenced to obtain the sequencing results.
[0231] Furthermore, when the multi-gene discovery network construction program is executed by the processor, it also performs the following operations:
[0232] The sequencing results are iterated through based on the start codons to generate a start table, and the sequencing results are iterated through based on the stop codons to generate a stop table. The start table is a list of the positions of all start codons in the genome, and the stop table is a list of the positions of all stop codons in the genome.
[0233] The sequencing results are decoded according to the start table and the end table to find the target gene located in the preset indirect sequence length range in the sequencing results;
[0234] Obtain the gene location information of the target gene, and determine whether the encoded gene exists in other genes besides the target gene based on the gene location information.
[0235] This embodiment, through the above-described scheme, acquires the genomic information input by the user, constructs a discovery network among all genes to be encoded based on the genomic information, and determines the storage information of each gene to be encoded; generates a DNA sequence based on the discovery network and the storage information, and sequences the DNA sequence to obtain sequencing results; decodes the sequencing results to obtain gene location information, and determines whether the encoded gene exists based on the gene location information. This allows for rapid and efficient locating of the corresponding gene during decoding, saving gene search time, quickly detecting gene reporting errors, and avoiding sequencing result errors caused by various variations and the sequencing process itself.
[0236] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0237] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0238] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A multi-gene discovery network construction method, characterized by, The multi-gene discovery network construction method comprises: obtaining user input genome information, constructing a discovery network between all to-be-encoded genes according to the genome information, and determining storage information of each to-be-encoded gene; generating a DNA sequence according to the discovery network and the storage information, and sequencing the DNA sequence to obtain a sequencing result; decoding the sequencing result to obtain gene position information, and determining whether there is an encoded gene according to the gene position information; wherein, the obtaining user input genome information, constructing a discovery network between all to-be-encoded genes according to the genome information, and determining storage information of each to-be-encoded gene comprises: obtaining user input genome information, and constructing a discovery network according to the genome information; obtaining interval sequences between each to-be-encoded gene, and forming a plurality of circular genomes according to the interval sequences; obtaining the number of genes and genome decoding information of each circular genome, obtaining associated genes of each to-be-encoded gene according to the number of genes and the genome decoding information, and determining storage information to be embedded by each to-be-encoded gene according to the associated genes; wherein, the decoding the sequencing result to obtain gene position information, and determining whether there is an encoded gene according to the gene position information comprises: traversing the sequencing result according to a start codon to generate a start table, and traversing the sequencing result according to a stop codon to generate a stop table, the start table being a list of positions of all start codons in the genome, and the stop table being a list of positions of all stop codons in the genome; decoding the sequencing result according to the start table and the stop table to find a target gene located in a preset interval sequence length interval in the sequencing result; obtaining gene position information of the target gene, and determining whether there is an encoded gene in other genes except the target gene according to the gene position information.
2. The multi-gene discovery network construction method of claim 1, wherein, The obtaining user input genome information, constructing a discovery network according to the genome information comprises: obtaining user input genome information, and abstracting each to-be-encoded gene as a node; constructing a directed acyclic graph with two in-degree and out-degree for each node, and generating a discovery network according to the directed acyclic graph.
3. The multi-gene discovery network construction method of claim 1, wherein, The obtaining interval sequences between each to-be-encoded gene, and forming a plurality of circular genomes according to the interval sequences comprises: identifying all to-be-encoded genes in a clockwise direction according to the genome information to obtain interval sequences between each to-be-encoded gene, and grouping each to-be-encoded gene into a plurality of circular genomes according to the interval sequences.
4. The multi-gene discovery network construction method of claim 1, wherein, The obtaining the number of genes and genome decoding information of each circular genome, obtaining associated genes of each to-be-encoded gene according to the number of genes and the genome decoding information, and determining storage information to be embedded by each to-be-encoded gene according to the associated genes comprises: obtaining the number of genes contained in each circular genome, solving each circular genome according to the number of genes to obtain corresponding genome decoding information; solving associated genes associated with each to-be-encoded gene according to the number of genes and the genome decoding information. Obtain a target natural number to be embedded with the associated genes, and take the target natural number as the storage information to be embedded with each gene to be encoded.
5. The multi-gene discovery network construction method of claim 1, wherein, The DNA sequence is generated according to the discovery network and the storage information, and the DNA sequence is sequenced to obtain a sequencing result, which includes: The position information of each gene to be encoded is obtained from the discovery network, and each gene to be encoded is embedded and coded according to the position information and the storage information to generate a DNA sequence. The DNA sequence is subjected to a biochemical process, and the processed DNA sequence is sequenced to obtain a sequencing result.
6. A multi-gene discovery network construction apparatus, characterized by comprising: The multi-gene discovery network construction device includes: An information acquisition module is configured to acquire genome information input by a user, construct a discovery network among all genes to be encoded according to the genome information, and determine storage information of each gene to be encoded; A sequencing module is configured to generate a DNA sequence according to the discovery network and the storage information, and sequence the DNA sequence to obtain a sequencing result; A decoding module is configured to decode the sequencing result to obtain gene position information, and determine whether there is a coded gene according to the gene position information. The information acquisition module is further configured to acquire genome information input by a user, construct a discovery network according to the genome information, obtain interval sequences between each gene to be encoded, form a plurality of circular genomes according to the interval sequences, acquire a number of genes and genome decoding information of each circular genome, acquire associated genes of each gene to be encoded according to the number of genes and the genome decoding information, and determine storage information to be embedded with each gene to be encoded according to the associated genes. The decoding module is further configured to traverse the sequencing result according to a start codon to generate a start table, traverse the sequencing result according to a stop codon to generate a stop table, the start table is a list of positions of all start codons in a genome, and the stop table is a list of positions of all stop codons in the genome; decode the sequencing result according to the start table and the stop table, find a target gene located in a preset interval sequence length range in the sequencing result; obtain gene position information of the target gene, and determine whether there is a coded gene in other genes except the target gene according to the gene position information.
7. A multi-gene discovery network construction apparatus, characterized by comprising: The multi-gene discovery network construction device includes a memory, a processor, and a multi-gene discovery network construction program stored on the memory and executable on the processor, and the multi-gene discovery network construction program is configured to implement the steps of the multi-gene discovery network construction method according to any one of claims 1 to 5.
8. A storage medium, characterized by The storage medium stores a multi-gene discovery network construction program, and the multi-gene discovery network construction program is executed by the processor to implement the steps of the multi-gene discovery network construction method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and apparatus for the access to bioinformatics data structured in access units
CA3040057A1
Method and system for selective access of stored or transmitted bioinformatics data
CA3040138A1