Methods and systems for assembling medicinal plant genomes based on high-throughput sequencing

By influencing sequences through screening, performing sequence clustering, and constructing association networks, the splicing errors and redundancy caused by repetitive sequences in the assembly of medicinal plant genomes during high-throughput sequencing were resolved, thereby improving the accuracy and integrity of the assembly.

CN121565257BActive Publication Date: 2026-04-03XIAN BOTANICAL GARDEN SHAANXI PROV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

High-throughput sequencing technology cannot traverse long repetitive sequence fragments in the assembly of medicinal plant genomes, leading to sequence splicing errors, fragment breaks, and redundant assembly, which reduces the integrity and accuracy of repetitive sequence regions.

Method used

By comparing the target sequencing sequence with a library of known functional sequences, sequences with functional influences are screened, sequence clustering and feature analysis are performed, non-repetitive sequence clusters and repetitive sequence clusters are divided, and an association network is constructed. Sequence splicing is performed based on the dynamic association strength and the differential assembly path.

Benefits of technology

It significantly improves the accuracy and integrity of medicinal plant genome assembly, and solves the splicing errors and redundancy problems caused by repetitive sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565257B_ABST
    Figure CN121565257B_ABST
Patent Text Reader

Abstract

This invention relates to the field of genome assembly technology, specifically to a method and system for assembling medicinal plant genomes based on high-throughput sequencing. The method includes: acquiring target sequencing sequences and screening sequences that affect function; classifying the target sequencing sequences into repetitive clusters and non-repetitive sequence clusters, further dividing the repetitive sequence clusters into functional repetitive sequences and non-functional repetitive sequences; determining the dynamic association strength by combining the expression abundance of functional genes, the content of medicinal components, and the base overlap characteristics with adjacent sequences; constructing an association network; and, based on the dynamic association strength in the association network, using differentiated assembly paths to splice sequences from functionally associated repetitive sequences, non-functionally associated repetitive sequences, and adjacent non-repetitive sequences, to obtain a complete genome assembly result. This invention can solve the problems of splicing errors, fragment breaks, and redundancy caused by repetitive sequences, significantly improving the accuracy and completeness of assembly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of genome assembly technology, and specifically to a method and system for assembling the genome of medicinal plants based on high-throughput sequencing. Background Technology

[0002] Medicinal plants are a core resource for traditional medicine and modern drug development. The analysis of their genome information is crucial for elucidating the synthetic pathways of medicinal components, identifying functional genes, promoting variety improvement, and facilitating industrial applications. The rapid development of high-throughput sequencing technology has provided an efficient and low-cost means of data acquisition for medicinal plant genome research. Genome assembly, the key step in piecing together short sequence fragments obtained from high-throughput sequencing into a complete genome sequence, directly affects the accuracy of medicinal plant genome analysis.

[0003] Short reads from high-throughput sequencing cannot span longer repetitive sequence fragments, which can lead to misidentification of repetitive sequences at different positions as the same region during assembly, resulting in sequence splicing errors, fragment breaks, or redundant assembly. At the same time, the large number of tandem repeats in the genomes of medicinal plants further increases the difficulty of sequence alignment and error correction, resulting in a significant reduction in the integrity and accuracy of repetitive sequence regions in the assembly results. Summary of the Invention

[0004] To address the technical problem in related technologies where short reads from high-throughput sequencing cannot span longer repetitive sequence fragments, resulting in a significant reduction in the integrity and accuracy of repetitive sequence regions in the assembly results, this invention provides a method and system for assembling medicinal plant genomes based on high-throughput sequencing. The specific technical solution adopted is as follows:

[0005] This invention proposes a method for assembling the genome of medicinal plants based on high-throughput sequencing, the method comprising:

[0006] Obtain the target sequencing sequence, and screen for functionally influential sequences based on the alignment of the target sequencing sequence with a library of known functional sequences;

[0007] By using sequence clustering and target sequencing sequence feature analysis, the target sequencing sequences are classified into repetitive sequences to obtain non-repetitive sequence clusters and repetitive sequence clusters. The repetitive sequence clusters are then compared with functionally affected sequences to determine functional and non-functional repetitive sequences.

[0008] By combining the expression abundance of functional genes, the content of medicinal components, and the base overlap characteristics with adjacent sequences, the dynamic association strength of functionally influential sequences was determined. Non-repetitive sequences, functional repetitive sequences, and non-functional repetitive sequences in the non-repetitive sequence cluster were used as network nodes to construct an association network. The edge weights between network nodes were determined based on the dynamic association strength, and the edge directions were determined based on the transcription direction and the sequence end overlap characteristics.

[0009] Based on the dynamic association strength in the association network, differential assembly paths are used to assemble functionally associated repetitive sequences, non-functionally associated repetitive sequences, and adjacent non-repetitive sequences to obtain complete genome assembly results.

[0010] Furthermore, the screening of functionally influential sequences based on the alignment of the target sequencing sequence with a library of known functional sequences includes:

[0011] The high-quality, effective sequence set is compared with the functional sequence library, and sequencing sequences that have a continuous, predetermined length of overlap with the functional elements of the functional sequence library are selected as functionally influential sequences.

[0012] Furthermore, the process of classifying the target sequencing sequences into repetitive clusters and repetitive sequence clusters through sequence clustering and target sequencing sequence feature analysis includes:

[0013] All target sequencing sequences are globally paired and compared to calculate the base matching probability. Multiple sequence clusters are formed through agglomerative hierarchical clustering.

[0014] The copy number and average matching probability of each sequence cluster are calculated. Sequence clusters with both copy number and base matching probability higher than the average are selected as repetitive sequence clusters, and the others are selected as non-repetitive sequence clusters. The copy number is the number of sequences contained in each sequence cluster.

[0015] Furthermore, the repetitive sequence clusters are compared with functionally affected sequences to identify functional and non-functional repetitive sequences, including:

[0016] Calculate the overlap probability between each target sequencing sequence and each functionally affected sequence within the repetitive sequence cluster;

[0017] Based on the overlap probability, all target sequencing sequences within a repetitive sequence cluster are classified into functionally associated repetitive sequences and non-functionally associated repetitive sequences.

[0018] Furthermore, the determination of the dynamic association strength of functionally influential sequences by combining the expression abundance of functional genes, the content detection of medicinal components, and the base overlap characteristics with adjacent sequences includes:

[0019] Acquire transcriptomic and metabolomic data;

[0020] The expression abundance of functional genes corresponding to functionally influential sequences were extracted from transcriptome data, and the average value was calculated as a functional weight parameter.

[0021] The content of target medicinal components was extracted from metabolomics data, and the Pearson correlation coefficient between the component and the abundance of functional gene expression was calculated and normalized as a functional confidence parameter.

[0022] For functionally related repeating sequences, the product of the functional weight parameter and the functional confidence parameter is calculated and normalized to obtain the dynamic association strength.

[0023] For non-functionally associated repetitive sequences, the dynamic association strength is determined by combining functional weight parameters, functional confidence parameters, and the distribution of the preliminary genome assembly results of medicinal plants.

[0024] Furthermore, the determination of dynamic association strength by combining functional weight parameters, functional confidence parameters, and the distribution of preliminary genome assembly results in medicinal plants includes:

[0025] The preliminary assembly results of the genome of medicinal plants are divided into a predetermined number of continuous gene fragments of equal length;

[0026] Calculate the distribution information entropy of non-functionally associated repeat sequences in different gene segments, and use the normalized information entropy as a correction parameter.

[0027] The product of the correction parameter, the functional weight parameter, and the functional confidence parameter is calculated as the dynamic association strength of the non-functionally associated repeating sequence.

[0028] Furthermore, the method for obtaining the edge directions between the network nodes includes:

[0029] Determine the DNA sequence from the 5' end to the 3' end during transcription;

[0030] Use any two adjacent network nodes as the first node and the second node.

[0031] When the 3' end of the first node points to the 5' end of the second node, the edge direction is determined to be from the first node to the second node;

[0032] There is a base pair overlap at the 3' end of the first node and the 5' end of the second node, and the overlap length is greater than the preset length threshold. The edge direction is determined to be from the first node to the second node.

[0033] Furthermore, methods for assembling functionally related repetitive sequences to obtain complete genome assembly results include:

[0034] Functionally associated repetitive sequences are spliced ​​together with non-repetitive sequences in clusters of non-repetitive sequences adjacent to the transcriptional sequence to form functionally anchored core fragments;

[0035] Using the core segment of functional anchoring as the anchor point, extract all functional repetitive sequences directly connected to the anchor point from the association network, sort them from high to low according to the dynamic association strength, and select the functional repetitive sequence with the highest strength as the first extension object.

[0036] Based on whether the repetitive base units in the functionally related repeat sequence are arranged in the forward direction during DNA transcription, the extension direction is determined and the assembly is performed. If the repetitive base units are arranged in the forward repeat direction, the extension is performed in the direction of aligning the 3' end of the anchor point with the 5' end of the functionally related repeat sequence. If the repetitive base units are arranged in the reverse repeat direction, the base matching probability between the reverse complementary sequence of the repetitive base unit and the anchor point is verified. When the matching probability exceeds the preset matching threshold, the extension is performed in the direction of aligning the 3' end of the anchor point with the reverse complementary sequence at the 3' end of the repeat sequence.

[0037] After the first functionally associated repeat sequence is extended, the anchor point plus the first repeat sequence is used as a new temporary anchor point, and the extension continues with the new temporary anchor point to achieve genome assembly.

[0038] Furthermore, the method for assembling non-functionally associated repetitive sequences to obtain a complete genome assembly includes:

[0039] From the association network, select non-repetitive sequences from multiple non-repetitive sequence clusters that have high association strength with non-functional association repeating sequences. Then, concatenate the non-repetitive sequences with non-functional association repeating sequences to obtain candidate concatenation segments. The high association strength is defined as a dynamic association strength greater than a preset association threshold.

[0040] Calculate the sequence consistency and network association consistency for each candidate splice segment, where the sequence consistency is the normalized value of the average base matching probability between any two nodes in the candidate splice segment, and the network association consistency is the normalized value of the coefficient of variation of the dynamic association strength between nodes in the candidate splice segment.

[0041] Calculate the difference between sequence consistency and network association consistency as a judgment index; select the candidate splicing segment with the largest judgment index for splicing.

[0042] On the other hand, a medicinal plant genome assembly system based on high-throughput sequencing is also provided. The system includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the method as described in any of the foregoing.

[0043] The present invention has the following beneficial effects:

[0044] This invention obtains target sequencing sequences and, based on alignment with known functional sequence libraries, screens sequences that influence function. Then, it performs sequence clustering and feature analysis to classify the target sequencing sequences into repetitive clusters, resulting in non-repetitive sequence clusters and repetitive sequence clusters. These repetitive sequence clusters are further functionally classified into functional repetitive sequences and non-functional repetitive sequences. Next, it determines the dynamic association strength of functionally influential sequences by combining features such as the expression abundance of functional genes and the content of medicinal components, and constructs an association network. Finally, based on the dynamic association strength in the association network, it uses differentiated assembly paths to assemble functionally associated repetitive sequences, non-functionally associated repetitive sequences, and adjacent non-repetitive sequences, obtaining a complete genome assembly result. This invention, by establishing a function-oriented repetitive sequence classification system and a dynamic association network optimization mechanism, effectively solves the splicing errors, fragment breaks, and redundancy problems caused by repetitive sequences in high-throughput sequencing of medicinal plant genomes, significantly improving assembly accuracy and completeness. Attached Figure Description

[0045] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a flowchart of a method for assembling the genome of medicinal plants based on high-throughput sequencing, provided as an embodiment of the present invention. Detailed Implementation

[0047] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a method and system for assembling medicinal plant genomes based on high-throughput sequencing proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0049] Medicinal plants are a core resource for traditional medicine and modern drug development. The analysis of their genome information is crucial for elucidating the synthetic pathways of medicinal components, identifying functional genes, promoting variety improvement, and facilitating industrial applications. The rapid development of high-throughput sequencing technology has provided an efficient and low-cost means of data acquisition for medicinal plant genome research. Genome assembly, the key step in piecing together short sequence fragments obtained from high-throughput sequencing into a complete genome sequence, directly affects the accuracy of medicinal plant genome analysis.

[0050] Short reads from high-throughput sequencing cannot span longer repetitive sequence fragments, which can lead to misidentification of repetitive sequences at different positions as the same region during assembly, resulting in sequence splicing errors, fragment breaks, or redundant assembly. At the same time, the large number of tandem repeats in the genomes of medicinal plants further increases the difficulty of sequence alignment and error correction, resulting in a significant reduction in the integrity and accuracy of repetitive sequence regions in the assembly results.

[0051] The analysis is based on the above-mentioned technical issues, and specific details can be found in the following embodiments.

[0052] The following detailed description, in conjunction with the accompanying drawings, illustrates a specific scheme for a high-throughput sequencing-based method for assembling the genome of medicinal plants provided by this invention.

[0053] Please see Figure 1 The diagram illustrates a flowchart of a method for assembling the genome of a medicinal plant based on high-throughput sequencing, according to an embodiment of the present invention. The method includes:

[0054] S101: Obtain the target sequencing sequence and screen for functionally influential sequences based on the alignment of the target sequencing sequence with a library of known functional sequences.

[0055] The short sequences obtained by the sequencer are very short. The task of genome assembly is to organize these short sequences, find complete segments (gene fragments) related to the synthesis of medicinal components, and identify recurring short sequences as repetitive sequences.

[0056] In this embodiment of the invention, a raw sequencing short sequence is first obtained using a sequencer. The raw sequencing short sequence is then subjected to base quality filtering and low-complexity sequence removal. This part is a well-known technique and will not be described in detail here. The target sequencing sequence is obtained after processing.

[0057] Furthermore, in some embodiments of the present invention, the screening of functionally influential sequences based on the comparison between the target sequencing sequence and a known functional sequence library includes: comparing a high-quality effective sequence set with a functional sequence library, and screening out sequencing sequences that have a continuous preset length of overlap with the functional elements of the functional sequence library as functionally influential sequences.

[0058] In this embodiment of the invention, the target sequencing sequence is compared with a known functional sequence library (a conserved sequence library of functional elements of medicinal plants, containing conserved domains of key enzyme genes in the synthesis pathways of medicinal components such as terpenes and alkaloids, promoter core elements, transcription factor binding sites, etc.) to screen out target sequencing sequences that have a continuous overlap of more than 20 bp (preset length) with functional elements in the library as functionally influential sequences.

[0059] Among them, the functionally influential sequence is the potentially related gene sequence obtained by comparing it with the gene sequences of reported functional medicinal components.

[0060] S102: Through sequence clustering and target sequencing sequence feature analysis, the target sequencing sequences are classified into repetitive sequences to obtain non-repetitive sequence clusters and repetitive sequence clusters. The repetitive sequence clusters are compared with functionally affected sequences to determine functional repetitive sequences and non-functional repetitive sequences.

[0061] Since the target sequencing sequence contains a large number of repetitive and non-repetitive fragments, further differentiation is required. In this embodiment of the invention, repetitive classification is performed to obtain non-repetitive sequence clusters and repetitive sequence clusters.

[0062] Furthermore, in some embodiments of the present invention, the target sequencing sequences are classified into repetitive sequences through sequence clustering and target sequencing sequence feature analysis to obtain non-repetitive sequence clusters and repetitive sequence clusters. This includes: performing global pairwise alignment of all target sequencing sequences, calculating the base matching probability, and forming multiple sequence clusters through agglomerative hierarchical clustering; calculating the copy number and average cluster matching probability of each sequence cluster, and selecting sequence clusters with both copy number and base matching probability higher than the average value as repetitive sequence clusters, and the others as non-repetitive sequence clusters, wherein the copy number is the number of sequences contained in each sequence cluster.

[0063] In this embodiment of the invention, for any two target sequencing sequences, the base matching probability is determined based on the number of identical bases. The specific calculation method is as follows: First, the two target sequencing sequences are optimally aligned using a sequence alignment tool. The number of identical bases at corresponding positions after alignment is counted. Then, the base matching probability is calculated as the ratio of the number of identical bases to the total length of the two target sequencing sequences after alignment.

[0064] After traversing the base matching probabilities between all target sequencing sequences, agglomerative hierarchical clustering can be performed on all target sequencing sequences: First, each target sequencing sequence is treated as an initial cluster. Then, the base matching probabilities between all initial clusters are calculated, and the group of target sequencing sequences with the highest base matching probabilities is selected and merged into a new cluster. After the first merge, the similarity between all clusters is calculated, that is, the average base matching probabilities between any two target sequencing sequences in any two clusters are taken, and the pair of clusters with the highest average probabilities are merged into a larger new cluster. Then, the merging is repeated continuously. Each time, the global profile coefficient of all target sequencing sequences is calculated. The global profile coefficient is a global indicator. The closer the value is to 1, the more reasonable the current clustering is.

[0065] The global silhouette coefficient is calculated as follows: the mean base matching probability of the target sequencing sequence within the current cluster is calculated, and the average of each pair is used as the current similarity. The mean base matching probability of the target sequencing sequence within the current cluster is calculated and the average of each pair is used as the similarity of other clusters. The difference between the current similarity and the similarity of each other cluster is used as the similarity difference index. Then, the average of the similarity difference indices between all pairs of clusters is used as the global silhouette coefficient.

[0066] In other words, the larger the global contour coefficient, the greater the difference in base matching probability between different clusters, which means the better the clustering effect. Therefore, when the global contour coefficient is at its maximum value, sequence clustering is stopped, and the final sequence clusters are obtained.

[0067] In this embodiment of the invention, the number of target sequencing sequences contained in each sequence cluster, i.e. the copy number, is counted. The average copy number of all sequence clusters and the average base matching probability between any two target sequencing sequences within all sequence clusters are calculated as the cluster average matching probability.

[0068] Specifically, the cluster average matching probability is calculated by first calculating the average base matching probability between each pair of target sequencing sequences within each sequence cluster, which is then used as the cluster matching mean of the sequence cluster. Finally, the average of the cluster matching means of all sequence clusters is used as the cluster average matching probability.

[0069] Sequence clusters with a copy number greater than the average copy number of all sequence clusters and a cluster matching mean greater than the average cluster matching probability are classified as repeating sequence clusters; other sequence clusters are classified as non-repeating sequence clusters.

[0070] In this embodiment of the invention, all target sequencing sequences in each repetitive sequence cluster can be further divided to obtain functionally associated repetitive sequences and non-functionally associated repetitive sequences.

[0071] Furthermore, in some embodiments of the present invention, the repetitive sequence clusters are compared with functionally influential sequences to determine functional repetitive sequences and non-functional repetitive sequences, including: calculating the overlap probability between each target sequencing sequence and each functionally influential sequence within the repetitive sequence cluster; and classifying all target sequencing sequences within the repetitive sequence cluster into functionally associated repetitive sequences and non-functionally associated repetitive sequences based on the overlap probability.

[0072] In this embodiment of the invention, each target sequencing sequence within the repetitive sequence cluster is compared with each functionally affected sequence to determine the number of target sequencing sequences within the repetitive sequence cluster that have ≥10bp overlap with the functionally affected sequence and a base matching probability ≥60%, which is taken as the functionally relevant number. The ratio of the functionally relevant number to the total number of target sequencing sequences in the repetitive sequence cluster is taken as the overlap probability.

[0073] In this embodiment of the invention, target sequencing sequences with an overlap probability greater than 0.5 are classified as functionally related repetitive sequences, while others are classified as non-functionally related repetitive sequences. In this embodiment, the functional element types of functionally related repetitive sequences can be further labeled. Specific functional element types include conserved regions of alkaloid synthase genes, promoter core elements, transcription factor binding sites, etc.

[0074] S103: Combining the expression abundance of functional genes, the content detection of medicinal components, and the base overlap characteristics with adjacent sequences, the dynamic association strength of functionally influential sequences is determined; non-repetitive sequences, functional repetitive sequences and non-functional repetitive sequences in the non-repetitive sequence cluster are used as network nodes to construct an association network. The edge weights between network nodes are determined based on the dynamic association strength, and the edge directions are determined based on the transcription direction and the sequence end overlap characteristics.

[0075] After classifying the sequences into different types, an association network can be constructed based on the classification results.

[0076] The dynamic association strength of functionally influential sequences was determined by combining the expression abundance of functional genes, the content of medicinal components, and the base overlap characteristics with adjacent sequences. This process included: acquiring transcriptomic and metabolomic data; extracting the expression abundance of functional genes corresponding to the functionally influential sequences from the transcriptomic data and calculating the average as a functional weight parameter; extracting the content of the target medicinal component from the metabolomic data, calculating its Pearson correlation coefficient with the expression abundance of functional genes, and normalizing it as a functional confidence parameter; for functionally associated repetitive sequences, calculating the product of the functional weight parameter and the functional confidence parameter, and normalizing it as the dynamic association strength; and for non-functionally associated repetitive sequences, determining the dynamic association strength by combining the functional weight parameter, the functional confidence parameter, and the distribution of the preliminary genome assembly results in medicinal plants.

[0077] In one embodiment of the present invention, the normalization process can be specifically, for example, maximum and minimum value normalization. Furthermore, the normalization in subsequent steps can all adopt maximum and minimum value normalization. In other embodiments of the present invention, other normalization methods can be selected according to the specific range of the numerical values, which will not be elaborated further.

[0078] It should be noted that, for ease of calculation, all indicator data involved in the calculation in this embodiment of the invention have undergone data preprocessing to eliminate the influence of dimensions. The specific methods for eliminating the influence of dimensions are well known to those skilled in the art and are not limited here.

[0079] Metabolomics data represents information about the metabolome of medicinal plants, including the content of all components within the plant, such as organic acids, amino acids, nucleotides, carbohydrates, and lipid molecules. It also includes secondary metabolites closely related to plant stress resistance, such as flavonoids, alkaloids, phenols, and terpenes. Transcriptome data represents the set of all mRNA genes in a plant, including key genes for the synthesis of specific medicinal components (such as alkaloid synthase genes). Both transcriptome and metabolome data were obtained using well-known techniques, which will not be elaborated further.

[0080] Functional influencing sequences correspond to key genes in the synthesis of medicinal components (such as alkaloid synthase genes). The higher the expression level of these genes, the more active they are under the current physiological state, and the greater their contribution to the synthesis of medicinal components. Therefore, we extract the expression abundance information of functional genes from transcriptome data, calculate the average expression abundance of functional genes corresponding to functional influencing sequences, and use it as a functional weight parameter.

[0081] The expression abundance can be calculated using the RPKM (Reads Per Kilobase Per Million Reads) algorithm, a well-known technique, and will not be elaborated further. Each functionally influential sequence corresponds to multiple functional genes, and the average expression abundance of the functional genes corresponding to the functionally influential sequence is used as the functional weight parameter. That is, the larger the value of the functional weight parameter, the greater the potential contribution to the synthesis of medicinal components.

[0082] In this embodiment of the invention, each functional influence sequence corresponds to multiple functional genes (a combination of base pairs), and each functional gene corresponds to a content detection value of a target medicinal component. This value is obtained from the metabolome data of medicinal plants. The expression abundance of functional genes is arranged according to the fixed order of functional genes to obtain an abundance sequence. The content detection values ​​are then arranged to obtain a content sequence. Subsequently, the Pearson correlation coefficient between the abundance sequence and the content sequence is calculated. The Pearson correlation coefficient is normalized and used as a functional confidence parameter to verify whether the functional influence sequence truly affects the synthesis of the target component.

[0083] In this embodiment of the invention, for functionally related repeating sequences, the product of the functional weight parameter and the functional confidence parameter is directly calculated and normalized to obtain the dynamic association strength.

[0084] For non-functionally associated repetitive sequences, it is necessary to avoid misjudgment of association due to the dispersed distribution of non-functional repetitive sequences. Therefore, the dynamic association strength is determined by combining functional weight parameters, functional confidence parameters, and the distribution of the preliminary genome assembly results of medicinal plants.

[0085] Furthermore, in some embodiments of the present invention, the dynamic association strength is determined by combining the functional weight parameter, the functional confidence parameter, and the distribution of the preliminary genome assembly results of medicinal plants. This includes: dividing the preliminary genome assembly results of medicinal plants into a predetermined number of continuous gene fragments of equal length; calculating the distribution information entropy of non-functionally associated repetitive sequences in different gene fragments, and normalizing the information entropy as a correction parameter; and calculating the product of the correction parameter, the functional weight parameter, and the functional confidence parameter as the dynamic association strength of the non-functionally associated repetitive sequences.

[0086] Specialized assembly software (such as Canu, Flye, HiCanu, Necat, etc.) can be used for initial assembly. The initial assembly result is then divided into a preset number of continuous gene fragments of equal length. The preset number is set according to the actual detection needs, for example, 10. That is, the initial assembly result is divided into 10 continuous gene fragments of equal length.

[0087] Then, for all non-functionally related repeating sequences, count the number of sequences in each of the above segments, that is, the number of non-functionally related repeating sequences contained in the i-th segment. The ratio of this value to the total number of all non-functionally related repeating sequences is denoted as the frequency pi of non-functionally related repeating sequences in the i-th segment.

[0088] For example: If there are a total of 100 non-functional repetitive sequences, and the third segment contains 15 sequences, then the p3 of that segment is 15 ÷ 100 = 0.15.

[0089] Based on the pi of the non-functionally associated repeat sequence in each initially assembled gene fragment, the information entropy of the non-functionally associated repeat sequence in all initially assembled fragments is calculated. This entropy, after normalization, serves as a correction parameter, characterizing the dispersion of its distribution across all initially assembled gene fragments. The product of the correction parameter, functional weight parameter, and functional confidence parameter is then calculated as the dynamic association strength of the non-functionally associated repeat sequence. This approach avoids misclassification of associations due to the dispersed distribution of non-functionally associated repeat sequences.

[0090] Subsequently, non-repetitive sequences, functional repetitive sequences and non-functional repetitive sequences in the non-repetitive sequence cluster were used as network nodes to construct an association network. The edge weights between network nodes were determined based on the dynamic association strength, and the edge directions were determined based on the transcription direction and sequence end overlap characteristics.

[0091] In other words, the dynamic correlation strength is directly used as the edge weight, while the edge direction needs further analysis.

[0092] Furthermore, in some embodiments of the present invention, the method for obtaining the edge direction between network nodes includes: determining that the DNA sequence is transcribed from the 5' end to the 3' end; taking any two adjacent network nodes as a first node and a second node; when the 3' end of the first node points to the 5' end of the second node, determining the edge direction as from the first node to the second node; and determining the edge direction as from the first node to the second node when there is base pair overlap between the 3' end of the first node and the 5' end of the second node, and the overlap length is greater than a preset length threshold.

[0093] In this embodiment of the invention, it is common knowledge that DNA sequences are transcribed from the 5' end to the 3' end. Two adjacent network nodes on the transcription strand are taken as the first node and the second node, with the first node being gene fragment X and the second node being gene fragment Y. For specific examples: if the 3' end of gene fragment X points to the 5' end of gene fragment Y, then the edge direction is X→Y; if the 3' end of gene fragment X overlaps with the 5' end of gene fragment Y (the preset length threshold can be specifically, for example, 20bp, such as the last 20bp of X matching the first 20bp of Y), then the edge direction is X→Y. Thus, edge direction analysis is achieved.

[0094] After determining the network nodes, the edge weights and directions between different network nodes, an interconnected network can be constructed.

[0095] S104: Based on the dynamic association strength in the association network, differential assembly paths are used to assemble functionally associated repetitive sequences, non-functionally associated repetitive sequences, and adjacent non-repetitive sequences to obtain complete genome assembly results.

[0096] Because functionally related repetitive sequences and non-functionally related repetitive sequences behave differently, different assembly approaches are needed for sequence splicing, i.e., differentiated assembly paths are used for sequence splicing.

[0097] A method for assembling functionally associated repetitive sequences to obtain a complete genome assembly includes: assembling the functionally associated repetitive sequence with non-repetitive sequences from a non-repetitive sequence cluster adjacent to the transcriptional sequence to form a functionally anchored core fragment; using the functionally anchored core fragment as an anchor point, extracting all functionally associated repetitive sequences directly connected to that anchor point from the association network, sorting them from high to low according to dynamic association strength, and preferentially selecting the functionally associated repetitive sequence with the highest strength as the first extension target; and determining whether the repetitive base units in the functionally associated repetitive sequence follow the DNA sequence transfer pattern. The genome is assembled by first determining the forward alignment during recording, then extending the sequence in the direction of aligning the anchor 3' end with the 5' end of the functionally associated repeat sequence. If the repeating base unit is in a forward repeating arrangement, it extends in the direction of aligning the anchor 3' end with the 5' end of the functionally associated repeat sequence. If it is in a reverse repeating arrangement, the base matching probability between the reverse complementary sequence of the repeating base unit and the anchor is verified. When the matching probability exceeds a preset matching threshold, it extends in the direction of aligning the anchor 3' end with the reverse complementary sequence of the repeat sequence. After the extension of the first functionally associated repeat sequence is completed, the anchor plus the first repeat sequence is used as a new temporary anchor, and the extension continues with the new temporary anchor to achieve genome assembly.

[0098] Among them, based on the edge formed by the two functionally associated repeat sequences corresponding to the highest dynamic association strength value in the association network, the functionally associated repeat sequences are spliced ​​with adjacent non-repetitive sequences to form a functional anchoring core fragment. The functional anchoring core fragment is directly associated with the key genes for the synthesis of medicinal components and is the functional core of the splicing.

[0099] Using the functional anchor core segment as a fixed starting point, extract all functionally related repetitive sequences directly connected to the anchor point from the association network, sort them from high to low according to the dynamic association strength, and prioritize the functionally related repetitive sequence with the highest dynamic association strength as the first extension object.

[0100] Based on whether the repetitive base units in functionally related repetitive sequences are arranged in the forward direction during DNA transcription, the extension direction is determined and the sequences are assembled.

[0101] It should be noted that in a DNA sequence, the 5' end is the head and the 3' end is the tail. If the sequence is arranged in the forward direction from 5' end to 3' end, the extension order must follow the direction of the 5' end of the anchor point docking with the functionally related repeat sequence (consistent with the side direction).

[0102] If it is an inverted repeat, the base matching probability between the inverted complementary sequence of the repeating base unit and the anchor point must be verified first. Inverted complementarity can be understood as "reversed and mirrored". For example, if the base of the repeating base unit is "ATCG", its inverted sequence is "GCTA". The complementary sequence of the inverted sequence is CG and AT. Therefore, the inverted complementary sequence is "CGAT". Here, the repeating base unit is a base segment of a functionally related repeating sequence.

[0103] Calculate the base matching probability between the reverse complementary sequence of the repeating base unit and the anchor point. When the probability is ≥80% (preset matching threshold), the verification is successful. Then extend the reverse complementary sequence of the 3' end of the functionally related repeating sequence according to the 3' end of the anchor point.

[0104] Secondly, after completing the extension of the first functionally related repeat sequence, the integrated fragment of "anchor point + first functionally related repeat sequence" is used as a new temporary anchor point. The next level of functionally related repeat sequences connected to it are selected again from the association network. They are also sorted according to the dynamic association strength. Combining the periodicity of their repeat units with the end overlap characteristics of the temporary anchor point, that is, calculating the base matching probability between the last 15 bp at the 3' end of the temporary anchor point and the first 15 bp at the 5' end of the candidate functionally related repeat sequence, as long as the base matching probability is ≥70% (which can be gradually reduced according to actual needs), the candidate functionally related repeat sequence is determined as the next splicing order. Similarly, each time it is extended, the end base sequence of the current integrated fragment is recorded as the matching benchmark for the next round of extension.

[0105] In some embodiments of the present invention, the association information of the integrated fragment after each extension with the medicinal components in the metabolomics data can be verified. If the arrangement of the repetitive sequences in the integrated fragment is positively correlated with the content of the corresponding medicinal components, the extension result is retained; otherwise, the extension direction and length are adjusted according to the association network to achieve dynamic adaptation.

[0106] Furthermore, in some embodiments of the present invention, a method for assembling non-functionally associated repetitive sequences to obtain a complete genome assembly result includes: screening non-repetitive sequences from multiple non-repetitive sequence clusters that have high association strength with non-functionally associated repetitive sequences in an association network; assembling the non-repetitive sequences with non-functionally associated repetitive sequences respectively as candidate assembly fragments, wherein high association strength is defined as dynamic association strength greater than a preset association threshold; calculating the sequence consistency and network association consistency of each candidate assembly fragment, wherein sequence consistency is the normalized value of the average base matching probability between any two nodes in the candidate assembly fragment, and network association consistency is the normalized value of the coefficient of variation of the dynamic association strength between nodes in the candidate assembly fragment; calculating the difference between sequence consistency and network association consistency as a judgment index; and selecting the candidate assembly fragment with the largest judgment index for assembly.

[0107] The preset association threshold is a threshold value for the dynamic association strength, and in this embodiment of the invention, it can be specifically 0.5, for example.

[0108] It is understandable that candidate splicing fragments are assembly schemes formed by temporary connections of multiple network nodes in an associative network. These may include network nodes with non-functionally associated repeating sequences and multiple network nodes with non-repetitive sequences. Sequence consistency represents the normalized value of the average base-matching probability between all pairs of nodes in the candidate splicing fragments; a larger value indicates a higher base-matching probability and higher consistency. Network association consistency is the normalized value of the coefficient of variation of the dynamic association strength between nodes in the candidate splicing fragments. The specific calculation of the coefficient of variation is common knowledge; it represents the dispersion of the probability distribution. A larger value indicates a more dispersed distribution among all nodes in the candidate splicing fragments. Therefore, the difference between sequence consistency and network association consistency is calculated as a judgment index; the candidate splicing fragment scheme with the highest judgment index is selected for splicing.

[0109] This invention obtains target sequencing sequences and, based on alignment with known functional sequence libraries, screens sequences that influence function. Then, it performs sequence clustering and feature analysis to classify the target sequencing sequences into repetitive clusters, resulting in non-repetitive sequence clusters and repetitive sequence clusters. These repetitive sequence clusters are further functionally classified into functional repetitive sequences and non-functional repetitive sequences. Next, it determines the dynamic association strength of functionally influential sequences by combining features such as the expression abundance of functional genes and the content of medicinal components, and constructs an association network. Finally, based on the dynamic association strength in the association network, it uses differentiated assembly paths to assemble functionally associated repetitive sequences, non-functionally associated repetitive sequences, and adjacent non-repetitive sequences, obtaining a complete genome assembly result. This invention, by establishing a function-oriented repetitive sequence classification system and a dynamic association network optimization mechanism, effectively solves the splicing errors, fragment breaks, and redundancy problems caused by repetitive sequences in high-throughput sequencing of medicinal plant genomes, significantly improving assembly accuracy and completeness.

[0110] On the other hand, a medicinal plant genome assembly system based on high-throughput sequencing is also provided. The system includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any of the methods described above.

[0111] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0112] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A method for assembling the genome of medicinal plants based on high-throughput sequencing, characterized in that, The method includes: Obtain the target sequencing sequence, and screen for functionally influential sequences based on the alignment of the target sequencing sequence with a library of known functional sequences; By using sequence clustering and target sequencing sequence feature analysis, the target sequencing sequences are classified into repetitive sequences to obtain non-repetitive sequence clusters and repetitive sequence clusters. The repetitive sequence clusters are then compared with functionally affected sequences to determine functional and non-functional repetitive sequences. By combining the expression abundance of functional genes, the content of medicinal components, and the base overlap characteristics with adjacent sequences, the dynamic association strength of functionally influential sequences was determined. Non-repetitive sequences, functional repetitive sequences, and non-functional repetitive sequences in the non-repetitive sequence cluster were used as network nodes to construct an association network. The edge weights between network nodes were determined based on the dynamic association strength, and the edge directions were determined based on the transcription direction and the sequence end overlap characteristics. Based on the dynamic association strength in the association network, differential assembly paths are used to assemble functionally associated repetitive sequences, non-functionally associated repetitive sequences, and adjacent non-repetitive sequences to obtain complete genome assembly results. Methods for obtaining non-repetitive sequence clusters and repetitive sequence clusters include: All target sequencing sequences are globally paired and compared to calculate the base matching probability. Multiple sequence clusters are formed through agglomerative hierarchical clustering. The copy number and average matching probability of each sequence cluster are calculated. Sequence clusters with both copy number and base matching probability higher than the average are selected as repetitive sequence clusters, and the others are selected as non-repetitive sequence clusters. The copy number is the number of sequences contained in each sequence cluster. Methods for determining the strength of dynamic correlation include: For functionally affected sequences, transcriptomic and metabolomic data were obtained; the expression abundance of the corresponding functional genes of the functionally affected sequences was extracted from the transcriptomic data, and the average value was calculated as a functional weight parameter; the content of the target medicinal component was extracted from the metabolomic data, and the Pearson correlation coefficient between it and the expression abundance of the functional genes was calculated and normalized as a functional confidence parameter. For functionally related repeating sequences, the product of the functional weight parameter and the functional confidence parameter is calculated and normalized to obtain the dynamic association strength. For non-functionally associated repetitive sequences, the preliminary genome assembly results of medicinal plants are divided into a predetermined number of continuous gene fragments of equal length; the distribution information entropy of the non-functionally associated repetitive sequences in different gene fragments is calculated, and the information entropy is normalized and used as a correction parameter; the product of the correction parameter, the functional weight parameter, and the functional confidence parameter is calculated as the dynamic association strength of the non-functionally associated repetitive sequences. Differentiated assembly paths include: Functionally associated repetitive sequences are spliced ​​together with non-repetitive sequences in clusters of non-repetitive sequences adjacent to the transcriptional sequence to form functionally anchored core fragments; For functionally related repetitive sequences, using the functionally anchored core segment as the anchor point, all functionally related repetitive sequences directly connected to the anchor point are extracted from the association network, sorted from high to low according to the dynamic association strength, and the functionally related repetitive sequence with the highest strength is selected as the first extension object. Based on whether the repetitive base units in the functionally related repeat sequence are arranged in the forward direction during DNA transcription, the extension direction is determined and the assembly is performed. If the repetitive base units are arranged in the forward repeat direction, the extension is performed in the direction of aligning the 3' end of the anchor point with the 5' end of the functionally related repeat sequence. If the repetitive base units are arranged in the reverse repeat direction, the base matching probability between the reverse complementary sequence of the repetitive base unit and the anchor point is verified. When the matching probability exceeds the preset matching threshold, the extension is performed in the direction of aligning the 3' end of the anchor point with the reverse complementary sequence at the 3' end of the repeat sequence. After the first functionally associated repeat sequence is extended, the anchor point plus the first repeat sequence is used as a new temporary anchor point, and the extension continues with the new temporary anchor point to achieve genome assembly. For non-functional repetitive sequences, non-repetitive sequences from multiple non-repetitive sequence clusters that have high correlation strength with non-functional repetitive sequences are selected from the association network. The non-repetitive sequences are then concatenated with the non-functional repetitive sequences to form candidate concatenation segments. The high correlation strength is defined as a dynamic correlation strength greater than a preset correlation threshold. Calculate the sequence consistency and network association consistency for each candidate splice segment, where the sequence consistency is the normalized value of the average base matching probability between any two nodes in the candidate splice segment, and the network association consistency is the normalized value of the coefficient of variation of the dynamic association strength between nodes in the candidate splice segment. Calculate the difference between sequence consistency and network association consistency as a judgment index; select the candidate splicing segment with the largest judgment index for splicing.

2. The method for assembling the genome of a medicinal plant based on high-throughput sequencing as described in claim 1, characterized in that, The process of screening functionally influential sequences based on the alignment of the target sequencing sequence with a library of known functional sequences includes: The high-quality, effective sequence set is compared with the functional sequence library, and sequencing sequences that have a continuous, predetermined length of overlap with the functional elements of the functional sequence library are selected as functionally influential sequences.

3. The method for assembling the genome of a medicinal plant based on high-throughput sequencing as described in claim 1, characterized in that, Clusters of repetitive sequences are aligned with functionally affected sequences to identify functionally repetitive sequences and non-functionally repetitive sequences, including: Calculate the overlap probability between each target sequencing sequence and each functionally affected sequence within the repetitive sequence cluster; Based on the overlap probability, all target sequencing sequences within a repetitive sequence cluster are classified into functionally associated repetitive sequences and non-functionally associated repetitive sequences.

4. The method for assembling the genome of a medicinal plant based on high-throughput sequencing as described in claim 1, characterized in that, The method for obtaining the edge directions between network nodes includes: Determine the DNA sequence from the 5' end to the 3' end during transcription; Use any two adjacent network nodes as the first node and the second node. When the 3' end of the first node points to the 5' end of the second node, the edge direction is determined to be from the first node to the second node; There is a base pair overlap at the 3' end of the first node and the 5' end of the second node, and the overlap length is greater than the preset length threshold. The edge direction is determined to be from the first node to the second node.

5. A medicinal plant genome assembly system based on high-throughput sequencing, the system comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method, device and terminal for detecting genome variations

    CN109074429A

  • Grape seed development gene VvPHERES1 and application thereof

    CN119685345A