Construction Method and Application of an Identification Model for Extrachromosomal Circular DNA in Plants
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-08-14
AI Technical Summary
当前植物eccDNA鉴定主要依托高通量测序结合序列比对算法实现,以 split read、discordant read 片段比对为核心(如 CircleMap 软件),但植物基因组普遍富含大量重复序列,极易造成测序片段错配,导致鉴定结果假阳性偏高,精准筛选困难
本发明突破了传统序列比对工具易受植物重复序列干扰、通用深度学习模型无法适配环状 DNA 拓扑结构、现有工具功能单一等多项技术瓶颈,在一方面定制严苛的数据筛选与预处理规则,结合对称截断策略完整保留成环断点信息,从源头降低假阳性;另一方面创新提出高斯分布注意力惩罚机制,为环状 DNA 引入结构先验,显式建模首尾相连的拓扑特征,精准捕获断点信号;同时采用卷积融合 Transformer 的定制化网络架构,兼顾序列局部特征与长程依赖,模型训练稳定、泛化能力强。本发明训练得到的鉴定模型不仅可高精度鉴定小麦、玉米等复杂植物基因组中的 eccDNA,框架还可拓展至环状 RNA 等其他环状核酸检测,具有良好的泛化能力。
Smart Images

Figure CN122436003B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, and in particular to a method for constructing and applying an identification model for extrachromosomal circular DNA in plants. Background Technology
[0002] Extrachromosomal circular DNA (eccDNA) is widely distributed in plant genomes and participates in regulating genome plasticity, environmental stress response, and gene expression diversity, making it an important area of research in plant genome variation. Current plant eccDNA identification mainly relies on high-throughput sequencing combined with sequence alignment algorithms, focusing on split read and discrete read alignment (such as the CircleMap software). However, plant genomes are generally rich in repetitive sequences, which easily lead to sequencing fragment mismatches, resulting in a high false positive rate and difficulty in accurate screening. Existing intelligent deep learning models such as DNABERT and DeepECD are mostly developed based on human genome data and have not been designed for the end-to-end topological structure of circular DNA, lacking targeted capture of breakpoint connectivity features, resulting in insufficient adaptability to complex plant genomes such as wheat and maize.
[0003] In addition, existing identification tools are limited to single-molecule recognition of eccDNA and cannot be compatible with circular RNA identification, resulting in poor generalization prediction performance across species and plant tissues.
[0004] In summary, the lack of a universal detection framework that relies on sequence information, can explicitly model circular structures, and can recognize multiple types of circular nucleic acids restricts the large-scale research of plant circular nucleic acids. Summary of the Invention
[0005] The purpose of this invention is to provide a method for constructing and applying an identification model for extrachromosomal circular DNA in plants. This model provides a basis for large-scale research on plant circular nucleic acids, allowing for the identification of circular structures based on sequence information and accommodating the recognition of multiple types of circular nucleic acids. It can accurately identify extrachromosomal circular DNA in plants.
[0006] To achieve the above objectives, this technical solution provides a method for constructing an identification model for extrachromosomal circular DNA in plants, comprising the following steps: Obtain plant eccDNA positive sequences labeled with binary classification tags and non-circular CDS negative sequences of the same species as training sequences; The training sequence is input into the identification framework and trained until the conditions are met to obtain the identification model. The identification framework includes a data processing layer, a CAM encoder module, a feedforward network layer, a linear layer and an activation layer connected in sequence. Each CAM encoder module includes a one-dimensional convolutional layer, a multi-head self-attention mechanism and a Gaussian distribution attention mechanism. Residual connections and layer normalization layers are introduced between each CAM encoder module. The input sequence is fed into the data processing layer and converted into a labeled sequence based on the word segmentation strategy. The labeled sequence is then subjected to feature encoding to generate a sequence embedding vector that incorporates effective positional information. The sequence embedding vector is fed into the CAM encoder module for encoding to obtain encoded features. The encoded features are then fed into the feedforward network layer to obtain mapped features. The mapped features are then fed into the linear layer and the activation layer to output the prediction results.
[0007] Secondly, this scheme provides a method for identifying circular DNA outside plant chromosomes, including the following steps: Obtain the test sequence to be identified; The sequence to be tested is input into the identification model as the input sequence, and the output prediction result is either eccDNA or a linear genome sequence.
[0008] Compared with existing technologies, this technical solution has the following characteristics and beneficial effects: This invention overcomes several technical bottlenecks of traditional sequence alignment tools, such as susceptibility to interference from repetitive plant sequences, the inability of general deep learning models to adapt to the topology of circular DNA, and the limited functionality of existing tools. On one hand, it customizes stringent data screening and preprocessing rules, combining a symmetric truncation strategy to fully preserve circular breakpoint information, reducing false positives from the source. On the other hand, it innovatively proposes a Gaussian distribution attention penalty mechanism, introducing structural priors to circular DNA, explicitly modeling the topological features of connected ends, and accurately capturing breakpoint signals. Simultaneously, it employs a customized network architecture combining convolutional and Transformer layers, balancing local sequence features and long-range dependencies, resulting in stable model training and strong generalization ability. The identification model trained by this invention can not only accurately identify eccDNA in the genomes of complex plants such as wheat and maize, but the framework can also be extended to the detection of other circular nucleic acids such as circular RNA, demonstrating excellent generalization ability. Attached Figure Description
[0009] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a framework diagram of a plant chromosome extracyclic circular DNA identification model according to embodiments of this application.
[0010] Figure 2 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0011] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0012] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this specification in other embodiments. In some other embodiments, the methods may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; and multiple steps described in this specification may be combined into a single step in other embodiments.
[0013] Example 1 This approach addresses the unique characteristics of extrachromosomal circular DNA (ECC) in plants by constructing a sequence-based identification model with explicit topological modeling capabilities for accurate identification. The model incorporates a CAM encoder to capture the head-to-tail dependencies within the ECC sequence and enhances end-breakpoint related signals through a Gaussian-guided attention penalty mechanism. Furthermore, this model is generalizable to the identification of various plant ECC sequences, providing a high-precision and highly generalizable research framework for circular nucleic acid (ECC) recognition and establishing a sequence-driven research paradigm for analyzing eccDNA and related structural variations in complex plant genomes.
[0014] Specifically, this scheme provides a method for constructing an identification model of extrachromosomal circular DNA in plants, including the following steps: Obtain plant eccDNA positive sequences labeled with binary classification tags and non-circular CDS negative sequences of the same species as training sequences; The training sequence is input into the identification framework and trained until the conditions are met to obtain the identification model. The identification framework includes a data processing layer, a CAM encoder module, a feedforward network layer, a linear layer and an activation layer connected in sequence. Each CAM encoder module includes a one-dimensional convolutional layer, a multi-head self-attention mechanism and a Gaussian distribution attention mechanism. Residual connections and layer normalization layers are introduced between each CAM encoder module. The input sequence is fed into the data processing layer and converted into a labeled sequence based on the word segmentation strategy. The labeled sequence is then subjected to feature encoding to generate a sequence embedding vector that incorporates effective positional information. The sequence embedding vector is fed into the CAM encoder module for encoding to obtain encoded features. The encoded features are then fed into the feedforward network layer to obtain mapped features. The mapped features are then fed into the linear layer and the activation layer to output the prediction results.
[0015] In order to construct a high-confidence training dataset, the plant eccDNA positive sequences in this scheme need to be strictly screened and filtered.
[0016] Specifically, this method obtains high-throughput sequencing data of plant eccDNA enriched by exonuclease digestion and rolling circle amplification from public databases. This approach allows for selective enrichment of circular molecules and effective removal of linear genomic DNA. CircleMap software is used to annotate candidate eccDNA sites in the high-throughput sequencing data, and positive plant eccDNA sequences are screened from the high-throughput sequencing data according to the following screening rules: Only candidate sites with a CircleMap confidence score of 5 or higher are retained, while sequences containing more than five ambiguous bases (N) are removed. Candidate sequences must be between 50 and 10,000 bp in length and have at least one split read and one discordant read to ensure the robustness of the circularization evidence.
[0017] Specifically, in this scheme, the CircleMap confidence score of candidate eccDNA sites in plant eccDNA positive sequences is not less than 5, the number of ambiguous bases (N) does not exceed 5, the sequence length is between 50 and 10000 bp, and it has at least one break read and an abnormal distance read.
[0018] In some embodiments, simple repetitive sequences in plant eccDNA positive sequences are removed, thereby eliminating sequencing alignment errors caused by genomic repetitive regions and reducing false positives in the model.
[0019] In addition, this scheme selects plant eccDNA positive sequences as positive samples and non-circular CDS negative sequences of the same species as negative samples. These regions are transcribed and translated in the chromosomal background and do not form circular structures, thus ensuring the consistency of the genomic background with the positive samples.
[0020] The framework diagram of the identification framework of this scheme is as follows: Figure 1As shown, the identification framework includes a data processing layer, a multi-layer stacked CAM encoder module, a feedforward network layer, a linear layer, and an activation layer connected in sequence. This scheme introduces a deep learning algorithm to replace the traditional sequence alignment, which allows the trained identification model to autonomously learn the higher-order biological patterns in plant eccDNA positive sequences, avoiding the defects of traditional sequence alignment, such as susceptibility to interference from genomic repetitive sequences and high false positive rate.
[0021] Regarding the data processing layer of this identification framework: The data processing layer performs input embedding and position embedding on the input sequence to obtain a sequence embedding vector with positional encoding information, similar to "words" in natural language. The advantage of this is that the CAM encoder module can be used to capture sequence patterns and long-range dependencies.
[0022] The input sequence is fed into the data processing layer. The overlapping k-mers with a set length of k and a configurable sliding step size step are divided into tokens. A k-mer vocabulary is constructed based on nucleotide combinations. The token is then mapped to a unique index based on the k-mer vocabulary to obtain the labeled sequence.
[0023] In some embodiments, a basic k-mer vocabulary is constructed based on the permutations and combinations of four nucleotides: A, G, C, and T. In addition, four special markers, [START], , [PAD], and [UNK], are added to the k-mer vocabulary.
[0024] It should be noted that [START] and are used to mark the beginning and end boundaries of the input sequence, matching the breakpoint characteristics at both ends of circular DNA, making it easier to capture key sites for circularization; [PAD] is used to complete input sequences of varying lengths; and [UNK] is used to match unrecognized unknown bases and mutated bases.
[0025] In some embodiments, when the configurable sliding step is greater than 1, segments with a length less than k at the end of the input sequence are discarded. That is, when step > 1, k-mers are truncated from the beginning of the input sequence at intervals step. If there are fewer than k bases remaining at the end, no token is generated, and the segments are discarded directly.
[0026] In some embodiments, the length k is set to 6 by default, and the configurable sliding step is set to 1 by default.
[0027] In addition, in order to preserve key information while processing variable-length sequences, this scheme adopts a symmetrical truncation strategy to ensure the preservation of key breakpoint information of eccDNA. The circularity features of eccDNA are concentrated at the beginning and end of the sequence. Ordinary unilateral truncation is prone to deleting the circularity sites at the beginning / end, while the symmetrical truncation of this scheme ensures that the base information at both ends of the sequence is completely preserved, which is adapted to the subsequent capture of the beginning and end connection features (circular topology).
[0028] In some embodiments, a maximum token length threshold is set. If the total number of tokens to be divided exceeds the maximum token length threshold, a symmetrical truncation strategy is adopted to equally truncate tokens from the head and tail of the input sequence and discard the middle segment.
[0029] Furthermore, half of the maximum length threshold of the tokens located at the beginning of the input sequence is retained, and half of the maximum length threshold of the tokens located at the end of the input sequence is retained.
[0030] In some embodiments, the maximum token length threshold is calculated based on the maximum sequence length and a set length.
[0031] For example, for the input sequence S = s1, s2, ..., s... n The corresponding label sequence can be represented as: ; in Let i represent the token sequence, k represent the i-th token, step represent the configurable sliding step size, and n represent the input length in the input sequence.
[0032] If the maximum sequence length is set to L max Then the maximum token length threshold is: ; Among them This is the maximum length threshold for the token. k is the maximum sequence length, where k is the length. If the total number of tokens to be divided exceeds the maximum token length threshold, only the tokens at the head position will be retained. and located at the tail position The token.
[0033] After obtaining the marker sequence, this scheme also requires the application of learnable positional encoding to reflect the bidirectional structural characteristics of plant eccDNA positive sequences. In some embodiments, tokens in a token sequence participate in location encoding calculations based on their positions, and a learnable embedding layer is used to generate a sequence embedding vector that incorporates valid location information.
[0034] In some embodiments, the token marked [PAD] does not participate in the positional encoding calculation, and a Boolean mask matrix is constructed to distinguish between valid positions and padding positions in the sequence embedding vector. In other words, the token with the [PAD] occupier will skip the positional encoding calculation and will be set to an invalid bit in the Boolean mask matrix.
[0035] In some embodiments, the sequence embedding vector is a 256-dimensional dense vector representation.
[0036] Regarding the CAM encoder module of this identification framework: The core architecture of this identification framework consists of multiple stacked CAM encoder modules that integrate convolutional operations. These CAM encoder modules enhance the identification framework's ability to focus on key sequence regions, and residual connections and layer normalization are introduced between the CAM encoder modules to improve training stability.
[0037] Of particular note is that this scheme introduces a Gaussian distribution attention mechanism in the CAM encoder module to incorporate structural prior information in the identification of plant eccDNA positive sequences, thereby explicitly guiding the identification framework to focus on the two ends of the input sequence.
[0038] eccDNA is a circular DNA with its first and last bases linked together to form a circle. After sequencing, it is broken down into a linear sequence. The circular breakpoints fall in the head and tail regions of this linear sequence. Since the determination of positive plant eccDNA sequences depends on the breakpoint signals, and these signals are usually located at both ends of the linearized sequence, there is a clear biological basis for strengthening the end positions.
[0039] In some embodiments, each CAM encoder module includes a one-dimensional convolutional layer, a multi-head attention mechanism, and a Gaussian distribution attention mechanism connected in sequence. The original attention features and original attention scores output by the multi-head self-attention mechanism are fed into the Gaussian distribution attention mechanism to obtain encoded features.
[0040] Specifically, the Gaussian distribution attention mechanism constructs a Gaussian function penalty term based on the sequence position of each token in the original attention features. The Gaussian function penalty has a symmetrical decay distribution centered on the current token position, which is used to apply position-related suppression weights to distant tokens. The constructed Gaussian function penalty term is then weighted element-wise with the original attention score to apply local prior constraints to the attention distribution, thereby obtaining a position-modulated attention weight matrix. Finally, a weighted aggregation operation is performed on the original attention features based on this attention weight matrix to output an encoded feature representation with local structure awareness.
[0041] That is, the Gaussian distribution attention mechanism maintains the symmetry at both ends of the sequence while tilting the attention distribution towards the end region, thereby enhancing the identification model's ability to perceive key breakpoint regions.
[0042] Furthermore, the formula for calculating the attention weights of the Gaussian distribution attention mechanism in this scheme is as follows: ; ; in The term represents the Gaussian penalty term, where i represents the normalized distance from position i to the nearest end of the sequence (range [0,1][0,1][0,1]). Controlling the diffusion range of the Gaussian distribution, Here, the penalty intensity hyperparameter is fixed, and L is the sequence length. The attention score is the result of Gaussian penalty, i.e., the attention weight matrix modulated by position. The original attention score. These are the original attention features.
[0043] As can be seen from the calculation formula, the Gaussian function penalty term is based on a Gaussian function, which is determined by the normalized distance from the sequence position to the nearest end.
[0044] Regarding the feedforward network layer of this scheme: The feedforward network layer of this scheme contains a global pooling layer and a linear layer. The encoded features are input into the global pooling layer and compressed into a fixed-length representation vector. The representation vector is then input into the linear layer and mapped to obtain the mapped features.
[0045] In some embodiments, the features input to the feedforward network layer and the features output by the feedforward network layer are subjected to residual connection and normalization.
[0046] Regarding the linear layers and activation layers in this scheme: In some embodiments, the classification head outputs prediction results and confidence scores for the linear layer and the activation layer, wherein the prediction results are eccDNA or non-eccDNA, and the confidence scores are probability distributions generated by the Softmax function.
[0047] In addition, the identification framework training in this scheme uses the binary cross-entropy loss function as the loss constraint, and constructs the loss value based on the deviation between the true label and the predicted result output by the model; the optimizer uses the AdamW optimizer to dynamically update the network parameters, and at the same time, it is combined with weight decay regularization and Dropout regularization strategies to suppress overfitting; the training process sets an early stopping mechanism to monitor the validation set loss, and terminates the iteration when the validation set loss does not decrease for several consecutive rounds, thus saving the optimal model weights.
[0048] To comprehensively evaluate the identification performance of the model trained using this method, multiple standard metrics were employed, including accuracy, recall, precision, F1 score, and area under the ROC curve (AUC). Accuracy measures the overall prediction correctness; recall and precision assess the ability to identify true positives and predict reliability, respectively; the F1 score is the harmonic mean of the two; and AUC is used to evaluate the model's overall discriminative ability through the true positive rate versus false positive rate curves. These metrics collectively provide a quantitative basis for evaluating the effectiveness of the identification model in eccDNA breakpoint identification and regional prediction tasks.
[0049] 1. Model performance evaluation and comparative analysis: To comprehensively evaluate the performance of the identification model in this scheme, this scheme systematically compares the CircFormer-Rice model trained in this study (i.e., the identification model trained above) with several state-of-the-art models such as DNABert, DNAGemma, DNAGPT, and the standard Transformer.
[0050] All pre-trained models used in this study were derived from PDLLM. In the eccDNA classification task, this approach introduces a binary classification head above the output layer of each pre-trained model to distinguish between eccDNA and linear genomic sequences. Differentiated sequence-level feature extraction strategies are employed for different types of pre-trained models: for models based on the Masked Language Model (MLM), the latent representation corresponding to the first marker ([CLS]) is extracted as the overall sequence representation; while for the Causal Language Model (CLM), the latent representation of the sequence end marker is used as the global feature. The extracted sequence-level features are then input into a fully connected layer, and the probability of belonging to eccDNA is output through a sigmoid activation function. During the comparison of models, the backbone and classification head parameters of the pre-trained models are jointly optimized. The models are trained on the training set, and performance is monitored on the validation set to prevent overfitting. Simultaneously, a 1:1 balanced sampling of positive and negative samples is performed in each training batch. Finally, the fine-tuned models are evaluated on an independent test set to measure their classification performance.
[0051] The results, shown in Table 1, demonstrate that CircFormer achieved the best performance in classification across all four metrics: accuracy, precision, recall, and F1 score. Its accuracy and F1 score both reached 0.961, significantly outperforming all benchmark models. Further evaluation of model performance using ROC curves and AUC values revealed that CircFormer achieved the highest AUC value of 0.989. These results indicate that CircFormer maintains an optimal balance between sensitivity and specificity across a wide range of decision thresholds. This performance improvement is primarily attributed to the circular attention mechanism. This module explicitly models circular dependencies in the eccDNA sequence and captures long-range interactions within the circular topology.
[0052] Table 1 Classification performance of different models .
[0053] 2. Assessment of cross-organizational and cross-species generalization ability: This study systematically evaluated the generalization ability of CircFormer through cross-organizational and cross-species prediction tasks.
[0054] In the cross-tissue setting, this approach directly applies the CircFormer-Rice model, trained on rice leaf data, to the identification of eccDNA in multiple organs, including reproductive organs, panicles, leaf sheaths, leaves, roots, and stems. The results, shown in Table 2, demonstrate consistently stable and excellent performance across all tissues with minimal fluctuations. The accuracy ranges from 0.959 to 0.966, and the F1 score ranges from 0.954 to 0.962.
[0055] For cross-species assessment, this approach employed a zero-sample transfer strategy. The model trained on rice was directly applied to seven representative plant species covering both dicotyledons and monocotyledons, including Arabidopsis thaliana, rapeseed, upland cotton, soybean, tomato, common wheat, and maize. The results, shown in Table 3, demonstrate that CircFormer maintained high predictive performance across species, with accuracies ranging from 0.935 to 0.972 and F1 scores ranging from 0.936 to 0.972. ROC curve analysis further confirmed these findings. CircFormer's AUC values exceeded 0.980 in all tested species. These results highlight the model's excellent cross-species generalization ability.
[0056] Table 2. CircFormer-Rice's performance across organizations. .
[0057] Table 3. CircFormer-Rice's performance across species. .
[0058] 3. Generalization ability on circular RNA sequences: Although circular RNA and eccDNA differ in their biological mechanisms and cellular functions, they exhibit similarities at the sequence level. The closed-loop topology induces consistent local sequence patterns around the junction sites, a characteristic that provides a theoretical basis for unified modeling of circularization signals. This study further evaluated the cross-task generalization ability of CircFormer-Rice, directly applying it to a circular RNA and linear RNA classification task across nine plant species. All circular RNA datasets were derived from published plant transcriptome studies (collected from the PlantcircBase database) and reprocessed using standardized bioinformatics procedures to ensure consistency and comparability.
[0059] CircFormer demonstrated consistently strong performance on the RNA task, as shown in Table 4. Accuracy ranged from 0.848 to 0.945 across nine species, with F1 scores ranging from 0.837 to 0.943. Despite not undergoing RNA-specific training, the model achieved an AUC of 0.969 on the RNA task, a small difference compared to 0.994 on the DNA task. This small performance gap indicates that the circular features learned by the model are highly shared between DNA and RNA. ROC analysis further evaluated the model's cross-species discrimination performance on the RNA task. AUC values remained high across all species, ranging from 0.899 to 0.976.
[0060] Table 4. Performance of CircFormer-Rice in sequencing circular RNA
[0061] In summary, CircFormer not only excels in cross-species eccDNA prediction but also effectively transfers to circular RNA recognition without requiring task-specific training. The circular attention mechanism facilitates the extraction of biologically relevant circularization patterns at the sequence level, supports robust generalization across species and molecular types, and provides a unified computational framework for the systematic analysis of circular nucleic acids in plant genomes and transcriptomes.
[0062] Example 2 This scheme provides a method for identifying circular DNA outside plant chromosomes, including the following steps: Obtain the test sequence to be identified; The sequence to be tested is input into the identification model constructed as in Example 1, and the predicted result is output, where the predicted result is eccDNA or a linear genome sequence.
[0063] The technical content in Embodiment 2 that is the same as that in Embodiment 1 will not be repeated.
[0064] Example 3 This embodiment also provides an electronic device, see reference. Figure 2 It includes a memory 402 and a processor 401, the memory 402 storing a computer program, and the processor 401 being configured to run the computer program to perform the steps in any of the above-described methods for constructing a model for identifying plant extrachromosomal circular DNA or methods for identifying plant extrachromosomal circular DNA.
[0065] Specifically, the processor 401 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0066] The memory 402 may include a large-capacity memory 402 for data or instructions. The memory 402 can be used to store or cache various data files that need to be processed and / or communicated, as well as possible computer program instructions executed by the processor 401.
[0067] The processor 401 reads and executes computer program instructions stored in the memory 402 to implement any of the methods for constructing a cross-condition health index prediction model for wind turbine gearbox components or a cross-condition health index prediction method for wind turbine gearbox components in the above embodiments.
[0068] Optionally, the electronic device may further include a transmission device 403 and an input / output device 404, wherein the transmission device 403 is connected to the processor 401 and the input / output device 404 is connected to the processor 401.
[0069] The transmission device 403 can be used to receive or send data via a network. Specific examples of the network described above may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 403 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0070] Input / output device 404 is used to input or output information. In this embodiment, the input information may be a sequence, etc., and the output information may be a prediction result, etc.
[0071] Optionally, in this embodiment, the processor 401 can be configured to perform the following steps via a computer program: Obtain plant eccDNA positive sequences labeled with binary classification tags and non-circular CDS negative sequences of the same species as training sequences; The training sequence is input into the identification framework and trained until the conditions are met to obtain the identification model. The identification framework includes a data processing layer, a CAM encoder module, a feedforward network layer, a linear layer and an activation layer connected in sequence. Each CAM encoder module includes a one-dimensional convolutional layer, a multi-head self-attention mechanism and a Gaussian distribution attention mechanism. Residual connections and layer normalization layers are introduced between each CAM encoder module. The input sequence is fed into the data processing layer and converted into a labeled sequence based on the word segmentation strategy. The labeled sequence is then subjected to feature encoding to generate a sequence embedding vector that incorporates effective positional information. The sequence embedding vector is fed into the CAM encoder module for encoding to obtain encoded features. The encoded features are then fed into the feedforward network layer to obtain mapped features. The mapped features are then fed into the linear layer and the activation layer to output the prediction results.
[0072] or: Obtain the test sequence to be identified; The sequence to be tested is input into the identification model constructed as in Example 1, and the predicted result is output, where the predicted result is eccDNA or a linear genome sequence.
[0073] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.
[0074] Generally, various embodiments can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of the invention can be implemented in hardware, while others can be implemented by firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, by way of non-limiting example, these blocks, apparatuses, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0075] Embodiments of the present invention can be implemented by computer software, which may be executable by a data processor of a mobile device, such as a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. A computer program product may include one or more computer-executable components configured to perform embodiments when the program is run. One or more computer-executable components may be at least one piece of software code or a portion thereof. Additionally, it should be noted that any block in the logical flow of the figures may represent a program step, or interconnected logical circuitry, blocks and functions, or a combination of program steps and logical circuitry, blocks and functions. The software may be stored on physical media such as memory chips or blocks of storage implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs, etc. The physical medium is a non-transient medium.
[0076] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0077] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for constructing an identification model for extrachromosomal circular DNA in plants, characterized in that, Includes the following steps: Obtain plant eccDNA positive sequences labeled with binary classification tags and non-circular CDS negative sequences of the same species as training sequences; The training sequence is input into the identification framework and trained until the conditions are met to obtain the identification model. The identification framework includes a data processing layer, a CAM encoder module, a feedforward network layer, a linear layer and an activation layer connected in sequence. Each CAM encoder module includes a one-dimensional convolutional layer, a multi-head self-attention mechanism and a Gaussian distribution attention mechanism. Residual connections and layer normalization layers are introduced between each CAM encoder module. The input sequence is input into the data processing layer and divided into tokens by overlapping k-mers with a set length of k and a configurable sliding step size step. A maximum token length threshold is set. If the total number of tokens exceeds the maximum token length threshold, a symmetrical truncation strategy is used to equally truncate tokens from the head and tail of the input sequence and discard the middle segment. A k-mer vocabulary is constructed based on nucleotide combinations. Each token is mapped to a unique index based on the k-mer vocabulary to obtain a labeled sequence. The labeled sequence is feature-encoded to generate a sequence embedding vector that incorporates effective positional information. The sequence embedding vector is input into the CAM encoder module for encoding to obtain encoded features. The encoded features are input into the feedforward network layer to obtain mapped features. The mapped features are input into the linear layer and the activation layer to output the prediction result. The formula for calculating the attention weights in the Gaussian distribution attention mechanism is as follows: ; ; in The term represents the Gaussian penalty term, where i represents the normalized distance from position i to the nearest end of the sequence. Controlling the diffusion range of the Gaussian distribution, Here, the penalty intensity hyperparameter is fixed, and L is the sequence length. This is the position-modulated attention weight matrix. The original attention score. These are the original attention features.
2. The method for constructing an identification model for extrachromosomal circular DNA in plants according to claim 1, characterized in that, The original attention features and original attention scores output by the multi-head self-attention mechanism are fed into the Gaussian distribution attention mechanism. The Gaussian distribution attention mechanism constructs a Gaussian function penalty term based on the sequence position of each token in the original attention features. The constructed Gaussian function penalty term is then weighted element-wise with the original attention score to obtain a position-modulated attention weight matrix. Based on this attention weight matrix, a weighted aggregation operation is performed on the original attention features to output an encoded feature representation with local structure awareness.
3. The method for constructing an identification model for extrachromosomal circular DNA in plants according to claim 1, characterized in that, In plant eccDNA positive sequences, the CircleMap confidence score of candidate eccDNA sites should be no less than 5, the number of ambiguous bases should not exceed 5, the sequence length should be between 50 and 10000 bp, and there should be at least one break read and an abnormal distance read.
4. The method for constructing an identification model for extrachromosomal circular DNA in plants according to claim 1, characterized in that, A basic k-mer vocabulary is constructed based on the permutations and combinations of four nucleotides: A, G, C, and T. Additionally, four special markers, [START], 5. The method for constructing an identification model for extrachromosomal circular DNA in plants according to claim 1, characterized in that, , [PAD], and [UNK], are added to the k-mer vocabulary.
6. A method for identifying circular DNA outside plant chromosomes, characterized in that, When the configurable sliding step is greater than 1, discard segments in the input sequence whose length is less than k at the end. Includes the following steps: Obtain the test sequence to be identified; The test sequence is input as an input sequence into the identification model constructed by the method of constructing an identification model of plant extrachromosomal circular DNA as described in any one of claims 1 to 5, and the output prediction result is an eccDNA or a linear genome sequence.
7. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method for constructing an identification model of plant extrachromosomal circular DNA as described in any one of claims 1 to 5.
Citation Information
Patent Citations
ScATAC-seq-based ecDNA structure prediction method, method for identifying cells carrying ecDNA and medium
CN120748478A
Methods of amplifying and sequencing nucleic acids
US20060040297A1