Phage dna sequence generation method and apparatus based on hierarchical modeling

By employing a collaborative modeling approach combining hierarchical modeling and layered generation models, the problem of existing technologies failing to balance global dependencies and local contextual information in phage DNA sequence generation is solved, resulting in the generation of high-quality, complete phage DNA sequences.

CN122337327APending Publication Date: 2026-07-03SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN UNIV
Filing Date
2026-03-23
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing deep learning-based phage DNA sequence generation methods struggle to simultaneously consider global dependencies and local contextual information, resulting in poor quality of the generated phage DNA sequences.

Method used

A hierarchical modeling approach is adopted to convert phage DNA sequences into element sequences and map them into digital sequences. A hierarchical generation model is used for collaborative modeling, including a global processing layer and a local processing layer, to capture the overall structural patterns and local element arrangement characteristics of phage DNA. The complete phage DNA sequence is then generated through a preset mapping relationship.

Benefits of technology

It achieves high-quality generation of phage DNA sequences, balancing global structural consistency and local sequence rationality, to meet practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122337327A_ABST
    Figure CN122337327A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of bioinformatics, and discloses a bacteriophage DNA sequence generation method and device based on hierarchical modeling, which comprises the following steps: obtaining an elementized characterization result of a bacteriophage DNA sequence to obtain an element sequence composed of elements arranged in sequence; converting the element sequence into a digital sequence composed of digital identifiers based on a preset mapping relationship; inputting the digital sequence into a trained hierarchical generation model, cooperatively modeling global dependency relationships and local context information in the digital sequence through the hierarchical generation model, and generating a target digital sequence; decoding the target digital sequence into a corresponding target element sequence based on the preset mapping relationship, and splicing each element in the target element sequence one by one according to a preset correspondence relationship between elements and DNA fragments to obtain a complete bacteriophage DNA sequence. The application can effectively generate a complete bacteriophage DNA sequence meeting characteristics, and improve the generation quality and integrity of the bacteriophage DNA sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics technology, specifically to a method and apparatus for generating bacteriophage DNA sequences based on hierarchical modeling. Background Technology

[0002] Artificial generation of phage DNA (Deoxyribonucleic acid) sequences has important applications in phage modification, antibacterial agent development, and other fields. With the development of deep learning technology, deep learning-based modeling methods have become the mainstream means of DNA sequence generation, which can complete various tasks such as sequence completion, local editing, and de novo generation. However, the task of de novo generation of long phage DNA sequences requires the model to accurately characterize the overall features and local details of the sequence, which places higher demands on sequence modeling capabilities.

[0003] Currently, existing deep learning-based DNA sequence generation methods struggle to simultaneously model global dependencies and local contextual information when generating long phage DNA sequences de novo. They fail to effectively capture both the overall correlation features between distant positions within the long sequence and the contextual features of local elements, resulting in phage DNA sequences that do not meet practical application requirements in terms of structural consistency and overall quality. In summary, existing DNA sequence generation methods cannot effectively characterize both global dependencies and local contextual information in the de novo generation of long phage DNA sequences, leading to poor quality phage DNA sequences.

[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention

[0005] This application provides a method and apparatus for generating phage DNA sequences based on hierarchical modeling, which can effectively generate complete phage DNA sequences that meet the characteristics, thereby improving the generation quality and integrity of phage DNA sequences.

[0006] In a first aspect, embodiments of this application provide a method for generating phage DNA sequences based on hierarchical modeling, including: The elemental characterization results of the phage DNA sequence were obtained, resulting in an element sequence composed of elements arranged in sequence; Based on a preset mapping relationship, the component sequence is converted into a number sequence composed of digital identifiers; The number sequence is input into a trained hierarchical generation model, which collaboratively models the global dependencies and local contextual information in the number sequence to generate the target number sequence; wherein, the hierarchical generation model includes a global processing layer and a local processing layer; Based on the preset mapping relationship, the target digital sequence is decoded into the corresponding target element sequence, and according to the preset correspondence between elements and DNA fragments, each element in the target element sequence is spliced ​​together one by one to obtain a complete phage DNA sequence.

[0007] Secondly, embodiments of this application provide a phage DNA sequence generation device based on hierarchical modeling, comprising: The acquisition module is used to acquire the elemental characterization results of the phage DNA sequence, and obtain the element sequence composed of elements arranged in sequence; A conversion module is used to convert the element sequence into a number sequence composed of numeric identifiers based on a preset mapping relationship; A generation module is used to input the number sequence into a trained hierarchical generation model, and to generate a target number sequence by co-modeling the global dependencies and local context information in the number sequence through the hierarchical generation model; wherein, the hierarchical generation model includes a global processing layer and a local processing layer; The splicing module is used to decode the target digital sequence into the corresponding target element sequence based on the preset mapping relationship, and splice each element in the target element sequence one by one according to the preset correspondence between elements and DNA fragments to obtain a complete phage DNA sequence.

[0008] This application provides a method and apparatus for generating phage DNA sequences based on hierarchical modeling. First, by converting the phage DNA sequence into element sequences and further mapping them into digital sequences, a suitable modeling input format is provided for the hierarchical generation model. This allows the model to perform targeted modeling based on the element-level features of the phage DNA, ensuring the accuracy and adaptability of the modeling process. Then, a hierarchical generation model containing global and local processing layers is adopted to collaboratively model the global dependencies and local contextual information of the digital sequences. This allows the model to simultaneously capture the overall structural regularity of the phage DNA sequence and the correlation features of the local element arrangement. The generated target digital sequence not only conforms to the global structural logic of the phage DNA sequence but also ensures the rationality of the local element sequences, avoiding the problem of missing sequence structure or local details caused by a single modeling dimension. Finally, the target digital sequence is decoded into element sequences through a preset mapping relationship, and splicing is completed according to the correspondence between elements and DNA fragments. This achieves accurate conversion from element-level modeling results to complete phage DNA sequences. The final generated phage DNA sequence can balance global structural consistency and local sequence rationality. As can be seen, this application achieves high-quality generation of phage DNA sequences by co-modeling the global dependencies and local contextual information of the phage DNA sequence through a hierarchical generation model. This effectively adapts to the actual needs of phage DNA sequence generation and solves the problem that existing technologies cannot simultaneously and effectively characterize the global dependencies and local contextual information of the sequence, resulting in poor quality of the generated phage DNA sequences. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is an application environment diagram of the phage DNA sequence generation method based on hierarchical modeling provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the phage DNA sequence generation method based on hierarchical modeling provided in this application embodiment; Figure 3 This is a schematic diagram of the hierarchical sequence modeling architecture provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the phage DNA sequence generation device based on hierarchical modeling provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0011] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of systems and methods consistent with those detailed in the appended claims or with some aspects of this application.

[0012] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover descriptions such as non-exclusive inclusion, so that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.

[0013] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0014] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.

[0015] To address the aforementioned technical problems and overcome the shortcomings of existing technologies, this application provides a method and apparatus for generating phage DNA sequences based on hierarchical modeling, which can effectively generate complete phage DNA sequences that meet the characteristics, thereby improving the quality and integrity of the generated phage DNA sequences.

[0016] Figure 1 This is a diagram illustrating the application environment of a hierarchical modeling-based phage DNA sequence generation method in one embodiment. (Refer to...) Figure 1This hierarchical modeling-based phage DNA sequence generation method is applied to a hierarchical modeling-based phage DNA sequence generation system. The system includes a terminal 110 and a server 120. The terminal 110 and server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal; a mobile terminal can be at least one of a mobile phone, tablet, or laptop. The server 120 can be a standalone server or a server cluster consisting of multiple servers. Server 120 is configured to execute the above-mentioned hierarchical modeling-based phage DNA sequence generation method, including: obtaining the elemental characterization results of the phage DNA sequence to obtain an element sequence composed of elements arranged in sequence; converting the element sequence into a digital sequence composed of digital identifiers based on a preset mapping relationship; inputting the digital sequence into a trained hierarchical generation model, and generating a target digital sequence by co-modeling the global dependencies and local context information in the digital sequence through the hierarchical generation model; wherein, the hierarchical generation model includes a global processing layer and a local processing layer; decoding the target digital sequence into the corresponding target element sequence based on the preset mapping relationship, and splicing each element in the target element sequence one by one according to the preset correspondence between elements and DNA fragments to obtain a complete phage DNA sequence.

[0017] Please see Figure 2 , Figure 2 This is a flowchart illustrating a hierarchical modeling-based phage DNA sequence generation method according to an embodiment of this application. This embodiment primarily uses the application of this hierarchical modeling-based phage DNA sequence generation method to a computer device as an example for illustration. Specifically, the hierarchical modeling-based phage DNA sequence generation method provided in this application may include the following steps: S1. Obtain the elemental characterization results of the phage DNA sequence to obtain the element sequence composed of elements arranged in sequence; Specifically, in step S1, the complete phage DNA sequence is first decomposed into multiple independent elements according to its inherent structural features and functional attributes using elemental characterization. Each element corresponds to a specific functional or structural fragment in the original phage DNA sequence and is a basic building block of the phage DNA sequence. Then, according to the natural order of these elements in the original phage DNA sequence, the decomposed elements are arranged in an orderly manner to form an element sequence that completely preserves the arrangement logic of the elements in the original phage DNA sequence. For example, after elemental characterization, a certain phage DNA sequence is decomposed into three elements: a coding element X, a non-coding element Y, and a repeating element Z. These three elements exist in the original DNA sequence in the order X, Y, Z, and the final element sequence is [X, Y, Z].

[0018] S2. Based on a preset mapping relationship, convert the component sequence into a number sequence composed of numeric identifiers; Specifically, for step S2, a unified and fixed mapping relationship between elements and numerical identifiers is first established in advance. Each different element obtained after the elemental characterization of the phage DNA sequence is assigned a unique corresponding numerical identifier. This mapping relationship provides a unified basis for the bidirectional conversion between elements and numerical identifiers. After the mapping relationship is established, each element is replaced sequentially with its corresponding numerical identifier in the preset mapping relationship, according to the order of the elements in the element sequence. The replacement process strictly follows the original order of the elements, ultimately forming a numerical sequence composed of a string of numerical identifiers in a specific order, thus realizing the digital conversion of the phage DNA element sequence. For example, if the mapping relationship is preset as follows: coding element X corresponds to the number 1, non-coding element Y corresponds to the number 2, and repeating element Z corresponds to the number 3, then the element sequence [X,Y,Z] will be replaced sequentially, ultimately converting it into the numerical sequence [1,2,3].

[0019] S3. Input the digit sequence into the trained hierarchical generation model, and generate the target digit sequence by co-modeling the global dependencies and local context information in the digit sequence; wherein, the hierarchical generation model includes a global processing layer and a local processing layer; Specifically, for step S3, the hierarchical generation model is a model trained on phage DNA sequence-related digital sequence samples, possessing the ability to model phage DNA sequence features. It comprises two functional layers: a global processing layer and a local processing layer, which cooperate to achieve collaborative modeling. After the converted digital sequence is input into the trained hierarchical generation model, the global processing layer performs an overall analysis of the digital sequence, capturing and modeling the relationships between distant positional identifiers, i.e., global dependencies, thereby grasping the overall structural patterns of the phage DNA sequence. Simultaneously, the local processing layer focuses on identifiers at adjacent or nearby positions in the digital sequence, mining and modeling the relationships between them, i.e., local contextual information, thereby accurately depicting the local detailed features of the phage DNA sequence. The modeling processes of the global and local processing layers are performed synchronously and collaboratively, achieving a fusion learning of global and local features of the digital sequence. Finally, based on this collaborative modeling result, a target digital sequence conforming to the inherent characteristics and arrangement patterns of the phage DNA sequence is generated. For example, when the number sequence [1,2,3] is input into a trained hierarchical generation model, the model generates a target number sequence [1,3,2,5] that conforms to the arrangement logic of phage DNA elements through collaborative modeling of global dependencies and local contexts, where the number 5 is a unique number identifier corresponding to another type of phage DNA element.

[0020] S4. Based on the preset mapping relationship, the target digital sequence is decoded into the corresponding target element sequence, and according to the preset correspondence between the elements and DNA fragments, each element in the target element sequence is spliced ​​together one by one to obtain the complete phage DNA sequence; Specifically, step S4 consists of two core stages: decoding and splicing, both following pre-defined rules. In the decoding stage, the established mapping relationship between elements and numerical identifiers is used to reverse-convert the generated target numerical sequence. Following the order of the numerical identifiers in the target sequence, each identifier is restored to its corresponding phage DNA element, ultimately forming a target element sequence arranged in a specific order. The mapping relationship used for decoding is completely consistent with the conversion mapping relationship, ensuring the accuracy of element restoration. In the splicing stage, based on the pre-defined one-to-one correspondence between elements and DNA fragments, where each phage DNA element uniquely associates with a specific DNA fragment sequence, the DNA fragments corresponding to each element are sequentially joined end-to-end according to the order of the elements in the target element sequence. The splicing process strictly follows the original order of the elements, ultimately integrating multiple independent DNA fragments into a complete phage DNA sequence. For example, the target numerical sequence [1,3,2,5] is decoded to obtain the target element sequence [X,Z,Y,W]. The pre-defined correspondence between the elements and DNA fragments is X corresponds to SeqX, Z corresponds to SeqZ, Y corresponds to SeqY, and W corresponds to SeqW. Then, by sequentially splicing SeqX, SeqZ, SeqY, and SeqW, the complete phage DNA sequence SeqX-SeqZ-SeqY-SeqW is obtained.

[0021] This embodiment converts the phage DNA sequence into an element sequence and further digitizes it. It then uses a hierarchical generation model with global and local processing layers to collaboratively model the global dependencies and local contextual information of the sequence. After decoding and element splicing, a complete DNA sequence is obtained. This approach takes into account both the global structural features and local detailed features of the phage DNA sequence, effectively improving the generation quality of the phage DNA sequence and enabling the efficient generation of complete phage DNA sequences that conform to the inherent characteristics of phages.

[0022] Furthermore, in some embodiments, step S2, "converting the element sequence into a number sequence composed of numeric identifiers based on a preset mapping relationship," may specifically include: S21. Count all element types appearing in the element sequence of all samples, and construct a mapping table between elements and identifiers. The mapping table contains the correspondence between filler markers and preset numeric identifiers. Specifically, for step S21, firstly, the element sequences corresponding to all phage DNA sequence samples used for modeling are collected. All elements appearing in all element sequences are fully analyzed and statistically analyzed. Deduplication is performed according to the element type characteristics to determine all unique element types that have appeared, ensuring that the mapping table fully covers all element types in the samples. Then, a unique preset numerical identifier is assigned to each unique element type. Simultaneously, a padding marker is included in the mapping table, and a corresponding numerical identifier is assigned to it separately, ultimately forming a complete mapping table. This mapping table clarifies the fixed one-to-one correspondence between the padding marker, various element types, and their corresponding numerical identifiers, serving as the core basis for subsequent conversion between element sequences and numerical sequences. For example, after statistically analyzing the element sequences of all samples, four unique element types are obtained: coding element X, coding element Y, non-coding element M, and repeating element N. The padding marker is set to [PAD], and a numerical identifier 0 is assigned to it. Coding element X is assigned 1, coding element Y is assigned 2, non-coding element M is assigned 3, and repeating element N is assigned 4. Based on this, an element-numerical identifier mapping table containing all the above correspondences is constructed.

[0023] S22. For each element sequence, according to the mapping table, replace each element in each element sequence with its corresponding numeric identifier in turn to obtain the numeric sequence; Specifically, in step S22, each element sequence of a phage DNA sequence sample is processed independently. Following the natural order of the elements in the sequence, the mapping table constructed in step one by one is consulted to find the corresponding numerical identifier. The element is then directly replaced with the corresponding numerical identifier. During the replacement process, the order of the numerical identifiers is strictly maintained to be completely consistent with the order of the elements in the original sequence, without changing the overall arrangement logic of the sequence. After replacing all elements in a single element sequence, a numerical sequence corresponding one-to-one with the original element sequence is obtained. The element sequences of all samples are converted according to this rule. For example, if the element sequence of a sample is [coding element X, non-coding element M, coding element Y, repeating element N], the replacements are performed sequentially according to the mapping table: coding element X is replaced with 1, non-coding element M with 3, coding element Y with 2, and repeating element N with 4, ultimately resulting in the numerical sequence [1, 3, 2, 4].

[0024] S23. Store the mapping table between element identifiers for use in decoding during the sequence generation stage; Specifically, in step S23, the constructed and verified element-digit identifier mapping table is stored in a stable, readable, and non-loss-prone format. During storage, the integrity and accuracy of all correspondences within the mapping table are guaranteed, with no omissions or mismatches. This mapping table is specifically retained and serves as the sole basis for reverse decoding the target digital sequence generated by the model into the target element sequence during the subsequent phage DNA sequence generation stage. This ensures that the decoding process in the sequence generation stage uses the exact same element-digit identifier correspondence as the previous element sequence conversion process, avoiding decoding deviations caused by inconsistent correspondences. For example, the constructed mapping table can be saved as a standardized, readable file. When decoding the target digital sequence in the subsequent generation stage, this file can be directly retrieved, and the digital identifiers can be restored to their corresponding elements according to the correspondences within the file.

[0025] This embodiment constructs an element-identifier mapping table containing padding tags by statistically analyzing all sample element types. It sequentially completes the conversion from element sequence to number sequence and saves the mapping table. This establishes a unified standard for the bidirectional conversion between elements and number identifiers, ensuring the accuracy and consistency of the conversion process, and providing a reliable foundation for subsequent model building and sequence decoding.

[0026] Furthermore, in some embodiments, after step S2 "converting the element sequence into a number sequence composed of numeric identifiers based on a preset mapping relationship", the method further includes: S301. Obtain the preset target length value; Specifically, in step S301, a pre-set target length value is first obtained. This value is a fixed value determined by comprehensively considering factors such as the model input requirements for subsequent sequence modeling, computational processing efficiency, and the characteristic distribution of phage DNA sequences. It serves as a unified reference standard for length standardization of all digital sequences. All digital sequences to be processed must ultimately be processed to this target length to adapt to the batch input and unified modeling requirements of subsequent models. For example, if the preset target length value is 256 based on modeling requirements, all subsequent digital sequences must be processed into fixed-length sequences of 256.

[0027] S302. For a number sequence whose length is less than the target length value, add a numeric identifier corresponding to the padding mark to the end of the sequence so that the length of the number sequence reaches the target length value, thus obtaining a fixed-length number sequence; Specifically, in step S302, the actual length of the number sequence is first compared with the preset target length. If the actual length of the number sequence is less than the target length, it is padded. The padding method involves adding numeric identifiers corresponding to the padding markers sequentially to the end of the number sequence, maintaining the original element order during the addition process. This continues until the total length of the number sequence exactly reaches the preset target length. The resulting number sequence with the same length as the target length is a fixed-length number sequence. For example, if the preset target length is 256, the padding marker's numeric identifier is 0, and the actual length of a certain number sequence is 180, which is less than 256, then 76 consecutive 0s are added to the end of the number sequence to make the total length of the sequence reach 256, forming the corresponding fixed-length number sequence.

[0028] S303. For a number sequence whose length is greater than the target length value, extract a continuous subsequence of length equal to the target length value from the number sequence to obtain a fixed-length number sequence; wherein, the extraction is performed using a random starting point extraction method; Specifically, in step S303, the actual length of the number sequence is first compared with the preset target length. If the actual length of the number sequence is greater than the target length, it is truncated. The core requirement of truncation is to extract a continuous subsequence from the original number sequence whose length is exactly equal to the target length. The truncation uses a random starting point. First, the effective position range of the original number sequence that can be used as the starting point is determined. This range is based on the principle of ensuring that a continuous subsequence of the target length can be obtained after truncation. Then, a position is randomly selected from this effective position range as the truncation starting point. Starting from this starting point, continuous digit identifiers with a length equal to the target length are truncated. The resulting continuous subsequence is the fixed-length number sequence that meets the requirements. For example, if the preset target length is 256 and the actual length of a certain number sequence is 400, its effective truncation starting point range is from the 1st to the 144th position. The 58th position is randomly selected from this range as the starting point. Starting from the 58th position, 256 consecutive digit identifiers are truncated, finally resulting in a fixed-length number sequence of length 256.

[0029] This embodiment achieves standardized length for digital sequences of different lengths by performing fixed-length processing, padding short sequences with markers, and truncating long sequences to a fixed length subsequence from a random starting point. This perfectly adapts to the batch input requirements of subsequent models. At the same time, the core features of the original sequence are preserved during the processing, avoiding the problem of feature homogenization, and allowing the model to learn more comprehensive phage DNA sequence features.

[0030] Furthermore, in some embodiments, step S3, "inputting the digit sequence into a trained hierarchical generation model, and using the hierarchical generation model to collaboratively model the global dependencies and local contextual information in the digit sequence to generate the target digit sequence," may specifically include: S31. Divide the numerical sequence input to the hierarchical generation model into multiple consecutive sequence blocks; Specifically, in step S31, the number sequence to be modeled is first segmented continuously and without overlap according to a preset fixed block length. The core requirement of the segmentation is to ensure the continuity of the subsequences and that all segmented sequence blocks can completely cover the original number sequence without any missing digit identifiers. The number of sequence blocks is determined by the total length of the original number sequence and the preset block length. Through this segmentation method, the overall number sequence is decomposed into multiple independent sequence block units, providing a structured processing basis for subsequent global and local level modeling. For example, if the total length of the input number sequence is 256 and the preset fixed block length is 16, then the number sequence is sequentially segmented into 16 consecutive sequence blocks. Each sequence block contains 16 consecutive digit identifiers. The first sequence block contains the identifiers of bits 1-16 of the original sequence, the second sequence block contains the identifiers of bits 17-32, and so on, until the segmentation of the entire sequence is completed.

[0031] S32. Through the global processing layer, modeling is performed on a per-sequence-block basis to capture long-range dependencies between blocks and obtain a global context representation; Specifically, in step S32, the global processing layer treats each segmented sequence block as an independent, holistic modeling unit. Instead of analyzing individual numerical identifiers within a block, it focuses on the relationships and arrangement patterns between sequence blocks. For all sequence blocks, it mines and models the intrinsic connections between distant sequence blocks, i.e., the long-range dependencies between blocks. This process captures the overall structural features and global arrangement logic of the phage DNA sequence corresponding to the entire digital sequence. The global processing layer extracts, integrates, and characterizes these long-range dependency features, ultimately forming a global context representation that comprehensively reflects the global structural information of the original digital sequence. This representation reflects the overall element arrangement patterns and structural features of the sequence. For example, for the 16 segmented sequence blocks, the global processing layer analyzes the relationships between distant sequence blocks such as block 1 and block 8, block 3 and block 12, and block 6 and block 15, capturing the global arrangement features of elements in different regions of the phage DNA sequence and integrating these features into a vector-based global context representation.

[0032] S33. Through the local processing layer, the elements are modeled within each sequence block, capturing the local contextual relationships of the elements within the block to obtain a local detail representation; Specifically, in step S33, the local processing layer analyzes each independent sequence block without cross-block feature modeling, focusing on the numeric identifiers within a single sequence block. For the sequentially arranged numeric identifiers within each sequence block, it mines and models the inherent relationships and combination patterns between adjacent or similar identifiers, i.e., the local contextual relationships corresponding to elements within the block. This process accurately depicts the detailed features and element combination rules of each local region of the original numeric sequence. The local processing layer extracts and represents the local contextual features of each sequence block, then integrates the local feature results of all sequence blocks to ultimately form a local detail representation that fully reflects the detailed information of each local region of the original numeric sequence. This representation can reflect the local combination features of elements in each part of the sequence. For example, for each of the 16 sequence blocks mentioned above, the local relationships between the 1st and 2nd, 5th and 6th, 10th and 15th identifiers within the block are analyzed, capturing the element combination patterns of each local region. A corresponding local feature vector is generated for each block, and then all vectors are integrated to obtain the overall local detail representation.

[0033] S34. Based on global context representation and local detail representation, predict and generate the target number sequence bit by bit in an autoregressive manner; Specifically, in step S34, the global context representation and local detail representation are first fused to form a comprehensive feature representation that combines the global structural features and local detail features of the original digit sequence. This comprehensive representation is the core basis for subsequent sequence generation, ensuring that the generated sequence conforms to the overall arrangement logic and meets the requirements of local element combination. Based on this comprehensive feature representation, an autoregressive generation method is used to construct the target digit sequence. Starting from the first digit of the sequence, based on the features of the one or more digit identifiers already generated, the next digit identifier that conforms to the phage DNA sequence features is predicted digit by digit. The predicted identifier is appended to the currently generated sequence, and the next digit identifier is predicted based on the updated sequence. This process is iteratively advanced until the entire target digit sequence is generated. For example, using the fused comprehensive feature representation as a basis, the first digit identifier of the target digit sequence is predicted to be 1. Then, based on the first digit 1, the second digit is predicted to be 3. Next, based on the already generated 1 and 3, the third digit is predicted to be 2, and so on, generating the complete target digit sequence digit by digit.

[0034] This embodiment divides the digital sequence into blocks, with a global processing layer capturing long-range dependencies between blocks and a local processing layer capturing local contextual relationships within blocks. After fusing these two types of features, a target digital sequence is generated in an autoregressive manner. This achieves collaborative modeling of the global and local features of the digital sequence, ensuring that the generated target digital sequence conforms to both the global structural arrangement of the phage DNA sequence and the combination characteristics of local elements, thereby improving the accuracy and rationality of the target digital sequence generation.

[0035] Furthermore, in some embodiments, step S34, "predicting and generating the target number sequence digit by digit in an autoregressive manner," may specifically include: S341. In each generation process, obtain the probability distribution vector output by the hierarchical generation model. Each element in the probability distribution vector represents the probability value that the next numeric identifier is the corresponding candidate identifier. Specifically, in step S341, in each iteration of the autoregressive generation of the target digit sequence, the hierarchical generation model outputs a probability distribution vector containing all candidate digit identifiers based on the overall features of the currently generated digit sequence fragments. The dimension of this vector is consistent with the total number of candidate digit identifiers. Each element in the vector has a value between 0 and 1, corresponding to the probability that the candidate digit identifier associated with that position will be the next digit identifier to be generated. The sum of the probabilities of all elements in the vector is 1. This probability distribution vector is the core data basis for selecting the next digit identifier. For example, in a certain generation step, if the candidate digit identifiers are 0, 1, 2, and 3, and the model outputs a probability distribution vector of [0.05, 0.70, 0.15, 0.10], it means that the probability of the next digit identifier being 0 is 5%, 1 is 70%, 2 is 15%, and 3 is 10%.

[0036] S342. Based on a preset threshold parameter, select a set of candidate identifiers whose sum of probability values ​​reaches the threshold parameter from the probability distribution vector to form a truncated candidate set; Specifically, for step S342, all probability values ​​in the probability distribution vector are first sorted in descending order, while retaining the candidate numeric identifiers corresponding to each probability value. Then, starting from the first sorted probability value, the values ​​are accumulated sequentially until the sum of the accumulated probability values ​​reaches a preset threshold parameter. At this point, all candidate numeric identifiers corresponding to the probability values ​​involved in the accumulation are filtered out and form a truncated candidate set. Low-probability candidate numeric identifiers that did not participate in the accumulation are eliminated. This process eliminates candidate identifiers with extremely low probabilities through threshold filtering, reducing invalid sampling while ensuring that the identifiers in the truncated candidate set have high generation rationality. For example, if the preset threshold parameter is 0.9, the sorted probability values ​​and corresponding identifiers in the above example are 1 (0.70), 2 (0.15), 3 (0.10), and 0 (0.05). The accumulated 0.70 + 0.15 = 0.85 does not reach the threshold. After accumulating 0.10, the sum is 0.95 ≥ 0.9, so identifiers 1, 2, and 3 are selected to form the truncated candidate set.

[0037] S343. Based on preset temperature parameters, the probability values ​​in the truncated candidate set are scaled to obtain the adjusted probability distribution; Specifically, for step S343, firstly, the original probability values ​​corresponding to each candidate numeric identifier in the truncated candidate set are extracted. Then, these original probability values ​​are scaled according to a pre-set temperature parameter. The temperature parameter is used to adjust the steepness of the probability distribution: when the temperature parameter is greater than 1, it flattens the gap between probability values ​​and increases the randomness of sampling; when the temperature parameter is less than 1, it amplifies the proportion of high probability values ​​and reduces the randomness of sampling; when the temperature parameter is equal to 1, the probability values ​​remain unchanged. After the scaling operation is completed, all processed probability values ​​are normalized so that the sum of the adjusted probability values ​​is 1 again, finally obtaining the adjusted probability distribution adapted to subsequent sampling. For example, the original probability values ​​of the above truncated candidate set are 0.70, 0.15, and 0.10. If the preset temperature parameter is 1.0, the probability values ​​remain unchanged after scaling. After normalization, the adjusted probability distribution is [0.70 / 0.95, 0.15 / 0.95, 0.10 / 0.95]≈[0.7368, 0.1579, 0.1053]. If the temperature parameter is set to 0.8, the relative probability value of the high-probability identifier 1 will be further increased.

[0038] S344. Based on the adjusted probability distribution, the next digital identifier is obtained by sampling from the truncated candidate set using the Gumbel sampling method; Specifically, for step S344, after obtaining the adjusted probability distribution, the next digit identifier is selected from the truncated candidate set using Gumbel sampling. This sampling method assigns a corresponding Gumbel noise value to the adjusted probability of each identifier in the truncated candidate set, combines the probability magnitude with random noise to calculate the sampling score of each identifier, and finally selects the candidate digit identifier with the highest sampling score as the next digit identifier for the current generation step. Gumbel sampling, while retaining the core principle that high-probability identifiers are more likely to be selected, introduces random noise to achieve appropriate random sampling, avoiding the generation process from falling into the repetitive generation of a fixed sequence. For example, if the truncated candidate set corresponding to the adjusted probability distribution is 1, 2, and 3, after assigning Gumbel noise values ​​to the three and calculating the sampling score, if the sampling score of identifier 1 is significantly higher than that of 2 and 3, then identifier 1 is selected as the next digit identifier generated in the current step.

[0039] In this embodiment, at each step of the autoregressive generation, the probability distribution vector is obtained, the candidate set is truncated by thresholding, the probability is scaled by temperature parameters, and the next identifier is selected by Gumbel sampling. This ensures that the selection of digital identifiers at each step is both consistent with the characteristics of the phage DNA sequence and regular, while also achieving diversity through appropriate randomness. At the same time, low-probability invalid candidates are eliminated, improving sampling efficiency and effectively ensuring the stability, rationality and richness of the target digital sequence generation.

[0040] Furthermore, in some embodiments, step S34, "predicting and generating the target number sequence digit by digit in an autoregressive manner," further includes: S345. During the generation process, determine whether the type combination between the element corresponding to the currently generated digital identifier and the previously generated element conforms to the preset biological sequence constraint rule; the preset biological sequence constraint rule prohibits the consecutive generation of two non-coding elements. Specifically, for step S345, in each iterative step of the autoregressive generation of the target digit sequence, after the current digit identifier is generated, it is first converted into a corresponding actual element based on the preset mapping relationship between elements and digit identifiers, and the type of the element is determined to be either a coded element or a non-coded element. Then, the element corresponding to the previously generated and confirmed digit identifier is retrieved, and it is confirmed that it is also a coded element or a non-coded element. Subsequently, the type combination of these two consecutively generated elements is checked to determine whether they conform to the preset biological sequence constraint rules. The core requirement of these rules is that consecutive combinations of "non-coded element + non-coded element" are not allowed, and all other type combinations conform to the rule requirements. This judgment process is executed immediately after each digit identifier is generated and is a pre-verification step for whether the generation result is retained. For example, if the previously generated numeric identifier corresponds to the non-coded element nonconding_3905, and the currently generated numeric identifier corresponds to the non-coded element nonconding_17896, the two are consecutive non-coded element combinations, which is determined to be inconsistent with the preset constraint rules; if the previous element is the coded element CPB0094#YP_009902727.1#NCBI#337#1104, the current element is the non-coded element nonconding_3905, or both the previous and current elements are coded elements, then it is determined to be consistent with the preset constraint rules.

[0041] S346. If it does not meet the requirements, the generated result of the current step shall be corrected or resampled; Specifically, for step S346, if, after verification, it is determined that the type combination of the currently generated element and the previously generated element does not conform to the preset biological sequence constraint rules, the numerical identifier generated in the current step is not included in the already generated target numerical sequence, nor is the subsequent sequence generation step continued. Instead, a result correction or resampling mechanism is initiated for the current generation step to process the generation result. The core purpose of this operation is to discard generation results that do not conform to the biological sequence rules, and obtain numerical identifiers that conform to the preset constraint rules through correction or resampling. This ensures that the generated sequence fragments always follow the inherent element arrangement rules of the phage DNA sequence. After obtaining the current generation result that conforms to the rules, the subsequent bit-by-bit generation steps are then executed. For example, if the previous element is a non-coding element nonconding_3905 and the currently generated element is a non-coding element nonconding_17896, after determining that it does not conform to the rules, the numerical identifier corresponding to nonconding_17896 is immediately discarded, and the current step is corrected or resampled until a numerical identifier corresponding to a coding element is generated, satisfying the constraint rules, before continuing generation.

[0042] In this embodiment, during the generation of the target digital sequence, the method determines in real time whether the combination of consecutive generated elements conforms to the biological sequence constraint rules that prohibit consecutive non-coding elements. If the result does not conform, it corrects or resamples in a timely manner. This directly avoids element combinations that violate the inherent arrangement rules of phage DNA from the generation stage, making the element sequence corresponding to the generated digital sequence more biologically reasonable and laying the biological characteristic foundation for the subsequent generation of high-quality phage DNA sequences.

[0043] Furthermore, in some embodiments, step S346, "correcting or resampling the generation result of the current step," may specifically include: When it is determined that the type combination of the currently generated element and the previously generated element does not conform to the preset biological sequence constraint rules, the generation result of the current step is marked as invalid; Specifically, during the autoregressive generation of the target numerical sequence, if, upon verification, it is found that the combination of the element corresponding to the currently generated numerical identifier with the previously confirmed generated element violates the biological sequence constraint rule prohibiting the consecutive generation of two non-coding elements, the numerical identifier and its corresponding element generated in the current step are immediately marked as invalid. This marking operation excludes this generated result from the valid sequence; it is neither appended to the already generated sequence fragment nor used to advance subsequent sequence generation steps. This provides a clear execution premise for subsequent resampling operations, defining invalid generated results from the source and preventing results that do not conform to biological laws from interfering with sequence generation. For example, if the previous generated element is the non-coding element nonconding_3905, and the element generated in the current step is the non-coding element nonconding_17896, their combination does not conform to the constraint rule, and the numerical identifier corresponding to nonconding_17896 generated in the current step is immediately marked as invalid.

[0044] Resample the current step to select the next numeric identifier from the candidate identifier set; Specifically, after marking the current generation result as invalid, a resampling process is initiated for that generation step. Using the same candidate identifier set from the original generation step, a new numeric identifier is selected from this set according to predetermined sampling rules, serving as the new generation candidate for the current step. The resampled candidate identifier set is consistent with the original step, and the sampling rules remain unchanged, ensuring that the resampled result still conforms to the overall characteristics of the phage DNA sequence. Only the erroneous generation result of the current step is replaced, without altering the fundamental sampling logic of the entire generation process. For example, in the aforementioned invalidated step, the original candidate identifier set contains various numeric identifiers corresponding to coded and non-coded elements. The numeric identifier corresponding to the coded element CPB0094#YP_009902727.1#NCBI#337#1104 is resampled from this set and used as the new generation candidate for the current step.

[0045] Determine whether the type combination of the element obtained after resampling and the previously generated element conforms to the preset biological sequence constraint rules. If it does, accept the generation result and continue to execute the subsequent generation steps. If it does not, repeat the sampling and judgment steps until the preset resampling limit is reached. Specifically, after converting the resampled numerical identifier into the corresponding element, the biological sequence constraint rule verification is performed again to determine whether the type combination of the new element and the previously generated element complies with the requirement of prohibiting the consecutive generation of two non-coding elements. If the verification result is compliant, the resampled numerical identifier is determined as the valid generation result of the current step, appended to the already generated sequence fragment, and the generation step of the next numerical identifier continues based on the updated valid sequence. If the verification result is still non-compliant, the resampled result is not accepted, and the resampling operation is performed again for the current step. Then, the above rule judgment steps are repeated. This sampling and judgment loop continues until a result that complies with the constraint rule is generated, or a preset resampling limit is triggered. For example, after resampling to obtain a coded element, if the combination with the previous non-coded element complies with the rule, the result is accepted and the next element is generated. If resampling still yields the non-coded element nonconding_5680, and the combination with the previous non-coded element still does not comply with the rule, then resampling and rule judgment continue.

[0046] If a generation result that meets the preset biological sequence constraint rules cannot be obtained after reaching the maximum number of resampling attempts, the current candidate sequence is discarded and the generation process is restarted. Specifically, during the sampling and judgment loop, the number of resampling operations is counted in real time. When the count reaches the pre-set upper limit for resampling, if no element and corresponding digital identifier conforming to the biological sequence constraint rules are generated, it is determined that the candidate sequence currently being generated cannot meet the biological requirements within a reasonable number of iterations. Therefore, all segments of the candidate sequence already generated are discarded, the current generation process is terminated, and the generation process of the target digital sequence is restarted from the beginning. This avoids falling into an infinite loop due to a single generation error, ensuring the efficiency of the sequence generation process. For example, if the preset upper limit for resampling is 5 times, and after 5 sampling and judgments, the result is still a non-coding element that does not meet the constraint rules, then all currently generated sequence segments are discarded, and the autoregressive generation of the entire target digital sequence is restarted.

[0047] This embodiment establishes a complete processing flow for generated results that do not conform to biological sequence constraint rules, including marking invalid, resampling, loop judgment, and discarding and regenerating after an upper limit on the number of iterations. This forms a closed loop for biological constraint verification, ensuring that the generated target digital sequence strictly follows the element arrangement rules of phage DNA, while avoiding the efficiency loss of infinite resampling. It ensures the biological rationality of the sequence while taking into account the execution efficiency of the generation process.

[0048] Furthermore, in some embodiments, step S4, "based on a preset mapping relationship, decoding the target digital sequence into a corresponding target element sequence, and according to the preset correspondence between elements and DNA fragments, splicing each element in the target element sequence one by one to obtain a complete phage DNA sequence," may specifically include: S41. Obtain the preset mapping relationship and decode the target number sequence into the corresponding target element name sequence; Specifically, for step S41, the pre-established mapping relationship between components and numeric identifiers is retrieved first. This mapping relationship is the same fixed correspondence used when converting a component sequence to a numeric sequence, possessing unique bidirectional conversion. Following the sequential arrangement of the numeric identifiers in the target numeric sequence, each numeric identifier is reverse-converted, querying its corresponding component name in the pre-defined mapping relationship. The conversion process strictly adheres to the original order of the numeric identifiers, without altering the overall arrangement logic of the sequence. After completing the reverse conversion of all numeric identifiers in the target numeric sequence, a target component name sequence composed of component names in their corresponding order is obtained, achieving accurate decoding from a numeric sequence to a component name sequence. For example, in the preset mapping relationship, the number 1 corresponds to the coded element name CPB0094#YP_009902727.1#NCBI#337#1104, the number 2 corresponds to the non-coded element name nonconding_3905, and the number 3 corresponds to the repeating element name CPB0094#repeat#NCBI#1#63. If the target number sequence is [1,2,3], then the target element name sequence obtained after sequential decoding is [CPB0094#YP_009902727.1#NCBI#337#1104, nonconding_3905, CPB0094#repeat#NCBI#1#63].

[0049] S42. Query the preset element DNA fragment library to obtain the DNA fragment sequence corresponding to each target element name; Specifically, in step S42, a pre-constructed element DNA fragment library is retrieved. This library stores a one-to-one correspondence between the names of various bacteriophage DNA elements and their corresponding DNA base fragment sequences. Each valid element name can be matched with a unique DNA fragment sequence in the library, forming the core data foundation for DNA sequence assembly. For the decoded target element name sequence, the element names are precisely queried and matched in the element DNA fragment library one by one according to their order, extracting the DNA fragment sequence corresponding to each element name. The query process ensures an accurate correspondence between the element name and the DNA fragment sequence, preventing mismatches or omissions. For example, for the above target element name sequence, the fragment library is queried sequentially, finding the DNA fragment sequence SeqA corresponding to the coding element name, the DNA fragment sequence SeqB corresponding to the non-coding element name, and the DNA fragment sequence SeqC corresponding to the repeating element name, thus obtaining the ordered DNA fragment sequences [SeqA, SeqB, SeqC]. S43. Following the order of the target element name sequence, the obtained DNA fragment sequences are spliced ​​end to end to obtain the complete bacteriophage DNA sequence; Specifically, for step S43, the original order of the target element name sequence is strictly followed. Using this order as the core basis, all the retrieved DNA fragment sequences are continuously spliced ​​end-to-end. During splicing, the terminal base of the previous DNA fragment sequence is directly linked to the starting base of the next DNA fragment sequence, maintaining the integrity of each DNA fragment sequence without adding, deleting, or modifying any bases. This process is repeated for all DNA fragment sequences. After all fragments are spliced, a continuous and complete DNA base sequence is formed, which is the final complete phage DNA sequence. For example, the sequentially obtained DNA fragment sequences SeqA, SeqB, and SeqC are spliced ​​end-to-end, with the end of SeqA linked to the start of SeqB, and the end of SeqB linked to the start of SeqC, ultimately resulting in the complete phage DNA sequence SeqA-SeqB-SeqC.

[0050] This embodiment uses a unified preset mapping relationship to accurately decode the target digital sequence into an element name sequence. After obtaining the corresponding DNA fragment by querying the element DNA fragment library, the fragments are spliced ​​together sequentially, achieving a precise and orderly conversion from the digitized target digital sequence to the actual complete phage DNA sequence. This ensures the accuracy of the decoding and splicing process and successfully completes the generation of the actual DNA sequence from the digitized sequence characterization.

[0051] Furthermore, in some embodiments, step S43, "joining the obtained DNA fragment sequences end-to-end according to the order of the target element name sequence to obtain a complete bacteriophage DNA sequence," may specifically include: If the target element name corresponding to the DNA fragment cannot be found during the query process, the element corresponding to the target element name will be skipped, and subsequent elements will continue to be assembled.

[0052] Specifically, during the DNA fragment search and assembly process based on the target element name sequence, the order of the element names in the sequence is strictly followed. Each element name is sequentially selected and precisely matched in a pre-defined element DNA fragment library to obtain a unique DNA fragment sequence corresponding to each element name. Simultaneously, real-time result verification is performed for each search operation to determine whether a matching DNA fragment sequence can be found in the element DNA fragment library. The verification result serves as the direct basis for whether to perform the subsequent assembly operation. This search and verification process proceeds synchronously without changing the original sequence search order. For example, if the target element name sequence is [Element X, Element Y, Element Z, Element W], element X is searched first, and if a corresponding DNA fragment is found, element Y is searched, and a verification operation to check for a corresponding fragment is performed simultaneously.

[0053] When the query results for a target element name are validated and no matching DNA fragment sequence is found in the element DNA fragment library, no DNA fragment retrieval or ligation operations are performed on that element. The element name is directly excluded from the element sequence to be assembled and not included in the subsequent assembly process. This skip operation only applies to elements for which there is currently no corresponding DNA fragment; it does not change the overall arrangement order of the original target element name sequence, nor does it trigger an interruption in the assembly process. For example, if the validation in the above sequence finds that element Y has no corresponding DNA fragment, element Y is skipped directly without any assembly processing related to that element.

[0054] After skipping elements without corresponding DNA fragments, the overall splicing process remains continuous without interrupting subsequent operations. Following the original sequence of the target element names, the process continues sequentially from the next element name after the skipped element, querying and verifying in the element DNA fragment library. For elements with corresponding DNA fragments, splicing is performed normally according to the rule of concatenating the first and last elements, until all subsequent elements in the target element name sequence have been queried and spliced, ultimately forming a complete DNA sequence. For example, after skipping element Y, the process continues sequentially to query element Z. If a corresponding DNA fragment is found, splicing is performed. Then, element W is queried, and if a corresponding DNA fragment is found, splicing continues, completing the subsequent splicing process for the entire sequence.

[0055] In this embodiment, during the phage DNA fragment splicing query stage, a skip operation is performed on target elements for which no corresponding DNA fragment exists, and subsequent elements are spliced. This sets up practical fault-tolerant rules for the splicing process, effectively avoiding the problem of the entire splicing process being interrupted due to the abnormality of individual elements, ensuring the continuity and integrity of the splicing process, and making the phage DNA sequence generation process more robust and practically feasible.

[0056] To facilitate understanding of the hierarchical modeling-based phage DNA sequence generation method provided in this embodiment, another implementation of the hierarchical modeling-based phage DNA sequence generation method is also provided, applicable to de novo phage DNA long sequence generation tasks. This involves digitally mapping the preprocessed sequence element characterization, combining it with a hierarchical Transformer generation model containing global and local layers for sequence modeling, and then decoding and assembling the generated sequence element-level results into a complete phage DNA sequence. Specifically, the steps include: Step S101: Receive / acquire elementation characterization results. Quality screening and elementation results can be provided by existing methods or third parties. In practice, this approach primarily begins with the step of "acquiring existing element sequence inputs and using them for model generation." The DNA sequence samples are 850 KP phage DNA sequences (fa or fasta format) provided by BGI Genomics in collaboration with BGI Genomics. The elementation results include a list of elements corresponding to each sequence (in order), as well as information such as the element's name / ID, start and end positions, and length.

[0057] Step S102: Digital mapping of sequence elements.

[0058] A "component-digit ID" mapping table is constructed based on the preprocessed component set. This table is used to convert the component sequence of each sample into a digit ID sequence, which serves as the input format for model training and generation. First, the mapping table is constructed by statistically analyzing all components appearing in the training data (based on component name / identifier), and assigning a unique integer ID to each component. A padding marker [PAD] is set, and its ID is specified as 0; the remaining valid components are numbered sequentially starting from 1. For each sample, the component sequence is replaced one by one with the corresponding integer ID according to the mapping table, resulting in a digit ID sequence used for model training and generation.

[0059] Deduplication rules for component statistics: When performing component statistics, the "component name" is used as the unique key for deduplication. If a component with the same name appears multiple times in different samples, it is only counted once.

[0060] Component name / identifier definition specifications: According to the provided componentization result fa file, the basic form is that each component is represented by a string, and fields are separated by #; different components are separated by English commas to form a component sequence.

[0061] Elements are divided into coding elements and non-coding elements. Coding elements are named as follows: Sample ID#Protein / Gene ID#Source Library#Start Coordinates#End Coordinates, for example, CPB0094#YP_009902727.1#NCBI#337#1104, and CPB0094#tr|A0A385E8F2|A0A385E8F2_9CAUD#Uniprotkb#1174#1398. In these examples, CPB0094 is the name of the phage to which they belong, YP_009902727.1 and tr|A0A385E8F2|A0A385E8F2_9CAUD are their identifiers in the corresponding databases, NCBI and Uniprotkb represent their source libraries, and the last two numbers represent the start and end coordinates, i.e., the interval (bp coordinates) of the element in the original DNA sequence. Non-coded elements are named in the form of nonconding_number, such as nonconding_3905 and nonconding_17896, where the number is a unique identifier for the non-coded element. Repeating elements are represented by sample ID#repeat#source library#start#end, used to mark repeated segments in a sequence, such as CPB0094#repeat#NCBI#1#63.

[0062] The priority principle for assigning numeric IDs is as follows: [PAD] is always assigned 0; the remaining components are numbered starting from 1. The numbering order is determined by the traversal order during the construction of the mapping table (e.g., the order in which the component list file is read / lexicographical order / first appearance order are all acceptable), as long as the same mapping table is used for training and generation.

[0063] The batch processing data format requirements for digitization mapping are as follows: Each sample input is a "sequence of elements arranged in order" (e.g., E1, E2, ..., Ek). During batch processing, mapping is performed sample by sample, and the output is a corresponding integer ID sequence (e.g., id(E1), id(E2), ..., id(Ek)). The mapped results are used for subsequent length alignment and model training / generation. The mapping table is saved as element_id_map.json at runtime, and the same mapping table is directly read for decoding and concatenation during the generation phase.

[0064] The mapping table must be updated and maintained using the same "element-ID mapping table" for both training and generation, and this table must be saved / loaded along with the model weights to avoid decoding errors caused by inconsistent numbering. If new data is added subsequently, resulting in new elements, the mapping table must be re-statistically analyzed and rebuilt based on the expanded data, and the model must be retrained.

[0065] Step S103: Training sample construction and length alignment.

[0066] Training samples are constructed based on the component sequences obtained in step S102, and the samples are processed to a fixed length to adapt to the model requirements. The target sequence length for model input is set to L (256). When the component ID sequence length of a certain sequence sample is less than L, PAD is appended to the end of the sequence until the length reaches L; when the component ID sequence length of a certain sequence sample is greater than L, a continuous subsequence of length L is extracted from the sequence as training input. The training adopts an autoregressive training method, shifting the input sequence one position to the right as the prediction target, so that the model learns the conditional probability distribution of "predicting the next component based on the previous component".

[0067] The selection criteria for the target length L are as follows: Currently, L=256 is fixed in the implementation. This length is used to balance the uniformity of input length and controllable computation during batch training; it can cover the component sequence fragments of most samples for learning local and global relationships. L is not limited to 256; it is an adjustable hyperparameter. In practice, it can be adjusted according to GPU memory, training speed, and sequence length distribution, for example, using {128, 256, 512}. A larger L can cover a longer context, but the training overhead is higher; a smaller L is more computationally efficient, but the context information is shorter.

[0068] Sliding step size for truncating ultra-long sequences: The current implementation truncates ultra-long sequences into continuous subsequences of length L, using a random starting point (randomly selecting a window from the available starting points each time). Therefore, this implementation does not have a fixed sliding step size; this is equivalent to random sampling within the range of [1, len-L]. If a deterministic window is required, a fixed step size (e.g., step size = L or L / 2) can be set, and multiple windows can be extracted as training samples using a sliding window method.

[0069] The input / output alignment method for autoregressive training is as follows: For each ID sequence of length L, x=[x1,xL], the input sequence is [x1,xL-1}], and the prediction target is [x2,xL]. That is, the next element is predicted using the previous element. The training loss ignores the [PAD] position (ID=0) in the target sequence.

[0070] Processing ratio limitations for samples of different lengths: For samples with insufficient length, [PAD] padding is used to fill to a fixed length L; for excessively long samples, a window of length L is obtained by truncating. The training loss ignores the calculation of the [PAD] position, so no supervision signal is generated for the filling position. No additional filtering or sampling weight adjustment rules for the "effective element ratio threshold" are set. If further control of the filling ratio is required, a minimum effective length threshold or a bucket sampling strategy can be added in the optional implementation.

[0071] Sample equalization strategy: This embodiment does not employ any additional sample equalization strategy. It only performs iterative training on the training samples in the conventional batch method. If it is necessary to further control the proportion of samples of different lengths, strategies such as sliding window augmentation, filtering of excessively short samples, or bucket sampling can be introduced in optional embodiments.

[0072] Step S104: Hierarchical Transformer Generate Model Construction.

[0073] A hierarchical Transformer model for component sequence generation is constructed, employing a two-level hierarchical modeling structure: first, global-level modeling, then local-level modeling, to simultaneously characterize long-range dependencies and local compositional patterns. The model structure is as follows: 1. Input and Embedding: Input a sequence of element IDs of length L into the model, map it into a vector sequence through the Embedding layer, and add position encoding.

[0074] 2. Sequence Blocking (Hierarchical Division): Divide the element sequence of length L into several blocks according to a fixed block length for subsequent global and local layer processing.

[0075] 3. Global Transformer layer: Models sequences in blocks, learns the dependencies between blocks, and captures global structural information and long-range dependencies.

[0076] 4. Local Transformer layer: Models the component sequence within each block, learns the local context and composition patterns of adjacent components, and supplements local detail information.

[0077] 5. Output layer: Maps the model's hidden states to the probability distribution (softmax) of the next element ID, used for autoregressive generation of element sequences.

[0078] This hierarchical structure enables the model to utilize both global structural information and local sequence patterns on longer element sequences, thereby enhancing the modeling capabilities in de novo generation of long phage DNA sequences.

[0079] The global layer's embedding dimension is 512, and the local layer's embedding dimension is 256.

[0080] Position encoding type: Learnable position encoding is used, corresponding to global and local sequence positions respectively.

[0081] Fixed block length values ​​and selection criteria for sequence partitioning: The input sequence length L = 256. The example implementation uses fixed partitioning, with both global and local block lengths being 16, i.e., 256 = 16 × 16. This is to ensure that the global layer can cover a relatively long range while the local layer models local composition patterns within blocks, and the computational overhead is controllable.

[0082] Number of layers in the global / local Transformer: The example implementation has a total of 6 layers, with 3 layers in the global layer and 3 layers in the local layer.

[0083] Attention count: 4.

[0084] The hidden layer dimension is the same as the embedding dimension (global 512, local 256).

[0085] The feedforward layer multiplier ff_mult = 4 (the intermediate dimension of the feedforward layer is 4 times the hidden dimension). The hidden dimension of the global layer is 512, corresponding to a feedforward layer dimension of 2048; the hidden dimension of the local layer is 256, corresponding to a feedforward layer dimension of 1024.

[0086] The relevant sampling parameters for the output layer are as follows: The output layer uses a linear mapping + softmax to map the model's hidden states to the probability distribution of the next element ID. The example implementation in the generation stage uses threshold truncation (top-k threshold filtering) + Gumbel sampling to obtain the next element ID. k is calculated based on the filtering threshold `filter_thres`, and Gumbel sampling is performed after retaining the top-k candidates. The generation length is controlled by preset parameters. After generation, the [PAD] (ID=0) in the output sequence is filtered to obtain the final candidate element ID sequence.

[0087] like Figure 3 As shown, Figure 3 This is a schematic diagram of the hierarchical sequence modeling architecture provided in this embodiment. First, the input sequence is initially encoded using pre-trained embedding. Then, the encoded sequence is divided into multiple consecutive sequence blocks (Patch 1, Patch 2, Patch 3, i.e., patching operation). Next, a global Transformer performs global modeling on all sequence blocks, extracting and outputting global context information that reflects the overall structural regularity of the sequence. Then, this global context information is input into multiple local Transformers, and each local Transformer performs local feature modeling within the corresponding sequence block, accurately capturing the local contextual relationships of the sequence within the block. Finally, each local Transformer outputs the corresponding element identifier sequence, realizing the collaborative modeling of global dependencies and local contextual information of the sequence, providing core feature support for the subsequent generation of target sequences that conform to the characteristics of phage DNA sequences.

[0088] Step S105: Model training.

[0089] The fixed-length element ID sequence constructed in step S3 is input into the hierarchical Transformer generative model constructed in step S4 for training, enabling the model to learn the conditional probability relationships between elements. The training method is as follows: using the first L-1 element IDs of the input sequence as input, the model predicts the next L-1 element IDs (i.e., shifting the sequence one position to the right as the label). The cross-entropy loss function is used to calculate the error between the predicted distribution and the true label, and the position corresponding to [PAD] (ID=0) is set as an ignored term and does not participate in the loss calculation. The example implementation trains for 500 epochs, and the model weight file is periodically saved for subsequent generation stages.

[0090] Suppose a training sample has the following element ID sequence: Where PAD's ID = 0. The model outputs the predicted distribution of the next element's ID at position t: The training loss is the average cross-entropy of the next-token across all non-PAD locations: Where N is the batch size (a batch has N sequences), and I(·) is an indicator function, which is 1 if the condition is true and 0 otherwise. It is the total number of non-PAD labels within the batch, meaning the loss is calculated only at positions where the label is not a PAD, and the average is taken for these positions.

[0091] Model training batch size: Mini-batch iteration is used, and the example implementation has a batch size of 8.

[0092] Initial learning rate and decay strategy: The initial learning rate is set to 1e-4. In this embodiment, the training script does not have a separate learning rate warmup or decay scheduler configured (cosine decay, step decay, or warmup + decay strategy can be added if needed).

[0093] Weight configuration rules for the cross-entropy loss function: Training uses next-token autoregressive cross-entropy loss, [PAD] is set to ignore_index and does not participate in the loss calculation, and no additional class weights are set for different component categories (default equal weight).

[0094] The selection criteria for training epochs and the quantification standards for selecting the optimal model are as follows: The training epochs are set to 500. During training, the training loss (and optional validation loss / perplexity) is used as the quantification criteria for convergence and model quality. The model loss continuously decreases and gradually stabilizes during training. To balance generation quality and training convergence during the generation phase, the model weights at the 150th epoch are selected as the checkpoint for generation, as this epoch corresponds to a relatively stable convergence state, while avoiding the risk of overfitting that may arise from excessively long training. This selection is an empirical setting in the example implementation; in practical applications, the number of generation epochs can be adjusted based on validation set metrics or generation quality evaluation results.

[0095] In summary, the key training configurations and saving strategies determined in this embodiment include: batch size=8, lr=1e-4, 500 training rounds, saving a checkpoint every 50 rounds, and using the weights from the 150th round in the generation phase.

[0096] Early stopping trigger conditions and model weight saving time / effect interval: Early stopping is not enabled; training and model selection are performed using a fixed number of training epochs (500 epochs) and checkpoints are saved at intervals. During the generation phase, appropriate checkpoints can be selected for generation based on validation metrics (such as validation loss / perplexity) or generation quality assessment results. If necessary, early stopping rules can be introduced in the optional implementation (such as stopping training if the validation metric does not improve for a certain number of consecutive times) to further improve training efficiency.

[0097] Step S106: Candidate element sequence generation.

[0098] The model parameters saved during the training phase are loaded as the generated model; the example implementation uses the model weights from the 150th epoch (model_epoch_150.pth). Using the initial input (which can be empty or a preset initial element sequence) as a condition, the model progressively predicts the next element ID and appends the prediction to the end of the sequence, iterating until the generated length reaches the preset length. The next element ID is obtained by sampling from the probability distribution of the next element ID output by the model (using strategies such as top-k or threshold truncation), generating a candidate element ID sequence. After generation, [PAD] (ID=0) in the sequence is filtered out to obtain the final candidate element ID sequence. Furthermore, a "continuous merging" rule is adopted for non-coding regions in the elementization representation: adjacent non-coding segments are merged into one non-coding element during the elementization stage. Therefore, under normal circumstances, two consecutive non-coding elements will not appear in the element sequence; a consistency constraint is added during generation: two consecutive non-coding elements (i.e., two adjacent elements are both nonconding_xxx) are not allowed to be generated. If consecutive non-coded elements appear during the generation process, resampling is performed on the current step. If the constraint cannot be met after multiple resamplings, the candidate sequence is discarded and regenerated.

[0099] The specific rules for selecting the preset starting element sequence are as follows: No starting element sequence needs to be provided during generation; the model will generate autoregressively from an empty prefix. If it is necessary to control the generation style, the first few elements of a real phage element sequence can be selected as the starting input (e.g., randomly selecting the first K elements of a sequence from the training set) to guide subsequent generation. The specific numerical parameters for top-k / threshold truncation are as follows: In this example, threshold filtering is used to control the candidate set; in the example implementation, filter_thres=0.9. The sampling temperature is also set to temperature=1.0 (adjustable). After threshold filtering, the next element ID is obtained by sampling using the Gumbel method.

[0100] Generation length preset threshold and adjustment criteria for adapting to different phages: The generation length is controlled by preset parameters. The example implementation uses a fixed generation length. In practical applications, the generation length can be adjusted according to the element number distribution of different phages or the target DNA length requirements. For example, the generation length range can be selected based on the statistical distribution of the number of elements in the training set.

[0101] Iteration termination criteria: In this embodiment, "reaching the preset generation length" is the primary termination condition. After generation, the [PAD] in the output sequence is filtered to obtain the final candidate element ID sequence. Optional termination conditions can be added if needed, such as stopping when a large number of [PAD] appear consecutively in the generation sequence.

[0102] Preliminary screening criteria for candidate sequences (e.g., percentage of effective components): Basic screening of candidate component sequences can be performed, for example, requiring the "percentage of effective components" to reach a certain threshold: Percentage of effective components = (Number of non-[PAD] components) / (Length of generated sequence). When the percentage of effective components is lower than the threshold, the candidate sequence can be discarded or regenerated. The threshold can be set according to actual needs (e.g., ≥0.9).

[0103] Step S107: Decoding and DNA fragment splicing.

[0104] Based on the element-digit ID mapping table, the candidate element ID sequence is decoded into the corresponding element name sequence. According to the pre-established element name-DNA fragment correspondence, the DNA fragment sequence corresponding to each element is obtained. The corresponding DNA fragments are then spliced ​​together in the order of element names to obtain the complete phage DNA sequence output.

[0105] The rules for constructing and updating the element name-DNA fragment correspondence are as follows: The correspondence between element names and DNA fragments is provided by the partner during the elementation stage, and each element name corresponds to a fixed DNA fragment sequence. This invention directly reads this correspondence for lookup and assembly during the generation stage. Training and generation must use the same version of the "element name" mapping. ID mapping table and component name The "DNA fragment library" and both should be version-compatible with the model weights to avoid decoding or splicing errors. The "component name" is also relevant. The "ID mapping table" is mainly used in the training / input / output generation stage, and the "component name" is... The DNA fragment library is mainly used in the decoding and splicing stages; if the element library is updated (adding / replacing elements), the mapping table and DNA fragment library need to be updated synchronously and retrained to ensure consistency.

[0106] Adapter sequence design for DNA fragment splicing (if applicable): Direct splicing is used, connecting the corresponding DNA fragments end-to-end according to the generated element order, without designing additional adapter / linking sequences. If subsequent applications require the introduction of linking sequences (e.g., a uniform linker or restriction enzyme site), a preset adapter sequence can be inserted between adjacent fragments in an optional implementation, but this is not a necessary step in this embodiment.

[0107] Methods for length verification and format standardization of the assembled complete sequence: After assembly, the output sequence length is recorded as the final length of the candidate sequence, and a list of the lengths of each element fragment is also recorded (for tracking the assembly structure). No additional "too short / too long" threshold filtering rules are set; only the length is recorded and output. The generated results are output in FASTA format: a unique sequence ID is assigned to each generated sequence (e.g., >gen_0001), followed by the corresponding DNA base sequence, with line breaks at a fixed line width (80 base characters per line in the example implementation) to facilitate subsequent reading and analysis by tools.

[0108] Handling strategy for anomalous elements without corresponding DNA sequences: If the decoded element name cannot find a corresponding sequence in the DNA fragment library, the element is considered an anomalous element, skipped, and its corresponding fragment is not spliced ​​into the final DNA sequence. Subsequent elements are then processed to ensure the splicing process can continue and output results. If necessary, a stricter strategy can be added to the optional implementation, such as discarding the entire candidate sequence and regenerating it if an anomalous element is found, or replacing it with a preset placeholder fragment and marking it.

[0109] In summary, compared with the prior art, the phage DNA sequence generation method based on hierarchical modeling provided in this embodiment adopts a hierarchical Transformer structure. Through the collaborative modeling of the global layer and the local layer, it can simultaneously characterize the long-range dependencies and local combination patterns between elements, thereby enhancing the comprehensive modeling ability of global structural information and local details in the long sequence de novo generation task, thus improving the quality of the generated sequence and enhancing its practicality.

[0110] To facilitate better implementation of the hierarchical modeling-based phage DNA sequence generation method of this application, this application also provides a hierarchical modeling-based phage DNA sequence generation device based on the above-described hierarchical modeling-based phage DNA sequence generation method. The meanings of the terms used are the same as in the hierarchical modeling-based phage DNA sequence generation method described above, and specific implementation details can be found in the descriptions in the method embodiments.

[0111] Please see Figure 4 , Figure 4 This is a schematic diagram of the hierarchical modeling-based phage DNA sequence generation device provided in an embodiment of this application. Specifically, the hierarchical modeling-based phage DNA sequence generation device may include an acquisition module 201, a conversion module 202, a generation module 203, and a splicing module 204, as follows: The acquisition module 201 is used to acquire the element characterization results of the phage DNA sequence and obtain the element sequence composed of elements arranged in sequence; The conversion module 202 is used to convert a sequence of components into a sequence of numbers consisting of digital identifiers based on a preset mapping relationship. The generation module 203 is used to input the digit sequence into the trained hierarchical generation model, and to generate the target digit sequence by co-modeling the global dependencies and local context information in the digit sequence; wherein, the hierarchical generation model includes a global processing layer and a local processing layer; The splicing module 204 is used to decode the target digital sequence into the corresponding target element sequence based on a preset mapping relationship, and splice each element in the target element sequence one by one according to the preset correspondence between the element and the DNA fragment to obtain a complete phage DNA sequence.

[0112] Furthermore, in some embodiments, the conversion module 202 is specifically used for: Count all element types appearing in the element sequences of all samples, construct a mapping table between elements and identifiers, and the mapping table contains the correspondence between filler markers and preset numeric identifiers; For each element sequence, according to the mapping table, each element in each element sequence is replaced with its corresponding numeric identifier in turn to obtain the numeric sequence; Store a mapping table between element identifiers for use during the sequence generation stage.

[0113] Furthermore, in some embodiments, the apparatus further includes an alignment module, specifically used for: Get the preset target length value; For a number sequence whose length is less than the target length value, add a padding marker corresponding to the number identifier at the end of the sequence to make the length of the number sequence reach the target length value, thus obtaining a fixed-length number sequence; For a number sequence whose length is greater than the target length, a continuous subsequence of length equal to the target length is extracted from the number sequence to obtain a fixed-length number sequence; wherein, the extraction is performed using a random starting point extraction method.

[0114] Furthermore, in some embodiments, the generation module 203 is specifically used for: The numerical sequence input to the hierarchical generation model is divided into multiple consecutive sequence blocks; By using a global processing layer, modeling is performed on a per-sequence-block basis to capture long-range dependencies between blocks and obtain a global context representation; By using a local processing layer, elements are modeled within each sequence block, capturing the local contextual relationships of elements within the block to obtain a local detail representation; Based on global context representation and local detail representation, the target number sequence is predicted and generated bit by bit in an autoregressive manner.

[0115] Furthermore, in some embodiments, the generation module 203 is specifically used for: In each generation process, the probability distribution vector output by the hierarchical generation model is obtained. Each element in the probability distribution vector represents the probability value that the next numeric identifier is the corresponding candidate identifier. Based on a preset threshold parameter, a set of candidate identifiers whose sum of probability values ​​reaches the threshold parameter is selected from the probability distribution vector to form a truncated candidate set. The probability values ​​in the truncation candidate set are scaled based on preset temperature parameters to obtain the adjusted probability distribution. Based on the adjusted probability distribution, the next digital identifier is obtained by sampling from the truncated candidate set using the Gumbel sampling method.

[0116] Furthermore, in some embodiments, the generation module 203 is specifically used for: During the generation process, it is determined whether the type combination between the element corresponding to the currently generated digital identifier and the previously generated element conforms to the preset biological sequence constraint rules; the preset biological sequence constraint rules prohibit the consecutive generation of two non-coding elements. If it does not meet the requirements, the result generated in the current step will be corrected or resampled.

[0117] Furthermore, in some embodiments, the generation module 203 is specifically used for: When it is determined that the type combination of the currently generated element and the previously generated element does not conform to the preset biological sequence constraint rules, the generation result of the current step is marked as invalid; Resample the current step to select the next numeric identifier from the candidate identifier set; Determine whether the type combination of the element obtained after resampling and the previously generated element conforms to the preset biological sequence constraint rules. If it does, accept the generation result and continue to execute the subsequent generation steps. If it does not, repeat the sampling and judgment steps until the preset resampling limit is reached. If the maximum number of resampling attempts is reached and a result that meets the preset biological sequence constraint rules is still not obtained, the current candidate sequence is discarded and the generation process is restarted.

[0118] Furthermore, in some embodiments, the splicing module 204 is specifically used for: Obtain the preset mapping relationship and decode the target number sequence into the corresponding target component name sequence; Query the preset element DNA fragment library to obtain the DNA fragment sequence corresponding to each target element name; The obtained DNA fragment sequences are spliced ​​together end-to-end according to the order of the target element names to obtain the complete bacteriophage DNA sequence.

[0119] Furthermore, in some embodiments, the splicing module 204 is specifically used for: If the target element name corresponding to the DNA fragment cannot be found during the query process, the element corresponding to the target element name will be skipped, and subsequent elements will continue to be assembled.

[0120] For specific limitations regarding the hierarchical modeling-based phage DNA sequence generation device, please refer to the limitations of the hierarchical modeling-based phage DNA sequence generation method above, which will not be repeated here. Each module in the aforementioned hierarchical modeling-based phage DNA sequence generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0121] The phage DNA sequence generation device based on hierarchical modeling provided in this embodiment uses a hierarchical generation model to collaboratively model the global dependencies and local contextual information of phage DNA sequences, thereby achieving high-quality generation of phage DNA sequences. This effectively adapts to the actual needs of phage DNA sequence generation and solves the problem that existing technologies cannot simultaneously and effectively characterize the global dependencies and local contextual information of sequences, resulting in poor quality of the generated phage DNA sequences.

[0122] Furthermore, embodiments of this application also provide an electronic device, such as... Figure 5 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically: The electronic device may include components such as a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, a power supply 303, and an input unit 304. Those skilled in the art will understand that... Figure 5 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 301 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 302, and by calling data stored in the memory 302, thereby providing overall monitoring of the electronic device. Optionally, the processor 301 may include one or more processing cores; preferably, the processor 301 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 301.

[0123] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and a hierarchical modeling-based phage DNA sequence generation method by running the software programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.

[0124] The electronic device also includes a power supply 303 that supplies power to various components. Preferably, the power supply 303 can be logically connected to the processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 303 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0125] The electronic device may also include an input unit 304, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0126] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 301 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 302 according to the following instructions, and the processor 301 runs the applications stored in the memory 302 to realize various functions, as follows: The component characterization results of the phage DNA sequence are obtained, resulting in a component sequence composed of components arranged in sequence. Based on a preset mapping relationship, the component sequence is converted into a numerical sequence composed of numerical identifiers. The numerical sequence is input into a trained hierarchical generation model, which collaboratively models the global dependencies and local context information in the numerical sequence to generate a target numerical sequence. The hierarchical generation model includes a global processing layer and a local processing layer. Based on the preset mapping relationship, the target numerical sequence is decoded into the corresponding target component sequence, and according to the preset correspondence between components and DNA fragments, each component in the target component sequence is spliced ​​together one by one to obtain the complete phage DNA sequence.

[0127] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0128] This application embodiment uses a hierarchical generation model to collaboratively model the global dependencies and local contextual information of phage DNA sequences, thereby achieving high-quality generation of phage DNA sequences. This effectively adapts to the actual needs of phage DNA sequence generation and solves the problem that existing technologies cannot simultaneously and effectively characterize the global dependencies and local contextual information of sequences, resulting in poor quality of the generated phage DNA sequences.

[0129] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0130] Therefore, embodiments of this application provide a storage medium storing multiple instructions that can be loaded by a processor to execute steps in any of the hierarchical modeling-based phage DNA sequence generation methods provided in this application. For example, the instructions can execute the following steps: The component characterization results of the phage DNA sequence are obtained, resulting in a component sequence composed of components arranged in sequence. Based on a preset mapping relationship, the component sequence is converted into a numerical sequence composed of numerical identifiers. The numerical sequence is input into a trained hierarchical generation model, which collaboratively models the global dependencies and local context information in the numerical sequence to generate a target numerical sequence. The hierarchical generation model includes a global processing layer and a local processing layer. Based on the preset mapping relationship, the target numerical sequence is decoded into the corresponding target component sequence, and according to the preset correspondence between components and DNA fragments, each component in the target component sequence is spliced ​​together one by one to obtain the complete phage DNA sequence.

[0131] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0132] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0133] Since the instructions stored in the storage medium can execute the steps in any of the hierarchical modeling-based phage DNA sequence generation methods provided in the embodiments of this application, the beneficial effects that any of the hierarchical modeling-based phage DNA sequence generation methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0134] The foregoing has provided a detailed description of a method and apparatus for generating phage DNA sequences based on hierarchical modeling, as provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A hierarchical modeling-based phage DNA sequence generation method, characterized by, include: The elemental characterization results of the phage DNA sequence were obtained, resulting in an element sequence composed of elements arranged in sequence; Based on a preset mapping relationship, the component sequence is converted into a number sequence composed of digital identifiers; The number sequence is input into a trained hierarchical generation model, which collaboratively models the global dependencies and local contextual information in the number sequence to generate the target number sequence; wherein, the hierarchical generation model includes a global processing layer and a local processing layer; Based on the preset mapping relationship, the target digital sequence is decoded into the corresponding target element sequence, and according to the preset correspondence between elements and DNA fragments, each element in the target element sequence is spliced ​​together one by one to obtain a complete phage DNA sequence.

2. The hierarchical modeling-based phage DNA sequence generation method of claim 1, wherein, The process of converting the element sequence into a number sequence composed of numeric identifiers based on a preset mapping relationship includes: All element types appearing in the element sequences of all samples are counted, and a mapping table between elements and identifiers is constructed. The mapping table contains the correspondence between filler markers and preset numeric identifiers. For each element sequence, according to the mapping table, each element in each element sequence is replaced sequentially with the corresponding numeric identifier to obtain the numeric sequence; A mapping table between the component identifiers is stored for decoding during the sequence generation stage.

3. The hierarchical modeling-based phage DNA sequence generation method of claim 1, wherein, After converting the element sequence into a number sequence composed of numeric identifiers based on a preset mapping relationship, the method further includes: Get the preset target length value; For a number sequence whose length is less than the target length value, add a numeric identifier corresponding to the padding marker to the end of the sequence so that the length of the number sequence reaches the target length value, thus obtaining a fixed-length number sequence; For a number sequence whose length is greater than the target length value, a continuous subsequence of length equal to the target length value is extracted from the number sequence to obtain a fixed-length number sequence; wherein, the extraction is performed using a random starting point extraction method.

4. The hierarchical modeling based phage DNA sequence generation method of claim 1, wherein, The step of inputting the digit sequence into a trained hierarchical generation model, and using the hierarchical generation model to collaboratively model the global dependencies and local context information in the digit sequence to generate the target digit sequence, includes: The input digital sequence to the hierarchical generation model is divided into multiple consecutive sequence blocks; The global processing layer models the sequence blocks as units, captures long-range dependencies between blocks, and obtains a global context representation. Through the local processing layer, the elements are modeled within each sequence block, capturing the local contextual relationships of the elements within the block to obtain a local detail representation; Based on the global context representation and the local detail representation, the target number sequence is predicted and generated bit by bit in an autoregressive manner.

5. The phage DNA sequence generation method based on hierarchical modeling according to claim 4, characterized in that, The step of predicting and generating the target number sequence digit by digit in an autoregressive manner includes: In each generation process, the probability distribution vector output by the hierarchical generation model is obtained, and each element in the probability distribution vector represents the probability value that the next numeric identifier is the corresponding candidate identifier. Based on a preset threshold parameter, a set of candidate identifiers whose sum of probability values ​​reaches the threshold parameter is selected from the probability distribution vector to form a truncated candidate set. The probability values ​​in the truncation candidate set are scaled based on preset temperature parameters to obtain an adjusted probability distribution. Based on the adjusted probability distribution, the next digital identifier is obtained by sampling from the truncated candidate set using the Gumbel sampling method.

6. The phage DNA sequence generation method based on hierarchical modeling according to claim 5, characterized in that, The step of predicting and generating the target number sequence digit by digit in an autoregressive manner further includes: During the generation process, it is determined whether the type combination between the element corresponding to the currently generated digital identifier and the previously generated element conforms to the preset biological sequence constraint rule; the preset biological sequence constraint rule prohibits the consecutive generation of two non-coding elements. If it does not meet the requirements, the result generated in the current step will be corrected or resampled.

7. The phage DNA sequence generation method based on hierarchical modeling according to claim 6, characterized in that, The step of correcting or resampling the generated result of the current step includes: When it is determined that the type combination of the currently generated element and the previously generated element does not conform to the preset biological sequence constraint rules, the generation result of the current step is marked as invalid; Resample the current step to select the next numeric identifier from the candidate identifier set; Determine whether the type combination of the element obtained after resampling and the previously generated element conforms to the preset biological sequence constraint rules. If it does, accept the generation result and continue to execute the subsequent generation steps. If it does not, repeat the sampling and judgment steps until the preset upper limit of resampling times is reached. If the generation result that meets the preset biological sequence constraint rules cannot be obtained after reaching the upper limit of the number of resampling attempts, the current candidate sequence is discarded and the generation process is restarted.

8. The phage DNA sequence generation method based on hierarchical modeling according to claim 1, characterized in that, Based on the preset mapping relationship, the target digital sequence is decoded into a corresponding target element sequence, and according to the preset correspondence between elements and DNA fragments, each element in the target element sequence is spliced ​​together one by one to obtain a complete phage DNA sequence, including: Obtain the preset mapping relationship and decode the target number sequence into the corresponding target element name sequence; Query the preset element DNA fragment library to obtain the DNA fragment sequence corresponding to each target element name; The obtained DNA fragment sequences are spliced ​​together end-to-end according to the order of the target element name sequence to obtain a complete bacteriophage DNA sequence.

9. The phage DNA sequence generation method based on hierarchical modeling according to claim 8, characterized in that, The step of splicing the obtained DNA fragment sequences end-to-end according to the order of the target element name sequence to obtain a complete bacteriophage DNA sequence includes: If the target element name corresponding to the DNA fragment cannot be found during the query process, the element corresponding to the target element name is skipped, and subsequent elements are spliced.

10. A phage DNA sequence generation device based on hierarchical modeling, characterized in that, include: The acquisition module is used to acquire the elemental characterization results of the phage DNA sequence, and obtain the element sequence composed of elements arranged in sequence; A conversion module is used to convert the element sequence into a number sequence composed of numeric identifiers based on a preset mapping relationship; A generation module is used to input the number sequence into a trained hierarchical generation model, and to generate a target number sequence by co-modeling the global dependencies and local context information in the number sequence through the hierarchical generation model; wherein, the hierarchical generation model includes a global processing layer and a local processing layer; The splicing module is used to decode the target digital sequence into the corresponding target element sequence based on the preset mapping relationship, and splice each element in the target element sequence one by one according to the preset correspondence between elements and DNA fragments to obtain a complete phage DNA sequence.