Multi-constraint-based anti-parallel DNA triplex design method and system

By constructing an mCGR initialization model and using a flow network strategy to screen sequences, the stability and design efficiency issues of antiparallel DNA triplexes in complex in vivo environments were solved, enabling the efficient generation of TFO sequences that meet clinical applications and promoting the development of DNA triplexes in the field of targeted gene therapy.

CN120977384APending Publication Date: 2025-11-18DALIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511225513.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Antiparallel DNA triplets are not stable enough in complex in vivo environments and are inefficient to design, making it difficult for traditional design methods to meet the needs of clinical applications.

Method used

A multi-constraint-based antiparallel DNA triplet design method was adopted. By constructing an mCGR initialization model and combining GC content, homopolymer and structured anti-crosstalk constraints, a flow network strategy was used to screen sequences and generate TFO sequences that meet the requirements of stability and target affinity.

Benefits of technology

It significantly improves the stability and targeting affinity of antiparallel DNA triple strands, enhances the efficiency of TFO sequence design and the scale of generated sequence sets, and is suitable for targeted gene therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977384A_ABST
    Figure CN120977384A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-constraint-based anti-parallel DNA triplex design method and system, and relates to the technical field of DNA triplex targeting genes, and the method comprises the following steps: constructing an mCGR initialization model, adjusting a matrix iteration rule according to a sequence length and a basic constraint, and generating all possible sequence sets satisfying the constraint; mapping the sequence set into flow network nodes, screening out approximate nodes through similarity search, constructing to-be-screened edges, then performing sequence comparison on the to-be-screened edges, only reserving the edges of the sequence within a preset stem length range, and deleting the edges beyond the range, so as to overcome the problem of high complexity of mCGR during combined constraint calculation; and finally, outputting a DNA triplex coding sequence meeting all constraint conditions. According to the method, the anti-parallel DNA triplex TFO sequence meeting multiple constraints can be efficiently generated, the sequence set scale is large, the stability is high, the crosstalk rate is low, and a key technical support is provided for research and development of nucleic acid drugs in the field of targeted gene therapy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of DNA triplex targeting gene technology, in particular to an anti-parallel DNA triplex design method and system based on multiple constraints. BACKGROUND

[0002] Anti-parallel DNA triplex (anti-parallel triplex DNA) forms by Hoogsteen base pairing between triplex-forming oligonucleotides (TFO) and the major groove of DNA double helix, and its unique sequence-specific binding ability makes it show great potential in the field of gene expression regulation. For example, TFO can specifically target oncogenes, inhibit their overexpression, and thus interfere with cancer cell proliferation, providing a new direction for the development of nucleic acid drugs for cancer treatment.

[0003] However, the current application of anti-parallel DNA triplex faces two major challenges:

[0004] 1. Insufficient stability: Although the stability of triplex can be improved to some extent by adjusting the ion concentration (such as K + , Mg 2+ ) in the physiological environment, its structure is easily disturbed by non-B-DNA secondary structures (such as mirror structure, slip strand structure) in the complex in vivo environment, making it difficult to maintain stable binding state;

[0005] 2. Low design efficiency: Traditional TFO design methods (such as TripDesign) not only have high computational complexity and long time-consuming when generating sequences that meet basic constraints (GC content, homopolymer), but also have small scale of generated sequence set, and cannot effectively avoid the problem of structured crosstalk, resulting in difficulty in meeting the clinical application requirements of targeting specificity and affinity.

[0006] Therefore, developing an anti-parallel DNA triplex design method that can simultaneously meet the basic biological constraints and structured anti-crosstalk constraints, and balance the design efficiency and sequence quality, has become a key to promote its application in the field of targeted gene therapy. SUMMARY

[0007] The purpose of the present application is to propose an anti-parallel DNA triplex design method and system based on multiple constraints, which can not only improve the stability and targeting affinity of anti-parallel DNA triplex, but also improve the design efficiency of TFO sequence and the scale of generated sequence set.

[0008] According to a first aspect of the embodiments of the present disclosure, an anti-parallel DNA triplex design method based on multiple constraints is provided, comprising the following steps:

[0009] An mCGR initialization model is constructed, and iteration rules of the adjustment matrix are adjusted according to the sequence length and the basic constraint to generate a set of all possible sequences satisfying the constraint;

[0010] The sequence set is mapped to a flow network node, approximate nodes are screened out through similarity search, and edges to be screened are constructed. Then, sequence alignment is performed on the edges to be screened, only the edges with the sequence within a preset stem length range are retained, and the edges exceeding the range are deleted, so as to overcome the high complexity problem of mCGR in combination constraint calculation;

[0011] Finally, the DNA triplex coding sequence meeting all the constraint conditions is output.

[0012] In one embodiment, an mCGR (chaos game representation) initialization model is established according to the sequence length of a triplex-forming oligonucleotide (TFO):

[0013] When the TFO sequence length n = 1, the initial matrix of the mCGR is directly composed of the four bases {A, T, C, G} of DNA;

[0014] When the TFO sequence length n > 1, a high-dimensional matrix is generated by recursive operation of Kronecker product (direct product) to realize the dimension expansion of the TFO sequence, and the expression is:

[0015]

[0016] Wherein, M n represents a high-dimensional chaos game matrix (mCGR) corresponding to a triplex-forming oligonucleotide (TFO) sequence with a length of n, and M represents an initial chaos game matrix.

[0017] In one embodiment, the basic constraint includes a GC content constraint and a homopolymer constraint:

[0018] The GC content constraint based on binary coding is implemented as follows: for the GC content constraint, the bases in the TFO sequence are first converted by binary coding (the rule is A = T → 0, G = C → 1); then the mCGR sequence matrix is mapped and induced into a single row form, and each base is mapped to a different column of the matrix; finally, the numerical sum of each column is calculated, and the sum is divided by the TFO sequence length to obtain the GC content of the corresponding sequence as:

[0019]

[0020] The homopolymer constraint implementation based on matrix operation is: creating mCGR separately in each quartile interval of a given length, marking the bases that can form homopolymer as 1, and marking the rest of the bases as 0, to serve as a generating matrix for homopolymer screening; first performing tiling and stretching on the generating matrix to make its dimensions adapt to the next sequence length; then adding the tiled matrix and the stretched matrix to obtain a new generating matrix that adapts to the next sequence length; repeating the above “tiling-stretching-adding” process until the sequence length corresponding to the generating matrix reaches the required TFO sequence length.

[0021] In one of the embodiments, on the basis of the mCGR initialization model, the GC content constraint and the homopolymer constraint are integrated, and the model is optimized as follows:

[0022] mCGR n =X n +Y n

[0023]

[0024]

[0025] wherein, X n , Y n represent the extension of all bases.

[0026] In one of the embodiments, for the structured anti-crosstalk constraint, the sequence generated by the optimized model mCGR is mapped to a node in the network, each node v i ∈V, represents a sequence generated by the mCGR matrix, and the expression of the node set is:

[0027] V={v1,v2,…,vq}

[0028] wherein, q is the total number of sequences generated by mCGR;

[0029] Similar sequences are mapped to the same “hash bucket”, and the edges to be screened are constructed in this way;

[0030] Check whether the reverse complementary sequences of the prefix v pre and the suffix v suf of each sequence in the sequence set are the same, if they are the same, the sequence is regarded as a mirror structure;

[0031] At the same time, check whether the sequences have the same prefix and suffix, if they are the same, the sequence is regarded as a slip chain structure;

[0032] Delete the sequences that meet the mirror structure and the slip chain structure, and keep the edges whose sequences are within the preset stem length range.

[0033] In one of the embodiments, the sequence alignment formula in the mirror structure is:

[0034] S (i) = s1,s2,s3,…,s i

[0035]

[0036] wherein S (i) represents the prefix part of the sequence, specifically the subsequence composed of the first i bases of the sequence; S (j) represents the suffix part of the sequence, which is the subsequence composed of the jth base to the Lth base (L is the total length of the sequence) of the sequence; m is a preset minimum stem length threshold; S mirror is an identification variable for judging whether the sequence is a mirror structure;

[0037] In one of the embodiments, the sequence alignment formula in the slip chain structure is:

[0038] S′ (j) = reverse(Sj)

[0039]

[0040] wherein S' (j) is the sequence obtained by reversing S (j) .

[0041] According to a second aspect of the embodiments of the present disclosure, a multiple constraint-based antiparallel DNA triplex design system is provided, comprising:

[0042] a sequence set generation module, which constructs an mCGR initialization model, adjusts the matrix iteration rule according to the sequence length and the basic constraint, and generates a sequence set satisfying the constraint;

[0043] a flow network screening optimization module, which maps the above sequence set into flow network nodes, screens out approximate nodes by similarity search and constructs edges to be screened, then performs sequence alignment on the edges to be screened, only retains the edges whose sequences are within the preset stem length range, and deletes the edges beyond the range, so as to overcome the high complexity problem of mCGR in combination constraint calculation;

[0044] an encoding sequence output module, which outputs the DNA triplex encoding sequence meeting all the constraint conditions.

[0045] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, which comprises a memory, a processor and a computer program stored in the memory and running on the memory, and the processor implements the multiple constraint-based antiparallel DNA triplex design method when executing the program.

[0046] According to a fourth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, and the computer readable storage medium has stored thereon a computer program which, when executed by a processor, implements the method for designing anti-parallel DNA triplex based on multiple constraints.

[0047] Compared with the prior art, the above technical scheme has the advantages that: 1. On the basis of homopolymer ≤ 3 and GC content constraints, the application innovatively introduces mirror constraints, slip chain constraints and other structured anti-crosstalk constraints, which can effectively identify and exclude sequences prone to form non-B-DNA structures, significantly reduce the interference probability of non-B-DNA structures on the formation of normal anti-parallel DNA triplex, and ensure the stability and reliability of the triplex structure formation.

[0048] 2. When designing a set of triplex-forming oligonucleotide (TFO) sequences that meet the constraints, the sequence direction is first filtered based on common biochemical constraints (homopolymer and GC content constraints), and then the matrix iteration and dimension expansion capabilities of the chaotic game matrix representation (mCGR) are used to systematically generate all possible sequences that meet the basic constraints, avoiding the sequence omission problem in traditional design methods, and providing a more comprehensive candidate sequence pool for subsequent screening.

[0049] 3. In view of the high complexity problem of mCGR in complex structured anti-crosstalk constraint calculation, the application integrates a flow network strategy, maps the sequences generated by mCGR to flow network nodes, and efficiently completes the anti-crosstalk screening through similarity search and sequence comparison, not only overcoming the calculation limitations of mCGR, but also obtaining a large number of high-quality TFO sequences that meet all constraints. Such TFO sequences have higher targeting affinity when targeting DNA triplex formation, which can effectively promote the practical application and development of DNA triplex in the field of targeted gene therapy. BRIEF DESCRIPTION OF DRAWINGS

[0050] The drawings accompanying the specification of this application are used to provide further understanding of the application, the illustrative embodiments of the application and the description thereof, and do not constitute an improper limitation on the application.

[0051] Figure 1 Flow chart of the method for designing anti-parallel DNA triplex based on multiple constraints. DETAILED DESCRIPTION

[0052] The present disclosure will be further described below in conjunction with the drawings and embodiments.

[0053] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as generally understood by those skilled in the art to which the application belongs.

[0054] It should be noted that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting of example embodiments in accordance with the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, components, and / or groups thereof, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.

[0055] It should be noted that the flow diagrams and block diagrams in the drawings are representations of the architectural, functional, and operational aspects of possible implementations of methods and systems according to various embodiments of the present disclosure. It should be noted that each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which includes one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the flow diagrams and / or block diagrams and combinations of blocks in the flow diagrams and / or block diagrams can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and

[0056] Embodiment one:

[0057] The embodiment is aimed at the design requirements of anti-parallel DNA triplex, adopts a design scheme based on chaotic game matrix representation (mCGR) and flow network, to generate triplex forming oligonucleotide (TFO) sequences meeting the basic constraints (GC content constraint, homopolymer constraint) and structured anti-crosstalk constraints (mirror constraint, sliding chain constraint) as the target, clearly defines the complete process from mCGR initialization, constraint screening to sequence output, and verifies the advantages of the scheme in sequence generation efficiency and quality.

[0058] The embodiment provides an anti-parallel DNA triplex design method based on multiple constraints, comprising the following steps:

[0059] Step one, construct an mCGR initialization model, generate all possible sequence sets meeting the constraints according to the sequence length and the matrix iteration rule of the basic constraints;

[0060] Specifically, initialize the mCGR according to the TFO sequence length n, when the sequence length n=1, use four bases A, T, C, G to represent, with the increase of the sequence length, generate higher dimension matrix through Kronecker product recursion of the matrix itself, to represent longer sequence, defined as follows:

[0061] mCGR n = X n + Y n (1)

[0062]

[0063] where formula X n , Y n represents the expansion for all bases, formula (2) gets the first base of all possible sequences, formula (3) processes the last base of the sequence in the same way by swapping the two matrices, and finally adds the two matrices to get the complete matrix mCGR n .

[0064] Different constraints are allowed to be obeyed by appending different rule matrices. For generating TFO sequences that satisfy GC content constraints, by converting the bases in the sequence into a binary pattern, A = T = 0, G = C = 1, mapping the elements in the matrix to a row, and mapping the bases therein to the corresponding column, the GC content of the corresponding sequence is obtained by calculating the sum of the columns and then dividing by the sequence length.

[0065] For the homopolymer constraint of anti-parallel DNA triplex, each sequence ends with the same base in one quarter of each matrix generated by mCGR. Therefore, for any given subsequence, there is a specific pattern of repetition. In order to exclude homopolymers, an mCGR can be created in each quartile of a given length, where the appearance of homopolymers is represented as 1 and the rest of the elements are represented as 0 as the generated matrix. The generated matrix is first used for tiling and stretching of the matrix to the next sequence length. Adding the tiled and stretched two matrices gets the next generated matrix, and this process is repeated until the desired TFO length is reached. The optimized model is defined as follows:

[0066] mCGR n = X n + Y n (4)

[0067]

[0068] Step two, map the above sequence set to a flow network node, filter out approximate nodes by similarity search and construct edges to be screened, then perform sequence alignment on the edges to be screened, only keep the edges whose sequences are within the preset stem length range, and delete the edges that exceed the range, so as to overcome the high complexity problem existing in the combination constraint calculation of mCGR;

[0069] Specifically, for the structured anti-crosstalk constraint, a fine and efficient sequence is needed, and the strategy of flow network is applied to mCGR, that is, the sequence generated by mCGR is mapped to a node in the network, each node v i ∈V, represents a sequence generated by an mCGR matrix:

[0070] V={v1,v2,…,vn} (7)

[0071] Where n is the total number of sequences generated by mCGR.

[0072] Local sensitive hashing (LSH) is used to reduce the complexity of high-dimensional data comparison, and similar data is mapped to a bucket, thereby accelerating similarity queries.

[0073] Check if the reverse complement of the prefix v pre and the suffix v suf of each sequence in the sequence set is the same. If it is the same, the sequence is considered to be a mirror structure, and the sequence alignment formula is:

[0074] S (i) =s1,s2,s3,…,s i (8)

[0075]

[0076] At the same time, check if the sequences have the same prefix and suffix, and if they are the same, consider them as a sliding chain structure, and the sequence alignment formula is:

[0077] S′ (j) =reverse(Sj) (11)

[0078]

[0079] Delete sequences that meet the mirror structure and sliding chain structure, and keep the edges within the preset stem length range.

[0080] Step three, finally output the DNA triplex coding sequence that meets all the constraints.

[0081] Example two

[0082] The present embodiment provides a multiple constraint-based antiparallel DNA triplex design system, comprising:

[0083] A sequence set generation module constructs an mCGR initialization model, adjusts the matrix iteration rule according to the sequence length and basic constraints, and generates all possible sequence sets that meet the constraints;

[0084] The flow network screening optimization module maps the sequence set to flow network nodes, screens approximate nodes through similarity search and constructs edges to be screened, then performs sequence alignment on the edges to be screened, only retains edges with sequences within a preset stem length range, and deletes edges beyond the range, thereby overcoming the high complexity problem of mCGR in performing combination constraint calculation.

[0085] The coding sequence output module outputs the DNA triplex coding sequence meeting all constraint conditions.

[0086] Embodiment three:

[0087] An electronic device comprising a memory, a processor and a computer program stored on the memory, wherein the processor executes the program to implement the above-mentioned multiple constraint-based antiparallel DNA triplex design method, comprising:

[0088] The mCGR initialization model is constructed, the matrix iteration rule is adjusted according to the sequence length and the basic constraint, and all possible sequence sets meeting the constraint are generated;

[0089] The sequence set is mapped to flow network nodes, approximate nodes are screened through similarity search and edges to be screened are constructed, then sequence alignment is performed on the edges to be screened, only edges with sequences within a preset stem length range are retained, and edges beyond the range are deleted, thereby overcoming the high complexity problem of mCGR in performing combination constraint calculation;

[0090] The DNA triplex coding sequence meeting all constraint conditions is finally output.

[0091] Embodiment four:

[0092] A computer readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement the above-mentioned multiple constraint-based antiparallel DNA triplex design method, comprising:

[0093] The mCGR initialization model is constructed, the matrix iteration rule is adjusted according to the sequence length and the basic constraint, and all possible sequence sets meeting the constraint are generated;

[0094] The sequence set is mapped to flow network nodes, approximate nodes are screened through similarity search and edges to be screened are constructed, then sequence alignment is performed on the edges to be screened, only edges with sequences within a preset stem length range are retained, and edges beyond the range are deleted, thereby overcoming the high complexity problem of mCGR in performing combination constraint calculation;

[0095] The DNA triplex coding sequence meeting all constraint conditions is finally output.

[0096] Application example one:

[0097] In the application example, the triplex-forming oligonucleotide (TFO) of the antiparallel DNA triplex is coded with a length of 20mers, and the core constraint parameters are set as follows: GC content accounts for 60% (GC%=60%), homopolymer length≤3, and the remaining constraints (mirror constraint, slip chain constraint) are executed according to the requirements of the foregoing scheme. The specific implementation steps are as follows:

[0098] Step 1: According to the TFO sequence length n=20, initialize the chaotic game matrix representation (mCGR). When the sequence length n=1, the initial matrix of mCGR is constructed with the four bases {A, T, C, G} of DNA; as the sequence length increases, the high-dimensional mCGR matrix corresponding to the longer sequence is generated through the Kronecker product (direct product) recursive operation of the matrix itself, until the matrix corresponding to the 20mers sequence is obtained.

[0099] Step 2: The antiparallel DNA triplex sequence generated in step 1 is substituted into the constraint condition of GC%=60% and homopolymer≤3 for screening, and the sequences with substandard GC content and existing continuous 3 or more same bases are removed, and the antiparallel TFO sequences meeting the quality requirements are retained.

[0100] Step 3: For the sequence set obtained in step 2, the local sensitive hashing (LSH) algorithm is used to cluster similar sequences, and sequences with similar characteristics are mapped to the same “hash bucket”, which greatly reduces the calculation amount of subsequent similarity search and improves the screening efficiency.

[0101] Step 4: Combined with the designed TFO sequence length n=20, the screening range of the repeat sequence (prefix and suffix) is defined as follows: wherein the length range of the prefix and suffix of the mirror constraint is set as (i.e. [8, 10]), and the length range of the prefix and suffix of the slip chain constraint is set as (i.e. [5, 10]), and all sequences are traversed according to the range.

[0102] Step 5: Check whether the reverse complementary sequences of the prefix s pre and the suffix s suf of the sequence are the same one by one, and if the reverse complementary sequences are the same, the sequence is determined as a mirror structure.

[0103] Step 6: At the same time, according to the range of the prefix and suffix of the slip chain constraint defined in step 4, the prefix and suffix of each sequence are rechecked whether they are completely the same, and if they are the same, the sequence is determined as a slip chain structure.

[0104] Step 7: Remove the sequences that do not meet the constraints determined as mirror structure and slip chain structure, and finally output the high-quality TFO sequence set meeting all core constraints.

[0105] Step 8: Based on the TFO sequence set generated in step 7, combined with the base pairing matching rules of AG type anti-parallel DNA triplex (such as Hoogsteen pairing principle), further filter out the sequence combination that can stably form AG type anti-parallel DNA triplex, and obtain the triplex set meeting the targeting demand.

[0106] Step 9: To verify the stability of the constructed triplex set, two key indicators are calculated: one is the number of undesired secondary structures (such as hairpin structure, stem loop structure) that may be formed in the set, and the other is the hydrogen bond interaction energy of the triplex; wherein the hydrogen bond interaction energy is defined as follows:

[0107] ΔE int = ΔE pair + ΔE Hoogsteen

[0108]

[0109] Wherein, ΔE is the bond energy of each layer of base, E is the total hydrogen bond energy, n is the sequence length, ΔE pair is the Watson-Crick hydrogen bond energy, ΔE Hoogsteen is the Hoogsteen hydrogen bond energy. According to the total bond energy E to evaluate the stability of the secondary structure: the larger |E| indicates that the DNA molecule is easy to produce stable secondary structure, which affects the formation of the final anti-parallel DNA triplex.

[0110] The experimental verification of the application is carried out in a specific hardware and software environment: in terms of hardware, Intel(R) Core(TM) i5-1050 3.10GHz CPU is adopted, the basic memory configuration is 16.00GB, and the operating system is Windows 11; in terms of software, the construction of flow network is realized through NetworkX library, and the memory is expanded to 60G in the flow path search stage to meet the high-dimensional sequence data processing demand. During the experiment, a large number of sequences not meeting the constraints are successfully screened out, verifying the effectiveness of the scheme, and the experimental results show that the performance of the method of the application is superior to other similar algorithms.

[0111] To further embody the advantages of the application in TFO sequence design efficiency and generated sequence scale, the application is compared with similar research work in the current field, and the specific comparison data is shown in Table 1 (the optimal result is marked in bold in the table).

[0112] Table 1: Size of DNA triplex code set and time spent

[0113]

[0114] From the results of Table 1, the present application designs TFO sequences with length of 16-22, and then compares with representative works in different dimensions. The present application is shorter in time in constructing the anti-parallel DNA triplex coding set, and the coding set is larger.

[0115] Table 2: Number of slip strand structures and corresponding hydrogen bond interaction energy designed by different triplex design methods

[0116]

[0117] Table 3: Number of slip strand structures and corresponding hydrogen bond interaction energy designed by different triplex design methods

[0118]

[0119] From the results of Table 2 and Table 3, the secondary structure is caused by the stem length of the larger mirror / slip strand structure, and for the mirror constraint, the average bond energy of the secondary structure of the other two methods is higher, and the absolute value is greater than 100, while the absolute value of the bond energy of the sequence generated by the present application is less than 100. Experimental results show that the secondary structure bond energy of the sequence set generated by MFN is lower than that of the other two methods, and the same phenomenon occurs for the slip strand constraint, indicating that the secondary structure of the sequence set generated by the present application has lower crosstalk rate.

[0120] Those skilled in the art should understand that each module or step of the present disclosure described above can be realized by a general computer device, and alternatively, they can be realized by program codes executable by a computing device, so that they can be stored in a storage device and executed by a computing device, or they can be respectively manufactured into each integrated circuit module, or a plurality of modules or steps thereof can be manufactured into a single integrated circuit module. The present disclosure is not limited to any specific combination of hardware and software.

[0121] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

[0122] The above describes the specific embodiments of the present disclosure in combination with the drawings, but is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications or changes made on the basis of the technical solutions of the present disclosure without creative labor are still within the protection scope of the present disclosure.

Claims

1. A method for designing antiparallel DNA triplets based on multiple constraints, characterized in that, Includes the following steps: Construct an mCGR initialization model, adjust the matrix iteration rules according to the sequence length and basic constraints, and generate a set of all possible sequences that satisfy the constraints; The above sequence set is mapped to flow network nodes. Similar nodes are selected through similarity search and edges to be selected are constructed. Then, the edges to be selected are sequence compared. Only edges whose sequences are within the preset stem length range are kept, and edges that exceed the range are deleted. This overcomes the high complexity problem of mCGR when performing combinatorial constraint calculation. The final output is a DNA triplet coding sequence that meets all constraints.

2. The antiparallel DNA triplet design method based on multiple constraints according to claim 1, characterized in that, Based on the sequence length of the triplet oligonucleotide, an mCGR initialization model is established: When the TFO sequence length n = 1, the initial matrix of mCGR is directly composed of the four bases of DNA {A, T, C, G}; When the length n>1 of the TFO sequence, a high-dimensional matrix is ​​generated through recursive operation of the Kronecker product, thereby expanding the dimension of the TFO sequence. Its expression is: Among them, M n Let M represent the high-dimensional chaotic game matrix corresponding to the oligonucleotide sequence formed by three strands of length n, and M represent the initial chaotic game matrix.

3. The antiparallel DNA triplet design method based on multiple constraints according to claim 1, characterized in that, The fundamental constraints include GC content constraints and homopolymer constraints: The implementation method of GC content constraint based on binary encoding is as follows: For GC content constraint, firstly, the bases in the TFO sequence are converted into binary encoding (rule A = T → 0, G = C → 1); then, the mCGR sequence matrix is ​​mapped and summarized into a single row form, and each base is mapped to a different column of the matrix; finally, by calculating the sum of the values ​​in each column and dividing the sum by the length of the TFO sequence, the GC content of the corresponding sequence is obtained. The homopolymer constraint implementation based on matrix operations is as follows: Create a separate mCGR within each quartile interval of a given length, marking bases that may form homopolymers as 1 and the rest as 0, using this as the generator matrix for homopolymer screening; first, flatten and stretch this generator matrix to adapt its dimensions to the next sequence length; then, add the flattened matrix and the stretched matrix to obtain a new generator matrix adapted to the next sequence length; repeat the above "flatten-stretch-add" process until the sequence length corresponding to the generator matrix reaches the required TFO sequence length.

4. The antiparallel DNA triplet design method based on multiple constraints according to claim 3, characterized in that, Based on the mCGR initialization model, GC content constraints and homopolymer constraints are incorporated to optimize the model, as shown below: mCGR n =X n +Y n Among them, X n Y n This indicates an expansion of all bases.

5. The antiparallel DNA triplet design method based on multiple constraints according to claim 1, characterized in that, To address the structured anti-crosstalk constraints, the sequences generated by the optimized model mCGR are mapped to nodes in the network, with each node v i ∈V, each represents a sequence generated by the mCGR matrix, and the expression for the node set is: V = {v1, v2, ..., vq} Where q is the total number of sequences generated by mCGR; Similar sequences are mapped to the same "hash bucket" to construct the edges to be filtered; Check the prefix v of each sequence in the sequence set pre and the suffix v suf If the reverse complementary sequences are the same, then the sequence is considered a mirror image. Simultaneously check whether the sequences have the same prefix and suffix; if they are the same, the sequence is considered a slip chain structure. Delete sequences that satisfy the mirror structure and slip chain structure, and retain the edges of sequences that are within the preset stem length range.

6. The antiparallel DNA triplet design method based on multiple constraints according to claim 5, characterized in that, The formula for sequence alignment in a mirror structure is: S (i) =s1,s2,s3,…,s i Among them, S (i) The prefix portion of the sequence is specifically the subsequence consisting of the first base to the i-th base; S (j) The suffix represents the subsequence from the j-th base to the L-th base (L is the total sequence length); m is the preset minimum stem length threshold; S mirror This is a flag variable used to determine whether a sequence is a mirror image.

7. The antiparallel DNA triplet design method based on multiple constraints according to claim 5, characterized in that, The formula for sequence alignment in a slip chain structure is: S′ (j) =reverse(Sj) Among them, S' (j) It is for S (j) The sequence obtained after performing the reversal operation.

8. A multi-constraint-based antiparallel DNA triplet design system, characterized in that, include: The sequence set generation module constructs the mCGR initialization model, adjusts the matrix iteration rules based on the sequence length and basic constraints, and generates a set of all possible sequences that satisfy the constraints. The flow network filtering and optimization module maps the above sequence set to flow network nodes, filters out similar nodes through similarity search and constructs edges to be filtered, and then performs sequence comparison on the edges to be filtered, keeping only the edges whose sequences are within the preset stem length range and deleting the edges that exceed the range, thereby overcoming the high complexity problem of mCGR when performing combined constraint calculations. The coding sequence output module outputs the DNA triplet coding sequence that meets all constraints.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running thereon, characterized in that, When the processor executes the program, it implements the antiparallel DNA triplet design method based on multiple constraints as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the antiparallel DNA triplet design method based on multiple constraints as described in any one of claims 1-7.