A method for calculating synthesis sequences of an in-situ DNA microarray synthesis system

Optimizing the synthetic sequence calculation of the in-situ DNA microarray synthesis system through the probability beam search algorithm, solving the local trap problem, improving the calculation efficiency and synthesis efficiency, and reducing redundant calculations and reagent use.

CN115881222BActive Publication Date: 2025-07-18XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211036695.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-26
Publication Date
2025-07-18
Estimated Expiration
2042-08-26

AI Technical Summary

Technical Problem

Existing enhanced beam search algorithms are prone to fall into local traps when calculating the shortest common supersequence problem, resulting in redundant calculations and inefficiency.

Method used

The probability beam search algorithm is used to construct a probability database and a probability index database, and combine branch pruning and perturbation mechanisms to optimize the node selection and amplification process, reduce the possibility of local traps and improve computational efficiency.

Benefits of technology

It effectively reduces the computation time complexity and redundant calculation of the in-situ DNA microarray synthesis system, improves the efficiency of the synthesis sequence, saves reagents and time, and shortens the synthesis rounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115881222B_ABST
    Figure CN115881222B_ABST
Patent Text Reader

Abstract

The invention discloses a method for calculating a synthesis sequence of an in-situ DNA microarray synthesis system, including: obtaining an integer sequence set; constructing a probability database and a probability index database; entering an iterative process, amplifying the sequences in the nodes, updating the matching positions of the amplified nodes, and updating the iterative node set; calculating the guiding values of all nodes in the iterative node set; selecting a dominant node set from the iterative node set, and performing branch pruning according to the dominant node set; clearing the candidate node set, selecting a specified number of candidate nodes from the iterative node set in descending order of the guiding values, and adding them to the candidate node set; randomly selecting a specified number of candidate nodes from the iterative node set and adding them to the candidate node set; checking the amplified integer sequences, and if they are the common super-sequences of the integer sequence set encoded by the sequence set, decoding them according to the encoding rules to obtain a string of characters. It can effectively improve the efficiency of the in-situ DNA microarray synthesis system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of computational biology, and in particular to a synthetic sequence calculation method in an in-situ DNA microarray synthesis system. Background Art

[0002] Deoxyribonucleic acid (DNA) is a macromolecular polymer composed of deoxynucleotides. Its in vitro synthesis is mainly carried out by in situ DNA microarray synthesis. The in situ DNA microarray synthesis system is a system that uses light-guided polymerization technology to synthesize DNA microarrays in a solid phase by using a pre-made light shield and four modified bases to activate them through light. Light-guided polymerization technology is a technology that integrates photolithography technology with solid-phase synthesis technology, computer technology, and molecular biology. Light-guided polymerization technology can quickly and easily synthesize a large number of oligonucleotides or polypeptide molecules at preset sites according to a predetermined sequence. As an efficient in vitro DNA synthesis method, the in situ DNA microarray synthesis system can quickly and high-throughput customize DNA synthesis and is now widely used in commercial synthesizers.

[0003] The working of the in situ DNA microarray synthesis system based on light-guided polymerization technology is described as follows: (1) Before synthesis, the glass slide is pre-aminated and the activated amino group is protected with a photolabile protective agent. The 5-end and 3-end activated of the nucleoside molecule are protected by a photosensitive protective agent; (2) An appropriate light-blocking plate is selected to allow light to pass through the site where polymerization is required and to shield the site where polymerization is not required; in this way, light is irradiated onto the support through the light-blocking plate, and the amino groups in the illuminated part are deprotected, thereby undergoing a coupling reaction with the nucleoside molecule; each reaction adds a specific base at tens of thousands of sites; (3) Since the site after the reaction is still protected by the protective agent, the purpose of synthesizing a large amount of DNA with a predetermined sequence at a specific site can be achieved by controlling the light-transmitting and light-shielding patterns of the light-blocking plate and the types of monomer molecules participating in each reaction.

[0004] The synthetic sequence refers to the character sequence composed of the base types added in each synthesis of the above steps. The synthetic sequence must be a common supersequence of the predetermined synthetic DNA sequence set. A common supersequence refers to the existence of a sequence such that all sequences in the sequence set can be obtained by deleting certain elements from this sequence. By calculating the shortest possible common supersequence of the predetermined synthetic DNA character sequence set and finding the shortest possible common supersequence of the sequence set, the number of synthesized times can be reduced.

[0005] There are currently various algorithms for solving the Shortest Common Supersequence Problem (SCSP), such as the Minimum Height algorithm (MH), the Deposition and Reconstruction algorithm (DR), the Ant Colony Optimization algorithm (ACO_SCS), the Improved Beam Search algorithm (IBS_SCS), etc. Among them, the algorithm with the best performance is the Improved Beam Search algorithm. The beam search algorithm is an incomplete tree search algorithm. The beam search algorithm selects multiple nodes at each step, and these nodes compete with each other during iteration and then output the node with the best performance. However, when calculating the shortest common supersequence problem, the Improved Beam Search algorithm always converges along a fixed path and is prone to falling into local traps. Therefore, the above algorithms have the problem of redundant calculation during the calculation process. Summary of the Invention

[0006] The present invention provides a method for calculating a synthesis sequence in an in-situ DNA microarray synthesis system in view of the above problems.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A method for calculating a synthesis sequence of an in-situ DNA microarray synthesis system includes:

[0009] S1. Obtain the path of a sequence file carrying sequence set information, traverse each character element in the sequence file, and encode the character elements into integer elements to obtain an integer sequence set;

[0010] S2. Construct a probability database and a probability index database;

[0011] S3. Enter the iterative process, amplify the sequences in the nodes, update the matching positions of the amplified nodes, and update the iterative node set; it includes: S3.1. Amplify the sequences of each node in the candidate node set, and add a total of |∑| integers 0, 1,..., |∑|-1 respectively; S3.2. Compare the amplified sequences with each sequence in the integer sequence set one by one, and judge whether the matching position is the length of the DNA sequence. If so, terminate the iteration and execute S8, otherwise execute S3.3; S3.3. Judge whether the amplified node is a useful node. If it is a useful node, add it to the iterative node set, otherwise discard the node;

[0012] S4. Calculate the guidance values of all nodes in the iterative node set;

[0013] S5. Select a dominant node set from the iterative node set, and perform branch pruning according to the dominant node set;

[0014] S6. Clear the candidate node set, and select a specified number of candidate nodes from the iterative node set in descending order of the guidance value and add them to the candidate node set;

[0015] S7. Randomly select a specified number of candidate nodes from the iterative node set, add them to the candidate node set, and then execute S3;

[0016] S8. Check the amplified integer sequence. If it is a common supersequence of the integer sequence set encoded by the sequence set, decode it according to the encoding rule to obtain a string of characters.

[0017] In one embodiment: The S1 includes:

[0018] S1.1. Read the strings in the sequence set file according to the file path, and each line is a character sequence;

[0019] S1.2. Perform integer encoding on the characters in the sequence file. The encoding method is that the i-th character traversed from left to right and top to bottom in the sequence file is encoded as i - 1.

[0020] In one embodiment: The S2 includes:

[0021] S2.1. Construct a probability database;

[0022] The probability database stores the probability that sequence y becomes a supersequence of sequence r. The probability calculation formula is as follows:

[0023]

[0024] Where: represents the probability that sequence y becomes a supersequence of sequence r; q is the length of sequence r; k is the length of sequence y; |∑| is the number of types of all characters in the sequence file; r[2…q] represents the subsequence of sequence r taken from the second character to the last character;

[0025] S2.2. Construct a probability index database;

[0026] Construct a probability index database. The mapping relationship between the index value and the maximum matching position is as follows:

[0027] target = (int)maxR * ln|∑| / ln 2

[0028] Where: maxR is the maximum value of the matching positions of all nodes in the iterative node set; when constructing the database, maxR = 1, 2…, m, and m is the length of the longest sequence in the sequence set; target is the second-dimensional index of the index probability database, that is, the length of sequence y.

[0029] In one embodiment: Each node is a custom structure, and the node contains the integer sequence x and the possibility g that the sequence becomes a common supersequence of the sequence set k(x), the positions where sequence x matches each sequence in the sequence set, i.e., a one-dimensional array P[n] of length n, where n is the number of sequences in the sequence set; S3 includes:

[0030] S3.1, traverse each node in the candidate node set, and respectively amplify |∑| integers 0, 1, …, |∑|-1 for the sequences in the nodes;

[0031] S3.2, compare the amplified sequences with each sequence in the integer sequence set one by one, and judge whether the matching position P[i], i ∈ 0, 1, 2…n is the length of the DNA series. If so, terminate the iteration and execute S8; otherwise, execute S3.3;

[0032] S3.3, judge whether the amplified node is a useful node. If so, the node is called a useful node and added to the iterative node set; otherwise, discard the node.

[0033] In one embodiment: S4 includes:

[0034] Traverse each node in the iterative node set, call the probability database, and calculate the guiding value of the node. The calculation formula is as follows:

[0035]

[0036] Where: r i (x) is a subsequence where sequence x in the node matches a certain sequence s in the DNA sequence set, n is the number of sequences in the DNA sequence set, and max is the index of the node with the largest matching position in the iterative node set; i

[0037] Perform amplification processing on the guiding value, and perform logarithmic processing on it. The calculation formula of the logarithmically processed guiding value is as follows:

[0038]

[0039] In one embodiment: In S5, if the sequences of nodes k and j are x k and x j respectively, and there is then it is said that node k is dominated by node j. In subsequent iterations, no matter how x k is amplified, its guiding value g k (x k ) is less than g k (x j ), then perform branch pruning and discard the dominated node k.

[0040] In one embodiment: The method for judging whether a node is dominated is as follows:

[0041] S5.1, select the top n_best nodes with the largest guidance values from the iterative node set to form the dominant node set, where: n_best satisfies n_best < beamSize, and beamSize is the beam width, that is, the number of candidate node sets;

[0042] S5.2, traverse all nodes in the iterative node set except the dominant node set;

[0043] S5.3, traverse all dominant nodes, compare the matching positions of the current node k and the dominant node j, if there is then determine that node k is dominated by node j and discard node k.

[0044] In one embodiment: the said S6 includes:

[0045] S6.1, clear the candidate node set;

[0046] S6.2, select the top n guide large nodes with the largest guidance values from the nodes in the iterative node set, and add the n guide nodes to the candidate node set. n guide is determined by the beam width beamSize and the perturbation rate p, and the three satisfy the following relationship:

[0047] n guide = beamSize * (1 - p)

[0048] where: p = 0.3.

[0049] In one embodiment: in the said S7, randomly select n pertu nodes from the remaining iterative node set, where n pertu = beamSize * p.

[0050] The present invention has achieved the following technical effects compared with the prior art:

[0051] The synthesis sequence calculation method of the in-situ DNA microarray synthesis system adds a perturbation component and a deduplication database on the basis of the probabilistic beam search algorithm. The former avoids falling into local traps and reduces the time complexity, and the latter reduces repeated calculations. While converging, it reduces the possibility of falling into local traps. Moreover, since the calculated common supersequence can be used to guide the determination of the types of deoxynucleotides synthesized by the in-situ DNA microarray synthesis system each time, and the shorter the length of the common supersequence, the fewer the synthesis rounds, and the more reagents and time are saved, which can effectively improve the efficiency of the in-situ DNA microarray synthesis system. Description of the Drawings

[0052] Figure 1It is a flowchart of the in-situ DNA microarray synthesis sequence method of the present invention.

[0053] Figure 2 It is a 3D surface plot of the average common supersequence length calculated by the present invention versus the perturbation rate and beam width. Detailed implementation manners

[0054] In the description of the detailed implementation manners of the present invention, a node is a structure that contains an integer sequence x; the possibility (guidance value) g k (x) that this sequence becomes a common supersequence of the sequence set; the positions where the sequence x matches each sequence in the sequence set, that is, a one-dimensional array P[n] of length n, where n is the number of sequences in the sequence set.

[0055] Figure 1 It is a flowchart of the synthesis sequence calculation method of the in-situ DNA microarray synthesis system of the present invention. Please refer to Figure 1 , the synthesis sequence calculation method of the in-situ DNA microarray synthesis system, includes:

[0056] S1. Obtain the path of the sequence file carrying sequence set information, traverse each character element in the sequence file, and encode the character elements into integer elements to obtain an integer sequence set; it includes:

[0057] S1.1. Read the string in the sequence set file according to the file path. Each line is a character sequence, and the sequence file carrying sequence set information is a text file, and the characters in the text file belong to the specified set ∑;

[0058] Specifically: Since DNA is composed of four deoxynucleotides, and it is the different arrangement orders of the four bases adenine (abbreviated as A), thymine (abbreviated as T), cytosine (abbreviated as C), and guanine (abbreviated as G) in the deoxynucleotides that determine the biodiversity, the symbols in the sequence file must belong to the set ∑={ ′ A ′,′ C ′,′ G ′ , 'T'}. Each line of characters represents a DNA sequence to be synthesized at a synthesis site, and each column of characters represents the types of DNA bases to be synthesized in one layer of the microarray chip.

[0059] S1.2. Perform integer encoding on the characters in the sequence file. The encoding method is that the i-th character traversed from left to right and top to bottom in the DNA sequence file is encoded as i - 1.

[0060] Specifically, for a sequence file storing the information of "GTGGC′\n′CGATGC′\n′GCTAA′\n'", the encoding method is as follows:

[0061] 'G′→0

[0062] ′T′→1

[0063] ′C′→2

[0064] ′A′→3

[0065] S2. Construct a probability database and a probability index database; it includes:

[0066] S2.1: Construct a probability database.

[0067] In order to evaluate the quality of each useful node expanded during the iteration, it is necessary to calculate the guiding value of the node during the iteration. In order to avoid repeated calculation of probabilities during the iteration, it is necessary to construct the probability database and the probability index database prior to entering the iteration. The probability database stores the probability that sequence y becomes a super-sequence of sequence r. The probability is determined by the lengths of sequences r and y. The probability calculation formula is as follows:

[0068]

[0069] where represents the probability that sequence r becomes a super-sequence of sequence y. Where q is the length of sequence r; k is the length of sequence y. Where |∑| is the number of types of all characters in the sequence file. For example, if calculating the common super-sequence of a DNA sequence set, |∑| is 4. Where r[2…q] represents the subsequence of sequence r taken from the second character to the last character.

[0070] S2.2: Probability index database.

[0071] One of the values of the probability index database is calculated based on the maximum position matched by all nodes. In order to avoid repeated calculation of the index value when indexing the probability database, the present invention constructs a probability index database, and the mapping relationship between the index value and the maximum matching position is as follows:

[0072] target=(int)maxRln|∑| / ln 2

[0073] where: maxR is the maximum value of the matching positions of all nodes in the iteration node set. When constructing the database, maxR = 1, 2…, m, and m is the length of the longest sequence in the sequence set. target is the second-dimensional index of the index probability database, that is, the length of sequence y.

[0074] S3. Enter the iterative process, amplify the sequences in the nodes, update the matching positions of the nodes after amplification, and update the iterative node set.

[0075] It includes: S3.1. Traverse each node in the candidate node set, amplify the sequences of each node in the candidate node set, and add a total of |∑| integers 0, 1, …, |∑| - 1 respectively.

[0076] For example, if calculating the common supersequence of a DNA sequence set, it is necessary to amplify the sequences of the nodes with the four integers 0, 1, 2, and 3 respectively to obtain four new nodes.

[0077] S3.2. Align the amplified sequences with each sequence in the integer sequence set one by one, and judge whether the matching position P[i], i ∈ 0, 1, 2…n is the length of the DNA series. If so, terminate the iteration and execute S8; otherwise, execute S3.3.

[0078] For example, if calculating the common supersequence of a sequence set such as {GTGGC, CGATGC, GCTAA}, the condition for terminating the iteration is that the matching positions of the nodes are P[0] = 4, P[1] = 5, and P[2] = 4.

[0079] S3.3. Judge whether the amplified node is a useful node, and move the matching position P[i], i ∈ 0, 1, 2…n to the end of the directed sequence. If so, the node is called a useful node and is added to the iterative node set; otherwise, the node is discarded. If the matching position P[i] does not change, the amplified node is discarded.

[0080] For example, if calculating the common supersequence of a sequence set such as {GTGGC, CGATGC, GCTAA}, and the sequence in the node is 0210, then the integer-encoded sequence set is {01002, 203102, 02133}, and the matching positions are P[0] = 2, P[1] = 1, and P[2] = 2.

[0081] S4. Calculate the guidance values of all nodes in the iterative node set.

[0082] Traverse each node in the iterative node set, call the probability database, and calculate the guidance value of the node. The specific calculation formula is as follows:

[0083]

[0084] Where: r i (x) is the subsequence of the sequence x in the node that matches a certain sequence s in the DNA sequence set. n is the number of sequences in the DNA sequence set. max is the index of the node with the largest matching position in the iterative node set. i

[0085] Since​ After n multiplications, the numerical guidance value is too small, and there is a possibility that the difference in the guidance value is less than the minimum precision that a double-precision real number can distinguish. Therefore, this method magnifies the guidance value and performs logarithmic processing on it. The calculation formula for the guidance value with logarithmic processing added is as follows:

[0086]

[0087] S5. Select the dominant node set from the iterative node set, and perform branch pruning according to the dominant node set; update the dominant node set, and judge whether the nodes in the iterative node set are dominated by it according to the matching positions of the nodes in the dominant node set; if they are dominated, discard the dominated nodes, that is, perform branch pruning, otherwise execute S6; it includes:

[0088] S5.1: Clear the dominant node set, and select the n_best nodes with the largest guidance values from the iterative node set to form the dominant node set. Among them, n_best should satisfy n_best < beamSize; where beamSize is the beam width, that is, the number of candidate node sets.

[0089] S5.2: Traverse all nodes in the iterative node set except the dominant node set.

[0090] S5.3: Traverse all dominant nodes. Compare the matching positions of the current node k and the dominant node j. If there is Then it is judged that node k is dominated by node j, and node k is discarded.

[0091] For example, if nodes k and j have sequences x k 、x j , there is Then it is said that node k is dominated by node j. In subsequent iterations, no matter how x k is amplified, its guidance value g k (x k ) is less than g k (x j ). In order to avoid redundant calculations, branch pruning is required to discard the dominated node k.

[0092] S6. Clear the candidate node set, and select a specified number of candidate nodes from the iterative node set in descending order of the guidance value, and add them to the candidate node set; it includes:

[0093] S6.1: Clear the candidate node set.

[0094] S6.2: Select the nodes with the largest guidance values from the nodes in the iterative node set according to the guidance value, and add the n guide nodes with the largest guidance values to the candidate node set. Among them, n guide ​guide It is determined by the beam width beamSize and the perturbation rate p. The three satisfy the following relationship:

[0095] n guide = beamSize * (1 - p)

[0096] Among them, in order to determine the value of p when the beneficial effect obtained by the present invention is the greatest, this method is used to implement the method multiple times for 10 DNA sequence sets with 50 sequences and a length of 100, and the following is obtained Figure 2 The 3D surface plot of the average common supersequence length calculated by the present invention shown in the figure with respect to the perturbation rate and the beam width. From Figure 2 It can be seen that when the beam width is set larger and larger, the perturbation rate with the shortest average common supersequence length approaches 0. Therefore, when calculating the synthesized sequence of the in-situ DNA microarray synthesis system, in the embodiment of the present invention, it is stipulated that when the beam width beamSize = 100, the perturbation rate p = 0.3.

[0097] S7. From the remaining iterative node set, randomly select n pertu nodes and add them to the candidate node set. Then enter the above S3 and enter the next round of iteration.

[0098] Among them

[0099] n pertu = beamSize * p

[0100] S8. Check the amplified integer sequence. If it is the common supersequence of the integer sequence set after encoding the sequence set, decode it according to the encoding rule to obtain a string of characters.

[0101] For example, for a sequence file storing the information of "GTGGC'\n'CGATGC'\n'GTGGC'\n″′, in an embodiment of the present invention, the decoding method is as follows:

[0102] 0 → ′G′

[0103] 1 → ′T′

[0104] 2 → ′C′

[0105] 3 → ′A′

[0106] As mentioned above, it is only a preferred embodiment of the present invention, so the scope of implementation of the present invention cannot be limited thereby. That is, equivalent changes and modifications made according to the scope of the present invention patent and the content of the specification should still fall within the scope covered by the present invention.

Claims

1. A method for calculating a synthesis sequence of an in-situ DNA microarray synthesis system, characterized in that: Including: S1. Obtain the path of the sequence file carrying the sequence set information, traverse each character element in the sequence file, and encode the character elements into integer elements to obtain an integer sequence set. S2. Construct a probability database and a probability index database. The S2 includes: S2.

1. Construct a probability database. The probability database stores the probability that the sequence y becomes a super-sequence of the sequence r. The probability calculation formula is as follows: Where: Pr(r < y) represents the probability that the sequence y becomes a super-sequence of the sequence r; q is the length of the sequence r; k is the length of the sequence y; |∑| is the number of types of all characters in the sequence file; r[2…q] represents the subsequence of the sequence r taken from the second character to the last character. S2.

2. Construct a probability index database. When constructing a probability index database, the mapping relationship between the index value and the maximum matching position is as follows: target = (int)maxR * ln|∑| / ln 2 Where: maxR is the maximum value of the matching positions of all nodes in the iterative node set; when constructing the database, maxR = 1, 2…, m, and m is the length of the longest sequence in the sequence set; target is the second-dimensional index of the index probability database, that is, the length of the sequence y. S3. Enter the iterative process, amplify the sequences in the nodes, update the matching positions of the amplified nodes, and update the iterative node set. It includes: S3.

1. Amplify the sequences of each node in the candidate node set, and add a total of |∑| integers 0, 1,…, |∑|-1 respectively; S3.

2. Compare the amplified sequences with each sequence in the integer sequence set one by one, and judge whether the matching position is the length of the DNA sequence. If so, terminate the iteration and execute S8, otherwise execute S3.3; S3.

3. Judge whether the amplified node is a useful node. If it is a useful node, add it to the iterative node set, otherwise discard the node. S4. Calculate the guiding values of all nodes in the iterative node set. S5. Select the dominant node set from the iterative node set, and perform branch pruning according to the dominant node set. S6. Clear the candidate node set, and select a specified number of candidate nodes from the iterative node set in descending order of the guiding values, and add them to the candidate node set. S7. Randomly select a specified number of candidate nodes from the iterative node set, add them to the candidate node set, and then execute S3. S8. Check the amplified integer sequence. If it is a common super-sequence of the integer sequence set after encoding the sequence set, decode it according to the encoding rule to obtain a string of characters.

2. The method for calculating a synthesized sequence of an in-situ DNA microarray synthesis system according to claim 1, wherein: The S1 includes: S1.

1. Read the string in the sequence set file according to the file path, and each line is a string of character sequences. S1.

2. Perform integer encoding on the characters in the sequence file. The encoding method is that the i-th character traversed from left to right and from top to bottom in the sequence file is encoded as i-1.

3. The synthesis sequence calculation method of an in-situ DNA microarray synthesis system according to claim 1, wherein: Each node is a custom structure, which contains an integer sequence x and the possibility g that this sequence becomes a common supersequence of the sequence set k (x), the positions where the sequence x matches each sequence in the sequence set, that is, a one-dimensional array P[n] of length n, where n is the number of sequences in the sequence set; the S3 includes: S3.

1. Traverse each node in the candidate node set, and amplify a total of |∑| integers 0, 1,…, |∑|-1 for the sequence in the node. S3.

2. Compare the amplified sequence with each sequence in the integer sequence set one by one, and determine whether the matching position P[i], where i ∈ 0, 1, 2…n, is the length of the DNA sequence. If so, terminate the iteration and execute S8; otherwise, execute S3.3; S3.

3. Determine whether the amplified node is a useful node. If so, the node is called a useful node and added to the iterative node set; otherwise, discard the node.

4. The synthesis sequence calculation method of an in-situ DNA microarray synthesis system according to claim 1, characterized in that: The said S4 includes: Traverse each node in the iterative node set, call the probability database, and calculate the guidance value of the node. The calculation formula is as follows: where: r i (x) is a subsequence of the sequence x in the node that matches a certain sequence s in the DNA sequence set i , n is the number of sequences in the DNA sequence set, and max is the index of the node with the largest matching position in the iterative node set; Perform amplification processing on the guidance value, and perform logarithmic processing on it. The calculation formula of the logarithmically processed guidance value is as follows:

5. The synthesis sequence calculation method of an in-situ DNA microarray synthesis system according to claim 4, characterized in that: In S5, if the sequences of nodes k and j are x k and x j respectively, and there exists Pr(r i (x k ) < s i ) ≤ Pr(r i (x j ) < s i ), then it is said that node k is dominated by node j. In subsequent iterations, no matter how x k expands, its guiding value g k (x k ) is less than g k (x j ), then branch pruning is performed to discard the dominated node k.

6. The synthesis sequence calculation method of an in-situ DNA microarray synthesis system according to claim 5, characterized in that: The method for determining whether a node is dominated is as follows: S5.

1. Select the nodes with the top n_best largest guidance values from the iterative node set to form a dominant node set, where: n_best satisfies n_best < beamSize, and beamSize is the beam width, that is, the number of candidate node sets; S5.

2. Traverse all nodes in the iterative node set except the dominant node set; S5.3, traverse all dominant nodes, compare the matching positions of the current node k and the dominant node j. If there exists P k [i] < P j [i], then determine that node k is dominated by node j and discard node k.

7. The synthesis sequence calculation method of an in-situ DNA microarray synthesis system according to claim 1, characterized in that: The said S6 includes: S6.

1. Clear the candidate node set; S6.

2. Select the top n nodes with the largest guidance values from the nodes in the iterative node set, and add the n guide nodes to the candidate node set. n guide is determined by the beam width beamSize and the perturbation rate p, and the three satisfy the following relationship: guide ​ n guide = beamSize * (1 - p) where: p = 0.

3.

8. The synthesis sequence calculation method of an in-situ DNA microarray synthesis system according to claim 1, characterized in that: In the S7, randomly select n pertu nodes from the remaining iterative node set, where n pertu = beamSize * p.

Citation Information

Patent Citations

  • Generating method of virtual mask for genetic chip in-situ synthesis

    CN102004833A

  • Code completion method based on structural features and sequence features

    CN114924741A