Capture probe optimization design method

By establishing a probe capture capability prediction model and optimizing probe design, the problem of being unable to a priori evaluate probe performance in existing technologies is solved, efficient probe combination optimization and target hit rate improvement are achieved, and design and detection costs are reduced.

CN120673846AInactive Publication Date: 2025-09-193D BIOMEDICINE SCI & TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202410311540.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies lack the ability to evaluate probe performance a priori, resulting in high probe design costs and the inability to effectively optimize the on-target rates of single probes and probe combinations. The on-target rates of probes cannot be predicted during design, and there is a lack of quantitative evaluation indicators for single probes, making fine-grained optimization impossible.

Method used

By establishing a probe capture ability prediction model, calculating the on-target rates of single probes and probe combinations, and using the training set to predict the probe capture ability, the probe design is optimized, including deleting or modifying probes that drag down the overall on-target rate, until the expected performance is achieved.

Benefits of technology

It enables the on-target rate of the probe combination to be evaluated during design, reduces design and testing costs, improves the actual on-target rate, avoids the waste of resources caused by repeated experiments, and enables optimization at the single probe level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673846A_ABST
    Figure CN120673846A_ABST
Patent Text Reader

Abstract

The invention relates to a method for optimizing the design of a capture probe. Specifically, the invention provides a method for predicting the target rate in a single probe and a probe combination in a liquid phase impurity capture scene by analyzing the nucleic acid sequence characteristics of the single probe, the probe combination and an input pool, so that the product design can be accelerated, the actual target rate of a liquid phase impurity capture high-throughput sequencing gene detection product can be effectively improved, and the method is suitable for popularization and application. And the design cost and the detection cost of the product are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of gene detection and specifically relates to a method for optimizing capture probe design. Background Art

[0002] High-throughput sequencing is an important tool for genomic research, playing a key role in clinical scenarios such as rare disease diagnosis and cancer drug guidance. Compared with whole-genome sequencing, enriching target regions for detection is a more cost-effective strategy, and liquid phase capture is one of the mainstream technical approaches to achieve this.

[0003] Because liquid-phase hybrid capture relies on the chemical principle of forming a double-stranded hybrid between the probe and target molecules through hydrogen bonding between bases, probes are typically designed to be complementary to the target sequence. The expectation is that non-target molecules introduced into the pool will remain in solution and not enter downstream experimental processes. However, this is not the case. First, the genome contains numerous homologous regions with highly similar or even identical sequences. A probe designed for one region will inevitably capture sequences from all other homologous regions. Second, binding ability is not a binary choice but a continuous process, ranging from strong to weak. Therefore, non-target molecules are captured by the probe to varying degrees. The capture of non-target molecules is called "off-target" and results in waste of materials and data. Controlling off-target effects or improving on-target rates while ensuring detection performance is a core objective in the design of liquid-phase hybrid capture high-throughput sequencing genetic testing products.

[0004] Currently, various methods for controlling off-target effects have been disclosed. Chinese Patent Publication No. CN106282352A discloses a method for achieving a balance between capture efficiency and coverage by avoiding designing probes over repetitive regions of the genome. However, completely avoiding designing probes over repetitive regions of the genome is sometimes not feasible, as repetitive regions may also have important biological functions and must be included in the product's coverage.

[0005] US Patent Publication No. US2010279883A1 discloses a method for generating a set of capture probes by analyzing the properties or parameters of the probes using a computer program.

[0006] Chinese Patent Publication No. CN109337956A discloses adding an additional sequence to a probe to form a stem-loop structure, thereby reducing affinity for non-target sequences. However, adding an additional sequence to the probe to form a stem-loop structure increases the probe length and the cost of probe synthesis. Furthermore, while reducing affinity for non-target regions, it also affects affinity for the target region.

[0007] These existing technologies judge whether the on-target rate has been improved "a posteriori" rather than "a priori". That is to say, although the designer adopts a probe design method that is "expected" to improve the on-target rate, the on-target rate of the probe cannot be predicted at the time of design. The improvement of the on-target rate can only be known by conducting a comparative experiment (comparing the actual on-target rate of the probe designed without using this method with the actual on-target rate of the probe designed with this method). Because probe synthesis brings actual costs, experiments without using this method are usually not conducted, but there is a rough expectation value for the actual on-target rate. If the experimental results do not meet expectations, no effective improvement strategy can be proposed, and one can only be forced to accept the performance that does not meet expectations. Therefore, the existing methods lack the ability to evaluate probe performance a priori and optimize probe sequences.

[0008] In addition, existing technologies can only improve the on-target rate of the probe combination as a whole, but cannot determine the contribution of individual probes that constitute the probe combination to the overall on-target rate, because they lack quantitative evaluation indicators for individual probes and lack the ability to perform fine-grained optimization at the level of individual probes. Summary of the Invention

[0009] The present invention establishes a method for predicting the on-target rates of single probes and probe combinations in a liquid-phase hybrid capture scenario by analyzing the characteristics of single probes, probe combinations, and nucleic acid sequences put into the pool. It also proposes a means to optimize probe design, which can accelerate product design, effectively improve the actual on-target rate of liquid-phase hybrid capture high-throughput sequencing gene detection products, and reduce product design and detection costs.

[0010] In one aspect, the present invention provides a method for optimizing probe design, which comprises: (1) calculating the on-target rate of a single probe in an initial probe combination and the on-target rate of the initial probe combination as a whole; (2) determining whether the on-target rate of the initial probe combination as a whole meets expectations; if not, (3) sequentially checking the single probe that has the greatest impact on the overall on-target rate; (4) after deleting, modifying, or retaining these probes, recalculating the on-target rate of the probe combination as a whole; (5) repeating steps (1), (2), (3), and (4) until the on-target rate of the probe combination as a whole meets expectations.

[0011] In another aspect, the present invention provides a method for predicting the on-target rate of a single probe, comprising: (1) using a training set, with the probe affinity score and the relative molecular number of the probe as independent variables and the relative capture ability of the probe as the dependent variable, to obtain a capture ability prediction model; (2) using the model to calculate the capture ability of the probe for each hit; and (3) predicting the on-target rate of a single probe based on the sequence information of the probe, the relative molecular number of the probe, the target region, the hit position, the similarity result, the capture ability, and whether it is on-target.

[0012] In another aspect, the present invention provides a method for predicting the on-target rate of a probe combination, the method comprising: (1) using a training set, with the probe affinity score and the relative molecular number of the probe as independent variables and the relative capture ability of the probe as the dependent variable, to obtain a capture ability prediction model; (2) using the model to calculate the capture ability of the probe for each hit; (3) predicting the on-target rate of the probe combination based on the sequence information of the probe, the relative molecular number of the probe, the target region, the hit position, the similarity result, the capture ability, whether it is on-target, and the product of the capture ability and the molar amount.

[0013] In another aspect, the present invention provides a method for determining the single probe that has the greatest impact on the overall on-target rate, the method comprising: (1) using a training set, with probe affinity scores and probe relative molecular numbers as independent variables, and probe relative capture capabilities as dependent variables, to obtain a capture capability prediction model; (2) using the model to calculate the capture capability of each probe for each hit; (3) predicting the degree of impact of a single probe on the on-target rate of a probe combination based on the probe's sequence information, probe relative molecular number, target region, hit position, similarity result, capture capability, whether it is on-target, and the product of capture capability and molar weight; (4) determining the probe with the largest single probe impact as the single probe that has the greatest impact on the overall on-target rate.

[0014] In another aspect, the present invention provides a method for establishing a probe capture ability prediction model, the method comprising: (1) obtaining a reference probe combination and its detection data; (2) obtaining an affinity score and a probe relative capture ability based on the probe sequence information, the genomic sequence corresponding to the fragment interval, and the number of genomic fragments; (3) constructing a training data set; and (4) using the training set, with the probe affinity score and the relative number of probe molecules as independent variables and the probe relative capture ability as the dependent variable, to obtain a capture ability prediction model.

[0015] In another aspect, the present invention provides a method for optimizing probe design, the method comprising:

[0016] - Establish a probe capture capacity prediction model: (1) obtain a reference probe combination and its detection data; (2) obtain affinity scores and probe relative capture capacities based on probe sequence information, the genomic sequence corresponding to the fragment interval, and the number of genomic fragments; (3) construct a training data set; (4) use the training set to obtain a capture capacity prediction model with probe affinity scores and the relative number of probe molecules as independent variables and probe relative capture capacities as dependent variables; (5) after adjusting the parameters of the nucleic acid sequence alignment tool, repeat the above steps (1)-(4) to obtain the optimal probe capture capacity prediction model as the final model, and record the alignment tool parameters used by the model;

[0017] - Calculate the on-target rate of a single probe in the initial probe combination and the on-target rate of the initial probe combination as a whole: (1) Calculate the capture capacity of each probe for each hit using the capture capacity prediction model and alignment tool parameters obtained above; (2) Predict the on-target rate of a single probe based on the probe sequence information, relative molecular number of probes, target region, hit position, similarity results, capture capacity, and whether it is on-target; (3) Predict the on-target rate of a probe combination based on the probe sequence information, relative molecular number of probes, target region, hit position, similarity results, capture capacity, whether it is on-target, and the product of capture capacity and molar weight;

[0018] - Optimize probe design: (1) Observe whether the on-target rate of the probe combination meets expectations. If not, check the single probe that has the greatest impact on the overall on-target rate in turn; (2) After deleting, modifying, or retaining these probes, recalculate the overall on-target rate of the probe combination; (3) Repeat steps (1) and (2) until the overall on-target rate of the probe combination meets expectations or is maximized.

[0019] In one embodiment, the method of the present invention is performed by a computer program. The computer program may include one or more modules for performing one or more steps of the method of the present invention. The computer program may be stored in a memory or a carrier, such as a magnetic tape, a disk, an optical disc, a ROM, a PROM, a VCD, a DVD or other computer-readable medium. Those skilled in the art will appreciate that computer programs themselves are within the conventional art of this area. It is within the capabilities of those skilled in the art to obtain and execute a computer program after reading this disclosure.

[0020] Therefore, in another aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the steps of the method of the present invention. In yet another aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by the processor, implements the steps of the method of the present invention. In yet another aspect, the present invention provides a computer program product comprising the computer program, wherein the computer program, when executed by the processor, implements the steps of the method of the present invention.

[0021] Compared with the prior art, the advantages of the present invention are:

[0022] (1) A calculation method for predicting the on-target rate of a single probe is proposed. Designers can use the information provided by this indicator to optimize the probe design at the level of a single probe.

[0023] (2) A calculation method for predicting the on-target rate of a probe combination is proposed, so that the overall capture performance of the product can be estimated in advance.

[0024] (3) A method for predicting the effectiveness of probe design is proposed. The method is reliable and the predicted on-target rate of the probe combination is highly consistent with the actual on-target rate. This method avoids the resource loss caused by repeated experiments (trial and error) in the traditional probe design process and enables product developers to have the ability to evaluate the product on-target rate before the experiment begins when designing probes, as well as the ability to optimize the probe combination at the level of a single probe. This allows developers to flexibly adjust the design plan, select or adjust the probe sequence without repeated experiments, and ultimately maximize the product on-target rate while ensuring the detection range. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 : Flowchart of the present invention scheme.

[0026] Figure 2 : Schematic diagram of NGS alignment results, showing the definition and relationship between genome and reads, genomic fragments, and fragment intervals.

[0027] Figure 3 : Relationship curve between the number of eliminated probes and the overall on-target rate of the remaining probe combinations. DETAILED DESCRIPTION

[0028] In order to fully understand the purpose, features and effects of the present invention, the present invention will be described in detail through the following specific embodiments. Unless otherwise specified, the technical terms involved in the present invention have the meanings commonly understood by those skilled in the art.

[0029] 1. Definition

[0030] Probe: A synthetic or semisynthetic single- or double-stranded oligonucleotide with a specific sequence and an affinity tag. Probes used for liquid-phase hybridization capture typically range from 60 to 300 base pairs in length. Whether single- or double-stranded, probes ultimately function as single strands during capture. The sequence of the probe is often artificially designed, and its sequence determines which nucleic acid molecules it will capture. In some contexts, probes are also figuratively referred to as "bait."

[0031] Input pool: A mixture of nucleic acid molecules (e.g., fragmented genomic DNA) from which probes capture target molecules. While the probe is called "bait," the nucleic acid molecules in the input pool are correspondingly called "prey," and the interaction between the probe and the input pool molecules is called "capture."

[0032] Template sequence: represents the sequence information of the nucleic acid molecule input into the pool, such as the human genome reference sequence hg19.

[0033] Target sequence: The sequence information of the nucleic acid molecule (target molecule) expected to be captured by the probe.

[0034] Target region: referred to as "target region", which is the coordinate interval of the template sequence corresponding to the target sequence.

[0035] Target molecule: A nucleic acid molecule that is desired to be captured by a probe and that at least partially overlaps with the target region at a position in the template sequence.

[0036] Liquid-phase hybridization capture (LPHC) is a process in which probe molecules and nucleic acid molecules in a pool form a double-stranded hybrid in a hybridization solution through hydrogen bonding caused by complementary base pairing. The double-stranded hybrid is then bound to a solid-phase matrix (e.g., magnetic beads with cross-linked streptavidin) that recognizes the probe's affinity tag (e.g., biotin), ultimately separating the nucleic acid molecules in the pool from the solution. The nucleic acid molecules in the pool that bind to the probe may be target molecules (which is the desired scenario) or non-target molecules (which is the undesirable scenario).

[0037] Single probe: A pure substance composed of probe molecules with identical sequences, used only to capture the target molecule of the probe. Probes for genetic testing products are generally not single probes.

[0038] Probe set: A mixture of multiple individual probes with different sequences, used to simultaneously capture multiple targets. Probes in genetic testing products are often combinations of probes. These probe sets are sometimes also called "panels."

[0039] High-throughput sequencing (HTS): also known as "massively parallel sequencing (MPS)" or "next-generation sequencing (NGS)", is a sequencing technology characterized by the ability to sequence hundreds of thousands to billions of nucleic acid molecules in parallel and short read lengths.

[0040] Effective sequencing data: The data remaining after removing duplicate reads and the overlapping regions of paired-end sequencing R1 / R2 from high-throughput sequencing data.

[0041] On-target rate: The probability that sequencing data falls within the target region and its flanking regions.

[0042] Actual on-target rate: The on-target rate calculated from the actual high-throughput sequencing data. Ideally, the actual on-target rate is typically at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100%.

[0043] Predicted on-target rate: The on-target rate theoretically estimated from the probe's sequence and other information, used to describe the probe's capture performance. The predicted on-target rate is typically at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100%.

[0044] Homologous regions: Regions on the genome with high sequence similarity are called homologous regions.

[0045] Relative molecule count: The relative molecule count of each probe in the probe set is the number of molecules of the probe divided by the minimum number of molecules of all probes in the probe set. If the number of molecules of each probe in the probe set is {n1,n2,…,n b}, then the relative molecular number of probe j

[0046] Genomic fragments and reads: Genomic fragments are fragments formed by the fragmentation of genomic DNA. They are the targets of probe capture and sequencing. In different sequencing modes, the single-end or double-end sequence of the genomic fragment is determined, and the sequence information obtained is called a read. Conversely, the paired-end read sequence can be used to determine the genomic coordinates corresponding to the read, and then the starting coordinates of the genomic location of the genomic fragment can be inferred.

[0047] Interval: The starting and ending intervals of the projections of several overlapping genomic fragments on the genome.

[0048] 2. Design Methods of Capture Probes

[0049] The following is a detailed description of the technical solution of the present invention, which is named SAPOTR (Sequence Alignment-based Probe On-Target Rate prediction). The following part uses concepts from relational algebra, which can be found in the reference Silberschatz, Abraham, Henry F. Korth, and S. Sudarshan. Database System Concepts. Seventh edition. New York, NY: McGraw-Hill, 2020. The following solution is presented in the form of a flow chart. Figure 1 .

[0050] The template sequence is denoted as G. G is a character string, generally a certain version (such as hg19) of the human genome sequence, but it can also be the genome sequence of other species or other non-genomic nucleic acid sequences.

[0051] The initial probe combination is denoted as relational table B. B contains b tuples, or b probes. The sequence information of each probe is denoted as s, and the relative number of molecules in the probe combination is m, i.e., B = prototype(s, m). The initial probe combination targets a portion of G and requires a predicted on-target rate. This represents the initial product probe design and often leaves room for optimization.

[0052] The reference probe combination is recorded as relational table A. A contains a tuples, i.e., a probes. The sequence information of each probe is recorded as S, and the relative number of molecules in the probe combination is M, i.e., A = reference(S, M). The reference probe combination captures a part of G and has been proven to have an ideal target rate for capturing G. The number of probes in A and B should be of the same order of magnitude, i.e.,

[0053] 1. Establish a model to predict probe capture ability

[0054] a) Randomly extract several test data of A from the public database or private historical data.

[0055] The public database refers to a publicly accessible database that stores high-throughput sequencing data and detailed experimental methods, especially probe information, including but not limited to NCBI's BioProject (https: / / www.ncbi.nlm.nih.gov / bioproject).

[0056] b) Merge the above test data and extract the position information of all genomic fragments from the genome alignment results of the merged data, merge them, and obtain several discontinuous fragment intervals (defined in Figure 2 ). Extract the genome sequence I corresponding to the fragment interval, and record the number of genome fragments F in the synthetic fragment interval to form the relationship table interval(I,F).

[0057] like Figure 2 As shown in Figure 1, during paired-end sequencing, genomic DNA is fragmented into genomic fragments, and sequence information is obtained at both ends of the genomic fragments. Based on this sequence information, the starting and ending positions of the genomic fragments on the genome are inferred. The coordinate information of multiple overlapping genomic fragments is combined to obtain a fragment interval. The starting coordinate of the fragment interval is the minimum value of the starting coordinates of these genomic fragments, and the ending coordinate of the fragment interval is the maximum value of the coordinates of these genomic fragments.

[0058] c) Assume the relational table capture(S,I,F)=π S (reference)×interval(I,F). Use the nucleic acid sequence alignment tool to calculate the similarity between S and I under the parameter table p(parameter,value). S may have multiple hits at different positions on I. The maximum value of the similarity results of all hits is represented by H, which is used as the affinity score between S and I. In addition, the number C of different S that hit I is counted, and the relationship table capture(S,I,F,H,C) is obtained. F is normalized and converted into relative capture ability E, And add the relative molecular number information of the probe, and then we get the relational table capture(S,I,F,H,C,E,M).

[0059] Nucleic acid sequence alignment tools are software tools used to perform nucleic acid alignments. The purpose of these tools is to compare the similarity between two or more nucleic acid sequences. Common nucleic acid sequence alignment tools include, but are not limited to, BLAT, BLAST, ClustalOmega, etc.

[0060] d) Construct a training data set (referred to as “training set”) T = Π E,M,H (σ C=1 (capture)).

[0061] e) Using the training set, with M and H as independent variables and E as the dependent variable, regression analysis and other methods are used to establish a prediction model of affinity score to relative capture ability under the alignment tool parameter p, that is, E = f p (M,H).

[0062] f) After adjusting the parameters of the nucleic acid sequence alignment tool (adjusting the tool parameter p), repeat the above steps to obtain multiple versions of prediction models. The best one is selected as the final probe capture ability prediction model, which is recorded as And record the parameters of the comparison tool used by the model, recorded as O.

[0063] The term "optimal" used here means that the relationship between M / H and E presented by the prediction model should conform to the principles of chemical kinetics. The more consistent the relationship, the better. For example, the relationship between M / H and E should be monotonic, meaning that the higher M, the higher E, and the higher H, the higher E.

[0064] 2. Calculate the on-target rate of a single probe in the initial probe combination and the on-target rate of the initial probe combination as a whole

[0065] a) Initial probe combination prototype (s, m), the target region of s is denoted as t. S and G are aligned using the above alignment tool under the setting of parameter set O, and a series of hits are obtained. The hit position is represented by g, and the corresponding similarity result is represented by h. The probe capture ability prediction model is used Calculate the capture ability of the probe for each hit, i.e. So we get the relationship alignment(s,m,t,g,h,e).

[0066] b) Comparing the information in t and g in alignment, we can easily determine whether the hit was on target and record it in field f. Convention: If it was on target, f = 1; if it was off target, f = 0. This gives the relationship alignment(s,m,t,g,h,e,f).

[0067] c) Define the predicted on-target rate of a single probe: For any tuple (s i ,m i ), which predicts the on-target rate So we get the relation prototype(s,m,r).

[0068] d) Expressing the product of probe capture capacity and molar mass in em, we obtain the relationship alignment(s,m,t,g,h,e,f,em), and the predicted on-target rate of the probe combination is

[0069] 3. Optimize probe design

[0070] a) Observe whether R meets expectations. If not, check the probes that have the heaviest impact on the overall on-target rate of the relationship prototype (s, m, r) in turn, that is, the impact of the single probe on the on-target rate of the probe combination according to the value of the drag degree. Arrange from large to small and check in sequence: (1) If the area covered by the probe can be discarded according to actual needs, then directly remove the probe, and then recalculate the overall on-target rate of the remaining probe combination according to the method 2-d above; or (2) If the area covered by the probe is indispensable according to actual needs, then modify the probe sequence, and then calculate the overall on-target rate of the new probe combination according to the methods 2-a to 2-d above; or (3) If the area covered by the probe is indispensable according to actual needs, and multiple modifications to the probe sequence do not improve the situation, then retain the probe, check the next probe in the sorting, and then calculate the overall on-target rate of the new probe combination according to the methods 2-a to 2-d above. Repeat the above steps until R meets the design expectations (or is maximized).

[0071] b) Order probes and conduct experimental verification to observe the consistency between the actual on-target rate and the predicted on-target rate.

[0072] The present invention is further illustrated by way of examples below, but the present invention is not limited to the scope of the examples. Experimental methods in the following examples where specific conditions are not specified were performed according to conventional methods and conditions, or selected according to the product specifications.

[0073] Example 1

[0074] 1. Select probe set P1 as the reference probe set. Specifically, after the target region is identified, a reference genome positive strand sequence covering the target region is randomly selected based on the genomic coordinates of the target region as the probe sequence, i.e., P1. P1 contains a total of 58,901 probes, designed for the human reference genome hg19, for tumor gene detection. The molar quantity of each probe in P1 is equal.

[0075] 2. Four detection data were selected from private historical data to establish a probe capture capability prediction model, as shown in Table 1.

[0076]

[0077] Table 1: Sequencing data information used to construct the probe capture capacity prediction model.

[0078] sample_id: sample ID. sex: gender. data_size: sequencing data size.

[0079] target_dup_rate: target region duplication rate. on_target_probe: ratio of valid data falling within the probe region. on_target_flanking: ratio of valid data falling within the probe flanking region. off_target: off-target rate.

[0080] 3. Combine the above data to form the relationship interval(I,F). Combine the probe sequence information of P1 to obtain the relationship alignment(S,I,F). BLAT was selected as the nucleic acid alignment tool (see Kent, W. James. “BLAT—the BLAST-like Alignment Tool.” Genome Research 12, no. 4 (April 2002): 656–64. https: / / doi.org / 10.1101 / gr.229202). Some of its parameter combinations are shown in Table 2.

[0081] parID tileSize stepSize minMatch minScore minIdentity maxGap repMatch M000 11 11 2 30 90 2 1024 M001 12 12 2 30 90 2 256 M002 11 5 2 30 90 2 2253 M003 11 11 3 30 90 2 1024 M004 11 11 2 20 90 2 1024

[0082] Table 2: Some parameter combinations of BLAT. parID: parameter combination number.

[0083] tileSize,stepSize,minMatch,minScore,minIdentity,maxGap,

[0084] repMatch: See the BLAT documentation (Blat Suite Program Specifications and UserGuide.

[0085] https: / / www.genome.ucsc.edu / goldenPath / help / blatSpec.html ).

[0086] 4. Under various BLAT parameter combinations, the S and I sets were aligned, and the training set was extracted from the alignment results to obtain the capture ability prediction model.

[0087] 5. The optimal model obtained from the above results is Record the optimal parameters.

[0088] 6. Using a similar method to the reference probe set, the initial probe set P2 was obtained, which contained 54,214 probes with equal molar weights. The probe sequences were aligned with the human genome sequence using BLAT under optimal parameters, and the results were substituted into the above function. Therefore, the on-target rate r of each probe and the overall on-target rate R of P2 can be obtained (R=64%, and the median of the actual on-target rate of P2 measured in subsequent experiments is 68%, which are very close).

[0089] 7. The above on-target rates do not meet expectations, so the on-target rates of all probes are arranged from small to large and removed from P2 one by one. The relationship curve between the number of removed probes and the overall on-target rate of the remaining probes is obtained, as shown in the figure: Figure 3 As shown, it can be seen that removing the 1805 probes with the smallest r value can ensure that the R of the remaining probes is 80%.

[0090] 8. After removing these probes, the probes were resynthesized for experimental testing, and the median value of the actual on-target rate was 77%, which is very close to 80%.

[0091] The above embodiments are preferred implementations of the present invention, but the implementation of the present invention is not limited to the above embodiments. Any other substitutions, modifications, combinations, changes, simplifications, etc. that do not deviate from the spirit and principles of the present invention should be considered equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A method for optimizing probe design, characterized in that: The method comprises: (1) calculating the on-target rate of a single probe in the initial probe combination and the on-target rate of the initial probe combination as a whole; (2) determining whether the on-target rate of the initial probe combination as a whole meets expectations; if not, (3) sequentially checking the single probes that have the greatest impact on the overall on-target rate; (4) deleting, modifying, or retaining these probes, and then recalculating the on-target rate of the probe combination as a whole; (5) repeating steps (1), (2), (3), and (4) until the on-target rate of the probe combination as a whole meets expectations.

2. The method according to claim 1, characterized in that The probes are designed for liquid phase hybridization capture.

3. The method according to claim 1, characterized in that The overall on-target rate of the probe combination that meets the expectations is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99% or 100%.

4. The method according to claim 1, wherein Calculating the on-target rate of a single probe in the initial probe combination includes: (1) using the training set, with the probe affinity score and the relative molecular number of the probe as independent variables and the relative capture ability of the probe as the dependent variable, to obtain a capture ability prediction model; (2) using the model to calculate the capture ability of the probe for each hit; (3) predicting the on-target rate of a single probe based on the sequence information of the probe, the relative molecular number of the probe, the target region, the hit position, the similarity result, the capture ability, and whether it is on-target.

5. The method according to claim 1, wherein Calculating the on-target rate of the initial probe combination as a whole includes: (1) using the training set, with the probe affinity score and the relative molecular number of the probe as independent variables and the relative capture ability of the probe as the dependent variable, to obtain a capture ability prediction model; (2) using the model to calculate the capture ability of the probe for each hit; (3) predicting the on-target rate of the probe combination based on the sequence information of the probe, the relative molecular number of the probe, the target region, the hit position, the similarity result, the capture ability, whether it is on-target, and the product of the capture ability and the molar weight.

6. The method according to claim 1, characterized in that Determining the single probe that has the greatest impact on the overall on-target rate includes: (1) using a training set, with probe affinity score and probe relative molecular number as independent variables, and probe relative capture ability as dependent variable, to obtain a capture ability prediction model; (2) using the model to calculate the capture ability of each probe for each hit; (3) based on the probe's sequence information, probe relative molecular number, target region, hit position, similarity result, capture ability, whether it is on-target, and the product of capture ability and molar weight, predicting the degree of impact of a single probe on the on-target rate of the probe combination; (4) determining the probe with the largest single probe drag value as the single probe that has the greatest impact on the overall on-target rate.

7. The method according to claim 4, 5 or 6, characterized in that The method further includes establishing a probe capture ability prediction model by the following steps: (1) obtaining a reference probe combination and its detection data; (2) obtaining a probe affinity score and a probe relative capture ability based on the probe sequence information, the genomic sequence corresponding to the fragment interval, and the number of genomic fragments; (3) constructing a training data set; and (4) using the training set, with the probe affinity score and the relative number of probe molecules as independent variables and the probe relative capture ability as the dependent variable, to obtain a capture ability prediction model.

8. The method according to claim 7, characterized in that The method further comprises adjusting the parameters of the nucleic acid sequence alignment tool, repeating steps (1) to (4), obtaining an optimal probe capture capability prediction model, and recording the alignment tool parameters used by the model.

9. The method according to claim 1, characterized in that After determining the single probe that has the greatest impact on the overall on-target rate, if the area covered by the probe can be discarded, the probe is directly removed, and the overall on-target rate of the remaining probe combination is recalculated; or if the area covered by the probe is indispensable, the probe sequence is modified, and the overall on-target rate of the new probe combination is calculated; or if the area covered by the probe is indispensable and multiple modifications to the probe sequence cannot improve the on-target rate, the probe sequence is retained, and the single probe that has the second greatest impact on the overall on-target rate is processed instead, and the overall on-target rate of the new probe combination is recalculated.

10. The method according to claim 1, characterized in that The method also includes synthesizing probes based on the expected probe design and performing experimental verification to determine the consistency between the actual on-target rate of the entire probe combination and the predicted on-target rate.

11. The method according to claim 10, characterized in that The actual on-target rate is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99% or 100%.

12. A method for predicting the on-target rate of a single probe, characterized in that: The method comprises: (1) using a training set, with probe affinity scores and probe relative molecular numbers as independent variables and probe relative capture capabilities as dependent variables, to obtain a capture capability prediction model; (2) using the model to calculate the capture capability of each probe for each hit; and (3) predicting the on-target rate of a single probe based on the probe's sequence information, probe relative molecular number, target region, hit position, similarity result, capture capability, and whether it is on-target.

13. A method for predicting the on-target rate of a probe combination, characterized in that: The method comprises: (1) using a training set, with probe affinity scores and probe relative molecular numbers as independent variables and probe relative capture capabilities as dependent variables, to obtain a capture capability prediction model; (2) using the model to calculate the capture capability of each probe for each hit; and (3) predicting the on-target rate of a probe combination based on the sequence information of the probe, the relative molecular number of the probe, the target region, the hit position, the similarity result, the capture capability, whether the probe is on-target, and the product of the capture capability and the molar weight.

14. A method for determining the single probe that has the greatest impact on the overall on-target rate, characterized in that: The method comprises: (1) using a training set, with probe affinity scores and probe relative molecular numbers as independent variables and probe relative capture capabilities as dependent variables, to obtain a capture capability prediction model; (2) using the model to calculate the capture capability of each probe for each hit; (3) based on the sequence information of the probe, the relative molecular number of the probe, the target region, the hit position, the similarity result, the capture capability, whether the probe is on target, and the product of the capture capability and the molar amount, predicting the degree of drag on the on-target rate of the probe combination by a single probe; (4) determining the probe with the largest single probe drag value as the single probe that has the heaviest drag on the overall on-target rate.

15. A method for establishing a probe capture ability prediction model, characterized in that: The method comprises: (1) obtaining a reference probe combination and its detection data; (2) obtaining a probe affinity score and a probe relative capture ability based on probe sequence information, the genomic sequence corresponding to the fragment interval, and the number of genomic fragments; (3) constructing a training data set; and (4) using the training set, with the probe affinity score and the relative number of probe molecules as independent variables and the probe relative capture ability as the dependent variable, to obtain a capture ability prediction model.

16. The method according to claim 15, characterized in that The reference probe combinations are obtained from public databases or private historical data.

17. A method for optimizing probe design, characterized in that: The method comprises: - Establish a probe capture capacity prediction model: (1) obtain a reference probe combination and its detection data; (2) obtain the probe affinity score and the probe relative capture capacity based on the probe sequence information, the genomic sequence corresponding to the fragment interval, and the number of genomic fragments; (3) construct a training data set; (4) use the training data set, with the probe affinity score and the relative number of probe molecules as independent variables and the probe relative capture capacity as the dependent variable, to obtain a capture capacity prediction model; (5) after adjusting the parameters of the nucleic acid sequence alignment tool, repeat the above steps (1)-(4) to obtain the optimal probe capture capacity prediction model as the final model, and record the alignment tool parameters used by the model; - Calculate the on-target rate of a single probe in the initial probe combination and the on-target rate of the initial probe combination as a whole: (1) Calculate the capture capacity of each probe for each hit using the capture capacity prediction model and alignment tool parameters obtained above; (2) Predict the on-target rate of a single probe based on the probe sequence information, relative molecular number of probes, target region, hit position, similarity results, capture capacity, and whether it is on-target; (3) Predict the on-target rate of a probe combination based on the probe sequence information, relative molecular number of probes, target region, hit position, similarity results, capture capacity, whether it is on-target, and the product of capture capacity and molar weight; and - Optimize probe design: (1) Observe whether the on-target rate of the probe combination meets expectations. If not, check the single probe that has the greatest impact on the overall on-target rate in turn; (2) Delete, modify, or retain these probes and recalculate the overall on-target rate of the probe combination; (3) Repeat steps (1) and (2) until the overall on-target rate of the probe combination meets expectations or is maximized.

18. The method according to any one of claims 1 to 17, characterized in that The method is performed by a computer program.

19. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 17.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 17 are implemented.

21. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 17 are implemented.

Citation Information

Patent Citations

  • Method for evaluating off-target risk of hybrid capture probe

    CN115101128A

  • Targeting sequence capture probe design strategy selection method and system for next-generation sequencing and terminal

    CN115713971A

  • Method and device for evaluating capture safety of probe in repeat region of genome

    CN116206684A

  • Method and system for screening siRNA sequences to reduce off-target effect

    CN116798513A

  • Integrated systems and methods for automated processing and analysis of biological samples, clinical information processing and clinical trial matching

    TW201816645A