An adaptive phage panning method based on real-time affinity monitoring
By monitoring phage affinity distribution data in real time and generating decision control instructions, the problem of insufficient or excessive screening in phage screening methods is solved, intelligent control is achieved, and the efficiency and quality of antibody discovery are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CAMOD BIOLOGICAL TECH TIANJIN CO LTD
- Filing Date
- 2025-11-27
- Publication Date
- 2026-04-14
AI Technical Summary
Existing phage screening methods cannot precisely control the screening rounds, which can easily lead to insufficient or excessive screening, wasting time and resources, and may also cause high-affinity clones to be eliminated.
By monitoring the affinity distribution data of bacteriophages in real time, decision control instructions are generated, including instructions to continue, terminate early, or optimize and adjust, and screening parameters are dynamically adjusted to achieve intelligent control.
It significantly shortens the antibody discovery cycle, improves R&D efficiency, enhances the quality and diversity of the final antibody library, and avoids the problems of insufficient or excessive screening.
Smart Images

Figure CN121211047B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biological data processing technology, specifically to an adaptive phage screening method based on real-time affinity monitoring. Background Technology
[0002] Phage display technology is a powerful in vitro screening technique widely used in the discovery and optimization of antibody drugs. Its core technical process involves multiple rounds of adsorption-elution-amplification to enrich clones that can bind to specific targets with high affinity from a large antibody library.
[0003] Existing phage screening methods, such as the preparation of a human CD56 immune phage display antibody library disclosed in publication number CN120209145A, represent a widely used technical approach. This method typically relies on a pre-defined, fixed-round screening process (e.g., 3-5 rounds), forcing researchers to rely solely on past experience to determine the total number of screening rounds. This easily leads to two disadvantages:
[0004] Insufficient screening: Prematurely ending the process before sufficient enrichment of optimal high-affinity clones can lead to screening failure or unsatisfactory results.
[0005] Over-screening: Continuing to perform unnecessary screening rounds after obtaining ideal clones. This not only wastes valuable R&D time (each screening round typically takes 1-2 weeks) and expensive reagent costs, but more seriously, it may cause some high-affinity but slow-growing clones to be masked or even eliminated in subsequent rounds by rapidly growing but slightly less affinity clones, resulting in "clone loss". Summary of the Invention
[0006] In view of the above-mentioned defects or deficiencies in the prior art, this application aims to provide an adaptive phage screening method based on real-time affinity monitoring for screening high-affinity antibodies; in order to control the biological process by processing biological data, thereby achieving real-time monitoring of the health status and enrichment effect of the screening process.
[0007] The method includes the following steps:
[0008] Perform the Nth round of selection, with N initially set to 1;
[0009] After the Nth round of screening, the following actions are performed in parallel: amplification of the eluted phages to obtain the Nth round enrichment library; and sampling of the eluted phages and high-throughput affinity testing of the sample solution to generate Nth round affinity distribution data.
[0010] The affinity distribution data of the Nth round is compared with the reference data, and control instructions are generated according to preset decision conditions. The control instructions include a continue instruction, an early termination instruction, or an optimization adjustment instruction.
[0011] Execute the control command:
[0012] If the control instruction is the continue instruction, then N is incremented, and the Nth round enrichment library is used as input for the next round of selection;
[0013] If the control instruction is the early termination instruction, then the selection process is terminated, and the Nth round enriched library is output as the final library.
[0014] If the control instruction is the optimization adjustment instruction, then after adjusting the screening parameters, N is incremented, and the Nth round enrichment library is used as input for the next round of screening.
[0015] According to the technical solution provided in this application, the step of performing high-throughput affinity testing on the sample liquid to generate Nth round affinity distribution data includes the following steps:
[0016] Phage DNA was extracted from the sample solution and subjected to high-throughput sequencing to obtain a sequencing data set containing multiple sequence reads; wherein each sequence read contains a preset DNA barcode, which was uniquely assigned to different antibody clones during the initial phage library construction;
[0017] Based on the sequencing data set, identify and count the number of times each preset DNA barcode sequence appears in the set;
[0018] The preset barcode-clone correspondence database is retrieved, and each preset DNA barcode sequence is mapped to the specific antibody clone it represents to obtain the analysis result;
[0019] Based on the analysis results, the relative frequency of each antibody clone in the sequencing dataset and / or its enrichment relative to historical rounds are calculated to generate the Nth round affinity distribution data.
[0020] According to the technical solution provided in this application, the preset decision conditions include stability conditions for generating the early termination instruction, which are configured to be triggered in the following manner:
[0021] From the Nth round of affinity distribution data, the antibody clones with the highest enrichment ranking among the top K are identified, forming the first clone set;
[0022] Calculate the total frequency of the first clone set in the Nth round of sequencing data set, and denote it as the first total frequency;
[0023] If the first total frequency exceeds the first preset threshold; or, the ratio of the first total frequency to the second total frequency is less than or equal to a second preset threshold close to 1, then the early termination instruction is generated: the second total frequency is the total frequency of the first clone set in the N-1 round of sequencing data set.
[0024] According to the technical solution provided in this application, the optimization and adjustment instruction is executed as follows: in subsequent screening rounds, competitive elution is added; the competitive elution includes the following steps:
[0025] After the phage binds to the immobilized target and before specific elution, a certain concentration of soluble target molecules is added to the system for competitive incubation.
[0026] The concentration of the soluble target molecule is dynamically adjusted based on the enrichment degree of non-specific clones in the affinity distribution data.
[0027] According to the technical solution provided in this application, before identifying the antibody clones with the highest enrichment in the top K positions from the Nth round of affinity distribution data to form the first clone set, the method further includes determining the K value.
[0028] Determining the value of K includes the following steps:
[0029] The antibody clones are sorted from high to low according to their enrichment, and an initial enrichment-ranking curve is plotted. The total number of all antibody clones before the inflection point with the largest slope change in the initial enrichment-ranking curve is taken as the K value.
[0030] The process of constructing the first clone set includes the following steps:
[0031] The first clone set consists of all antibody clones preceding the inflection point with the largest change in slope in the initial enrichment-ranking curve.
[0032] According to the technical solution provided in this application, after plotting the enrichment-ranking curve, the following steps are also included:
[0033] Calculate the sequence diversity index corresponding to the antibody clone that is ranked in the top P position in the initial enrichment-ranking curve;
[0034] The step of using the total number of all antibody clones before the inflection point with the largest slope change in the initial enrichment-ranking curve as the K value includes the following steps:
[0035] If the sequence diversity index is higher than or equal to the preset diversity threshold, then the total number of all antibody clones before the inflection point with the largest slope change in the initial enrichment-ranking curve is taken as the K value.
[0036] According to the technical solution provided in this application, after calculating the sequence diversity index corresponding to the antibody clone located in the top P position in the initial enrichment-ranking curve, the method further includes the following steps:
[0037] If the sequence diversity index is lower than the preset diversity threshold, then starting from the antibody clone with the highest current enrichment ranking, the process is traversed downwards sequentially, and based on its amino acid sequence, the traversed antibody clones are clustered using a preset sequence similarity threshold; wherein, during the clustering process, when the sequence similarity of a newly traversed antibody clone with all antibody clones in any existing cluster is lower than the preset sequence similarity threshold, it is classified into a new independent cluster;
[0038] The number of independent clusters formed is monitored in real time, and the traversal stops when the number of independent clusters reaches the preset maximum number of clusters for the first time.
[0039] All antibody clones that have been traversed at this point and belong to different clusters are defined as the candidate clone group;
[0040] A second enrichment-ranking curve is plotted based on the candidate clonal population, and the first clonal set is composed of all antibody clones before the inflection point with the largest slope change in the second enrichment-ranking curve.
[0041] According to the technical solution provided in this application, the preset decision conditions further include potential prediction conditions for generating the continuation instruction: the potential prediction conditions are configured as follows:
[0042] When the triggering conditions of the early termination instruction are not met, an enrichment growth trend model is constructed based on the affinity distribution data of the Nth round and multiple historical rounds.
[0043] If the enrichment growth trend model predicts that the first total frequency may exceed the first preset threshold in the next 1 to 2 rounds, then the continue instruction is generated; otherwise, the optimization adjustment instruction for adjusting the screening strategy is generated.
[0044] According to the technical solution provided in this application, the first preset threshold and / or the second preset threshold are dynamically set through a threshold optimization model.
[0045] According to the technical solution provided in this application, the method further includes constructing a threshold optimization model, comprising the following steps:
[0046] Model training phase: After system initialization or completion of a specific round of screening, the historical database of historical screening items is retrieved; the historical database stores the affinity distribution data sequence of each historical screening item in the multi-round screening process, the final decision round, and the information of the high affinity antibody sequence finally obtained;
[0047] Feature extraction and target definition: The first total frequency and the ratio of the first total frequency to the second total frequency generated in each round of the historical screening projects are used as input features; the probability that the final antibody library can be included if the screening is terminated in this round is used as the optimization target.
[0048] Model training: Based on machine learning regression algorithms, the input features are fitted and trained with the optimization objective to obtain the threshold optimization model.
[0049] Compared with the prior art, the beneficial effects of this application are as follows:
[0050] I. Achieved intelligent and precise control of the screening process: This invention upgrades the screening process from experience-based blind screening to data-driven intelligent screening by introducing a closed-loop feedback mechanism of real-time affinity assessment and decision control commands. The system can automatically judge the screening status based on objective data and make optimal decisions, completely avoiding the problems of insufficient or excessive screening.
[0051] II. Significantly shortens the antibody discovery cycle and improves R&D efficiency: By introducing an early termination command, once the system determines that high-affinity clones have been sufficiently enriched (i.e., reached a stable state), subsequent unnecessary screening rounds can be terminated immediately. Compared to the traditional method of performing 3-5 rounds, this invention can save an average of 1-2 rounds of screening time, equivalent to shortening the entire antibody discovery process by 1-2 weeks, while simultaneously saving corresponding manpower and reagent costs.
[0052] III. Effectively improve the quality and diversity of the final antibody library: By optimizing and adjusting instructions, the system can dynamically adjust screening parameters (such as introducing competitive elution) when non-specific binding is detected or screening pressure is insufficient, thereby actively optimizing the screening path, more effectively enriching true high-affinity clones, and helping to maintain clone diversity and avoid the overgrowth of dominant clones. Attached Figure Description
[0053] Figure 1 The flowchart illustrates the steps of the adaptive phage screening method based on real-time affinity monitoring provided in this application. Detailed Implementation
[0054] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0055] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0056] Example 1
[0057] As mentioned in the background section, to address the problems in the prior art, this application proposes an adaptive phage screening method based on real-time affinity monitoring for screening high-affinity antibodies; such as Figure 1 As shown, the method includes the following steps:
[0058] S1. Perform the Nth round of selection, with N initially set to 1;
[0059] S2. After the Nth round of screening, the following actions are performed in parallel: amplification of the eluted phages to obtain the Nth round enrichment library; and sampling of the eluted phages and high-throughput affinity testing of the sample solution to generate Nth round affinity distribution data.
[0060] S3. Compare the Nth round affinity distribution data with the reference data, and generate control instructions based on preset decision conditions. The control instructions include a continue instruction, an early termination instruction, or an optimization adjustment instruction.
[0061] S4. Execute the control command:
[0062] S41. If the control instruction is the continue instruction, then let N increment and use the Nth round enrichment library as input for the next round of selection;
[0063] S42. If the control instruction is the early termination instruction, then terminate the selection and output the Nth round enriched library as the final library.
[0064] S43. If the control instruction is the optimization adjustment instruction, then after adjusting the screening parameters, let N increase, and use the Nth round enrichment library as input for the next round of screening.
[0065] Specifically, first, a first round (N=1) of routine phage panning is performed, which involves binding the phage library to an immobilized target protein (e.g., CD56), washing away unbound phages, and then eluting the specifically bound phages. After each round of panning, two independent follow-up processes are immediately initiated in parallel:
[0066] Amplification Procedure: A majority (e.g., 99%) of the eluted phages are mixed with a logarithmic growth phase E. coli TGI bacterial culture, followed by infection and amplification culture to prepare a phage library for potential next round of screening (Nth round enrichment library). This procedure follows standard molecular cloning procedures.
[0067] Monitoring Procedure: Simultaneously, a small portion (e.g., 1%) of the same eluent is taken and, without amplification, directly enters the high-throughput affinity testing process. The purpose of this process is to quickly generate a quantitative report on the phage population binding characteristics in the current eluent (Nth round affinity distribution data).
[0068] Next, the system (usually a decision-making algorithm executed by a computer) compares and analyzes the affinity distribution data generated in this round with the corresponding data from the previous round (round N-1). The core of the analysis is to determine whether the current screening state meets the preset decision conditions.
[0069] Decision-making and execution are key aspects of the process:
[0070] Continue command: When analysis indicates that high-affinity clones are being effectively enriched but have not yet reached the optimal concentration or steady state, the system automatically issues a continue command. At this time, the N value is increased by 1 (N=N+1), and the Nth round enrichment library prepared in the amplification process is used as input to automatically start the next round of screening.
[0071] Early termination command: When analysis indicates that high-affinity clones have become dominant in the library and growth has plateaued, the system issues a termination command. The process stops immediately, and the current Nth round of enrichment library is output as the final result, thus saving unnecessary subsequent screening rounds.
[0072] Optimization and Adjustment Commands: When analysis detects problems, such as a large enrichment of non-specific clones, the system issues adjustment commands. Before starting the next round (N=N+1), the system automatically modifies the screening parameters according to preset rules (e.g., increasing the number of washes, adding competing agents, etc.), and then performs optimized screening using the Nth round enrichment library as input.
[0073] This implementation method introduces a real-time feedback closed loop of "monitoring-decision-execution," transforming the traditional open-loop, fixed-round screening process into an intelligent, adaptive one. Its technical principle is similar to an automatic control system, dynamically adjusting control signals (screening instructions and parameters) by continuously monitoring the system output (affinity distribution), ensuring that the system output (final antibody library) quickly and stably approaches the set target (high-affinity antibodies). This fundamentally solves the problem of insufficient or excessive screening caused by the fixed rounds in traditional methods, significantly improving screening efficiency and success rate.
[0074] In a preferred embodiment, the step of performing high-throughput affinity testing on the sample liquid to generate Nth round affinity distribution data includes the following steps:
[0075] Phage DNA was extracted from the sample solution and subjected to high-throughput sequencing to obtain a sequencing data set containing multiple sequence reads; wherein each sequence read contains a preset DNA barcode, which was uniquely assigned to different antibody clones during the initial phage library construction;
[0076] Based on the sequencing data set, identify and count the number of times each preset DNA barcode sequence appears in the set;
[0077] The preset barcode-clone correspondence database is retrieved, and each preset DNA barcode sequence is mapped to the specific antibody clone it represents to obtain the analysis result;
[0078] Based on the analysis results, the relative frequency of each antibody clone in the sequencing dataset and / or its enrichment relative to historical rounds are calculated to generate the Nth round affinity distribution data.
[0079] Specifically, first, phage genomic DNA is extracted from the sampling fluid used for monitoring. This DNA contains antibody gene information carried by all successfully eluted phages, as well as unique DNA barcode sequences that were pre-inserted during library construction and replicated along with the phages (each barcode uniquely corresponds to one antibody clone).
[0080] Next, using this DNA as a template, PCR amplification was performed with primers to specifically enrich the sequence regions containing the DNA barcode. The amplification products were then sent to a high-throughput sequencing platform such as Illumina or MGI for sequencing. After the sequencing run was completed, a raw data file containing millions to hundreds of millions of short sequence reads was obtained.
[0081] The sequencing data is then parsed using bioinformatics scripts (e.g., written in Python). The script iterates through all sequence reads, identifies each unique DNA barcode sequence, and counts its frequency in the total data. This frequency directly reflects the relative abundance of the clone represented by that barcode in the elution buffer.
[0082] Next, the barcode-clone mapping database established during the initial library construction is retrieved. This is a spreadsheet or database file that records a one-to-one correspondence between each DNA barcode sequence and the antibody clone it represents (defined by its variable region amino acid sequence or gene sequence). By querying the database, the barcode sequences obtained from sequencing are mapped to specific antibody clone identities.
[0083] Finally, calculations are performed based on the above analysis results:
[0084] Calculate the relative frequency: For each specific antibody clone, its relative frequency = (the sequence reads of that clone / the total sequence reads of all clones) × 100%. This gives the proportion of that clone in the current round.
[0085] Enrichment is calculated as follows: For each specific antibody clone, its enrichment = (relative frequency of the clone in round N) / (relative frequency of the clone in round N-1). This yields the fold increase of the clone after one round of screening.
[0086] Finally, the relative frequency and / or enrichment data of all clones are organized into a structured list or graph, which generates the Nth round of affinity distribution data for decision-making.
[0087] In a preferred embodiment, the preset decision condition includes a stability condition for generating the early termination instruction, which is configured to be triggered in the following manner:
[0088] From the Nth round of affinity distribution data, the antibody clones with the highest enrichment ranking among the top K are identified, forming the first clone set;
[0089] Calculate the total frequency of the first clone set in the Nth round of sequencing data set, and denote it as the first total frequency;
[0090] If the first total frequency exceeds the first preset threshold; or, the ratio of the first total frequency to the second total frequency is less than or equal to a second preset threshold close to 1, then the early termination instruction is generated: the second total frequency is the total frequency of the first clone set in the N-1 round of sequencing data set.
[0091] Specifically, after obtaining the affinity distribution data for round N, the system first extracts the enrichment data of all clones and sorts them from highest to lowest enrichment. Then, the system selects the top K antibody clones (e.g., K=20) and defines them as a first clone set. This set represents the best-performing and most promising candidate group in the current round.
[0092] Next, the system performs key calculations:
[0093] The sum of the frequencies of all clones in the first clonal set in the Nth round of sequencing data is calculated and denoted as the first total frequency (FN). This value represents the collective strength share of the elite clonal population in the entire phage library.
[0094] Retrieve data from round N-1, calculate the sum of frequencies of the same first clone set (i.e., the K identical clones) in round N-1, and denote it as the second total frequency (FN-1). This value represents their collective strength in the previous round.
[0095] Subsequently, the system executes a judgment logic, which contains two independent triggering conditions. Meeting either condition will generate an early termination instruction:
[0096] Condition 1 (Strength Threshold Judgment): The system checks whether the first total frequency (FN) exceeds a preset first threshold (e.g., T1 = 70%). If FN > T1, termination is triggered. This indicates that the elite group has gained an absolute advantage, and the selection target has been basically achieved.
[0097] Condition 2 (Growth Stability Judgment): The system calculates the ratio of the first total frequency (FN / FN-1) and checks whether this ratio is less than or equal to a second threshold close to 1 (e.g., T2 = 1.05). If (FN / FN-1) ≤ T2, termination is triggered. This indicates that the growth momentum of the elite group has slowed significantly, entered a growth plateau, and the marginal benefit of continuing screening is very low.
[0098] The technical principle of this implementation is based on population dynamic convergence judgment. It intelligently determines whether the screening process has reached its profit maximization point or has stabilized by monitoring two macroscopic indicators: the first total frequency (FN) and the growth acceleration (FN / FN-1) of the dominant population (first clone set). This avoids common problems caused by fixed rounds in traditional methods: premature stopping (insufficient screening) when condition one is not triggered; and meaningless continued input (over-screening) when condition two is triggered. It enables the screening process to stop precisely at the right moment, which is the core decision-making mechanism for improving efficiency.
[0099] In a preferred embodiment, the optimization adjustment instruction is executed by adding competitive elution in subsequent screening rounds; the competitive elution includes the following steps:
[0100] After the phage binds to the immobilized target and before specific elution, a certain concentration of soluble target molecules is added to the system for competitive incubation.
[0101] The concentration of the soluble target molecule is dynamically adjusted based on the enrichment degree of non-specific clones in the affinity distribution data.
[0102] Specifically, when the system determines that the screening stringency needs to be increased based on affinity distribution data analysis (e.g., the discovery of abnormal enrichment of certain known non-specific binding sequences or low-quality clones), a competitive elution step will be embedded in the next round of screening.
[0103] In the next round of selection, after the phages have bound to the target protein immobilized on the ELISA plate or magnetic beads and undergone routine washing to remove unbound and weakly bound phages, conventional acid / alkali elution is not performed immediately. Instead, a solution containing soluble target molecules is added to the reaction system for incubation (e.g., incubation at 37°C for 30–60 minutes). These soluble target molecules are identical to the targets immobilized on the solid phase, and they will competitively bind the antibodies exhibited by the phages in solution.
[0104] The key to this implementation is that the concentration of soluble target molecules is dynamically set. The system calculates the required concentration based on the enrichment level of non-specific clones in the previous affinity distribution data.
[0105] For example, the system can preset a baseline concentration (e.g., 10 μg / mL). Analyzing the previous round of data, if the total frequency of identified non-specific clones exceeds a certain warning threshold (e.g., 15%), the system will automatically increase the concentration of the competing agent in this round by a certain proportion (e.g., increasing it to 200% of the baseline concentration, i.e., 20 μg / mL). Conversely, if the level of non-specific clones is very low, a lower concentration (e.g., 5 μg / mL) is used. After the competitive incubation is completed, a traditional specific elution step is performed, and the eluent is collected for subsequent amplification and monitoring procedures.
[0106] The technical principle of this implementation is to apply selection pressure using competitive binding kinetics. High-affinity antibodies bind very strongly to the target, with a slow dissociation rate, requiring higher competitive agent concentrations or longer times to be displaced. In contrast, low-affinity or non-specifically binding antibodies bind weakly, dissociate quickly, and are effectively eluted at lower competitive agent concentrations. By dynamically and purposefully increasing the competitive agent concentration, this implementation essentially sets an automatically adjustable threshold in the screening process. This threshold effectively blocks weakly binding, noisy clones from the next round, thereby forcing the screening process towards enriching higher-affinity, more specific clones, significantly improving the quality of the final antibody library and the success rate of screening.
[0107] In a preferred embodiment, before identifying the antibody clones with the highest enrichment from the Nth round of affinity distribution data and forming the first clone set, the method further includes determining the value of K.
[0108] Determining the value of K includes the following steps:
[0109] The antibody clones are sorted from high to low according to their enrichment, and an initial enrichment-ranking curve is plotted. The total number of all antibody clones before the inflection point with the largest slope change in the initial enrichment-ranking curve is taken as the K value.
[0110] The process of constructing the first clone set includes the following steps:
[0111] The first clone set consists of all antibody clones preceding the inflection point with the largest change in slope in the initial enrichment-ranking curve.
[0112] Specifically, after obtaining the enrichment data of all antibody clones in the Nth round, the system first sorts the clones from highest to lowest enrichment. Then, the system plots an "initial enrichment-rank curve" with the rank number on the x-axis and the corresponding enrichment value on the y-axis. This curve visually illustrates the trend of enrichment decreasing with rank, typically showing a rapid initial decrease followed by a gradual flattening.
[0113] The next crucial step in determining the K value is finding the inflection point in the curve. This inflection point is defined as the point on the curve where the rate of change of slope is greatest, i.e., the turning point where a steep descent transitions to a gentler descent. In practice, this can be achieved by calculating the slope between every two adjacent points on the curve and then finding the region with the largest absolute value of the continuous slope change. A commonly used algorithm is the elbow rule, which calculates the vertical distance from each potential inflection point to the line connecting its preceding high point and its trailing low point (chord), and selects the point with the largest distance as the inflection point.
[0114] Once the inflection point is found, the system automatically reads the corresponding rank number on the horizontal axis. This rank number is the dynamically determined K value. For example, if the inflection point is located at the 25th ranked clone, then K=25. Finally, the system automatically assigns all antibody clones ranked in the top K positions (i.e., 1 to 25) to the "first clone set" for subsequent total frequency calculations.
[0115] The technical principle of this implementation is based on dynamic clustering of data distribution. It avoids the rigidity of manually setting a fixed K value (e.g., always selecting the top 20), because the size of the population of high-quality clones naturally changes dynamically depending on the target and the round of screening. By automatically identifying natural inflection points in the enrichment distribution, this method can intelligently distinguish significantly high-enrichment clones from ordinary-enrichment clones, ensuring that the first clone set used for decision-making always represents the most advantageous clones in the current round and those with significant influence as a group, thus making the termination decision more scientific, adaptive, and precise.
[0116] In a preferred embodiment, after plotting the enrichment-ranking curve, the following step is further included:
[0117] Calculate the sequence diversity index corresponding to the antibody clone that is ranked in the top P position in the initial enrichment-ranking curve;
[0118] The step of using the total number of all antibody clones before the inflection point with the largest slope change in the initial enrichment-ranking curve as the K value includes the following steps:
[0119] If the sequence diversity index is higher than or equal to the preset diversity threshold, then the total number of all antibody clones before the inflection point with the largest slope change in the initial enrichment-ranking curve is taken as the K value.
[0120] Specifically, after plotting the "initial enrichment-ranking curve" as described above, the system does not immediately look for the inflection point, but instead performs a rapid diversity assessment first:
[0121] The system starts with the highest-ranked clone and proceeds downwards, selecting the top P clones (e.g., P=30) as a population to be evaluated.
[0122] The system retrieves the amino acid sequences of the antibody variable regions (such as VH and VL) of these P clones.
[0123] Based on these sequences, the system calculates a sequence diversity index (D). This index quantifies the degree of sequence difference among the top-ranking clones. A common implementation is to calculate Shannon entropy: first, the sequences are divided into different families through multiple sequence alignment and clustering (e.g., using CDR3 region sequences for clustering), and then the entropy value is calculated based on the number of clones in each family. The higher the entropy value, the more uniform the sequence distribution and the better the diversity.
[0124] The system compares the calculated diversity index D with a preset diversity threshold (D0). If D ≥ D0, it indicates that the currently enriched top clones are sequence-diverse and have not fallen into clonal bias. The system then follows the original process described above, which involves finding the inflection point based on the initial enrichment-ranking curve and using the number of clones before that inflection point as the K value. This means that the system considers the current enrichment results to be healthy and of high quality, and can be directly used for decision-making.
[0125] This implementation assesses the intrinsic quality (diversity) of the elite group before dynamically determining its boundaries. This acts like a security check, ensuring the data foundation used for subsequent decision-making is reliable. If the top clone itself possesses good diversity, indicating the selection direction is correct and the results are ideal, the system trusts the initial data and continues to use efficient direct analysis methods. This step effectively prevents decisions from being made based on biased data when diversity is already insufficient, adding an important quality assurance layer to the entire adaptive process.
[0126] In a preferred embodiment, after calculating the sequence diversity index corresponding to the antibody clone ranked in the top P position in the initial enrichment-ranking curve, the method further includes the following step:
[0127] If the sequence diversity index is lower than the preset diversity threshold, then starting from the antibody clone with the highest current enrichment ranking, the process is traversed downwards sequentially, and based on its amino acid sequence, the traversed antibody clones are clustered using a preset sequence similarity threshold; wherein, during the clustering process, when the sequence similarity of a newly traversed antibody clone with all antibody clones in any existing cluster is lower than the preset sequence similarity threshold, it is classified into a new independent cluster;
[0128] The number of independent clusters formed is monitored in real time, and the traversal stops when the number of independent clusters reaches the preset maximum number of clusters for the first time.
[0129] All antibody clones that have been traversed at this point and belong to different clusters are defined as the candidate clone group;
[0130] A second enrichment-ranking curve is plotted based on the candidate clonal population, and the first clonal set is composed of all antibody clones before the inflection point with the largest slope change in the second enrichment-ranking curve.
[0131] Specifically, when the system determines that the sequence diversity index is below a threshold, it initiates this diversity optimization and relocation process:
[0132] Initialization and traversal: The system starts with the clone ranked 1st in enrichment and visits the next clone in sequence (2nd, 3rd, ...).
[0133] Online clustering: For the currently accessed clone, the system compares its amino acid sequence with all existing clusters. Sequence comparisons can be performed by calculating the percentage of sequence identity or using a scoring matrix such as BLOSUM. The system presets a sequence similarity threshold (e.g., 85% identity).
[0134] If the similarity between the clone and all representative clones in an existing cluster is below the threshold, it cannot be added to any existing cluster, and the system will create a new independent cluster for it.
[0135] Conversely, it is assigned to the existing cluster with the highest similarity.
[0136] Termination condition check: After classifying each clone, the system checks the number of independent clusters formed in real time. Once this number reaches a preset maximum number of clusters (M, e.g., M=10), the traversal stops immediately.
[0137] Constructing a candidate clone population: At this point, the system collects all clones that have been traversed (regardless of which cluster they belong to) to form a "candidate clone population". The key characteristic of this population is that it guarantees to contain representatives of at least M different sequence families (clusters).
[0138] Reassessment and Decision-Making: The system ignores all unvisited clones and, based solely on this candidate clone group, recalculates the enrichment degree of each clone and plots a new "second enrichment degree-ranking curve." Subsequently, an inflection point finding method is executed on this new curve, and all clones before the found inflection point are ultimately determined as the first clone set.
[0139] The technical principle behind this implementation is a sampling strategy that prioritizes enforced diversity. When natural enrichment leads to clonal bias, it no longer blindly follows enrichment rankings. Instead, it proactively and strategically selects a group from the top-ranked clones that, within a limited quota (controlled by M), can cover the maximum sequence diversity. Then, their enrichment performance is reassessed on this new, high-quality, and diverse foundation. This ensures that even under adverse conditions, the initial clone set used for termination decisions is an elite group that achieves an optimal balance between affinity and diversity, significantly enhancing the application value and development potential of the final antibody library.
[0140] In a preferred embodiment, the preset decision condition further includes a potential prediction condition for generating the continuation instruction: the potential prediction condition is configured as follows:
[0141] When the triggering conditions of the early termination instruction are not met, an enrichment growth trend model is constructed based on the affinity distribution data of the Nth round and multiple historical rounds.
[0142] If the enrichment growth trend model predicts that the first total frequency may exceed the first preset threshold in the next 1 to 2 rounds, then the continue instruction is generated; otherwise, the optimization adjustment instruction for adjusting the screening strategy is generated.
[0143] Specifically, data preparation: The system collects affinity distribution data for the current round (round N) and several previous rounds (e.g., rounds N-2, N-1), especially the historical sequence (FN-2, FN-1, FN) of the calculated first total frequency (F).
[0144] Trend Modeling: The system constructs an enrichment growth trend model using the round number as the independent variable and the first total frequency F as the dependent variable. This is a time series forecasting problem. In practice, simple linear regression or exponential fitting can be used, or more complex time series forecasting algorithms, such as the ARIMA (Autoregressive Integrated Moving Average) model, can be employed. The model learns from existing F-value sequences to fit their growth trajectory.
[0145] Prediction and Judgment: The system uses a trained model to predict the first total frequency for the next 1 to 2 rounds (i.e., round N+1 and / or round N+2). Then, the predicted value is compared with the first preset threshold (T1) used for decision-making in this round.
[0146] If the model predicts a greater likelihood of exceeding T1 in the next 1 to 2 rounds (e.g., the lower bound of the predicted value's confidence interval is higher than T1, or the probability exceeds a certain threshold such as 70%), the system generates a continue instruction. This indicates that although the target has not yet been met, the growth momentum is strong, making it worthwhile to invest in another 1 to 2 rounds to obtain better results.
[0147] Otherwise (i.e., the forecast indicates that the target cannot be met or growth will stagnate in the short term), the system generates optimization and adjustment instructions. This suggests that the current screening strategy may be inefficient and that parameter changes are needed to break the deadlock, rather than simply continuing.
[0148] The technical principle of this implementation is to introduce decision optimization based on time series forecasting. It expands the decision-making basis from the current state to future trends, giving the system forward-looking decision-making capabilities. This resolves a ambiguity in traditional methods: when the current situation is unsatisfactory, should we continue or adjust? This implementation provides a data-driven answer to this question by quantitatively predicting future returns. It avoids premature abandonment (erroneous adjustment) when potential has not yet materialized, and also avoids ineffective persistence (erroneous continuation) when stuck in a plateau, thus guiding the entire selection process to always stay on the optimal path, further improving the system's intelligence and overall efficiency.
[0149] In a preferred embodiment, the first preset threshold and / or the second preset threshold are dynamically set through a threshold optimization model.
[0150] Specifically, during the operation of the system of the present invention, the two thresholds (the first preset threshold T1 and the second preset threshold T2) used to determine whether to terminate prematurely are not fixed. They are dynamically output by an algorithm module called the threshold optimization model.
[0151] The operation process of this model is as follows:
[0152] Initialization: When the system is first put into use or when historical data is lacking, a set of empirically conservative default thresholds can be set (e.g., T1=60%, T2=1.10).
[0153] Real-time invocation: After performing the Nth round of screening and generating affinity distribution data, the system prepares the input features for the current round. These features include: the first total frequency (FN) of the current round, and the ratio of the first total frequency to the second total frequency (FN / FN-1); Dynamic calculation: The system passes these input features to a pre-trained threshold optimization model. Based on the specific situation exhibited by the current screening process, the model calculates in real time and outputs a set of optimal threshold suggestions tailored to this round.
[0154] For example, the model might determine that the current screening process is going very smoothly and the clone quality is generally high, so it might output a more promising T1 (e.g., 75%) and a more stringent T2 (e.g., 1.03) to incentivize the system to continue screening to obtain a top-level library with higher purity.
[0155] Conversely, if the model's judgment process is slow or the library quality is complex, it may output a more lenient T1 (such as 55%) and T2, aiming to suggest that the system stop when it reaches a relatively reasonable enrichment level, thus avoiding the risk of over-screening.
[0156] In a preferred embodiment, the method further includes constructing a threshold optimization model, comprising the following steps:
[0157] Model training phase: After system initialization or completion of a specific round of screening, the historical database of historical screening items is retrieved; the historical database stores the affinity distribution data sequence of each historical screening item in the multi-round screening process, the final decision round, and the information of the high affinity antibody sequence finally obtained;
[0158] Specifically, in the model training phase—data preparation: the system retrieves a database storing multiple completed historical screening items. Each item's data package must contain:
[0159] A complete affinity distribution data sequence: records detailed data for each round of screening in this project.
[0160] Final decision round: This records in which round the project was terminated.
[0161] The final high-affinity antibody sequence information: This is the crucial standard answer, indicating which effective antibodies were ultimately successfully isolated from the project and validated.
[0162] Feature extraction and target definition: The first total frequency and the ratio of the first total frequency to the second total frequency generated in each round of the historical screening projects are used as input features; the probability that the final antibody library can be included if the screening is terminated in this round is used as the optimization target.
[0163] Specifically, feature extraction and target definition:
[0164] Feature engineering: For each item in the historical database and for each round of screening, the following calculations are performed:
[0165] Based on the data from this round, the first total frequency (F) is simulated and calculated.
[0166] Calculate the ratio of the first total frequency to the previous round (F / F_previous).
[0167] These two values constitute the input feature vector of the machine learning model.
[0168] Objective Definition: For this specific project-round data point, an optimization objective needs to be defined. This objective is: if the screening terminates in this round, the probability that the final antibody library will include known high-affinity antibodies.
[0169] The specific calculation involves determining how many of the ultimately successfully validated high-affinity antibodies were already present in the first clone set of this round. For example, if there are ultimately 5 effective antibodies, and 4 of them are in the first clone set of this round, then the coverage probability is 80%. This probability value is the learning target of the model.
[0170] Model training: Based on machine learning regression algorithms, the input features are fitted and trained with the optimization objective to obtain the threshold optimization model.
[0171] Specifically, the large number of feature vector-target probability data pairs prepared above are input into a machine learning regression algorithm (e.g., Gradient Boosting Decision Tree (GBDT) or Random Forest). The algorithm iteratively learns to find a complex mapping function that can predict the quality (coverage probability) of the outcome at the current epoch based on the input FN and FN / FN-1. After training, this mapping function is the so-called threshold optimization model.
[0172] The technical principle of this implementation is supervised learning based on historical experience. It distills the successes and failures of countless past screening experiments into a computable mathematical model. The core capability of this model lies in its ability to predict the expected future output (antibody library quality) when terminating an experiment at a specific screening state (represented by F and F / F_previous). In the application of the threshold optimization model, the system essentially asks the model: "Based on historical experience, at what level should the termination threshold be set in the current state to maximize the final benefit?" The model provides scientific suggestions by analyzing massive historical patterns. This allows the system to move away from relying on subjective human experience and instead become data-driven based on collective wisdom, automating and optimizing the decision-making process. This represents a crucial step in the advancement of screening technology from "automation" to "intelligence."
[0173] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. The above descriptions are only preferred embodiments of this application. It should be noted that due to the limitations of written expression, while there are objectively infinite specific structures, those skilled in the art can make several improvements, modifications, or changes without departing from the principles of this invention, and can also combine the above technical features in an appropriate manner. These improvements, modifications, changes, or combinations, or the direct application of the inventive concept and technical solution to other situations without modification, should all be considered within the scope of protection of this application.
Claims
1. An adaptive phage screening method based on real-time affinity monitoring for screening high-affinity antibodies, characterized in that, Includes the following steps: Perform the Nth round of selection, with N initially set to 1; After the Nth round of screening, the following actions are performed in parallel: amplification of the eluted phages to obtain the Nth round enrichment library; and sampling of the eluted phages and high-throughput affinity testing of the sample solution to generate Nth round affinity distribution data. The affinity distribution data of the Nth round is compared with the reference data, and control instructions are generated according to preset decision conditions. The control instructions include a continue instruction, an early termination instruction, or an optimization adjustment instruction. Execute the control command: If the control instruction is the continue instruction, then N is incremented, and the Nth round enrichment library is used as input for the next round of selection; If the control instruction is the early termination instruction, then the selection process is terminated, and the Nth round enriched library is output as the final library. If the control instruction is the optimization adjustment instruction, then after adjusting the screening parameters, N is incremented, and the Nth round enrichment library is used as input for the next round of screening; The preset decision conditions include stability conditions for generating the early termination instruction, which are configured to be triggered as follows: From the Nth round of affinity distribution data, the antibody clones with the highest enrichment ranking (K) are identified to form the first clone set. Calculate the total frequency of the first clone set in the Nth round of sequencing data set, and denote it as the first total frequency; If the first total frequency exceeds the first preset threshold; or, the ratio of the first total frequency to the second total frequency is less than or equal to a second preset threshold close to 1, then the early termination instruction is generated: the second total frequency is the total frequency of the first clone set in the N-1th round of sequencing data set; The preset decision conditions also include potential prediction conditions for generating the continuation instruction: the potential prediction conditions are configured as follows: When the triggering conditions of the early termination instruction are not met, an enrichment growth trend model is constructed based on the affinity distribution data of the Nth round and multiple historical rounds. If the enrichment growth trend model predicts that the first total frequency may exceed the first preset threshold in the next 1 to 2 rounds, then the continue instruction is generated; otherwise, the optimization adjustment instruction for adjusting the screening strategy is generated.
2. The adaptive phage screening method based on real-time affinity monitoring according to claim 1, characterized in that: The process of performing high-throughput affinity testing on the sampled liquid to generate Nth round affinity distribution data includes the following steps: Phage DNA was extracted from the sample solution and subjected to high-throughput sequencing to obtain a sequencing data set containing multiple sequence reads; wherein each sequence read contains a preset DNA barcode, which was uniquely assigned to different antibody clones during the initial phage library construction; Based on the sequencing data set, identify and count the number of times each preset DNA barcode sequence appears in the set; The preset barcode-clone correspondence database is retrieved, and each preset DNA barcode sequence is mapped to the specific antibody clone it represents to obtain the analysis result; Based on the analysis results, the relative frequency of each antibody clone in the sequencing dataset and / or its enrichment relative to historical rounds are calculated to generate the Nth round affinity distribution data.
3. The adaptive phage screening method based on real-time affinity monitoring according to claim 1, characterized in that: The optimization and adjustment instructions are executed as follows: competitive elution is added in subsequent selection rounds; the competitive elution includes the following steps: After the phage binds to the immobilized target and before specific elution, a certain concentration of soluble target molecules is added to the system for competitive incubation. The concentration of the soluble target molecule is dynamically adjusted based on the enrichment degree of non-specific clones in the affinity distribution data.
4. The adaptive phage screening method based on real-time affinity monitoring according to claim 1, characterized in that: Before identifying the antibody clones with the highest enrichment from the Nth round of affinity distribution data and forming the first clone set, the method further includes determining the K value. Determining the value of K includes the following steps: The antibody clones are sorted from high to low according to their enrichment, and an initial enrichment-ranking curve is plotted. The total number of all antibody clones before the inflection point with the largest slope change in the initial enrichment-ranking curve is taken as the K value. The process of constructing the first clone set includes the following steps: The first clone set consists of all antibody clones preceding the inflection point with the largest change in slope in the initial enrichment-ranking curve.
5. The adaptive phage screening method based on real-time affinity monitoring according to claim 4, characterized in that: After plotting the enrichment-ranking curve, the following steps are also included: Calculate the sequence diversity index corresponding to the antibody clone that is ranked in the top P position in the initial enrichment-ranking curve; The step of using the total number of all antibody clones before the inflection point with the largest slope change in the initial enrichment-ranking curve as the K value includes the following steps: If the sequence diversity index is higher than or equal to the preset diversity threshold, then the total number of all antibody clones before the inflection point with the largest slope change in the initial enrichment-ranking curve is taken as the K value.
6. The adaptive phage screening method based on real-time affinity monitoring according to claim 5, characterized in that: After calculating the sequence diversity index corresponding to the antibody clone ranked in the top P position in the initial enrichment-ranking curve, the method further includes the following steps: If the sequence diversity index is lower than the preset diversity threshold, then starting from the antibody clone with the highest current enrichment ranking, the process is traversed downwards sequentially, and based on its amino acid sequence, the traversed antibody clones are clustered using a preset sequence similarity threshold; wherein, during the clustering process, when the sequence similarity of a newly traversed antibody clone with all antibody clones in any existing cluster is lower than the preset sequence similarity threshold, it is classified into a new independent cluster; The number of independent clusters formed is monitored in real time, and the traversal stops when the number of independent clusters reaches the preset maximum number of clusters for the first time. All antibody clones that have been traversed at this point and belong to different clusters are defined as the candidate clone group; A second enrichment-ranking curve is plotted based on the candidate clonal population, and the first clonal set is composed of all antibody clones before the inflection point with the largest slope change in the second enrichment-ranking curve.
7. The adaptive phage screening method based on real-time affinity monitoring according to claim 1, characterized in that: The first preset threshold and / or the second preset threshold are dynamically set through a threshold optimization model.
8. The adaptive phage screening method based on real-time affinity monitoring according to claim 7, characterized in that: The method also includes constructing a threshold optimization model, comprising the following steps: Model training phase: After system initialization or completion of a specific round of screening, the historical database of historical screening items is retrieved; the historical database stores the affinity distribution data sequence of each historical screening item in the multi-round screening process, the final decision round, and the information of the high affinity antibody sequence finally obtained; Feature extraction and target definition: The first total frequency and the ratio of the first total frequency to the second total frequency generated in each round of the historical screening projects are used as input features; the probability that the final antibody library can be included if the screening is terminated in this round is used as the optimization target. Model training: Based on machine learning regression algorithms, the input features are fitted and trained with the optimization objective to obtain the threshold optimization model.
Citation Information
Patent Citations
Preparation of human CD56 immune bacteriophage display antibody library
CN120209145A
High throughput monoclonal antibody generation by b cell panning and proliferation
CN107428827A
Method for high throughput peptide-MHC affinity screening for TCR ligands
CN113195529A