Artificial intelligence-based phage library screening and management method and application

By using AI-based multidimensional data processing and model training, the problems of data silos and static data in phage screening have been solved, enabling intelligent and precise phage screening and improving screening speed and resource utilization efficiency.

CN121963887APending Publication Date: 2026-05-01LIAONING AGRI COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIAONING AGRI COLLEGE
Filing Date
2025-11-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing phage screening technologies suffer from problems such as data silos, reliance on one-sided decision-making, high costs due to static nature, inability to provide real-time feedback and optimization, lack of predictive capabilities, and low resource utilization, resulting in low efficiency, high cost, and poor accuracy in phage development.

Method used

By employing an artificial intelligence-based approach, a phage-host matching prediction model is established through multi-dimensional data collection, feature modeling, AI model training, scoring and recommendation, and feedback updates, enabling dynamic optimization and intelligent recommendation.

Benefits of technology

It significantly improves the speed and accuracy of phage screening, reduces costs, enables dynamic optimization of library management and efficient utilization of resources, and supports accurate recommendations in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963887A_ABST
    Figure CN121963887A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of biological artificial intelligence, and discloses an artificial intelligence-based phage library screening and management method and application. Comprising the steps of data acquisition, feature modeling, AI model training, scoring and recommendation, and feedback updating. According to the method, the phage screening speed and accuracy are remarkably improved, pure dependence on trial and error experiments is avoided, the cost is saved, a dynamically updated library management mechanism is provided, and the optimal combination of the phages in the library is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

An Artificial Intelligence-Based Method for Phage Library Screening and Management and Its Application Technical Field

[0001] This invention belongs to the field of bioartificial intelligence technology, specifically relating to an artificial intelligence-based method and application for screening and managing phage libraries. Background Technology

[0002] Currently, the development and application of bacteriophages heavily rely on high-throughput biological screening of bacteriophage libraries in laboratories. While this traditional paradigm has achieved results in discovering specific bacteriophages, its core process has inherent efficiency bottlenecks and decision-making blind spots, making it difficult to meet the growing demand for precise and rapid phage therapy. The limitations of existing technologies are specifically reflected in the following aspects: 1. Limited data dimensions and the problem of "data silos": Existing screening processes mainly rely on phenotypic observation (such as plaque formation), and the resulting data is often fragmented and isolated.

[0003] Lack of multi-source data integration: A single screening experiment may simultaneously generate multidimensional data, including host genomic information (such as resistance genes and receptor genes), phage genome sequences, infection dynamics parameters (such as latency, outbreak size, and optimal multiplicity of infection (MOI), environmental stability data (such as tolerance to pH and temperature), and synergistic / antagonistic effects of different phage combinations. However, current technologies lack a unified data architecture to systematically collect, store, and correlate this heterogeneous data. This valuable data is often scattered across different experimental records, Excel spreadsheets, or independent databases, forming "data silos" that cannot be comprehensively utilized.

[0004] Reliance on one-sided decision-making: Due to the inability to integrate analysis, researchers often make decisions based on only one or two core indicators (such as plaque size or clarity), while ignoring other key factors that may affect the effectiveness of practical applications (such as the environmental stability or combination effect of phages). This results in phages that perform well under laboratory conditions but are not effective in complex real-world application scenarios (such as in vivo therapy and industrial fermentation).

[0005] 2. Static Screening Process and High Costs: Traditional screening is a trial-and-error intensive process, and its static and linear characteristics lead to huge resource consumption.

[0006] High throughput but low intelligence: Although automated equipment achieves "high throughput" screening, the screening strategy itself lacks intelligent guidance. It is essentially an "exhaustive" or "random" traversal of the phage library, with each experiment being an independent cost center, including manpower, reagents, consumables, and time.

[0007] Lack of real-time feedback and optimization: The experimental design is pre-defined, making it impossible to dynamically adjust subsequent screening directions based on the experimental results. For example, after initially screening a batch of promising phages, the system cannot immediately and intelligently predict and recommend other complementary or synergistic phages for the next round of testing based on the genomic characteristics of these phages. This "open-loop" process results in long learning cycles and slow iteration speeds.

[0008] 3. Lack of predictive ability and difficulty in knowledge discovery: Existing technology is essentially a "descriptive" tool rather than a "predictive" tool.

[0009] Establishing genotype-phenotype associations is challenging: Phage-host interactions are determined by complex molecular mechanisms. While we have access to vast amounts of genomic data, traditional methods lack effective computational models to uncover the deep-seated rules governing the association between phage tail fimbriae, host surface receptors, and infection outcomes. We know the "what" (whether infection occurs), but not the "why" (why infection occurs), severely hindering our ability to accurately predict host range.

[0010] Highly dependent on experience: The success of screening depends heavily on the researcher's personal experience and intuition. For a novel or rare pathogen, due to the lack of prior knowledge, screening work is often like "finding a needle in a haystack," with poor reproducibility, and different laboratories or operators may obtain drastically different results, which is not conducive to the standardization and promotion of the technology.

[0011] 4. Rigid and inefficient phage library management and low resource utilization: The large phage library lacks dynamic "value" assessment. Some phages with unique host ranges or excellent stability may be "shelved" for a long time because they were not selected in a routine screening, and cannot be given priority in subsequent screenings for new pathogens.

[0012] Inconvenient information retrieval: Existing database management systems typically only have basic archiving and query functions (such as searching by name or host), and cannot support complex multi-condition intelligent searches, such as: "find all bacteriophages that can infect Escherichia coli ST131 strains, are stable in the pH range of 3-5, and have a synergistic lysis effect with bacteriophage A".

[0013] In summary, the core contradiction of the existing technological system lies in the fact that we live in a multi-dimensional and highly complex world of biological interactions, yet we are using a single-dimensional, linear decision-making tool that relies on human experience. This leads to prominent problems such as high cost, long development cycle, low accuracy, and poor predictability in phage development. Therefore, there is an urgent need in this field for a management system that can deeply integrate biological experimental data with artificial intelligence algorithms to achieve a paradigm shift from "blind screening" to "intelligent prediction and precise recommendation," thereby greatly improving the efficiency and intelligence level of phage screening and management. Summary of the Invention

[0014] To overcome the shortcomings of existing technologies, this invention provides an artificial intelligence-based method and application for phage library screening and management, which is used to quickly predict the optimal phage or phage combination corresponding to target bacteria, thereby improving the control effect and reducing experimental costs.

[0015] The above-mentioned objective of this invention is achieved through the following technical solution: a method for screening and managing a phage library based on artificial intelligence, comprising the following steps: 1. Data acquisition: acquiring multidimensional data related to phages and target bacteria; 2. Feature modeling: standardizing the data obtained in step 1 and converting it into a vectorized feature matrix, extracting key features for training the AI ​​model; 3. AI model training: establishing a phage-host matching prediction model using machine learning algorithms or deep learning networks; 4. Scoring and recommendation: after inputting the target bacteria, the AI ​​model outputs a matching score for the phage or a combination of phages and generates a recommendation list; 5. Feedback update: feeding back the actual application results to the model to achieve dynamic optimization.

[0016] Furthermore, the data obtained in step 1 includes basic phage properties, host range experimental results, lysis kinetic parameters, environmental stability (i.e., pH and temperature), and combined therapeutic efficacy data.

[0017] Further, step 2 specifically includes: 2.1. Standardization by feature category. Differentiated standardization rules are adopted based on the data attribute type: continuous physical and dynamic parameters are standardized using Z-score; proportional or probability data are first converted to a 0–1 interval, then min-max normalization or arcsin-sqrt transformation is used, with the standardization benchmark being the 1%–99th percentile range of historical samples; categorical or binary features are represented using one-hot encoding or embedding vectors, with the standardization benchmark being a fixed set of categories or an embedding model version; environmental condition features use the mean and standard deviation of a sliding time window as the standardization benchmark, and undergo dynamic Z-score normalization; 2.2. Constructing a sub-matrix architecture: The standardized data are constructed into a multi-sub-matrix structure according to category, forming an overall feature matrix; 2.3. Weighted fusion: The feature vectors of each sub-matrix are weighted and fused. The initial weights are set based on domain knowledge. During model training, the weights are dynamically adjusted using statistical correlation (such as mutual information or Pearson correlation coefficient) and through an attention mechanism; the specific fusion form is as follows:

[0018] Where v i Let W be the eigenvector of the i-th submatrix. i Let α be a linear mapping matrix. i The weighting coefficients of the submatrix satisfy the following condition: The fusion results are used for model scoring and recommendation. The system adjusts each weight and standardized benchmark in real time based on subsequent experimental feedback, realizing the model's self-learning and dynamic optimization.

[0019] Furthermore, the Z-score standardization process involves calculating x′=(x-μ) / σ based on the mean μ and standard deviation σ of the population sample to eliminate the influence of dimensions; when the parameter distribution is skewed, a log(x+ε) transformation is performed before standardization.

[0020] Further, step 3 specifically includes: 3.1. Dataset Construction: The system first aggregates phage-host association data from phage databases, host strain experimental records, and third-party published literature to form a unified dataset; after data cleaning, the system treats each "phage-host pairing relationship" as a sample instance, with the sample label being "match / mismatch" or "lysis probability value"; subsequently, it divides the dataset into training, validation, and test sets in an 8:1:1 ratio to ensure the objectivity and repeatability of model evaluation. During the construction process, the system automatically detects and removes duplicate entries and obvious outliers to ensure that each sample has a complete feature vector input; 3.2. Model Training: Based on the feature matrix input, machine learning and deep learning algorithms are used for modeling; during training, the model optimization uses Adam or SGD optimization algorithms, with an initial learning rate of 0.001 to 0.01, and the loss function is selected according to the task type as binary cross-entropy (match prediction) or mean squared error (probabilistic regression). Model training iterates 100–300 times in a GPU environment. The system continues training in rounds until the validation set loss converges or the early stopping condition is triggered. 3.3. Feedback Update and Model Reproducibility Mechanism: To ensure model reproducibility and continuous optimization, the system records all training parameters and stores them in a model version control library. Each training process can be fully reproduced based on the version number. Simultaneously, the system has a built-in "feedback update" module: when the results of lysis experiments or animal trials in actual applications are returned, the system automatically appends new data to the sample library, re-standardizes it, and incorporates it into incremental training to achieve dynamic optimization. Model weight updates follow the exponential decay rule (learning rate decay), ensuring a moderate impact from new data without destroying existing learning results. Model output includes predicted matching probabilities and feature importance rankings, used for subsequent recommendation and interpretable analysis.

[0021] Furthermore, the specific matching and scoring rules in step 4 are as follows: After receiving the feature input of the target bacteria, the AI ​​model outputs the matching probability value between each candidate phage and the bacteria. The model output value is normalized to the 0–1 range using the Sigmoid or Softmax function. For easier intuitive comparison, it is further linearly mapped to the 0–100 range. The calculation formula is: Scoresi = round(Pi × 100), where Pi is the phage-host matching probability predicted by the model. When a single phage score is ≥ 0.7 (or ≥ 70 points), it is determined to be a recommended candidate. When multiple phage scores are in the 0.5–0.7 range, if their combined score (calculated by product or synergistic effect factor weighting) is ≥ 0.75, then combination therapy is recommended. The system sorts the scores from high to low, outputs the top N candidate phages or phage combinations, and displays the reasons for recommendation on the interface (e.g., "high lysis rate", "good stability", "significant synergistic effect with the target strain").

[0022] Furthermore, step 5 specifically involves using the difference between the model's predicted matching score and the experimental verification results as the basis for adjustment. When the experimental lysis rate is ≥80% and the error with the predicted value is ≤10%, the model prediction is confirmed to be effective, and the current parameters are maintained; when the difference between the experimental result and the prediction is >10%, the feature weights are updated; when a certain phage performs worse than expected in multiple verifications, its weight in the feature matrix is ​​reduced.

[0023] Another objective of this invention is to protect the application of the aforementioned AI-based phage library screening and management method, specifically for the rapid screening of suitable phage combinations in livestock and poultry farms; the rapid recommendation of personalized phage therapies for clinically drug-resistant strains in hospitals or testing centers; and the matching of suitable phage combinations for prevention and control in food safety testing.

[0024] The beneficial effects of this invention compared with the prior art are: (1) It significantly improves the speed and accuracy of phage screening, avoids relying solely on trial-and-error experiments, and saves costs.

[0025] (2) Provides a dynamically updated library management mechanism to ensure the optimal combination of phages in the library; (3) Can be integrated with the experimental platform in a closed loop (AI prediction → experimental verification → data feedback → AI optimization) to achieve virtuous iteration. Attached Figure Description

[0026] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Figure 1 is a heat map of the phage-bacteria relationship matrix in Embodiment 2 of the present invention; Figure 2 is a scatter plot of the correlation between AI prediction scores and experimental verification in Embodiment 2 of the present invention; Figure 3 is a curve of AI model iteration accuracy improvement in Embodiment 3 of the present invention. Detailed Implementation

[0027] The present invention is described in detail below through specific embodiments, but this does not limit the scope of protection of the present invention. Unless otherwise specified, the experimental methods used in the present invention are all conventional methods, and the experimental equipment, materials, reagents, etc. used can all be obtained commercially.

[0028] Example 1: An AI-based method for screening and managing phage libraries, used to quickly predict the optimal phage or phage combination corresponding to target bacteria, thereby improving the control effect and reducing experimental costs.

[0029] Methodological steps: 1. Data acquisition. Obtain multidimensional data related to bacteriophage and target bacteria, including but not limited to: basic bacteriophage properties, host range experimental results, lysis kinetic parameters (such as latency period, outbreak size), environmental stability (pH, temperature), and combined therapeutic efficacy data.

[0030] II. Feature modeling. Standardize the above data and convert it into a vectorized feature matrix; extract key features for training the AI ​​model: 1. Standardization method and benchmark according to feature category. Differentiated standardization rules are adopted according to different data attribute types: (1) Continuous physical and dynamic parameters (latency, burst, lysis rate, etc.) are standardized using Z-score, that is, the mean μ and standard deviation σ of the sample population are used as benchmarks to calculate x′=(x-μ) / σ to eliminate the influence of dimensions; when the parameter distribution is skewed, log(x+ε) transformation is performed first and then standardized.

[0031] (2) For proportional or probability data (lysis rate, antibacterial rate, etc.), the percentage is first converted into a 0–1 interval, and then minimum-maximum normalization or arcsin-sqrt transformation is used. The standardization benchmark is the 1%–99% percentile range of historical samples.

[0032] (3) Category or binary features (whether the host is sensitive or whether the gene exists) are represented by one-hot encoding or embedding, and the standardization benchmark is a fixed set of categories or an embedding model version.

[0033] (4) Environmental conditions (temperature, pH, etc.) are normalized by using the mean and standard deviation of the sliding time window as the standardization benchmark and dynamic Z-score normalization.

[0034] All standardized benchmarks (mean, variance, extrema, and embedding model version, etc.) are stored in the "feature dictionary" database to support subsequent model iterations and backtracking.

[0035] To improve the efficiency of feature structure expression, the standardized data is constructed into a multi-sub-matrix structure according to categories, forming an overall feature matrix. This includes, but is not limited to: a host spectrum sub-matrix (M1): describing the lysis intensity or infection probability of bacteriophages on different host strains; a kinetic sub-matrix (M2): containing biological kinetic parameters such as latency, burst size, and lysis rate; a stability sub-matrix (M3): reflecting the stability indicators of bacteriophages under different pH, temperature, and salinity conditions; a genomic functional domain sub-matrix (M4): composed of genomic or protein domain embedding vectors; a combinatorial effect sub-matrix (M5): recording the synergistic and antagonistic effects between bacteriophages and other preparations or bacteriophages; and a metadata matrix (M-meta): including experimental confidence levels, data sources, and batch information.

[0036] Each submatrix can be independently reduced in dimensionality or have key features extracted to form a unified input structure for AI models to learn from.

[0037] 3. The weighted fusion benchmark is to integrate the multi-source information of different sub-matrices. The system performs weighted fusion on the feature vectors of each sub-matrix.

[0038] The initial weights are set based on domain knowledge, for example: host spectrum matrix weight 0.4, dynamic matrix weight 0.2, genome matrix weight 0.2, stability matrix weight 0.1, and combined effect matrix weight 0.1.

[0039] During the model training phase, the weights are dynamically adjusted using statistical correlations (such as mutual information or Pearson correlation coefficient) and through an attention mechanism (attention layer).

[0040] The specific form of integration is as follows:

[0041] Where v i Let W be the eigenvector of the i-th submatrix. i Let α be a linear mapping matrix. i The weighting coefficients of the submatrix satisfy the following condition: .

[0042] The fusion results are used for model scoring and recommendation. The system adjusts each weight and standardized benchmark in real time based on subsequent experimental feedback, realizing the model's self-learning and dynamic optimization.

[0043] AI Model Training. A phage-host matching prediction model is established using machine learning algorithms (such as random forests and support vector machines) or deep learning networks (such as neural networks and attention mechanism models). 1. Dataset Construction: To train the machine learning model, the system first aggregates phage-host association data from phage databases, host strain experimental records, and third-party published literature to form a unified dataset.

[0044] Data sources include, but are not limited to: basic properties and genomic characteristics of bacteriophages (sequence, GC content, functional domain annotation, etc.); taxonomic and resistance profile characteristics of host strains; experimental validation data (lysis test results, infection rate, inhibition zone diameter, etc.); environmental and culture condition records (temperature, pH, salinity, etc.); and bacteriophage combination or synergistic experimental data.

[0045] After data cleaning, the system treats each "phage-host pairing relationship" as a sample instance, with the sample label being "match / non-match" or "lysis probability value".

[0046] The model was then divided into training, validation, and test sets in an 8:1:1 ratio to ensure the objectivity and reproducibility of the model evaluation. During the model construction process, the system automatically detected and removed duplicate entries and obvious outliers to ensure that each sample had a complete feature vector input. To enhance generalization ability, model stability was evaluated using cross-validation (k-fold, k=5).

[0047] 2. Model Training Methods: Based on the feature matrix input, various machine learning and deep learning algorithms are used for modeling, including but not limited to: Random Forest model: classifies and learns phage features and host features through multi-tree ensemble methods, and outputs the lysis probability; SVM model: uses kernel functions to map high-dimensional space to determine the matching boundary between phage and host; Deep Neural Network (DNN) model: inputs multi-dimensional feature vectors and obtains high-order feature representations through multi-layer nonlinear transformations; Graph Neural Network (GNN) model: uses phage and host as nodes and experimental verification relationships as edges, and uses Graph Convolution (GCN) or Graph Attention Network (GAT) to capture complex topological interactions; Attention Mechanism model: automatically learns the important weights of different features in the features after multi-submatrix fusion to achieve dynamic feature aggregation.

[0048] During training, the model is optimized using either the Adam or SGD optimization algorithm, with an initial learning rate of 0.001–0.01. The loss function is selected based on the task type: binary cross-entropy (matching prediction) or mean squared error (probabilistic regression). Model training iterates for 100–300 rounds in a GPU environment until the validation set loss converges or the early stopping condition is triggered.

[0049] 3. Feedback Update and Model Reproducibility Mechanism: To ensure model reproducibility and continuous optimization, the system records all training parameters (learning rate, batch size, weight initialization seed, network structure, training epochs, etc.) and stores them in the model version control library. Each training process can be fully reproduced based on the version number.

[0050] Meanwhile, the system incorporates a built-in "feedback update" module: when the results of lysis experiments or animal trials in actual applications are returned, the system automatically appends the new data to the sample database, re-standardizes it, and incorporates it into incremental training to achieve dynamic optimization. Model weight updates follow an exponential decay rule (learning rate decay), ensuring that the influence of new data is moderate without destroying existing learning results. Model outputs include predicted matching probabilities and feature importance rankings, used for subsequent recommendation and interpretable analysis.

[0051] IV. Scoring and Recommendation. After inputting the target bacteria, the AI ​​model outputs a matching score (0–1 or 0–100 range) for the phage or phage combination, and generates a recommendation list: After receiving the feature input of the target bacteria, the AI ​​model outputs the matching probability value between each candidate phage and the bacteria. ① Scoring Rules: The model output value is normalized to the 0–1 range using the Sigmoid or Softmax function; for easier intuitive comparison, it can be further linearly mapped to the 0–100 range. The calculation formula is: Score i =round(P i ×100) where P i ① The phage-host matching probability predicted by the model. ② Recommendation criteria: When a single phage score is ≥ 0.7 (or ≥ 70), it is considered a recommended candidate; when multiple phage scores are in the range of 0.5–0.7, if their combined score (calculated by product or synergistic effect factor weighting) is ≥ 0.75, then combination therapy is recommended.

[0052] ③ Recommendation list generation: The system sorts the candidates from highest to lowest score, outputs the top N candidate phages or combinations of phages, and displays the reasons for the recommendation on the interface (e.g., "high lysis rate", "good stability", "significant synergistic effect with the target strain").

[0053] V. Feedback Updates. Actual application results (in vitro or animal experimental data) are fed back to the model for dynamic optimization: ① Feedback Data Types: These include in vitro phage lysis rate, daily weight gain, feed conversion ratio, diarrhea rate, and mortality rate in animal experiments. ② Feedback Baseline: The difference between the model's predicted matching score and the experimental validation results is used as the adjustment basis. When the experimental lysis rate is ≥80% and the error between the predicted value and the actual value is ≤10%, the model prediction is confirmed as effective, and the current parameters are maintained. When the difference between the experimental results and the prediction is >10%, the feature weights are updated. When a phage performs worse than expected in multiple validations (e.g., lysis rate below 50%), its weight in the feature matrix is ​​reduced. ③ Update Method: Incremental learning or online learning is used to gradually correct the model parameters to ensure that the prediction performance optimizes with data accumulation.

[0054] The model was trained and scored using the above method: Input data: an interaction data matrix of 50 bacteriophages × 20 target bacteria; Feature parameters: host spectrum, latency, outbreak size, pH stability index; Model: random forest + graph neural network; Output: matching score table.

[0055] Table 1 Model Training and Scoring

[0056] Example 2: Dynamically updated AI-predicted recommended phage combination: Phage-1 + Probiotic-A. Animal experiments verified that the effect was significantly better than using either combination alone. The system automatically updated the weight parameters to improve the accuracy of future predictions (Figures 1 and 2).

[0057] Example 3: Animal Experiment Verification. One hundred weaned piglets were randomly divided into four groups (control group, single phage group, single probiotic group, and AI-recommended combination group), with 25 piglets in each group. The feeding period was 42 days. Daily weight gain, feed conversion ratio, diarrhea rate, and mortality rate were recorded.

[0058] The results showed that the AI-recommended combination group was significantly better than the single group in all indicators (P<0.05).

[0059]

[0060] Note: * indicates a significant difference from the control group (P<0.05). Example 4: Model Iterative Optimization. The model was updated by feeding back experimental data from five consecutive rounds, and the accuracy of AI predictions improved with each round (Figure 3). The results show that the prediction performance is significantly enhanced with the increase of the number of iterations.

[0061] Table 1. Schematic Data

[0062] Example 5: Effect of Reducing Drug-Resistant Bacteria. Fecal samples were collected from a large-scale farm to detect the proportion of drug-resistant strains. A comparison was made between standard feeding and AI-recommended feeding.

[0063]

[0064] Note: The AI-recommended combination group showed a significant reduction in drug resistance rate at 28 and 42 days (P<0.05). Example 6: Multi-scenario intelligent recommendation. The system connects the phage library with different livestock and poultry disease scenarios to output recommended combinations:

[0065] The results indicate that the system is effective not only in common livestock and poultry such as pigs and chickens, but can also be extended to ruminants such as cattle and sheep.

[0066] The embodiments described above are merely preferred embodiments of the present invention, and not all feasible embodiments of the present invention. Any obvious modifications made by those skilled in the art without departing from the principles and spirit of the present invention should be considered to be included within the scope of protection of the claims of the present invention.

Claims

1. A method for screening and managing phage libraries based on artificial intelligence, characterized in that, The steps include: S1. Data Acquisition: Obtain multidimensional data related to bacteriophages and target bacteria; S2. Feature Modeling: Standardize the data obtained in step S1 and convert it into a vectorized feature matrix, extracting key features for training the AI ​​model; S3. AI Model Training: Build a bacteriophage-host matching prediction model using machine learning algorithms or deep learning networks; S4. Scoring and Recommendation: After inputting the target bacteria, the AI ​​model outputs a matching score for the bacteriophage or a combination of bacteriophages and generates a recommendation list; S5. Feedback Update: Feed back the actual application results to the model to achieve dynamic optimization.

2. The method for screening and managing phage libraries based on artificial intelligence according to claim 1, characterized in that, The data obtained in step S1 includes basic phage properties, host range experimental results, lysis kinetic parameters, environmental stability (pH and temperature), and combined therapeutic data.

3. The method for screening and managing phage libraries based on artificial intelligence according to claim 1, characterized in that, Step S2 specifically includes: S2.

1. Standardizing by feature category. Differentiated standardization rules are adopted based on the data attribute type: continuous physical and dynamic parameters are standardized using Z-score; proportional or probability data are first converted to a 0–1 interval, then min-max normalization or arcsin-sqrt transformation is used, with the standardization benchmark being the 1%–99th percentile range of historical samples; categorical or binary features are represented using one-hot encoding or embedding vectors, with the standardization benchmark being a fixed set of categories or an embedding model version; environmental condition features use the mean and standard deviation of a sliding time window as the standardization benchmark, and undergo dynamic Z-score normalization; S2.

2. Constructing a sub-matrix architecture: the standardized data are constructed into a multi-sub-matrix structure according to category, forming the overall feature matrix; S2.

3. Weighted Fusion: The eigenvectors of each sub-matrix are weighted and fused. The initial weights are set based on domain knowledge. During model training, the weights are dynamically adjusted using statistical correlation (such as mutual information or Pearson correlation coefficient) and through an attention mechanism. The specific fusion form is as follows: Where v i Let W be the eigenvector of the i-th submatrix. i Let α be a linear mapping matrix. i The weighting coefficients of the submatrix satisfy the following condition: The fusion results are used for model scoring and recommendation. The system adjusts each weight and standardized benchmark in real time based on subsequent experimental feedback, realizing the model's self-learning and dynamic optimization.

4. The method for screening and managing phage libraries based on artificial intelligence according to claim 1, characterized in that, Step S3 specifically involves: S3.

1. Constructing a dataset: The system first aggregates phage-host association data from phage databases, host strain experimental records, and third-party published literature to form a unified dataset; after data cleaning, the system treats each "phage-host pairing relationship" as a sample instance, with the sample label being "match / mismatch" or "lysis probability value"; subsequently, it divides the dataset into training, validation, and test sets in an 8:1:1 ratio to ensure the objectivity and repeatability of model evaluation. During the construction process, the system automatically detects and removes duplicate entries and obvious outliers to ensure that each sample has a complete feature vector input; S3.

2. Model Training: Based on the feature matrix input, machine learning and deep learning algorithms are used for modeling; During training, the model is optimized using either Adam or SGD algorithms, with an initial learning rate of 0.001–0.

01. The loss function is selected based on the task type, using either binary cross-entropy or mean squared error. Model training iterates for 100–300 rounds in a GPU environment until the validation set loss converges or the early stopping condition is triggered. S3.

3. Feedback Update and Model Reproducibility Mechanism: To ensure model reproducibility and continuous optimization, the system records all training parameters and stores them in a model version control library. Each training process can be fully reproduced based on the version number. Simultaneously, the system has a built-in "feedback update" module: when the results of actual application lysis experiments or animal experiments are returned, the system automatically appends new data to the sample library, re-standardizes it, and incorporates it into incremental training to achieve dynamic optimization. Model weight updates follow an exponential decay rule, and the model output includes predicted matching probabilities and feature importance ranking.

5. The method for screening and managing phage libraries based on artificial intelligence according to claim 1, characterized in that, The specific matching and scoring rules for step S4 are as follows: After receiving the feature input of the target bacteria, the AI ​​model outputs the matching probability value between each candidate phage and the bacteria. The model output value is normalized to the 0–1 range using the Sigmoid or Softmax function; it is further linearly mapped to the 0–100 range, and the calculation formula is: Scoresi = round(Pi × 100), where Pi is the phage-host matching probability predicted by the model. When a single phage score is ≥ 0.7 or ≥ 70, it is determined to be a recommended candidate. When multiple phage scores are in the 0.5–0.7 range, if their combined score is ≥ 0.75, then combination therapy is recommended. The system sorts the scores from high to low, outputs the top N candidate phages or phage combinations, and displays the reasons for the recommendation on the interface.

6. The method for screening and managing phage libraries based on artificial intelligence according to claim 1, characterized in that, Step S5 specifically involves using the difference between the model's predicted matching score and the experimental verification results as the basis for adjustment. When the experimental lysis rate is ≥80% and the error with the predicted value is ≤10%, the model prediction is confirmed to be effective, and the current parameters are maintained; when the difference between the experimental results and the prediction is >10%, the feature weights are updated; when a certain phage performs worse than expected in multiple verifications, its weight in the feature matrix is ​​reduced.

7. The application of the artificial intelligence-based phage library screening and management method as described in claim 1, characterized in that, Specifically, it can be used for rapid screening of suitable phage combinations in livestock and poultry farms; rapid recommendation of personalized phage therapy for clinically resistant strains in hospitals or testing centers; and matching suitable phage combinations for prevention and control in food safety testing.