Training methods, prediction methods, devices, terminal equipment, and storage media for bacterial pathogenicity probability prediction models based on whole-genome features.

By using a machine learning model based on whole-genome features, combined with the gene prevalence and intake dose of bacterial serotype samples, a dose-response relationship is constructed, which solves the problem of inaccurate prediction of pathogenicity probability in existing technologies and enables accurate risk assessment in the absence of data.

CN122493930APending Publication Date: 2026-07-31CHINA NAT CENT FOR FOOD SAFETY RISK ASSESSMENT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA NAT CENT FOR FOOD SAFETY RISK ASSESSMENT
Filing Date
2026-03-18
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, the pathogenicity probability prediction models for foodborne pathogens lack serotype and region specificity, leading to inaccurate risk assessments. In particular, it is difficult to accurately predict the pathogenicity probability of different serotypes and regions in the absence of feeding or outbreak data.

Method used

By acquiring genomic data from bacterial serotype samples, calculating gene prevalence, training models using machine learning methods, and combining gene prevalence with ingested doses, a dose-response relationship is constructed to form a disease probability prediction model at the individual and population levels.

Benefits of technology

In the absence of serotype and region-specific data, it can accurately predict the pathogenicity probability of bacteria, improve the accuracy and consistency of risk assessment, and reduce systematic bias.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493930A_ABST
    Figure CN122493930A_ABST
Patent Text Reader

Abstract

This invention relates to a training method, prediction method, apparatus, terminal device, and storage medium for a bacterial pathogenicity probability prediction model based on whole-genome features. The method includes: acquiring bacterial serotype samples, the serotype samples including a preset intake dose, a preset pathogenicity probability, and genomic data; calculating the gene prevalence of the genomic data of the serotype samples; determining training sample data based on the gene prevalence and the preset intake dose; training an initial training model based on the training sample data to obtain a bacterial pathogenicity probability prediction model based on whole-genome features. This model jointly models genomic features extracted from whole-genome information with the intake dose, and uses machine learning methods to obtain the dose-response relationship that varies with serotype / region. Even in the absence of strain feeding / outbreak data, a usable dose-response curve is obtained from the machine learning model, outputting the pathogenicity probability at the individual and population levels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of bioinformatics technology, and in particular relates to a training method, prediction method, device, terminal equipment and storage medium for a bacterial pathogenicity probability prediction model based on whole genome features. Background Technology

[0002] Currently, foodborne pathogens (such as Salmonella) Salmonella enterica Microbial risk assessment (QMRA) commonly employs a "farm-to-table" link modeling: contamination levels are propagated and transformed at stages such as raw materials, transportation and storage, preparation and processing, cooking and consumption, and the ingested dose (in CFU or log) is factored into the dose-response (DR) model at the consumption end. 10 CFU (Chronic Fuel Intake) is mapped to the probability of pathogenicity / onset of disease. Existing DR models mostly use exponential or β-Poisson models, and their parameters are usually derived from historical data from human feeding experiments or outbreak surveys. In actual assessments, common parameters are often used to approximate the average response of different strains / serotypes in the population.

[0003] DR parameters are derived by fitting pathogenicity data from a few typical serotypes, assuming that different serotypes / lineages have the same pathogenicity. However, in reality, there are significant differences in pathogenicity, genomics, and other aspects among different serotypes / lineages. Moreover, numerous studies have shown that genomic differences such as virulence factors (VFs) exist significantly between different serotypes and regions, leading to significant differences in pathogenicity of different serotypes and strains in different regions. Using universal DR parameters may systematically overestimate or underestimate the risk for specific serotypes or regions. Furthermore, emerging or low-frequency serotypes and strains from specific regions often lack corresponding feeding or outbreak dose data, making it difficult to establish or calibrate serotype / region-specific DR models. Therefore, how to predict the pathogenicity probability of a serotype based on known genomic information in the absence of specific feeding / outbreak data for other serotypes is a technical problem that needs to be solved. Summary of the Invention

[0004] This application aims to provide a training method, prediction method, device, terminal equipment, and storage medium for a bacterial pathogenicity probability prediction model based on whole-genome features, in order to overcome the shortcomings of the prior art. The technical problem to be solved by this application is achieved through the following technical solutions.

[0005] In a first aspect, embodiments of this application provide a method for training a bacterial pathogenicity probability prediction model based on whole-genome features, the method comprising:

[0006] Obtain a serotype sample, the serotype sample including a preset intake dose of bacterial serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and the genomic data of the serotype sample; Calculate the gene prevalence of the genomic data of the serotype samples; Training sample data are determined based on the gene prevalence of the genomic data of the serotype samples and the preset intake dose; Based on the training sample data, the initial training model is trained to obtain the output pathogenic probability; The preset pathogenicity probability and the output pathogenicity probability are compared; If the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value, the initial training model corresponding to the output pathogenicity probability is determined as a prediction model based on the whole genome pathogenicity probability.

[0007] Optionally, determining the training sample data based on the gene prevalence of the genomic data of the serotype samples and the preset intake dose includes: Identify candidate genes in each of the serotype samples; Calculate the gene prevalence of candidate genes in each serotype sample; The training sample data are determined based on the gene prevalence of the candidate genes and the preset intake dose.

[0008] Optionally, the initial training model includes at least a neural network or a gradient boosting-based tree model.

[0009] Optionally, training the initial training model based on the training sample data to obtain the output pathogenicity probability includes: The training sample data is divided into K-fold cross-partitions. For each first-level learner, at the k-th fold, the remaining k... The parameters are fitted using 1-fold data, and the output pathogenicity probability is generated using the training sample data of the k-th fold. After completing all folds, the first-level out-of-fold vector of each training sample data is obtained; At the next higher level, the out-of-fold prediction from the previous layer and the feature training of the training sample data are combined as input to train a new stacked model. The top layer uses a greedy ensemble algorithm to weight and combine candidate models from the penultimate layer to obtain the output pathogenic probability.

[0010] Secondly, embodiments of this application provide a training device for a bacterial pathogenicity probability prediction model based on whole-genome features, the device comprising: An acquisition module is used to acquire bacterial serotype samples, wherein the serotype samples include a preset intake dose of the serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and the genomic data of the serotype sample; A calculation module is used to calculate the gene prevalence of the genomic data of the serotype sample; The sample generation module is used to determine training sample data based on the gene prevalence of the genomic data of the serotype sample and the preset intake dose; The training module is used to train the initial training model based on the training sample data to obtain the output pathogenicity probability. The comparison module is used to compare the preset pathogenicity probability and the output pathogenicity probability; The prediction module is used to determine the initial training model corresponding to the output pathogenicity probability as a prediction model based on the whole genome pathogenicity probability when the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value.

[0011] Optionally, the sample generation module is used for: Identify candidate genes in each of the serotype samples; Calculate the gene prevalence of candidate genes in each serotype sample; The training sample data are determined based on the gene prevalence of the candidate genes and the preset intake dose.

[0012] Optionally, the initial training model includes at least a neural network or a gradient boosting-based tree model.

[0013] Optionally, the training module is used for: The training sample data is divided into K-fold cross-partitions. For each first-level learner, at the k-th fold, the remaining k... The parameters are fitted using 1-fold data, and the output pathogenicity probability is generated using the training sample data of the k-th fold. After completing all folds, the first-level out-of-fold vector of each training sample data is obtained; At the next higher level, the out-of-fold prediction from the previous layer and the feature training of the training sample data are combined as input to train a new stacked model. The top layer uses a greedy ensemble algorithm to weight and combine candidate models from the penultimate layer to obtain the output pathogenic probability.

[0014] Thirdly, embodiments of this application provide a method for predicting bacterial pathogenicity probability based on whole-genome characteristics, the method comprising: Obtain genomic data of the bacterial serotype to be predicted; The pathogenicity probability corresponding to the genomic data of the serotype to be predicted is determined according to the prediction model based on whole-genome pathogenicity probability as described in any of the first aspects.

[0015] Fourthly, embodiments of this application provide a terminal device, including: at least one processor and a memory; The memory stores a computer program; the at least one processor executes the computer program stored in the memory to implement the training method for the bacterial pathogenicity probability prediction model based on whole-genome features provided in the first aspect.

[0016] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed, implements the training method for the bacterial pathogenicity probability prediction model based on whole-genome features provided in the first aspect.

[0017] The embodiments of this application have the following advantages: The present application provides a training method, apparatus, prediction method, terminal device, and storage medium for a bacterial pathogenicity probability prediction model based on whole-genome features. This involves acquiring a serotype sample, which includes a preset intake dose of the serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and the genomic data of the serotype sample; calculating the gene prevalence of the genomic data of the serotype sample; determining training sample data based on the gene prevalence of candidate genes in the bacterial serotype sample and the preset intake dose; training an initial training model based on the training sample data to obtain an output pathogenicity probability; and then adjusting the preset pathogenicity probability and the... The output pathogenicity probability is compared; if the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value, the initial training model corresponding to the output pathogenicity probability is determined as a prediction model based on the whole genome pathogenicity probability. In this embodiment, bacterial genome features (such as gene prevalence, serotype, etc.) extracted from genome information are jointly modeled with the intake dose-pathogenicity rate, and the dose-response relationship with serotype / region is obtained through machine learning methods. The goal is to construct a usable dose-response curve through a genome feature-driven machine learning model even in the absence of strain feeding / outbreak data, and output the pathogenicity probability at the individual and population levels in a form compatible with the existing QMRA framework. Attached Figure Description

[0018] To more clearly illustrate the embodiments of this application or the existing technical solutions, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart of a training method for a bacterial pathogenicity probability prediction model based on whole-genome features in one embodiment of this application; Figure 2 This application demonstrates the predictive performance of a genome-dose response model based on 6354 strains of Salmonella from China for six serotypes in one embodiment of the present application. Figure 3 This is a structural block diagram of an embodiment of a training device for a bacterial pathogenicity probability prediction model based on whole-genome features, as described in this application. Figure 4 This is a schematic diagram of the structure of a terminal device according to this application. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] One embodiment of this application provides a training method for a bacterial pathogenicity probability prediction model based on whole-genome features, used to predict the bacterial pathogenicity probability of serotypes. The execution entity of this embodiment is a training device for the bacterial pathogenicity probability prediction model based on whole-genome features, which is installed on a terminal device, such as a computer terminal.

[0022] Reference Figure 1 The diagram illustrates a step-by-step flowchart of an embodiment of a training method for a bacterial pathogenicity probability prediction model based on whole-genome features, which specifically includes the following steps: S101. Obtain a bacterial serotype sample, wherein the serotype sample includes a preset intake dose of the serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and the genomic data of the serotype sample. The terminal device acquires Salmonella dose-response data and corresponding serotypes. The genomic data is a set of strain genomes of a certain serotype s in a certain region / origin.

[0023] S102. Calculate the gene prevalence of the genomic data of the serotype sample; Specifically, serotyping tools (such as SeqSero2) are used to analyze the serotype of each strain, BLAST is used to determine the presence or absence of candidate genes in each strain, and the prevalence of each candidate gene in each serotype is calculated.

[0024] S103. Based on the gene prevalence of the genomic data of the serotype samples and the preset intake dose, determine the training sample data; S104. Based on the training sample data, train the initial training model to obtain the output pathogenicity probability; S105. Compare the preset pathogenicity probability and the output pathogenicity probability; S106. If the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value, the initial training model corresponding to the output pathogenicity probability is determined as a prediction model based on the whole genome pathogenicity probability.

[0025] Specifically, in this embodiment, Seqsero2 is used to analyze the serotype of each bacterial strain (based on different typing of antigen genes), and BLAST is used to determine the presence or absence of candidate genes for each strain; the prevalence of each candidate gene for each bacterial serotype is calculated. Training sample data are determined based on the gene prevalence of the genomic data of the serotype samples and a preset intake dose.

[0026] The initial training model is trained based on the training sample data to obtain the output pathogenicity probability. If the difference between the preset pathogenicity probability and the output pathogenicity probability is less than the preset value, the initial training model corresponding to the output pathogenicity probability is determined as the prediction model.

[0027] The training method for a bacterial pathogenicity probability prediction model based on whole-genome features provided in this application embodiment involves: acquiring bacterial serotype samples, wherein the serotype samples include a preset intake dose of the serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and the genomic data of the serotype sample; calculating the gene prevalence of the genomic data of the serotype sample; determining training sample data based on the gene prevalence of the genomic data of the serotype sample and the preset intake dose; training an initial training model based on the training sample data to obtain an output bacterial pathogenicity probability; and adjusting the preset pathogenicity probability and the output pathogenicity probability. A comparison is made; if the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value, the initial training model corresponding to the output pathogenicity probability is determined as a prediction model based on the whole genome pathogenicity probability. In this embodiment, the genomic features (such as gene prevalence, serotype, etc.) extracted from bacterial genomic information are jointly modeled with the intake dose-pathogenicity rate, and the dose-response relationship with serotype / region is obtained through machine learning methods. The goal is to construct a usable dose-response curve through a genomic feature-driven machine learning model even in the absence of strain feeding / outbreak data, and output the pathogenicity probability at the individual and population levels in a form compatible with the existing QMRA framework.

[0028] Another embodiment of this application further supplements the description of the training method for the bacterial pathogenicity probability prediction model based on whole genome features provided in the above embodiments.

[0029] To address the issue that limited feeding / outbreak data for only a few serotypes makes it impossible to establish separate dose-response (DR) models for all serotypes, this invention proposes a dose-response (DR) modeling method based on the interaction characteristics of gene prevalence and dose (logCFU). Specifically, the prevalence (proportion between 0 and 1) of candidate genes (virulence genes in this example) in a strain set (which can be stratified by region, etc.) of a specific serotype is quantified and multiplied by the logarithm of the exposure dose (logCFU) to form a unified independent variable vector, i.e., training sample data. A neural network model is then jointly trained on all serotypes, enabling the prediction of pathogenicity probability based on genomic information even in the absence of specific feeding / outbreak data for other serotypes.

[0030] This approach, which uses genomic features to drive curve parameters, allows for the construction of usable DR relationships and smooth integration with existing QMRAs even when direct feeding / outbreak data for the target serotype (or even the region) is lacking.

[0031] Optionally, training sample data are determined based on the gene prevalence of the genomic data of the serotype samples and a preset intake dose, including: Identify candidate genes in each serotype sample; that is, obtain candidate genes from the genomic data of serotype samples. Calculate the gene prevalence of candidate genes in each serotype sample; Training sample data are determined based on the gene prevalence of candidate genes and the preset intake dose.

[0032] Specifically, the terminal device acquires a set of candidate genes (such as virulence genes) as g = {g1,…,gp} [p=N(total number of candidate genes)]. For serotypes s = {S1,…, SX} [X=N(total number of serotypes)] (which can be calculated separately for each country / region), the prevalence of each gene is calculated. p g (s) ∈ [0,1].

[0033]

[0034] in Indicates serotype Number of genomes, 1 ( ) is an indicator function (1 if it exists, 0 otherwise). Given an intake dose d (in CFU), take l = log10(d). Define the independent variable vector z( s , d ) = ( p g1 (s) l, p g1 (s) l,…, p gp (s) l); In this model, the dose-response information consists of the intake dose for a specific serotype and its pathogenicity rate. The input to this predictive model is a matrix, where each column is the product of gene prevalence and dose, and the number of columns represents the number of candidate genes. The number of rows represents the number of dose-pathogenicity data points.

[0035] For example, the terminal device collects two dose-response information entries: 1. Salmonella Typhimurium, intake dose d is 2.5 × 10 3 Under CFU, the morbidity rate was 0.83. 2. Salmonella enteritidis, intake dose is 4 × 10⁻⁶ per day. 2 Under CFU, the morbidity rate was 0.95. Obtain candidate genes for each serotype, for example, A1-A10.

[0036] The specific training process is as follows: 1. Calculate the prevalence of the A1-A10 genes in Salmonella Typhimurium / Salmonella Enteritidis, expressed as P0. 鼠1 -P 鼠10 P 肠1 -P 肠10 express: 2. The training data is as follows: [First row of training data] P 鼠伤寒沙门氏菌-A1 ×2.5×10 3 , P 鼠伤寒沙门氏菌-A2 ×2.5×10 3 , P 鼠伤寒沙门氏菌-A3 ×2.5×10 3 …P 鼠伤寒沙门氏菌-A10 ×2.5×10 3 [First row of predicted data] 0.83; [Second row of training data] P 肠炎沙门氏菌-A1 ×4×10 2 , P 肠炎沙门氏菌-A2 ×4×10 2 , P 肠炎沙门氏菌-A3 ×4×10 2 …P 肠炎沙门氏菌-A10 ×4×10 2 [Second row prediction data] 0.95; Optionally, the initial training model may include at least a neural network or a gradient boosting-based tree model.

[0037] Optionally, the initial training model is trained based on the training sample data to obtain the output pathogenicity probability, including: Perform K-fold cross-multiplication on the training sample data; For each first-level learner, at the k-th fold, the remaining k... The parameters are fitted using 1-fold data, and the output pathogenicity probability is generated using the training sample data of the k-th fold. After completing all folds, the first-level out-of-fold vector of each training sample data is obtained; At the next higher level, the out-of-fold prediction from the previous layer and the feature training of the training sample data are combined as input to train a new stacked model. The top layer uses a greedy ensemble algorithm to weight and combine candidate models from the penultimate layer to obtain the output pathogenicity probability.

[0038] Specifically, the terminal device constructs a stacked neural network (SNN), which is the initial neural network. The SNN is composed of multiple stacked modules arranged sequentially and outputs the morbidity rate. P illness(d,s)=fθ(z(s,d)). In this example, the network consists of L stacked units fθ, taking the interaction feature vector z(s,d) as input and the pathogenicity probability as output. The output layer uses a sigmoid function to map the values ​​to (0,1).

[0039] This application further employs model-level stacking ensemble to enhance robustness and generalization ability. Using the same input vector z(s,d) as the independent variable, multiple first-level base learners are trained to obtain out-of-fold (OOF) predictions. First, the set of first-level learners is set as M={m1,…,mM}, including but not limited to: table neural networks (MLP), gradient boosting-based tree models (such as LightGBM, XGBoost, CatBoost), random forests, etc. The training set is then partitioned into K-fold cross-multiplications; for each first-level learner mj, at the k-th fold, the remaining K... Parameters are fitted using 1-fold data, and the probability is generated using the Kth fold sample. P illness (d,s). After completing all folds, the first-order OOF vector ui=[ P (1),…, P (M)]. At the next higher layer, the OOF out-of-fold predictions from the previous layer are used together with the original features as input to train a new stacked model; the top layer employs a Greedy Weighted Ensemble algorithm, which assigns a different weight to each model and then combines the various models for output, achieving better results. The candidate models from the penultimate layer are weighted and combined to obtain the final probability output. (Number of layers) The hidden unit dimension, normalization type, activation function, and Dropout probability are determined by automatic hyperparameter search.

[0040] The training method for a bacterial pathogenicity probability prediction model based on whole-genome features provided in this application embodiment involves obtaining bacterial serotype samples, which include a preset intake dose of the serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and the genomic data of the serotype sample; calculating the gene prevalence of the genomic data of the serotype sample; determining training sample data based on the gene prevalence of the genomic data of the serotype sample and the preset intake dose; training an initial training model based on the training sample data to obtain an output pathogenicity probability; comparing the preset pathogenicity probability and the output pathogenicity probability; and determining the initial training model corresponding to the output pathogenicity probability as a prediction model based on whole-genome pathogenicity probability if the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value. In this application embodiment, by jointly modeling genomic features (such as gene prevalence, serotype, etc.) extracted from genomic information with intake dose-pathogenicity rate, and obtaining the dose-response relationship with serotype / regional variations through machine learning methods, the goal is to construct a usable dose-response curve through a genomic feature-driven machine learning model even in the absence of strain feeding / outbreak data, and to compare it with existing QMRA (Quantitative Methods for Responsive Microbiome Research). The framework-compatible format outputs the probability of disease at the individual and population levels.

[0041] This application provides a method for predicting pathogenicity probability based on the whole genome, including: Step 1: Obtain the genomic data of the serotype to be predicted; Specifically, the terminal device acquires the serotype to be predicted. Serotype is a method of classifying microorganisms (such as bacteria and viruses), primarily distinguishing them based on differences in their surface antigens (such as proteins, polysaccharides, and lipopolysaccharides). This classification method is based on differences in the immune responses (such as agglutination and neutralization reactions) between microorganisms and specific antibodies, thereby classifying different strains or viral strains of the same species into different subtypes. Humans determine their serotype by detecting the reaction between antigens and specific antibodies through serological experiments (such as agglutination tests and enzyme-linked immunosorbent assays (ELISA)) or by different typing of antigen genes such as polysaccharides.

[0042] Step 2: Based on the pre-trained prediction model, determine the pathogenicity probability corresponding to the serotype to be predicted based on the dosage. The pre-trained prediction model is obtained by determining the training sample data based on the gene prevalence of candidate genes and the intake dose-pathogenicity, and then using the training sample data to train the initial training model. Specifically, the terminal device collects serotype samples, calculates the gene prevalence of different candidate genes based on the candidate genes of the serotype samples, and the intake dose of each serotype sample to obtain training sample data. The initial training model is then trained based on the training sample data to obtain the output pathogenicity probability. Finally, through training, a prediction model is obtained.

[0043] The terminal device inputs the serotype to be predicted into a pre-trained prediction model to obtain the pathogenic probability corresponding to the serotype to be predicted.

[0044] This application example uses 6354 non-typhoid strains from China in the NCBI database. Salmonella enterica The data were derived from whole-genome data, with virulence genes used as candidate genes.

[0045] First, the serotypes of the strains were analyzed using Sequero2 software, and then the presence or absence of virulence genes was determined using BLAST software. The prevalence of virulence genes (VFs) was calculated by summarizing the results by serotype; the list and annotations of virulence genes were obtained from the VFDB database.

[0046] Subsequently, based on available dose-outcome data, determine Newport , Derby , Typhimurium , Enteritidis , Meleagridis For the training set, Bareilly External validation set. A stacked neural network was built using AutoGluon (v1.5.2) and compared to a baseline β-Poisson dose-response model [with parameters α=0.1324, β=51.45 (WHO framework)].

[0047] These data and baseline selections ensure the reproducibility of the methodology: genome and VF annotations can be stably obtained from NCBI and VFDB; the functional form and parameterization of the β-Poisson model are clearly defined and implemented in the WHO report, as shown in Table 1.

[0048] Table 1

[0049] (a) Predictive accuracy of external serotypes In those who did not participate in training BareillyRegarding serotypes, the model of this invention predicted 15.25% (16.67% observed) and 27.00% (33.33% observed) for the two independent observations, respectively; at the same dose, the traditional β-Poisson model gave 64.38% and 71.61%, respectively. In terms of mean absolute error (MAE), this method was 3.88%, while the β-Poisson model was 43.00%. This method maintains considerable consistency with observed results for serotypes not learned by the model, significantly reducing the systematic overestimation of dose-response models in traditional methods.

[0050] (ii) Availability of serotypes without feeding / outbreak data The top 6 most prevalent serotypes in China (among which) I 4,[5],12:i:-, Infantis (Given the lack of data from human feeding experiments or outbreak surveys for serotypes), the predictive models in this application can output monotonic, reasonably shaped serotype-specific dose-response curves, distinguishing the differences in threshold dose and slope between different serotypes. After outputting the serotype-specific dose-response curves, the threshold dose (ID50, median pathogenic dose) and local slope (derivative at ID50) are extracted for each serotype as two summary indicators. Figure 2 As shown, the combined distribution of these two indicators exhibits a separable pattern: one serotype shows a lower ID50 and a steeper slope (high pathogenicity serotype); the other shows a higher ID50 and a gentler slope (low pathogenicity serotype). Typhimurium , Enteritidis It has a high pathogenicity.

[0051] Another embodiment of this application provides a training device for a bacterial pathogenicity probability prediction model based on whole-genome features, used to execute the training method for the bacterial pathogenicity probability prediction model based on whole-genome features provided in the above embodiment.

[0052] Reference Figure 3 The diagram illustrates a structural block diagram of an embodiment of a training device for a bacterial pathogenicity probability prediction model based on whole-genome features, according to this application. The device may specifically include the following modules: an acquisition module 301, a calculation module 302, a sample generation module 303, a training module 304, a comparison module 305, and a prediction module 306, wherein: The acquisition module 301 is used to acquire bacterial serotype samples, the serotype samples including a preset intake dose of the serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and the genomic data of the serotype sample; The calculation module 302 is used to calculate the gene prevalence of the genomic data of the serotype sample; The sample generation module 303 is used to determine training sample data based on the gene prevalence of the genomic data of the serotype sample and the preset intake dose; The training module 304 is used to train the initial training model based on the training sample data to obtain the output pathogenicity probability. The comparison module 305 is used to compare the preset pathogenicity probability and the output pathogenicity probability; The prediction module 306 is used to determine the initial training model corresponding to the output pathogenicity probability as a prediction model based on the whole genome pathogenicity probability when the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value.

[0053] The training device for a bacterial pathogenicity probability prediction model based on whole-genome features provided in this application embodiment acquires serotype samples, which include a preset intake dose of the serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and genomic data of the serotype sample; calculates the gene prevalence of the genomic data of the serotype sample; determines training sample data based on the gene prevalence of the genomic data of the serotype sample and the preset intake dose; trains an initial training model based on the training sample data to obtain an output pathogenicity probability; compares the preset pathogenicity probability and the output pathogenicity probability; if the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value, the initial training model corresponding to the output pathogenicity probability is determined as a prediction model based on whole-genome pathogenicity probability. In this application embodiment, by jointly modeling genomic features (such as gene prevalence, serotype, etc.) extracted from genomic information with intake dose-pathogenicity rate, and obtaining the dose-response relationship with serotype / regional variations through machine learning methods, the goal is to construct a usable dose-response curve through a genomic feature-driven machine learning model even in the absence of strain feeding / outbreak data, and to compare it with existing QMRA (Quantitative Methods for Pathogen Response). The framework-compatible format outputs the probability of disease at the individual and population levels.

[0054] Another embodiment of this application further illustrates the training device for the bacterial pathogenicity probability prediction model based on whole-genome features provided in the above embodiments.

[0055] Optionally, the sample generation module is used for: Identify candidate genes in each of the serotype samples; Calculate the gene prevalence of candidate genes in each serotype sample; The training sample data are determined based on the gene prevalence of the candidate genes and the preset intake dose.

[0056] Optionally, the initial training model includes at least a neural network or a gradient boosting-based tree model.

[0057] Optionally, the training module is used for: The training sample data is divided into K-fold cross-partitions. For each first-level learner, at the k-th fold, the remaining k... The parameters are fitted using 1-fold data, and the output pathogenicity probability is generated using the training sample data of the k-th fold. After completing all folds, the first-level out-of-fold vector of each training sample data is obtained; At the next higher level, the out-of-fold prediction from the previous layer and the feature training of the training sample data are combined as input to train a new stacked model. The top layer uses a greedy ensemble algorithm to weight and combine candidate models from the penultimate layer to obtain the output pathogenic probability.

[0058] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0059] The training device for a bacterial pathogenicity probability prediction model based on whole-genome features provided in this application embodiment acquires serotype samples, which include a preset intake dose of the serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and genomic data of the serotype sample; calculates the gene prevalence of the genomic data of the serotype sample; determines training sample data based on the gene prevalence of the genomic data of the serotype sample and the preset intake dose; trains an initial training model based on the training sample data to obtain an output pathogenicity probability; compares the preset pathogenicity probability and the output pathogenicity probability; if the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value, the initial training model corresponding to the output pathogenicity probability is determined as a prediction model based on whole-genome pathogenicity probability. In this application embodiment, by jointly modeling genomic features (such as gene prevalence, serotype, etc.) extracted from genomic information with intake dose-pathogenicity rate, and obtaining the dose-response relationship with serotype / regional variations through machine learning methods, the goal is to construct a usable dose-response curve through a genomic feature-driven machine learning model even in the absence of strain feeding / outbreak data, and to compare it with existing QMRA (Quantitative Methods for Responsive Microbiome Research). The framework-compatible format outputs the probability of disease at the individual and population levels.

[0060] In another embodiment of this application, a terminal device is provided for executing the training method of the bacterial pathogenicity probability prediction model based on whole genome features provided in the above embodiments.

[0061] Figure 4 This is a structural schematic diagram of a terminal device according to this application, such as... Figure 4 As shown, the terminal device includes: at least one processor 401 and a memory 402; The memory stores a computer program; at least one processor executes the computer program stored in the memory to implement the training method for the bacterial pathogenicity probability prediction model based on whole-genome features provided in the above embodiments.

[0062] The terminal device provided in this embodiment acquires serotype samples, which include a preset intake dose of the serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and genomic data of the serotype sample; calculates the gene prevalence of the genomic data of the serotype sample; determines training sample data based on the gene prevalence of the genomic data of the serotype sample and the preset intake dose; trains an initial training model based on the training sample data to obtain an output pathogenicity probability; compares the preset pathogenicity probability and the output pathogenicity probability; if the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value, the initial training model corresponding to the output pathogenicity probability is determined as a prediction model based on whole-genome pathogenicity probability. In this embodiment, by jointly modeling genomic features (such as gene prevalence, serotype, etc.) extracted from genomic information with intake dose-pathogenicity, the dose-response relationship varying with serotype / region is obtained through machine learning methods. The goal is to construct a usable dose-response curve through a genomic feature-driven machine learning model even in the absence of strain feeding / outbreak data, and output the pathogenicity probability at the individual and population levels in a form compatible with the existing QMRA framework.

[0063] Another embodiment of this application provides a computer-readable storage medium storing a computer program, which, when executed, implements the training method for the bacterial pathogenicity probability prediction model based on whole-genome features provided in any of the above embodiments.

[0064] According to the computer-readable storage medium of this embodiment, by acquiring a serotype sample, the serotype sample includes a preset intake dose of the serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and genomic data of the serotype sample; calculating the gene prevalence of the genomic data of the serotype sample; determining training sample data based on the gene prevalence of the genomic data of the serotype sample and the preset intake dose; training an initial training model based on the training sample data to obtain an output pathogenicity probability; comparing the preset pathogenicity probability and the output pathogenicity probability; if the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value, determining the initial training model corresponding to the output pathogenicity probability as a prediction model based on whole-genome pathogenicity probability. In this embodiment, by jointly modeling genomic features (such as gene prevalence, serotype, etc.) extracted from genomic information with intake dose-pathogenicity, and obtaining the dose-response relationship with serotype / region variation through machine learning methods, the goal is to construct a usable dose-response curve through a genomic feature-driven machine learning model even in the absence of strain feeding / outbreak data, and to compare with existing QMRA. The framework-compatible format outputs the probability of disease at the individual and population levels.

[0065] It should be noted that the above detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0066] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0067] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that the embodiments of this application described herein can be implemented in sequences other than those illustrated or described herein.

[0068] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or apparatus.

[0069] For ease of description, spatial relative terms such as "above," "on top of," "on the upper surface of," "above," etc., are used herein to describe the spatial positional relationship of a device or feature as shown in the figures to other devices or features. It should be understood that spatial relative terms are intended to encompass different orientations in use or operation beyond the orientation of the device as described in the figures. For example, if the device in the figures were inverted, a device described as "above" or "on top of" other devices or structures would subsequently be positioned as "below" or "under" other devices or structures. Thus, the exemplary term "above" can include both "above" and "below." The device may also be positioned in other different ways, such as rotated 90 degrees or in other orientations, and the spatial relative descriptions used herein will be interpreted accordingly.

[0070] In the detailed description above, reference has been made to the accompanying drawings, which form part of this document. In the drawings, similar symbols typically identify similar parts unless the context otherwise indicates otherwise. The illustrated embodiments described in the detailed specification, drawings, and claims are not intended to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein.

[0071] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for training a bacterial pathogenic probability prediction model based on whole genome features, characterized in that, The method includes: Obtain bacterial serotype samples, wherein the serotype samples include a preset intake dose of the serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and genomic data of the serotype sample; Calculate the gene prevalence of the genomic data of the serotype samples; Training sample data are determined based on the gene prevalence of the genomic data of the serotype samples and the preset intake dose; Based on the training sample data, the initial training model is trained to obtain the output pathogenic probability; The preset pathogenicity probability and the output pathogenicity probability are compared; If the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value, the initial training model corresponding to the output pathogenicity probability is determined as a bacterial pathogenicity probability prediction model based on whole genome features. 2.The method of training a bacteria pathogenic probability prediction model based on whole genome features according to claim 1, wherein, The step of determining training sample data based on the gene prevalence of the genomic data of the serotype samples and the preset intake dose includes: Identify candidate genes in each of the serotype samples; Calculate the gene prevalence of candidate genes in each serotype sample; The training sample data are determined based on the gene prevalence of the candidate genes and the preset intake dose. 3.The method of claim 1, wherein the method further comprises: determining a feature set of the bacteria based on a plurality of genomic features; and determining a feature set of the bacteria based on a plurality of genomic features. The initial training model includes at least a neural network or a gradient boosting-based tree model. 4.The method of claim 1, wherein the method further comprises: determining a feature set of the bacteria based on a plurality of genomic characteristics of the bacteria; and determining a feature set of the bacteria based on a plurality of genomic characteristics of the bacteria. The step of training the initial training model based on the training sample data to obtain the output pathogenicity probability includes: The training sample data is divided into K-fold cross-partitions. For each first-level learner, at the k-th fold, the remaining k... The parameters are fitted using 1-fold data, and the output pathogenicity probability is generated using the training sample data of the k-th fold. After completing all folds, the first-level out-of-fold vector of each training sample data is obtained; At the next higher level, the out-of-fold prediction from the previous layer and the feature training of the training sample data are combined as input to train a new stacked model. The top layer uses a greedy ensemble algorithm to weight and combine candidate models from the penultimate layer to obtain the output pathogenic probability.

5. A training device for a bacterial pathogenicity probability prediction model based on whole-genome features, characterized in that, The device includes: An acquisition module is used to acquire bacterial serotype samples, wherein the serotype samples include a preset intake dose of the serotype sample, a preset pathogenicity probability corresponding to the preset intake dose, and the genomic data of the serotype sample; A calculation module is used to calculate the gene prevalence of the genomic data of the serotype sample; The sample generation module is used to determine training sample data based on the gene prevalence of the genomic data of the serotype sample and the preset intake dose; The training module is used to train the initial training model based on the training sample data to obtain the output pathogenicity probability. The comparison module is used to compare the preset pathogenicity probability and the output pathogenicity probability; The prediction module is used to determine the initial training model corresponding to the output pathogenicity probability as a bacterial pathogenicity probability prediction model based on whole genome features when the difference between the preset pathogenicity probability and the output pathogenicity probability is less than a preset value.

6. The training device for the bacterial pathogenicity probability prediction model based on whole-genome features according to claim 5, characterized in that, The sample generation module is used for: Identify candidate genes in each of the serotype samples; Calculate the gene prevalence of candidate genes in each serotype sample; The training sample data are determined based on the gene prevalence of the candidate genes and the preset intake dose.

7. The training device for the bacterial pathogenicity probability prediction model based on whole-genome features according to claim 5, characterized in that, The training module is used for: The training sample data is divided into K-fold cross-partitions. For each first-level learner, at the k-th fold, the remaining k... The parameters are fitted using 1-fold data, and the output pathogenicity probability is generated using the training sample data of the k-th fold. After completing all folds, the first-level out-of-fold vector of each training sample data is obtained; At the next higher level, the out-of-fold prediction from the previous layer and the feature training of the training sample data are combined as input to train a new stacked model. The top layer uses a greedy ensemble algorithm to weight and combine candidate models from the penultimate layer to obtain the output pathogenic probability.

8. A method for predicting bacterial pathogenicity probability based on whole-genome characteristics, characterized in that, The method includes: Obtain genomic data for the serotype to be predicted; According to the bacterial pathogenicity probability prediction model based on whole-genome characteristics as described in any one of claims 1-4, the pathogenicity probability corresponding to the genomic data of the serotype to be predicted is determined based on the dosage.

9. A terminal device, characterized in that, include: At least one processor and memory; The memory stores a computer program; the at least one processor executes the computer program stored in the memory to implement the training method of the bacterial pathogenicity probability prediction model based on whole-genome features as described in any one of claims 1-4.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains a computer program that, when executed, implements the training method for the bacterial pathogenicity probability prediction model based on whole-genome features as described in any one of claims 1-4.