A method and system for quantitatively assessing the risk of pathogen transmission.
By constructing a pathogen transmission risk assessment model and combining genome resequencing and multi-dimensional data, the problem of the single dimension of pathogen transmission risk assessment in existing technologies has been solved, and accurate prediction and assessment of pathogen transmission risk has been achieved.
Patent Information
- Application Number
- CN202511699322.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-11-19
AI Technical Summary
Current technologies lack multi-dimensional assessment models for the risk of pathogen transmission, making accurate prediction impossible.
By using genome resequencing technology to obtain virulence genes and mutation sites of pathogen samples, and combining environmental factors and host susceptibility risk indices, a virulence index model and a transmission risk assessment model are constructed for multi-dimensional evaluation.
It enables accurate prediction of pathogen transmission risk, improving the accuracy of assessment and prediction precision.
Smart Images

Figure CN121171645B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of public health safety technology, and in particular to a method and system for quantitatively assessing the risk of pathogen transmission. Background Technology
[0002] Currently, pathogen monitoring mainly relies on traditional bacteriological detection methods, including bacterial isolation and culture, biochemical identification, and serotyping. These methods have shortcomings in terms of timeliness, accuracy, and risk prediction capabilities. Although molecular biology techniques such as PCR and gene chips have improved detection efficiency to some extent, they still cannot achieve quantitative assessment and dynamic early warning of transmission risks. With the development of high-throughput sequencing technology, whole-genome resequencing provides a new technical means to deeply understand the genetic characteristics, virulence mechanisms, and transmission patterns of pathogens. However, existing technologies lack a systematic risk assessment model and method that can integrate multi-dimensional information from the genome, environment, epidemiology, and host.
[0003] Therefore, the technical problem that this application needs to solve is: how to achieve accurate prediction of the risk of pathogen transmission. Summary of the Invention
[0004] The main objective of this application is to provide a quantitative assessment method for the transmission risk of pathogens, aiming to address the problems of existing technologies that rely on a single dimension for risk assessment and cannot provide quantitative predictions. This method uses genome resequencing technology to obtain the virulence genes and corresponding variant sites of pathogen samples, calculates the virulence index using a virulence index model, and then combines environmental factors, transmission coefficients, and host susceptibility risk indices into the transmission risk assessment model, thereby assessing the transmission risk of pathogen samples from multiple dimensions. This approach enables accurate prediction of pathogen transmission.
[0005] In addition, a quantitative assessment system for the risk of pathogen transmission is also provided.
[0006] To achieve the above objectives, the technical solution adopted in this application is as follows:
[0007] A method for quantitatively assessing the risk of pathogen transmission includes the following steps:
[0008] Step 1: Obtain the virulence gene, reference virulence gene, environmental parameters, transmission factors, and host risk factor parameters of the pathogen sample;
[0009] Step 2: Use multiple mutation detection algorithms to detect mutations in the virulence gene, obtain all mutation sites, process all mutation sites, and obtain effective mutation sites;
[0010] Step 3: Construct a virulence index model. Use multiple sequence comparison to compare the effective variant sites with the reference virulence gene to obtain the variation distance between the effective variant sites and the reference virulence gene. Input the variation distance into the virulence index model to calculate the virulence index of the pathogen sample.
[0011] Step 4: Construct a transmission risk assessment model. Calculate the environmental factors, transmission coefficient, and host susceptibility risk index based on environmental parameters, transmission factors, and host risk factor parameters. Input the virulence index, environmental factors, transmission coefficient, and host susceptibility risk index into the transmission risk assessment model, and the model will output a risk value.
[0012] Preferably, step 2 includes the following sub-steps:
[0013] Step A1: Use multiple mutation detection algorithms to detect mutations in the virulence gene and obtain all mutation sites;
[0014] Step A2: Perform the first screening on all variant sites, and retain variant sites that are detected simultaneously by at least two variant detection algorithms as candidate variant sites;
[0015] Step A3: Perform a second screening of candidate variant sites, calculate the variant confidence of candidate variant sites, and compare it with the variant confidence threshold. Retain candidate variant sites with a variant confidence greater than the variant confidence as valid variant sites.
[0016] Preferably, the mutation detection algorithm includes the GATK HaplotypeCaller algorithm, the FreeBayes algorithm, and the SAMtools mpileup algorithm;
[0017] In step A3, the formula for calculating the confidence level of variation is:
[0018] .
[0019] in, This is an index for mutation detection algorithms. For the first Weights of mutation detection algorithms For the first The confidence score of the mutation detection algorithm after normalization.
[0020] Preferably, in step 3, the formula for the toxicity index model is:
[0021] .
[0022] in, The virulence index is denoted by n, where n is the total number of virulence genes. Let i be the weight of the i-th virulence gene. Let be the mutation distance of the i-th virulence gene. is the attenuation coefficient of the mutation on the i-th virulence gene.
[0023] Preferably, a correction function is constructed based on the phylogenetic tree, and the weights of virulence genes are calculated based on the correction function;
[0024] The correction function is: .
[0025] in, This represents the evolutionary distance between the pathogen sample and known highly virulent pathogens. The distance attenuation constant is This is the maximum gain coefficient;
[0026] The weighting formula for the virulence gene is as follows: .
[0027] in, Let be the initial weight of the i-th virulence gene.
[0028] Preferably, step 4 includes the following sub-steps:
[0029] Step B1: Use fuzzy logic algorithm to calculate environmental parameters and obtain environmental factors;
[0030] Step B2: Calculate the transmission factors based on classical epidemiological principles to obtain the transmission coefficient;
[0031] Step B3: Calculate the host risk factor parameters using the risk index formula to obtain the host susceptibility risk index. The risk index formula is as follows:
[0032] .
[0033] in As a host susceptibility risk index, The total number of host risk factors, For the first Risk score of each risk factor To carry the first The proportion or prevalence of each risk factor in the entire target population;
[0034] Step B4: Construct a transmission risk assessment model. Input the virulence index, environmental factors, transmission coefficient, and host susceptibility risk index into the transmission risk assessment model. The transmission risk assessment model outputs a risk value.
[0035] Preferably, the transmission risk assessment model is as follows:
[0036] .
[0037] in, It is a risk value. The toxicity index after normalization. As environmental factors, The normalized propagation coefficient is... This is the normalized host susceptibility risk index. , , , All are weighting coefficients.
[0038] Preferably, the method further includes step 5: classifying risk levels based on risk values and outputting early warning information and prevention and control recommendations.
[0039] It should be noted that:
[0040] GATK HaplotypeCaller Algorithm: GATK HaplotypeCaller is a genome variation detection tool based on local assembly technology. Its core algorithm identifies genetic variations such as SNPs (single nucleotide polymorphisms) and INDELs (insertions and deletions) by constructing candidate haplotypes and comparing them with a reference genome.
[0041] FreeBayes Algorithm: FreeBayes is a variant detection tool based on a Bayesian statistical framework, primarily used to accurately identify genomic variations, such as SNPs (single nucleotide polymorphisms) and INDELs (insertions and deletions), from high-throughput sequencing data. FreeBayes assesses the reliability of site variations using Bayesian formulas, prioritizing the exclusion of variations that occur only in a single strand or at a single location, as these are more likely sequencing errors than genuine biological variations. Its algorithm combines local recombination rate models and error models to improve the detection accuracy in complex genomic regions.
[0042] SAMtools mpileup algorithm: SAMtools mpileup is a tool in bioinformatics used to generate deep data of genomic loci. It generates pileup data based on BAM files (binary alignment file format) and is mainly used for genetic variation detection and genome analysis.
[0043] In addition, a quantitative assessment system for pathogen transmission risk is provided to implement the aforementioned quantitative assessment method for pathogen transmission risk, including the following modules:
[0044] Data acquisition module: used to acquire virulence genes, reference virulence genes, environmental parameters, transmission factors, and host risk factor parameters of pathogen samples;
[0045] The mutation detection and analysis module is used to perform mutation detection on the virulence gene using multiple mutation detection algorithms, obtain all mutation sites, process all mutation sites, and obtain effective mutation sites.
[0046] Virulence assessment module: Used to construct a virulence index model. It uses multiple sequence comparison to compare effective variant sites with reference virulence genes to obtain the variation distance between effective variant sites and reference virulence genes. The variation distance is then input into the virulence index model to calculate the virulence index of the pathogen sample.
[0047] The transmission risk calculation module is used to construct a transmission risk assessment model. Based on environmental parameters, transmission factors, and host risk factor parameters, it calculates environmental factors, transmission coefficients, and host susceptibility risk indices, respectively. The virulence index, environmental factors, transmission coefficients, and host susceptibility risk indices are then input into the transmission risk assessment model, which outputs a risk value.
[0048] Preferably, it also includes an early warning decision module: used to classify risk levels according to risk values and output early warning information and prevention and control suggestions.
[0049] Compared with the prior art, this application has the following beneficial effects:
[0050] The quantitative assessment method of this application can quickly obtain the virulence genes and corresponding mutation sites of pathogen samples through gene resequencing technology. Then, the virulence index is calculated through a virulence index model. The environmental factors, transmission coefficient and host susceptibility risk index are then input into the transmission risk assessment model to obtain the risk value. By combining multi-dimensional data to assess the transmission risk of pathogens, the accuracy of assessing the transmission risk of pathogen samples is greatly improved.
[0051] In the mutation detection and analysis process, multiple mutation detection algorithms were used to detect virulence genes, and mutation sites were screened according to relevant screening methods to make the detected effective mutation sites more representative and accurate. Then, the virulence index model calculated the virulence index of the pathogen sample based on the mutation distance. In this way, the virulence of the pathogen sample can be quantitatively evaluated.
[0052] Furthermore, weights are assigned to reference virulence genes, and a correction function is constructed using a phylogenetic tree. This correction function dynamically adjusts the weights of the reference virulence genes. The core idea behind this design is that the closer the phylogenetic relationship between the tested bacterium and a highly virulent bacterium, the greater the threat posed by its homologous virulence genes, and the corresponding weight should be increased. This approach allows for a more accurate assessment of the virulence of pathogen samples, improving prediction precision. Attached Figure Description
[0053] Figure 1 This is a flowchart of the variant detection and analysis in Example 1;
[0054] Figure 2 This is a flowchart of the method for quantitatively assessing the risk of pathogen transmission in Example 1;
[0055] Figure 3 This is a block diagram of the system for quantitatively assessing the risk of pathogen transmission in Example 2. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of this application implemented as described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0057] Example 1
[0058] refer to Figures 1-2 A method for quantitatively assessing the risk of pathogen transmission includes the following steps:
[0059] Step 1: Obtain the virulence gene, reference virulence gene, environmental parameters, transmission factors, and host risk factor parameters of the pathogen sample;
[0060] The pathogen is Streptococcus suis, which is a zoonotic pathogen. This example uses Streptococcus suis for illustration.
[0061] 1. Pathogen sample data: Nasal swabs were collected from suspected infected pigs on site. DNA extraction, quality inspection, and sequencing library construction were completed according to standard operating procedures. Whole genome resequencing data (FASTQ format) of the samples were obtained through a high-throughput sequencing platform (such as Illumina NovaSeq).
[0062] To obtain the virulence gene of *Streptococcus suis* for risk assessment, suspected infected biological samples (such as nasal swabs) must first be collected at the farm under strict aseptic procedures and immediately transported to the laboratory at low temperatures to ensure the activity of the pathogen and the integrity of the genetic material. In the laboratory, following standardized molecular biology procedures, bacterial cells are lysed using a bacterial genomic DNA extraction kit to isolate and purify the complete genomic DNA from the complex sample matrix. Subsequently, the extracted DNA must undergo rigorous quality control. Its purity and concentration are precisely detected using a micro-spectrophotometer and quantitative fluorescence analyzer, and its integrity is verified by gel electrophoresis to ensure that the DNA quality meets the stringent requirements of high-throughput sequencing. The quality-controlled DNA is then processed into a "library" suitable for sequencer reading: it is randomly fragmented into segments of specific lengths using physical or enzymatic methods, and synthetic adapters with known sequences are ligated to both ends of the fragments. Finally, the constructed sequencing library is placed on a high-throughput sequencing platform (such as Illumina NovaSeq) for deep sequencing to obtain a massive amount of DNA sequence fragment data. This raw data then undergoes professional bioinformatics analysis, including quality filtering, sequence assembly, and alignment with the reference genome, ultimately allowing for the precise location and extraction of complete virulence genes, laying a solid data foundation for subsequent variant analysis and risk assessment.
[0063] 2. Reference virulence genes: A database of Streptococcus suis virulence genes was established. Through a literature survey in PubMed and a search of the NCBI public database, known reference virulence genes of Streptococcus suis were collected and input into the database. Based on the importance of the reference virulence genes in the pathogenesis process, each reference virulence gene was assigned a corresponding initial weight.
[0064] The initial weights were determined using the analytic hierarchy process (AHP) combined with expert ratings and literature reports. Reference virulence genes were divided into three levels: key genes (initial weight 0.8–1), important genes (initial weight 0.5–0.7), and general genes (initial weight 0.2–0.4).
[0065] In this embodiment, the reference virulence genes were obtained through literature review and public database searches, and include at least the following genes, which were assigned initial weights by experts using the analytic hierarchy process (AHP). :
[0066] Lysozyme release protein gene (mrp): (Key genes);
[0067] Extracellular factor gene (epf): (Key genes);
[0068] hemolysin gene (sly): (Key genes);
[0069] Serum type 2 capsular polysaccharide synthesis gene (cps2J): (Important genes).
[0070] 3. Environmental parameters: Environmental parameters of the farm and its surroundings were collected, including real-time temperature, humidity, stocking density, ventilation, and disinfection frequency. In this embodiment, the real-time temperature was 38℃, humidity was 75%, stocking density was 1.5 heads / square meter, ventilation score was 0.8, and disinfection frequency was once a week.
[0071] 4. Transmission Factors: Estimated through farm records and epidemiological surveys, including infection source intensity I (detected by random sampling PCR), effective contact rate C, average infectious period length D, number of susceptible individuals N, and recovery rate. In this embodiment, infection source intensity I is 5%, effective contact rate C is twice daily, average infectious period length D is 7 days, number of susceptible individuals N is 1000 heads in total on the farm, immunization rate is 90%, so N=100, and recovery rate is 1 / D=1 / 7.
[0072] 5. Host risk factor parameters: Statistical analysis of pig herd structure and health status, see Table 3 in step B3 for details.
[0073] Step 2: Use multiple mutation detection algorithms to detect mutations in the virulence gene, obtain all mutation sites, process all mutation sites, and obtain effective mutation sites;
[0074] Preferably, step 2 includes the following sub-steps:
[0075] Step A1: Three mutation detection algorithms (GATK HaplotypeCaller, FreeBayes, SAMtoolsmpileup) are used to detect mutations in the virulence gene to obtain all mutation sites;
[0076] First, the virulence genes of the *Streptococcus suis* samples were preprocessed: FastQC software was used to perform a comprehensive quality assessment of the raw data, generating a detailed diagnostic report that systematically analyzed several key indicators, including the phred score for each base, the degree of adapter contamination, and the level of sequence repetition. Based on this quality assessment report, Trimmomatic software was then used for targeted data cleaning and filtering. This cleaning process included: accurately identifying and removing residual adapter sequences; trimming low-confidence bases from both ends of the sequences according to a preset quality threshold (e.g., a phred quality value below 20); and removing sequence fragments that were too short after trimming and unsuitable for alignment analysis.
[0077] Then, the virulence genes of the Streptococcus suis sample were sequence-aligned with the standard reference genome of Streptococcus suis (such as strain P1 / 7) using the GATK HaplotypeCaller algorithm, the FreeBayes algorithm, and the SAMtools mpileup algorithm, respectively, to obtain the variant sites detected by the GATK HaplotypeCaller algorithm, the FreeBayes algorithm, and the SAMtools mpileup algorithm, thus obtaining all variant sites;
[0078] It should be noted that the GATK HaplotypeCaller algorithm, FreeBayes algorithm, and SAMtools mpileup algorithm all output a confidence score after each mutation detection.
[0079] Step A2: Perform the first screening on all variant sites, and retain variant sites that are detected simultaneously by at least two variant detection algorithms as candidate variant sites;
[0080] All variant sites are first screened, and variant sites that are detected by only one variant detection algorithm are removed. Variant sites that are detected by two or three variant detection algorithms are retained as candidate variant sites.
[0081] Step A3: Perform a second screening of candidate variant sites, calculate the variant confidence of candidate variant sites, and compare it with the variant confidence threshold. Retain candidate variant sites with a variant confidence greater than the variant confidence as valid variant sites.
[0082] First, the confidence scores output by each algorithm for candidate variant sites are normalized. For example, the quality scores output by each algorithm (such as the QUAL value of GATK) are mapped to the [0, 1] interval using the max-min normalization method, resulting in... The maximum-minimum normalization method is as follows: subtract the minimum value from the original data, and then divide by the difference between the maximum and minimum values.
[0083] Based on historical data performance, corresponding weight values are assigned to the respective mutation detection algorithms. In this embodiment, the weight values of the GATK HaplotypeCaller algorithm The weight value of the FreeBayes algorithm is 0.4. The weight value of the SAMtools mpileup algorithm is 0.35. The value is 0.25. Further calculation of the confidence level of variation is performed using the following formula:
[0084] .
[0085] in, This is an index for mutation detection algorithms. For the first Weights of mutation detection algorithms For the first The confidence score of the mutation detection algorithm after normalization.
[0086] When only two mutation detection algorithms detect a valid mutation site, the weight value of the mutation detection algorithm that did not detect that mutation site is... and confidence level , and assign a specific value of 0.
[0087] A confidence threshold for the variant was set; in this embodiment, the confidence threshold was 0.75. This threshold was obtained by performing ROC curve analysis on known true positive and false positive variant datasets.
[0088] After obtaining the variation confidence score for each candidate variant site, it is compared with the variation confidence score threshold. Candidate variant sites with variation confidence scores lower than the variation confidence score threshold are eliminated, and candidate variant sites with variation confidence scores higher than the variation confidence score threshold are retained as valid variant sites.
[0089] Step 3: Construct a virulence index model. Use a phylogenetic tree to obtain the evolutionary distance between the test strain and the high-virulence reference strain, then calculate the weight of the virulence gene. Compare the effective variant sites with the reference virulence gene to obtain the variation distance between the effective variant sites and the reference virulence gene. Input the variation distance into the virulence index model to calculate the virulence index of the pathogen sample.
[0090] A core genome phylogenetic tree was constructed using standard bioinformatics procedures (RAxML software and GTRGAMMA model) to compare the test strain with a known highly virulent reference strain (such as SC19), and the evolutionary distance between the two strains was extracted from the tree file. (Branch length).
[0091] Based on evolutionary distance, construct a correction function: .
[0092] in, The maximum gain coefficient is a preset normal value, representing the maximum gain when the tested strain is completely identical to the highly virulent reference strain (i.e., ...). When ), the highest increase ratio that its initial weight can obtain. The distance decay constant is a preset positive coefficient used to adjust the decay rate of the weight gain as the evolution distance increases. The larger the value, the faster the weight gain decreases with increasing evolutionary distance.
[0093] Based on the correction function and initial weights, the weight of each virulence gene is calculated. The formula for calculating the weight of a virulence gene is as follows: .
[0094] in, The initial weights for virulence genes. This is a correction function.
[0095] The core idea of this design is that the closer the phylogenetic relationship between the Streptococcus suis to be tested and the highly virulent Streptococcus suis, the greater the threat posed by the homologous virulence genes it carries, and the corresponding weight should be increased accordingly. Conversely, the closer the phylogenetic relationship between the Streptococcus suis to be tested and the low-virulence Streptococcus suis, the smaller the threat posed by the homologous virulence genes it carries, and the corresponding weight should be decreased accordingly. In this way, the weight of each reference virulence gene can be dynamically adjusted to improve the reliability of the prediction.
[0096] Then, taking into account the weight of each reference virulence gene, the mutation distance, and the degree of influence of the mutation on gene function, the virulence index of the tested Streptococcus suis is calculated using a virulence index model. The formula for the virulence index model is as follows:
[0097] .
[0098] in, The virulence index is denoted by n, where n is the total number of virulence genes. Let be the weight of the i-th virulence gene.
[0099] Let be the variation distance between the effective variant site and the reference virulence gene on the i-th virulence gene. The effective variant site and the reference virulence gene are compared using Muscle, MAFFT, ClustalW, or T-Coffee multiple sequence alignment methods to obtain the variation distance between the effective variant site and the reference virulence gene. Among them, Muscle, MAFFT, ClustalW, and T-Coffee are all multiple sequence alignment software.
[0100] To ensure the standardization and reproducibility of the calculation process, this application defines the following three specific calculation rules for different types of gene variations:
[0101] 1. For point mutations (SNPs): This is determined by calculating the density of effective variant sites in the gene sequence. The calculation formula is as follows:
[0102] .
[0103] in, The total number of valid point mutations detected on the i-th virulence gene; The total length of the i-th virulence gene (unit: base pairs bp).
[0104] 2. For insertions / deletions (Indels): The impact of Indels is related to both the length and location of the mutation; therefore, a positional weighting factor P is introduced for weighted calculation. The calculation formula is as follows:
[0105] .
[0106] Where m is the total number of valid Indels detected on the i-th gene; Let be the absolute value of the length of the j-th Indel; The total length of the gene; Let j be the position weighting factor for the j-th Indel. If the Indel is located within a known protein active site or key functional domain... ;otherwise, The functional domain information is determined by comparing the protein sequence with public authoritative databases (e.g., Pfam version 35.0, PROSITE 2023_01).
[0107] 3. For structural variations (SVs): For structural variations with significant impact, a rule-based assignment method is used: if an SV (e.g., a deletion >50 bp) occurs within a single gene, di is calculated according to the proportion of the sequence it affects. The specific calculation formula is as follows:
[0108] .
[0109] in, It is the absolute value of the length of the structural variation detected on the i-th virulence gene; This represents the total length of the gene. If the breakpoint of an SV (such as an inversion or translocation) is located inside the gene, causing damage to the gene structure, it is directly assigned a value. If the entire virulence gene is detected to have been deleted, assign a value. .
[0110] When multiple types of effective variants exist on the same virulence gene, the total variation distance The mean distance calculated for each type of variation, i.e. .
[0111] This represents the attenuation coefficient of the i-th virulence gene. The attenuation coefficient is determined based on the degree of influence of the variation on protein function, as shown in Table 1.
[0112] Table 1 Attenuation Coefficient Standard
[0113]
[0114] To ensure consistent scaling of all factors in the subsequent transmission risk assessment model, the virulence index needs to be normalized to obtain a normalized virulence index. ,in, The toxicity index is calculated using the toxicity index model. It is the theoretical maximum virulence index, which is the sum of the integrals when all genes have not mutated. , Let be the weight of the i-th virulence gene.
[0115] The method of the present invention will be described below using a specific Streptococcus suis sample as an example.
[0116] In this embodiment, Set the maximum gain coefficient =0.5, distance attenuation constant K=10, can be calculated as follows: .
[0117] In this embodiment, the weights of each virulence gene are as follows:
[0118] ;
[0119] ;
[0120] ;
[0121] ;
[0122] Gene mrp (532bp in length): One missense mutation was found, located in a non-critical region.
[0123] ;
[0124] (According to Table 1);
[0125] Gene sly (495bp in length): No valid variants were found.
[0126] ;
[0127] No need to consider it.
[0128] No effective mutations were found in other virulence genes (such as epf, cps2J). All are 0.
[0129] Substituting into the formula, the result is:
[0130] ;
[0131] ;
[0132] .
[0133] Step 4: Construct a transmission risk assessment model. Calculate the environmental factors, transmission coefficient, and host susceptibility risk index based on environmental parameters, transmission factors, and host risk factor parameters. Input the virulence index, environmental factors, transmission coefficient, and host susceptibility risk index into the transmission risk assessment model, and the model will output a risk value.
[0134] Preferably, step 4 includes the following sub-steps:
[0135] Step B1: Use fuzzy logic algorithm to calculate environmental parameters and obtain environmental factors;
[0136] Environmental factors aim to quantify the comprehensive impact of external environmental conditions on the survival, reproduction, and transmission of Streptococcus suis. Since the relationship between environmental parameters and pathogen transmission is often complex and nonlinear, and the measurement of environmental parameters themselves is uncertain, this application employs a fuzzy logic algorithm to construct an evaluation model for environmental factors.
[0137] The specific formula for the fuzzy logic algorithm is as follows:
[0138] .
[0139] in, This is an environmental factor with a value range of [0, 1]. The total number of environmental parameters considered. This represents the measured value of the k-th environmental parameter. Let be the weight of the k-th environmental parameter, and This weight can be determined by combining the analytic hierarchy process with expert scoring. The membership function for the k-th parameter is used to assign the actual measured value. It is mapped to a risk membership value in the interval [0, 1].
[0140] By inputting the specific values of the environmental parameters obtained in step 1 into the formula of the fuzzy logic algorithm above, the environmental factors for this operation can be calculated.
[0141] It should be noted that: membership function The specific form needs to be determined based on the biological significance of each parameter:
[0142] (1) For parameters with an optimal range (such as temperature and humidity): use trapezoidal or triangular membership functions. When the parameter value is in the range that is unfavorable to the survival of pathogens, the risk membership value is close to 1; when the parameter is in the optimal survival range, the risk membership value is close to 0.
[0143] (2) For parameters with monotonic effects (such as stocking density): use the sigmoid function or piecewise linear membership function.
[0144] In this embodiment, the selected environmental parameters and their weights and membership functions are defined as shown in Table 2:
[0145] Table 2. Definition of Environmental Parameters, Their Weights, and Membership Functions
[0146]
[0147] The numerical values obtained in step 1 are converted, and the results are as follows:
[0148] Temperature (38℃): >35℃, within the highest risk range. .
[0149] Humidity (75%): Between 70% and 80%, the risk increases linearly. .
[0150] Stocking density (1.5 heads / m²): Within the risk-increasing range. .
[0151] Ventilation condition (0.8): .
[0152] Disinfection frequency (once a week): .
[0153] The environmental factors for this assessment can be obtained by weighting and summing the weights calculated from Table 2.
[0154] .
[0155] Step B2: Calculate the transmission factors based on classical epidemiological principles to obtain the transmission coefficient;
[0156] The transmission coefficient is a core epidemiological parameter that quantifies the actual transmission capacity or efficiency of *Streptococcus suis* within a specific host population and environmental context. The transmission coefficient is not a static constant but a dynamically changing indicator that reflects the immediate potential for epidemic development. This embodiment will use a mathematical model based on classical epidemiological principles to analyze the transmission coefficient. The mathematical model is as follows:
[0157] .
[0158] in, For the propagation coefficient, In terms of the intensity of the infection source, For contact rate, The length of the infectious period. For the number of susceptible individuals, This refers to the recovery rate.
[0159] To ensure consistent scaling of all factors in the final model, the original propagation coefficients need to be normalized and mapped to the interval [0, 1], resulting in the normalized propagation coefficients:
[0160] .
[0161] in, This is a normalization constant representing a coefficient value for a high-risk transmission scenario determined based on historical data, used to calibrate the assessment scale. This constant can be calculated using parameters from a typical historical outbreak event that reaches a "high-risk" level, resulting in a Tcoeff value, known as Tcoeff_norm. In this embodiment, Tcoeff_norm is defaulted to 0.5, representing a relatively high transmission efficiency sufficient to cause rapid spread of the epidemic in this type of aquaculture environment.
[0162] Input the propagation factors obtained in step 1 into the mathematical model.
[0163] 1. Calculate the propagation coefficient: .
[0164] 2. Calculate the normalized propagation coefficient: .
[0165] It should be noted that:
[0166] Infection source strength This reflects the number or proportion of infected individuals in the environment that are present and able to shed pathogens. This data can be obtained through regular serological monitoring or etiological surveys of pig herds.
[0167] Effective contact rate The number of contacts between a susceptible individual and an infection source per unit of time that are sufficient to cause transmission is highly dependent on factors such as stocking density, pen design, mixed-pens behavior of pigs, and management of personnel and material flow. It can be estimated indirectly through animal behavior tracking and analysis using a sensor network deployed on the farm, or through detailed production management records.
[0168] Length of infectious period This refers to the average duration for which a single infected individual can shed pathogens and infect other individuals. This parameter is usually determined through clinical observation and laboratory animal infection experiments.
[0169] Total number of susceptible individuals This refers to the number of individuals in a specific area (such as a farm or an administrative village) that do not have effective immunity to the currently prevalent Streptococcus suis strain. It can be calculated by subtracting the number of individuals with immunity (through immunization or recovery from previous infection) from the total number of animals in stock.
[0170] recovery rate , which represents the rate at which infected individuals leave the infected population, including both individuals who have recovered and acquired immunity, and individuals removed due to death or culling. In simplified models, it is usually regarded as the reciprocal of the length of the infectious period.
[0171] Step B3: Calculate the host susceptibility risk index The host susceptibility risk index aims to quantify the overall health level and susceptibility to pathogens of a target host population. The risk status of a population is not determined by a single individual, but rather by the distribution of different risk factors within the population. This application calculates this index using a weighted risk model. The risk index formula is:
[0172] .
[0173] in, As a host susceptibility risk index; The total number of predefined host risk factors; For the first The risk score for a host risk factor is a semi-quantitative value (e.g., in the range of 1 to 10) set according to the severity of the impact of the risk factor on an individual's susceptibility. This score is determined primarily based on authoritative veterinary epidemiological research findings and broad expert consensus. To carry the first The proportion or prevalence of an individual host risk factor in the entire target population is empirical data that needs to be obtained through actual surveys and data statistics (such as reviewing breeding records and conducting laboratory tests).
[0174] To ensure the comparability of this index with other risk factors (such as virulence index), the host susceptibility risk index needs to be normalized and mapped to the [0, 1] interval. The normalized host susceptibility risk index is: .in, It is the theoretical maximum risk score, calculated as the sum of the hazard scores of all risk factors, i.e. .
[0175] By inputting the host risk factor parameters obtained in step 1 into the risk index formula, the host susceptibility risk index can be calculated. The host susceptibility risk index directly reflects the overall health level and disease resistance of the host population. The higher the host susceptibility risk index, the greater the possibility and severity of an outbreak when the population faces pathogen challenges.
[0176] It should be noted that:
[0177] For the first A risk factor risk score is a semi-quantitative or quantitative value set according to the severity of the host risk factor's impact on an individual's susceptibility. Its setting is primarily based on authoritative veterinary epidemiological research findings and broad expert consensus. For example, in pig herds, age is an extremely important host risk factor. Weaned piglets are far more susceptible than adult pigs due to the decline of maternal antibodies and their immature immune systems, thus they can be assigned a higher risk score. Other important risk factors include: the herd's immune status (e.g., unvaccinated or untimely vaccinated individuals have significantly increased risk), the pig's genetic background (certain breeds or strains may show higher susceptibility to specific serotypes of Streptococcus suis), nutritional status (malnutrition weakens the immune response), and the presence of concurrent infections (which increase the burden on the organism and the chance of secondary infections).
[0178] Representative carrying the first This refers to the proportion or prevalence of an individual risk factor within the entire target population. This is empirical data that requires actual surveys and statistical analysis. For example, to determine the risk associated with age structure, it is necessary to determine the proportion of piglets and nurseries in the entire pig herd. To assess the risk of immunization status, it is necessary to accurately calculate the proportion of unvaccinated or improperly immunized pigs based on immunization records.
[0179] In this embodiment, the host risk factors and their parameters are shown in Table 3:
[0180] Table 3 Host risk factors and their parameters
[0181] Host risk factor (i) <![CDATA[Hazard score (h i )]]> Prevalence (p i ) Acquisition mode Age: weaned piglets 10 Statistical pig farm inventory data, weaned piglets accounted for 20% Immune status: unvaccinated 8 Check the immunization records, unvaccinated accounted for 10% Concurrent infection: porcine reproductive and respiratory syndrome 7 Clinical diagnosis, prevalence rate in the field 5%
[0182] Substitute the data from Table 3 into the formula and perform the calculation:
[0183] 1. Calculate the host susceptibility risk index: ;
[0184] 2. Calculate the theoretical maximum host susceptibility risk index: ;
[0185] 3. Calculate the normalized host susceptibility risk index: .
[0186] Step B4: Construct a transmission risk assessment model. Input the virulence index, environmental factors, transmission coefficient, and host susceptibility risk index into the transmission risk assessment model. The transmission risk assessment model outputs a risk value.
[0187] The transmission risk assessment model is as follows:
[0188] .
[0189] in, The toxicity index after normalization. As environmental factors, The normalized propagation coefficient is... This is the normalized host susceptibility risk index. , , , These are all weighting coefficients. To determine these weights, a dataset containing historical Streptococcus suis outbreaks needs to be collected. Each event includes strain genome, environmental, transmission, and host data, and is ultimately labeled by veterinary experts with an actual risk level (integer from 1 to 5) based on the size and severity of the outbreak, serving as labels for model training. This embodiment... , , , The result was obtained by training a random forest model on historical data. , , , The calculated , , , Substitute the values to obtain the final risk value. .
[0190] The factors obtained above are input into the mathematical model to calculate the risk factor value R. total =0.4×0.998+0.2×0.6401+0.2×0.098+0.2×0.126=0.57202.
[0191] Preferably, the method further includes step 5: classifying risk levels based on risk values and outputting early warning information and prevention and control recommendations.
[0192] The risk value ranges from [0, 1], and the risk level is divided into five levels: extremely low risk ( Range is 0-0.2), low risk ( The range is 0.2 to 0.4, and the risk is medium. Range is 0.4–0.6), high risk ( Range of 0.6–0.8) and extremely high risk ( (0.8–1.0).
[0193] Extremely low risk (green alert): Maintain regular monitoring frequency and continue to implement standard biosafety measures. Routine testing is recommended monthly, with a focus on high-risk areas and vulnerable groups.
[0194] Low risk (blue alert): Increase monitoring density appropriately and strengthen biosecurity measures. It is recommended to conduct testing every two weeks, implement strict disinfection procedures at farms, and restrict the movement of people and vehicles.
[0195] Medium risk (yellow alert): Activate the key monitoring mode and implement preventive intervention measures. It is recommended to conduct weekly testing, perform genomic tracing analysis on high-risk strains, and prepare emergency prevention and control supplies.
[0196] High risk (orange alert): Activate the emergency response mechanism and implement strict prevention and control measures. Daily monitoring is recommended, along with sealing off and isolating affected areas, conducting epidemiological investigations, and developing a vaccination plan.
[0197] Extremely high risk (red alert): Activate the highest level of emergency response and implement comprehensive prevention and control measures. Real-time monitoring is recommended, along with a complete lockdown of the affected area, initiating emergency vaccination efforts, and coordinating multi-departmental joint prevention and control efforts.
[0198] The calculations obtained in this embodiment are as follows If the risk level is "medium risk," the system will automatically trigger a yellow alert and output prevention and control recommendations: "Activate key monitoring mode and implement preventive intervention measures. It is recommended to conduct weekly testing, perform genomic tracing analysis on high-risk strains, and prepare emergency prevention and control supplies."
[0199] This method allows for multi-dimensional quantitative assessment of the transmission risk of pathogen samples, enabling accurate prediction of pathogen transmission.
[0200] Example 2
[0201] refer to Figure 3 A system for quantitatively assessing the risk of pathogen transmission, used to implement the aforementioned method for quantitatively assessing the risk of pathogen transmission, includes the following modules:
[0202] Data acquisition module: used to acquire virulence genes, reference virulence genes, environmental parameters, transmission factors, and host risk factors of pathogen samples;
[0203] The mutation detection and analysis module is used to perform mutation detection on the virulence gene using multiple mutation detection algorithms, obtain all mutation sites, process all mutation sites, and obtain effective mutation sites.
[0204] Virulence assessment module: Used to construct a virulence index model. It uses multiple sequence comparison to compare effective variant sites with reference virulence genes to obtain the variation distance between effective variant sites and reference virulence genes. The variation distance is then input into the virulence index model to calculate the virulence index of the pathogen sample.
[0205] The transmission risk calculation module is used to construct a transmission risk assessment model. Based on environmental parameters, transmission factors, and host risk factors, it calculates environmental factors, transmission coefficients, and host susceptibility risk indices, respectively. The virulence index, environmental factors, transmission coefficients, and host susceptibility risk indices are then input into the transmission risk assessment model, which outputs a risk value.
[0206] Preferably, it also includes an early warning decision module: used to classify risk levels according to risk values and output early warning information and prevention and control suggestions.
[0207] In this embodiment, the specific workflow of the system for assessing the transmission risk of pathogens is as follows: the data acquisition module acquires the virulence gene of the pathogen sample to be tested, the reference virulence gene of the pathogen, environmental parameters, transmission factors, and host risk factor parameters, and inputs the virulence gene of the pathogen sample to be tested into the variation detection and analysis module, inputs the reference virulence gene of the pathogen into the virulence assessment module, and inputs the environmental parameters, transmission factors, and host risk factor parameters into the transmission risk calculation model;
[0208] After receiving the virulence gene of the pathogen sample to be detected, the mutation detection and analysis module uses the GATKHaplotypeCaller algorithm, the FreeBayes algorithm, and the SAMtools mpileup algorithm to perform sequence alignment between the virulence gene of the Streptococcus suis sample and the standard reference genome of Streptococcus suis (such as P1 / 7 strain). The module obtains the mutation sites detected by the GATKHaplotypeCaller algorithm, the FreeBayes algorithm, and the SAMtools mpileup algorithm, respectively. Based on retaining mutation sites detected by at least two mutation detection algorithms simultaneously, and the mutation confidence of the mutation site is greater than the mutation confidence threshold, the module selects effective mutation sites and inputs them into the virulence assessment module.
[0209] The virulence assessment module first constructs a virulence index model. After receiving the effective variant sites and reference virulence genes, it uses a multiple sequence comparison method to compare the effective variant sites with the reference virulence genes to obtain the variation distance between the effective variant sites and the reference virulence genes. Based on the weight of each reference virulence gene, the variation distance, and the degree of influence of the variation on gene function, the virulence index of the pathogen sample to be tested is calculated, and the virulence index is input into the transmission risk calculation module.
[0210] After receiving the virulence index, environmental parameters, transmission factors, and host risk factor parameters, the transmission risk calculation module uses a fuzzy logic algorithm to calculate the environmental parameters to obtain the environmental factors. Based on classical epidemiological principles, it calculates the transmission factors to obtain the transmission coefficient. It then uses the risk index formula to calculate the host risk factor parameters to obtain the host susceptibility risk index. The virulence index, environmental factors, transmission coefficient, and host susceptibility risk index are then input into the transmission risk assessment model. The transmission risk assessment model outputs a risk value, which is then input into the early warning decision module.
[0211] The early warning decision module divides the risk of transmission into five risk levels: extremely low risk (risk value range 0-0.2), low risk (risk value range 0.2-0.4), medium risk (risk value range 0.4-0.6), high risk (risk value range 0.6-0.8), and extremely high risk (risk value range 0.8-1.0). After receiving the risk value, the early warning decision module determines which risk level the risk value falls into, and automatically generates corresponding early warning information and prevention and control suggestions for different risk levels.
[0212] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
Claims
1. A method for quantitatively assessing the risk of pathogenic bacteria transmission, characterized by, The method comprises the following steps: Step 1: obtaining the virulence gene, reference virulence gene, environmental parameter, transmission factor and host risk factor parameter of the pathogenic bacteria sample; Step 2: performing variation detection on the virulence gene by using multiple variation detection algorithms to obtain all variation sites, and processing the all variation sites to obtain effective variation sites; Step 3: constructing a virulence index model, comparing the effective variation sites with the reference virulence gene by using a multiple sequence comparison method to obtain a variation distance between the effective variation sites and the reference virulence gene, inputting the variation distance into the virulence index model, and calculating the virulence index of the pathogenic bacteria sample; Step 4: constructing a transmission risk assessment model, respectively calculating an environmental factor, a transmission coefficient and a host susceptibility risk index according to the environmental parameter, the transmission factor and the host risk factor parameter, inputting the virulence index, the environmental factor, the transmission coefficient and the host susceptibility risk index into the transmission risk assessment model, and outputting a risk value by the transmission risk assessment model.
2. The method of quantifying the risk of pathogenic bacteria transmission according to claim 1, characterized in that, The step 2 comprises the following sub-steps: Step A1: performing variation detection on the virulence gene by using multiple variation detection algorithms to obtain all variation sites; Step A2: performing first screening on the all variation sites, retaining variation sites detected by at least two variation detection algorithms at the same time as candidate variation sites; Step A3: performing second screening on the candidate variation sites, calculating a variation confidence of the candidate variation sites, comparing the variation confidence with a variation confidence threshold, and retaining the candidate variation sites greater than the variation confidence threshold as effective variation sites.
3. The method of quantifying the risk of pathogenic bacteria transmission according to claim 2, characterized in that, The variation detection algorithms comprise a GATK HaplotypeCaller algorithm, a FreeBayes algorithm and a SAMtools mpileup algorithm; In the step A3, the calculation formula of the variation confidence is: ; in, This is an index for mutation detection algorithms. For the first Weights of mutation detection algorithms For the first The confidence score of the mutation detection algorithm after normalization.
4. The method of quantifying the risk of pathogenic bacteria transmission according to claim 1, wherein, In the step 3, the formula of the virulence index model is: ; wherein, is the virulence index, n is the total number of virulence genes, is the weight of the ith virulence gene, is the mutation distance of the ith virulence gene, is the decay coefficient of the mutation on the ith virulence gene.
5. The method of quantifying the risk of pathogenic bacteria transmission according to claim 4, characterized in that, A correction function is constructed based on a phylogenetic tree, and the weight of the virulence gene is calculated according to the correction function; The correction function is: ; wherein, is the evolutionary distance between the pathogenic sample and the known highly virulent pathogenic bacteria, is the distance decay constant, is the maximum gain coefficient; The calculation formula of the weight of the virulence gene is: ; wherein, is the initial weight of the ith virulence gene.
6. The method of quantifying the risk of pathogenic bacteria transmission according to claim 1, wherein, The step 4 comprises the following sub-steps: Step B1: calculating the environmental parameter by using a fuzzy logic algorithm to obtain an environmental factor; Step B2: calculating the transmission factor based on a classical epidemiology principle to obtain a transmission coefficient; Step B3: calculating the host risk factor parameter by using a risk index formula to obtain a host susceptibility risk index, and the risk index formula is: ; wherein is the host susceptibility risk index, is the total number of host risk factors, is the risk score of the risk factor, is the prevalence or proportion of individuals carrying the risk factor in the entire target population; Step B4: constructing a transmission risk assessment model, inputting the virulence index, the environmental factor, the transmission coefficient and the host susceptibility risk index into the transmission risk assessment model, and outputting a risk value by the transmission risk assessment model.
7. The method of quantifying the risk of pathogenic bacteria transmission according to claim 6, characterized in that, The transmission risk assessment model is: ; wherein, is a risk value, is a normalized virulence index, is an environmental factor, is a normalized transmission coefficient, is a normalized host susceptibility risk index, , , , are all weight coefficients.
8. The method of quantifying the risk of pathogenic bacteria transmission according to claim 1, wherein, The method further comprises the following step 5: dividing a risk level according to the risk value, and outputting early warning information and prevention and control suggestions.
9. A system for quantitatively assessing the risk of pathogen transmission, characterized by, The method for quantitatively evaluating the transmission risk of pathogenic bacteria according to any one of claims 1-8 comprises the following modules: A data acquisition module for acquiring the virulence gene, reference virulence gene, environmental parameter, transmission factor and host risk factor parameter of the pathogenic bacteria sample; The variation detection analysis module is configured to detect variations of the virulence genes by using multiple variation detection algorithms, obtain all variation sites, and process the all variation sites to obtain effective variation sites. The virulence evaluation module is configured to construct a virulence index model, compare the effective variation sites with reference virulence genes by using a multiple sequence comparison method to obtain variation distances between the effective variation sites and the reference virulence genes, and input the variation distances into the virulence index model to calculate a virulence index of the pathogenic bacteria sample. The transmission risk calculation module is configured to construct a transmission risk evaluation model, calculate an environmental factor, a transmission coefficient and a host susceptibility risk index according to environmental parameters, transmission factors and host risk factor parameters, input the virulence index, the environmental factor, the transmission coefficient and the host susceptibility risk index into the transmission risk evaluation model, and output a risk value by the transmission risk evaluation model.
10. The system for quantitative assessment of the risk of pathogenic bacteria dissemination according to claim 9, characterized in that, The early warning decision module is further included and is configured to divide risk levels according to the risk value and output early warning information and prevention and control suggestions.
Citation Information
Patent Citations
Method and system for assessing risk of bacterial propagation in chicken house
CN120068732A
Biohazard big data analysis and monitoring early warning system
CN120452553A