Pathogen gene sequence detection and analysis method based on nanopore sequencing
By combining nanopore sequencing technology and graph diffusion modeling with Bayesian causal models, the problem of on-site deployment of pathogen detection in customs quarantine has been solved, enabling rapid, full-spectrum monitoring and risk assessment, and generating structured reports.
Patent Information
- Application Number
- CN202511366941.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-24
AI Technical Summary
Existing technologies have long detection cycles and are difficult to deploy on-site in customs quarantine, port control and entry pathogen monitoring. They only support limited target identification, lack transmission map modeling mechanisms, and have simple risk identification methods that cannot assess transmission potential.
Nanopore sequencing technology is used for plug-and-play sequencing, sample classification and mutation monitoring, propagation maps are constructed and graph diffusion modeling is performed, and risk is assessed by combining Bayesian causal models to generate risk level labels.
It enables the completion of pathogen whole-genome sequencing and analysis within hours, supports the identification of multiple species and subtypes, reconstructs pathogen transmission relationships, quantifies transmission potential, and generates interpretable risk reasoning paths and early warning information.
Smart Images

Figure CN120853677A_ABST
Abstract
Description
Technical Field
[0001] This specification pertains to the field of customs quarantine; more specifically, this application relates to a method for detecting and analyzing pathogen gene sequences based on nanopore sequencing. Background Technology
[0002] Currently, in scenarios such as customs quarantine, port control, and monitoring of imported pathogens, traditional pathogen detection and analysis methods mainly rely on fluorescent PCR, specific primer amplification, enzyme-linked immunosorbent assay (ELISA), or high-throughput sequencing technology. These methods generally have the following limitations: 1. Long testing cycle and difficulty in on-site deployment: High-throughput sequencing relies on laboratory environment and centralized sequencing platform, which cannot meet the requirements of real-time testing and rapid response in mobile scenarios such as ports and high-speed channels.
[0003] 2. Only supports limited target identification and cannot monitor the entire spectrum: Traditional PCR and antigen detection require pre-setting detection targets, which cannot realize the discovery of unknown pathogens and the parallel monitoring of multiple types of mutations, and the ability to identify variant strains is limited.
[0004] 3. Lack of transmission map modeling mechanism: Most existing solutions focus on pathogen detection or genome annotation, lacking structured modeling of potential transmission relationships between samples, making it difficult to reconstruct their transmission paths or correlations.
[0005] 4. Risk identification methods are relatively simple and cannot assess transmission potential: Most current assessment methods assess risk based only on pathogen type or specific mutation markers, and fail to systematically model transmission mechanisms, transmission locations and the synergistic effects of mutations.
[0006] Therefore, there is an urgent need for a pathogen detection and analysis solution that supports rapid detection, structural mapping, intelligent modeling, and risk identification. Summary of the Invention
[0007] The summary section introduces a series of simplified concepts, which will be further explained in detail in the detailed description section. This summary section is not intended to limit the key and essential technical features of the claimed technical solution, nor is it intended to determine the scope of protection of the claimed technical solution.
[0008] Firstly, this application proposes a method for pathogen gene sequence detection and analysis based on nanopore sequencing, including: Nanopore sequencing was performed on the collected samples to obtain raw sequencing data; The raw sequencing data was subjected to quality screening to obtain the screened sequencing data; The above-screened sequencing data were subjected to species classification and mutation site monitoring operations to obtain pathogen classification and mutation information. Based on the above pathogen classification and mutation information, a transmission map is constructed to obtain a dynamic transmission structure diagram of the pathogen; Graph diffusion modeling is performed on the above dynamic propagation structure diagram to obtain potential transmission path scoring information of the pathogen; Based on the above mutation information and the above potential transmission path scoring information, a Bayesian causal model is constructed to obtain the posterior probability of the pathogen causing a high-risk event. Based on the aforementioned posterior probabilities, risk level labels are generated, and pathogen analysis reports and early warning information are output.
[0009] In one feasible implementation, the above-mentioned construction of a transmission map based on the pathogen classification information and mutation information to obtain a dynamic transmission structure map of the pathogen includes: Extract the metadata information for each sample, which includes the pathogen classification information, mutation information, sample collection time, sample collection location, and sequencing timestamp. Based on the above pathogen classification information, samples with the same or similar species are grouped into the same transmission map construction unit; Based on the above mutation information, a mutation feature vector is generated for each sample, wherein the mutation feature vector is encoded by mutation sites throughout the genome; Each sample is represented as a graph node, wherein the graph node includes the pathogen classification information, the mutation information, the sample collection time, the sample collection location, and the sequencing timestamp. The genomic similarity score between sample nodes is calculated based on the distance metric between the mutation feature vectors mentioned above; Based on the similarity score and sampling time and geographical location information, a propagation edge connection is established between the graph nodes if at least one preset condition is met. The preset conditions include a similarity score higher than a set threshold, a sample collection time interval less than a preset window, and the sample collection location being in a potential propagation path. Edge weights are set for the above-mentioned propagation edge connections, wherein the edge weights are calculated from similarity scores, time differences, and geographical distances; Using a time-sliding window mechanism, the dynamic propagation structure graph described above is established based on the graph nodes, propagation edge connections, and edge weights.
[0010] In one feasible implementation, the above-mentioned graph diffusion modeling operation on the dynamic propagation structure diagram to obtain potential transmission path scoring information of the pathogen includes: The graph nodes in the above dynamic propagation structure graph are encoded to construct a graph node feature matrix; A time series modeling mechanism is introduced into the above dynamic propagation structure graph. The graph structure is sorted by time based on the sample collection time to form a time-aware dynamic propagation graph. Based on the aforementioned time-aware dynamic propagation graph and the aforementioned graph node feature matrix, a graph diffusion model is used to perform graph diffusion modeling to obtain the graph node propagation embedding vector. Based on the above graph node propagation embedding vectors and the above time-aware dynamic propagation graph, the propagation potential is extracted to obtain the node potential propagation vectors; Calculate the similarity information between the potential propagation vectors of every two nodes in the above propagation path to obtain the potential propagation path score information.
[0011] In one feasible implementation, a Bayesian causal model is constructed based on the aforementioned mutation information and the aforementioned potential transmission path scoring information to obtain the posterior probability of a pathogen triggering a high-risk event, including: High-risk mutation feature identification is performed on the above mutation feature vectors to construct the first causal variable of the above Bayesian causal model; The propagation potential identification operation is performed on the above potential propagation path scoring information to construct the second causal variable of the above Bayesian causal model; Based on the first causal variable and the second causal variable, the Bayesian causal graph model is constructed. The Bayesian causal model includes posterior variable nodes representing high-risk events and directed edges from the causal variable nodes to the posterior variable nodes. Based on the prior conditional probability and the joint distribution inference rule, the conditional probability of the posterior variable node in the above Bayesian causal graph model is calculated to obtain the posterior probability of the above pathogen causing high-risk events.
[0012] In one feasible implementation, the above-mentioned high-risk mutation feature identification operation on the mutation feature vector to construct the first causal variable of the Bayesian causal model includes: Obtain a reference library of high-risk mutation sites; The above mutation feature vectors are compared with the above high-risk mutation site reference library to identify high-risk mutation-related sites and obtain the comparison results. Based on the above comparison results and the preset mutation risk determination rules, obtain statistical information on the number of mutation feature matches; Based on the above statistical information, the first causal variable of the Bayesian causal model is determined.
[0013] In one feasible implementation, the determination steps of the aforementioned preset mutation risk assessment rule include: A sliding time window sample set is constructed based on the mutation risk score, wherein the sliding time window sample set is all pathogen samples received within the preset time window, and the mutation risk score is calculated from the mutation feature vector. The mutation risk scores in the above sliding time window sample set are statistically analyzed, and the high quantile value of the distribution is calculated as the dynamic threshold. The mutation risk score of the current sample is compared with the above dynamic threshold to determine whether to update the above preset mutation risk judgment rule.
[0014] In one feasible implementation, the above-mentioned operation of identifying the propagation potential of the potential propagation path scoring information to construct the second causal variable of the Bayesian causal model includes: Determine the threshold information for determining the propagation potential; Based on the aforementioned threshold information for determining propagation potential and the aforementioned scoring information for potential propagation paths, information on the number of propagation paths, information on the maximum propagation probability, and an index of structural centrality are obtained. Based on the aforementioned information on the number of propagation paths, the aforementioned maximum propagation probability, the aforementioned structural centrality index, and the preset propagation potential determination rule, the second causal variable of the aforementioned Bayesian causal model is determined.
[0015] In one feasible implementation, the determination of the second causal variable in the Bayesian causal model based on the aforementioned propagation path quantity information, the aforementioned maximum propagation probability information, the aforementioned structure centrality index, and the preset propagation potential determination rule includes: Based on the above information on the number of propagation paths, the above information on the maximum propagation probability, and the above structural centrality index, we determine the propagation probability driver, the propagation depth adjustment term, the target node importance weighting term, and the spatial reachability adjustment term. The propagation potential scoring function is determined based on the above-mentioned propagation probability driving term, the above-mentioned propagation depth adjustment term, the above-mentioned target node importance weighting term, and the above-mentioned spatial accessibility adjustment term. The second causal variable of the Bayesian causal model is determined based on the aforementioned propagation potential scoring function and the aforementioned preset propagation potential determination rule.
[0016] In one feasible implementation, the risk level label is generated based on the aforementioned posterior probability, and a pathogen analysis report and early warning information are output, including: Obtain dynamic level threshold information; The risk level label is determined based on the above posterior probability and the above dynamic level threshold information. Based on the aforementioned risk level labels and sample analysis results, the above pathogen analysis report and early warning information will be output.
[0017] In one feasible implementation, the specific steps for determining the aforementioned dynamic level threshold information include: Construct a historical pathogen sample set containing the above posterior probability values and corresponding true risk labels; For the aforementioned historical sample set, construct receiver operating characteristic curves with the aforementioned posterior probabilities as predictors and the aforementioned risk labels as true classification criteria; Calculate the sensitivity and specificity indices corresponding to different posterior probability thresholds, and form the Youden index based on the sum of sensitivity and specificity minus one. The maximum posterior probability value of the Youden index is used as the dynamic judgment threshold for classifying the first-level risk level. By combining the macro-average AUC index and the classification accuracy evaluation index, the above posterior probability values are optimized with multiple thresholds to form probability intervals corresponding to high, medium, high and low risk levels.
[0018] In summary, the method proposed in this application employs nanopore sequencing technology, featuring plug-and-play sequencing, read length support, and on-the-fly analysis. It supports deployment at border crossings and in the field, enabling the completion of pathogen whole-genome sequencing and analysis within hours. Through species classification and mutation monitoring algorithms based on reference libraries or deep models, multiple species, subtypes, and their mutation combinations can be identified simultaneously, facilitating rapid tracing of imported pathogens. By modeling sequencing samples as graph nodes and constructing propagation maps based on mutation similarity, sampling time and space, and propagation reachability, the potential propagation relationships of pathogens across different individuals / regions can be reconstructed, supporting propagation path-level modeling and visualization. A graph diffusion algorithm is used to model the propagation map and calculate the propagation path scoring matrix between sample nodes, achieving a joint assessment of propagation potential from both structural and behavioral perspectives, addressing the weakness of traditional scoring models in structural expression. Pathogen mutation characteristics and propagation potential are introduced as causal variables. A Bayesian network is used to calculate the posterior probability of triggering high-risk events, constructing interpretable risk reasoning paths and supporting quantitative modeling and inferential identification of potential risk samples. By combining posterior probability with dynamic judgment thresholds, risk level labels are generated and structured reports are output for border inspection, epidemiological investigation and other systems to access, forming an integrated early warning mechanism from detection to disposal.
[0019] Other advantages, objectives and features of this application will be partly apparent from the description below, and partly understood by those skilled in the art through study and practice of this application. Attached Figure Description
[0020] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit this specification. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a flowchart illustrating a pathogen gene sequence detection and analysis method based on nanopore sequencing, provided in an embodiment of this application. Detailed Implementation
[0021] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus. The technical solutions of the embodiments of this application will now be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.
[0022] Please see Figure 1 This is a schematic flowchart of a pathogen gene sequence detection and analysis method based on nanopore sequencing, provided in an embodiment of this application. Specifically, it may include: S110. Perform nanopore sequencing on the collected samples to obtain raw sequencing data; For example, suspected samples can be collected from inbound personnel, cargo surfaces, transport vehicles, and logistics or cold chain transportation links, and nanopore sequencing technology can be used for real-time sequencing to obtain raw signal data containing complete pathogen genome sequences. Nanopore sequencing has advantages such as long read lengths, online analysis capability, and suitability for rapid on-site detection, making it particularly suitable for deployment in scenarios such as ports, highways, and airports.
[0023] S120. Perform quality screening on the above raw sequencing data to obtain the screened sequencing data; For example, to improve the accuracy of downstream analysis, the raw sequencing data needs to be processed. Quality screening may include, but is not limited to, removing low-quality reads, filtering out excessively short sequences, correcting sequencing errors (such as base call errors), and selecting high-quality fragments (usually evaluated by Q-value or Phred value). The sequencing data obtained after screening should have high confidence and coverage, and be suitable for subsequent alignment, variant identification, and other operations.
[0024] S130. Perform species classification and mutation site monitoring operations on the above-screened sequencing data to obtain pathogen classification and mutation information. For example, alignment-based classification algorithms (such as Minimap2+Kraken2) or deep learning models are used to identify species in sequencing fragments and extract corresponding pathogen classification information (genus, species, subtype, etc.). Simultaneously, base differences between the sequence and the reference genome are detected, mutation sites are recorded, including SNPs, InDels, and variant regions, and mutation feature vectors are constructed to provide a foundation for subsequent graph modeling and risk analysis.
[0025] S140. Construct a transmission map based on the above pathogen classification and mutation information to obtain a dynamic transmission structure map of the pathogen; For example, firstly, metadata information (species category, mutation information, sampling time, sampling location, etc.) is extracted for each sample, and each sample is represented as a graph node. Gene similarity between samples is calculated, and temporal / spatial reachability conditions are introduced to establish propagation edge connections. A sliding time window mechanism is used to update the graph structure time-by-time, constructing a dynamic propagation graph G containing node features, edge weights, and time series. t To depict the transmission network of pathogens.
[0026] S150. Perform graph diffusion modeling on the above dynamic transmission structure diagram to obtain potential transmission path scoring information of pathogens. For example, the above propagation graph G t With node feature matrix X t The input graph diffusion modeling algorithm performs propagation potential learning through a graph neural network structure, generating a propagation embedding vector for each sample node. Furthermore, by concatenating a scoring function or a similarity function, the potential propagation path score between any two nodes is calculated, forming a propagation path score matrix for subsequent causal analysis.
[0027] S160. Based on the above mutation information and the above potential transmission path scoring information, construct a Bayesian causal model to obtain the posterior probability of the pathogen causing a high-risk event. For example, mutation risk characteristics and transmission potential scores are used together as input variables for a causal graph. Conditional probability inference is performed by constructing a well-structured Bayesian network. The model output is the posterior probability that the sample triggers a high-risk event (such as large-scale transmission or a highly pathogenic outbreak), which reflects the potential threat to public safety posed by the pathogen represented by the sample.
[0028] S170. Based on the above posterior probabilities, generate risk level labels and output pathogen analysis reports and early warning information.
[0029] For example, by combining dynamic risk level determination thresholds, posterior probabilities are mapped to risk level labels (such as high, medium, and low). Subsequently, sample identification, mutation information, transmission scores, and risk inference results are integrated to output a structured pathogen analysis report for automatic collection and early warning by border inspection, epidemiological investigation, and epidemic monitoring departments.
[0030] In summary, the method proposed in this application employs nanopore sequencing technology, featuring plug-and-play sequencing, read length support, and on-the-fly analysis. It supports deployment at border crossings and in the field, enabling the completion of pathogen whole-genome sequencing and analysis within hours. Through species classification and mutation monitoring algorithms based on reference libraries or deep models, multiple species, subtypes, and their mutation combinations can be identified simultaneously, facilitating rapid tracing of imported pathogens. By modeling sequencing samples as graph nodes and constructing propagation maps based on mutation similarity, sampling time and space, and propagation reachability, the potential propagation relationships of pathogens across different individuals / regions can be reconstructed, supporting propagation path-level modeling and visualization. A graph diffusion algorithm is used to model the propagation map and calculate the propagation path scoring matrix between sample nodes, achieving a joint assessment of propagation potential from both structural and behavioral perspectives, addressing the weakness of traditional scoring models in structural expression. Pathogen mutation characteristics and propagation potential are introduced as causal variables. A Bayesian network is used to calculate the posterior probability of triggering high-risk events, constructing interpretable risk reasoning paths and supporting quantitative modeling and inferential identification of potential risk samples. By combining posterior probability with dynamic judgment thresholds, risk level labels are generated and structured reports are output for border inspection, epidemiological investigation and other systems to access, forming an integrated early warning mechanism from detection to disposal.
[0031] In one feasible implementation, the above-mentioned construction of a transmission map based on the pathogen classification information and mutation information to obtain a dynamic transmission structure map of the pathogen includes: Extract the metadata information for each sample, which includes the pathogen classification information, mutation information, sample collection time, sample collection location, and sequencing timestamp. Based on the above pathogen classification information, samples with the same or similar species are grouped into the same transmission map construction unit; Based on the above mutation information, a mutation feature vector is generated for each sample, wherein the mutation feature vector is encoded by mutation sites throughout the genome; Each sample is represented as a graph node, wherein the graph node includes the pathogen classification information, the mutation information, the sample collection time, the sample collection location, and the sequencing timestamp. The genomic similarity score between sample nodes is calculated based on the distance metric between the mutation feature vectors mentioned above; Based on the similarity score and sampling time and geographical location information, a propagation edge connection is established between the graph nodes if at least one preset condition is met. The preset conditions include a similarity score higher than a set threshold, a sample collection time interval less than a preset window, and the sample collection location being in a potential propagation path. Edge weights are set for the above-mentioned propagation edge connections, wherein the edge weights are calculated from similarity scores, time differences, and geographical distances; Using a time-sliding window mechanism, the dynamic propagation structure graph described above is established based on the graph nodes, propagation edge connections, and edge weights.
[0032] For example, embodiments of this application construct a dynamic propagation map to characterize the transmission relationships of pathogen samples by combining their classification and mutation information. This propagation map reflects the potential transmission links between different pathogen samples at the temporal, spatial, and genetic levels, providing a structural basis for subsequent propagation path modeling and risk inference.
[0033] First, metadata information for each sample is extracted from the completed nanopore sequencing data. This metadata information includes basic attributes such as pathogen species classification results, whole-genome mutation information, sample collection time and location, and sequencing completion timestamp.
[0034] Based on the classification information of the samples, samples belonging to the same species or variant subtype are grouped into the same propagation map construction unit to ensure the rationality and comparability of propagation relationships. Next, based on mutation information, a mutation feature vector is constructed for each sample. This feature vector can be generated by encoding mutation sites across the entire genome and is usually expressed in the form of Boolean encoding, site indexing, or sparse vectors to facilitate subsequent similarity analysis.
[0035] Subsequently, each sample is abstracted as a node in the graph. Each node contains attributes such as its classification label, mutation feature vector, collection time, and location. Then, based on the distance metric of mutation feature vectors between nodes, a genomic similarity score is calculated between samples. Similarity metrics can include Jaccard distance, Mash distance, or SNP site overlap rate.
[0036] Based on this, the system determines whether to establish a propagation edge between two sample nodes according to the similarity score, the time difference between collection, and the reachability of the geographical location. Specifically, if the similarity score between samples is higher than a set threshold, or the collection time interval is within a preset time window, or the sample collection location is in a propagation path (such as being located at the same port or adjacent areas), then a propagation edge is established between the two nodes.
[0037] Furthermore, weights are assigned to the propagation edges to quantify the reliability of the propagation path. These edge weights are generated using a weighted calculation method, taking into account factors such as sample similarity, time intervals, and geographical distance. Finally, a sliding time window mechanism is employed to update and reconstruct the propagation map structure at different time intervals, forming a time-aware dynamic propagation map sequence to reflect the propagation and evolution trend of the pathogen in a real-world scenario.
[0038] Through the above steps, the constructed dynamic propagation map not only retains multidimensional information about the pathogen's genetic variation, temporal distribution, and geographical distribution, but also provides structured data support for subsequent graph diffusion modeling and causal inference.
[0039] In one feasible implementation, the above-mentioned graph diffusion modeling operation on the dynamic propagation structure diagram to obtain potential transmission path scoring information of the pathogen includes: The graph nodes in the above dynamic propagation structure graph are encoded to construct a graph node feature matrix; A time series modeling mechanism is introduced into the above dynamic propagation structure graph. The graph structure is sorted by time based on the sample collection time to form a time-aware dynamic propagation graph. Based on the aforementioned time-aware dynamic propagation graph and the aforementioned graph node feature matrix, a graph diffusion model is used to perform graph diffusion modeling to obtain the graph node propagation embedding vector. Based on the above graph node propagation embedding vectors and the above time-aware dynamic propagation graph, the propagation potential is extracted to obtain the node potential propagation vectors; Calculate the similarity information between the potential propagation vectors of every two nodes in the above propagation path to obtain the potential propagation path score information.
[0040] For example, based on the constructed dynamic propagation structure graph, a graph diffusion modeling operation is performed to obtain potential transmission path scoring information for the pathogen.
[0041] For each node in the propagation graph Extracting three types of information to form an initial feature vector : ; in, Represents the mutation feature vector of a pathogen; Indicates the sampling timestamp; This indicates the sampling location code (such as area code, latitude and longitude vector).
[0042] Combine the feature vectors of all nodes into a graph node feature matrix: ; Based on the sampling time order, the propagation graph Introducing a time sliding window Construct a time-segmented plot sequence: ; in, To represent a time window A subset of nodes within; This represents the set of edges that satisfy the propagation condition within the window.
[0043] Using graph diffusion models (such as GAT, TGN, DCRNN), time-aware graph structures With node feature matrix Input the model to obtain the propagation representation vector of each node. : ; The graph diffusion modeling function represents the embedding features of each sample node in the propagation structure.
[0044] For any two nodes in the graph, a propagation path score is calculated based on their propagation embedding vectors. The scoring method employs a concatenated neural network scoring model. ; in: Denotes the L2 norm; This indicates vector concatenation; It is the sigmoid activation function; These are the trainable parameters of the model. Represents a node The graph diffusion embedding vector, For nodes The graph diffusion embedding vector.
[0045] The score values between all node pairs are combined into a propagation score matrix. : ; This matrix represents the potential propagation probability between any two samples in the graph, and is used for subsequent risk reasoning and source tracing.
[0046] Node feature encoding takes into account gene mutation, sampling time, and geographical location simultaneously, making propagation modeling more comprehensive and enhancing the model's ability to characterize complex propagation paths.
[0047] This embodiment constructs a dynamic graph by introducing a time-sliding window mechanism and utilizes a time-aware graph diffusion model to effectively reflect the temporal evolution characteristics of pathogen transmission. By representing transmission potential with graph embedding vectors and calculating the transmission score matrix between node pairs, the transmission probability between any samples can be quantified, improving the accuracy of transmission chain mining and source tracing analysis.
[0048] In one feasible implementation, a Bayesian causal model is constructed based on the aforementioned mutation information and the aforementioned potential transmission path scoring information to obtain the posterior probability of a pathogen triggering a high-risk event, including: High-risk mutation feature identification is performed on the above mutation feature vectors to construct the first causal variable of the above Bayesian causal model; The propagation potential identification operation is performed on the above potential propagation path scoring information to construct the second causal variable of the above Bayesian causal model; Based on the first causal variable and the second causal variable, the Bayesian causal graph model is constructed. The Bayesian causal model includes posterior variable nodes representing high-risk events and directed edges from the causal variable nodes to the posterior variable nodes. Based on the prior conditional probability and the joint distribution inference rule, the conditional probability of the posterior variable node in the above Bayesian causal graph model is calculated to obtain the posterior probability of the above pathogen causing high-risk events.
[0049] For example, firstly, a high-risk mutation identification operation is performed on the mutation feature vector of each sample. This operation includes the following sub-steps: obtaining a high-risk mutation reference library containing mutation sites known to have enhanced propagation or immune evasion capabilities; comparing the mutation feature vector of the sample with the reference library to identify whether it contains high-risk mutations; based on the comparison results, determining whether the sample hits a key risk mutation, and outputting the first causal variable. If the number of high-risk mutation sites hit exceeds a set threshold, or if a specific combination of mutations is hit, then it is considered to possess high-risk mutation characteristics. Otherwise, set .
[0050] Subsequently, based on the potential propagation ability of the samples in the propagation scoring map, a second causal variable was constructed. The specific operation is as follows: Extract the following indicators from the propagation path scoring matrix for the sample, such as: the number of outward propagation paths, the maximum propagation score, and structural centrality indicators in the propagation graph (e.g., out-degree, betweenness centrality). Based on the above indicators and preset rules, determine the propagation potential. If the sample's propagation ability exceeds a set threshold, then... Otherwise set .
[0051] Based on the two causal variables mentioned above, a simplified Bayesian causal graph structure is constructed: ; in: High-risk mutation characteristics (Boolean variable); For propagation potential characteristics (Boolean variable); This represents whether a high-risk event is triggered (Boolean posterior variable). The graph represents high-risk events. The occurrence of this disease is influenced by a combination of high-risk mutations and transmission potential.
[0052] Calculate the posterior variables according to Bayesian inference rules. The conditional probability, i.e.: ; This can be achieved through the following steps: Construct a joint probability table or a CPT (Conditional Probability Table) based on training samples or historical data. Obtain the prior probabilities using maximum likelihood estimation or Bayesian estimation. For the current sample, substitute it into... Calculate the corresponding value. ; Output the posterior probability of this sample triggering a high-risk event, as the basis for risk assessment.
[0053] This embodiment enables a dual causal explanation of pathogen samples from both the mutation and transmission levels, and utilizes Bayesian graph modeling to perform quantitative inferences of high-risk events, providing a basis for subsequent early warning classification and key monitoring.
[0054] In one feasible implementation, the above-mentioned high-risk mutation feature identification operation on the mutation feature vector to construct the first causal variable of the Bayesian causal model includes: Obtain a reference library of high-risk mutation sites; The above mutation feature vectors are compared with the above high-risk mutation site reference library to identify high-risk mutation-related sites and obtain the comparison results. Based on the above comparison results and the preset mutation risk determination rules, obtain statistical information on the number of mutation feature matches; Based on the above statistical information, the first causal variable of the Bayesian causal model is determined.
[0055] For example, a reference library of high-risk mutation sites can be constructed or loaded. This reference library can be derived from highly pathogenic and highly transmissible mutation sites reported in publicly available databases at home and abroad, as well as from epidemiological statistical analysis results or laboratory functional validation data.
[0056] Reference Library The form can be: ; Each of them It represents a key mutation site, including meta-information such as site number, gene name, and risk level.
[0057] The mutation feature vector of the target sample Perform a positional comparison with the reference library. If the sample contains a mutation at a given site, it is considered a high-risk mutation match. Output the matching results between the sample and the reference library, such as the number of matches. It can also record the hit combinations.
[0058] Based on the comparison results, the following information is compiled: Total number of mutations detected. Risk level weighted sum of mutations Whether a specific mutation combination is hit (such as the "E484K+N501Y" combined mutation) or a known domain sensitive site is hit (such as the RBD region).
[0059] The risk score can be expressed as: ; in, It is a mutation site Risk weights in the reference library.
[0060] Based on the pre-defined mutation risk assessment rules, the sample is determined as "whether it has high-risk mutation characteristics" as the first causal variable. .
[0061] For example, the decision rules may include: 1. If Then let .
[0062] 2. If Then let Otherwise, let .in, The minimum number of hits threshold (e.g., 3). This is a risk scoring threshold (such as the 90th percentile or an experience threshold).
[0063] In one feasible implementation, the determination steps of the aforementioned preset mutation risk assessment rule include: A sliding time window sample set is constructed based on the mutation risk score, wherein the sliding time window sample set is all pathogen samples received within the preset time window, and the mutation risk score is calculated from the mutation feature vector. The mutation risk scores in the above sliding time window sample set are statistically analyzed, and the high quantile value of the distribution is calculated as the dynamic threshold. The mutation risk score of the current sample is compared with the above dynamic threshold to determine whether to update the above preset mutation risk judgment rule.
[0064] For example, in one feasible implementation, in order to improve the flexibility and timeliness of high-risk mutation identification, a sliding time window mechanism can be introduced to achieve dynamic updating and adaptive optimization of mutation risk determination rules, specifically including the following steps: Set a sliding time window (e.g., 7 days, 14 days, 1 month), denoted as This window contains all pathogen sample data received by the system during that time period.
[0065] For each sample within the window Through its mutation feature vector Calculate its corresponding mutation risk score ,For example: ; in: For the sample The number of high-risk mutation sites hit; Risk weight for each site in the reference library; It is the total mutation risk score of the current sample.
[0066] Constructing a window of sample ratings : ; For the sliding time window sample set All risk scores are statistically analyzed for distribution, and their high quantiles (such as the 90th quantile, 95th quantile, etc.) are calculated as dynamic judgment thresholds. ; in: Set parameters for quantiles; This is the dynamic risk scoring threshold for the current time period.
[0067] This dynamic threshold reflects the high-ranking abnormality of the current mutation risk score in the full sample distribution, which helps to adaptively respond to the rising trend of new variant strains.
[0068] For the newly arrived samples Calculate its mutation risk score and with the current dynamic threshold Compare. If If the sample is considered to have high-risk mutation characteristics, then a first causal variable can be set. If it is below the threshold, then set This information can also be used to update the preset judgment rules for the next round of identification.
[0069] The mechanism proposed in this embodiment can continuously adjust the model's definition boundary for "high-risk mutations" to adapt to changes in mutation spectra at different stages and in different regions. The model no longer relies on fixed empirical thresholds and can automatically adjust its judgment rules according to the virus's evolutionary trend. The sliding time window design allows the model to respond in real-time to changes in sample risk. Judgment criteria are established through statistical distribution, reducing the impact of extreme individual values. Frequent manual intervention is unnecessary, which is beneficial for the long-term operation of the system.
[0070] In one feasible implementation, the above-mentioned operation of identifying the propagation potential of the potential propagation path scoring information to construct the second causal variable of the Bayesian causal model includes: Determine the threshold information for determining the propagation potential; Based on the aforementioned threshold information for determining propagation potential and the aforementioned scoring information for potential propagation paths, information on the number of propagation paths, information on the maximum propagation probability, and an index of structural centrality are obtained. Based on the aforementioned information on the number of propagation paths, the aforementioned maximum propagation probability, the aforementioned structural centrality index, and the preset propagation potential determination rule, the second causal variable of the aforementioned Bayesian causal model is determined.
[0071] For example, a set of threshold indicators for determining "high propagation potential" can be set or dynamically generated. One such threshold is the number of propagation paths. This represents the minimum number of effective paths for the sample to diffuse outwards. Maximum propagation probability threshold. This is the lower limit of the maximum single-path score corresponding to the sample. Structure centrality threshold. This serves as the lower bound for determining graph structure centrality (such as out-degree, betweenness centrality, and PageRank). The threshold can be a static empirical value or generated using a dynamic sliding window mechanism (e.g., adaptive adjustment using the 90th percentile).
[0072] For the target sample From the propagation path scoring matrix Extract the following indicators: Number of transmission paths : ; in, Determine the threshold for path scoring. This is an indicator function.
[0073] Maximum Propagation Probability : ; Structural centrality Optional metrics include out-degree centrality, PageRank value, K-shell value, etc., which reflect the "strategic position" of the node in the propagation graph.
[0074] Based on the extracted indicators and a set of preset or dynamically generated thresholds The following strategy is used to determine this: like: ; This sample is then considered to have high transmission potential, and a second causal variable is defined. ; Otherwise, set The aforementioned causal variables are used to model the risk diffusion capacity of pathogen samples at the transmission pathway level.
[0075] This embodiment allows for the systematic analysis of the diffusion intensity and centrality of samples within a propagation spectrum. The propagation scoring results are quantified into logical variables. This facilitates integration into Bayesian causal graph modeling. It improves the efficiency of identifying "potential super-spreading nodes" or "key network propagators," providing a model foundation for accurate source tracing and key early warning.
[0076] In one feasible implementation, the determination of the second causal variable in the Bayesian causal model based on the aforementioned propagation path quantity information, the aforementioned maximum propagation probability information, the aforementioned structure centrality index, and the preset propagation potential determination rule includes: Based on the above information on the number of propagation paths, the above information on the maximum propagation probability, and the above structural centrality index, we determine the propagation probability driver, the propagation depth adjustment term, the target node importance weighting term, and the spatial reachability adjustment term. Based on the aforementioned propagation probability driving term, the aforementioned propagation depth adjustment term, and the aforementioned target node, the second causal variable of the aforementioned Bayesian causal model is determined according to the aforementioned propagation potential scoring function and the aforementioned preset propagation potential judgment rule.
[0077] For example, in one feasible implementation, to achieve a quantitative assessment of the potential spread capacity of pathogen samples in a transmission network, this invention constructs a transmission potential scoring function based on information on the number of transmission paths, maximum transmission probability information, and structural centrality indicators. Combined with preset transmission potential determination rules, a second causal variable in the Bayesian causal model is determined. Specifically, this includes the following steps: (1) Constructing sub-indicators of transmissibility: For each pathogen sample node The following three types of core indicators were extracted from the transmission map and scoring matrix: 1. Number of transmission paths : Represents a node The total number of transmission paths sent (i.e.) The number of edges exceeding a set threshold); 2. Maximum propagation probability : Represents a node Send the edge with the highest score in the path; 3. Structural centrality index :node Centrality measures in a propagation graph, such as PageRank or betweenness centrality.
[0078] In addition, spatial relevance information can be extended to include "spatial accessibility adjustment items", such as weighted scores based on the number of areas covered by the propagation path or geographical distance.
[0079] (2) Constructing a propagation potential scoring function: Based on the extracted multiple indicators, a fusion-based propagation potential scoring function is constructed to comprehensively evaluate the propagation capability of nodes. This scoring function is expressed as follows: ; in: : Represents the Sigmoid activation function, ensuring the final score falls within . interval; : These are the weight parameters for the five features; To adjust the exponential coefficient for nonlinear enhancement effect; In the above scoring function This is a propagation depth adjustment term, which measures the number of paths a node can reach in the propagation graph, reflecting the breadth or coverage of the propagation. The more paths, the wider the propagation's influence, and the higher the score.
[0080] The propagation probability driving term represents the probability value of the highest-scoring path among the outward propagation paths of this node, reflecting its potential as the "strongest propagation path". The higher the score, the stronger the potential for efficient propagation at a single point.
[0081] Weighting terms for the importance of the target node, where This indicates the structural centrality of a node, such as PageRank or betweenness centrality. This value describes the importance of a node's position in the graph structure; higher centrality indicates that the node is in a critical position and its transmission is more strategic.
[0082] The term representing the intersection of breadth and strength indicates that the node possesses both a sufficient number of paths (breadth) and strong high-scoring paths (strength). The product of these two factors enhances its ability to spread widely but strike precisely in the transmission network.
[0083] The intersection of structure and strength indicates that a node not only has a strong path but also occupies an important position in the propagation graph, further reinforcing the "super propagation point" characteristic. Nodes possessing both high strength and centrality will receive higher scores.
[0084] In one feasible implementation, the risk level label is generated based on the aforementioned posterior probability, and a pathogen analysis report and early warning information are output, including: Obtain dynamic level threshold information; The risk level label is determined based on the above posterior probability and the above dynamic level threshold information. Based on the aforementioned risk level labels and sample analysis results, the above pathogen analysis report and early warning information will be output.
[0085] For example, a criterion for risk level classification can be constructed based on the posterior probability values of multiple pathogen samples output by a Bayesian causal model, combined with known historical label information (such as cluster transmission events, case pathogenicity levels, etc.).
[0086] Specifically, the posterior probability value corresponding to the largest Youden index can be selected as the threshold for judging the first-level risk level by constructing the receiver operating characteristic curve (ROC). In scenarios where multi-level risk classification is required, the posterior probability intervals corresponding to the multi-level risk levels can be dynamically determined based on the macro average AUC index or distribution statistics (such as the 90th, 75th, and 50th percentiles).
[0087] The risk level ranges can be set as follows: high risk is defined as a posterior probability ≥ 0.85; medium-high risk is defined as a posterior probability < 0.85 (≤ 0.70); medium risk is defined as a posterior probability < 0.70 (≤ 0.50); and low risk is defined as a posterior probability < 0.50. The thresholds can be updated in real time based on the distribution of sample data within the sliding time window.
[0088] For the current sample to be evaluated, the posterior probability value output by the Bayesian causal model is matched with the currently determined dynamic level threshold to assign a corresponding risk level label. For example, if the posterior probability is 0.91, it is labeled "high risk"; if the posterior probability is 0.68, it is labeled "medium risk". This labeling information will serve as the basis for the classification of downstream pathogen management strategies.
[0089] Combining metadata information obtained during sample preprocessing (such as classification results, mutation characteristics, transmission score vectors, transmission potential sub-items, causal variable values, etc.), a structured output template is constructed to generate pathogen analysis reports and early warning summaries containing the following: sample number, collection time and sampling location; taxonomic species and their confidence levels; information on hit mutation sites and identification of risk mutations; obtained transmission potential scores and their sub-item scores; Bayesian posterior probabilities and their key causal variable values; automatically generated risk level labels; and early warning suggestions generated based on rules (such as review, isolation, reporting, early warning prompts, etc.).
[0090] The pathogen analysis report can be output as a PDF with images and text through a visualization component, or as a structured output such as JSON, XML, or tabular files, for use in border inspection monitoring systems, epidemiological investigation platforms, emergency response systems, etc.
[0091] In one feasible implementation, the specific steps for determining the aforementioned dynamic level threshold information include: Construct a historical pathogen sample set containing the above posterior probability values and corresponding true risk labels; For the aforementioned historical sample set, construct receiver operating characteristic curves with the aforementioned posterior probabilities as predictors and the aforementioned risk labels as true classification criteria; Calculate the sensitivity and specificity indices corresponding to different posterior probability thresholds, and form the Youden index based on the sum of sensitivity and specificity minus one. The maximum posterior probability value of the Youden index is used as the dynamic judgment threshold for classifying the first-level risk level. By combining the macro-average AUC index and the classification accuracy evaluation index, the above posterior probability values are optimized with multiple thresholds to form probability intervals corresponding to high, medium, high and low risk levels.
[0092] For example, firstly, a historical sample set is constructed based on pathogen sample data that has already been evaluated. Each sample in the sample set includes a posterior probability value calculated using a Bayesian causal model and a corresponding true risk label (e.g., whether it triggered a transmission event, whether it is highly pathogenic, etc.). This sample set is used to train and optimize the correspondence between the posterior probability and the risk level.
[0093] Using the posterior probability as the predictor variable and historical true risk labels as the classification standard, model performance metrics, including sensitivity and specificity, were calculated at different probability thresholds. Receiver operating characteristic (ROC) curves were plotted based on these results to comprehensively evaluate the model's classification ability at various probability thresholds.
[0094] For each candidate's posterior probability, calculate its corresponding Youden index, defined as: ; Sensitivity represents the model's ability to identify positive samples (such as high-risk pathogens), while Specificity represents the model's ability to identify negative samples (such as low-risk pathogens).
[0095] The posterior probability interval that maximizes the Youden index is chosen as the dynamic threshold for classifying the first-level wind and hazard levels. This value offers optimal balance and discriminative power under the current state of the sample distribution.
[0096] When further refining the risk levels to high, medium, high-medium, and low, multi-classification model evaluation metrics are employed, such as macro-AUC, weighted accuracy, and F1-score. By analyzing the distribution density of posterior probabilities and classification performance across different intervals, probability boundary points corresponding to multiple risk levels are determined. These interval divisions can be regenerated in each system update based on the latest samples within the sliding window, thus forming an adaptive risk level determination mechanism.
[0097] The dynamic classification mechanism proposed in this embodiment makes risk level classification more timely and adaptable to different scenarios. The introduction of ROC and Youden index judgment methods ensures the statistical optimality of the classification threshold. Through multi-level classification optimization, risk management strategies can be refined, providing more accurate decision support for subsequent early warning and response systems.
[0098] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for detecting and analyzing pathogen gene sequences based on nanopore sequencing, characterized in that, include: Nanopore sequencing was performed on the collected samples to obtain raw sequencing data; The raw sequencing data is subjected to quality screening to obtain screened sequencing data; The selected sequencing data are subjected to species classification and mutation site monitoring operations to obtain pathogen classification and mutation information. Based on the pathogen classification and mutation information, a propagation map is constructed to obtain a dynamic propagation structure diagram of the pathogen; Perform graph diffusion modeling on the dynamic propagation structure diagram to obtain potential transmission path scoring information of the pathogen; Based on the mutation information and the potential transmission path scoring information, a Bayesian causal model is constructed to obtain the posterior probability of the pathogen causing a high-risk event. Based on the posterior probability, risk level labels are generated, and pathogen analysis reports and early warning information are output.
2. The method for pathogen gene sequence detection and analysis based on nanopore sequencing according to claim 1, characterized in that, The step of constructing a transmission map based on the pathogen classification and mutation information to obtain a dynamic transmission structure map of the pathogen includes: Extract metadata information for each sample, including pathogen classification information, mutation information, sample collection time, sample collection location, and sequencing timestamp; Based on the pathogen classification information, samples with the same or similar species are grouped into the same transmission map construction unit; Based on the mutation information, a mutation feature vector is generated for each sample, wherein the mutation feature vector is encoded by mutation sites throughout the genome; Each sample is represented as a graph node, wherein the graph node includes the pathogen classification information, the mutation information, the sample collection time, the sample collection location, and the sequencing timestamp; The genomic similarity score between sample nodes is calculated based on the distance metric between the mutation feature vectors; Based on the similarity score and sampling time and geographical location information, a propagation edge connection is established between the graph nodes if at least one preset condition is met. The preset conditions include a similarity score higher than a set threshold, a sample collection time interval less than a preset window, and the sample collection location being in a potential propagation path. An edge weight is set in the propagation edge connection, wherein the edge weight is calculated from the similarity score, time difference and geographical distance; The dynamic propagation structure graph is established based on the graph nodes, the propagation edge connections, and the edge weights using a time sliding window mechanism.
3. The method for pathogen gene sequence detection and analysis based on nanopore sequencing according to claim 2, characterized in that, The step of performing graph diffusion modeling on the dynamic propagation structure diagram to obtain potential transmission path scoring information for the pathogen includes: The graph nodes in the dynamic propagation structure graph are encoded to construct a graph node feature matrix; A time series modeling mechanism is introduced into the dynamic propagation structure graph to sort the graph structure by time based on the sample collection time, thereby forming a time-aware dynamic propagation graph. Based on the time-aware dynamic propagation graph and the graph node feature matrix, graph diffusion modeling is performed to obtain graph node propagation embedding vectors; Based on the graph node propagation embedding vector and the time-aware dynamic propagation graph, the propagation potential is extracted to obtain the node potential propagation vector; Calculate the similarity information between the potential propagation vectors of every two propagation path nodes to obtain the potential propagation path score information.
4. The method for pathogen gene sequence detection and analysis based on nanopore sequencing according to claim 3, characterized in that, The step of constructing a Bayesian causal model based on the mutation information and the potential transmission path scoring information to obtain the posterior probability of the pathogen causing a high-risk event includes: The mutation feature vector is subjected to a high-risk mutation feature identification operation to construct the first causal variable of the Bayesian causal model; The potential propagation path scoring information is used to perform a propagation potential identification operation to construct the second causal variable of the Bayesian causal model; Based on the first causal variable and the second causal variable, the Bayesian causal graph model is constructed, wherein the Bayesian causal model includes posterior variable nodes representing high-risk events and directed edges from the causal variable nodes to the posterior variable nodes; Based on the prior conditional probability and the joint distribution inference rule, the conditional probability of the posterior variable node in the Bayesian causal graph model is calculated to obtain the posterior probability of the pathogen causing a high-risk event.
5. The method for pathogen gene sequence detection and analysis based on nanopore sequencing according to claim 4, characterized in that, The step of performing high-risk mutation feature identification on the mutation feature vector to construct the first causal variable of the Bayesian causal model includes: Obtain a reference library of high-risk mutation sites; The mutation feature vector is compared with the high-risk mutation site reference library to identify high-risk mutation-related sites and obtain the comparison results. Based on the comparison results and the preset mutation risk determination rules, obtain statistical information on the number of mutation feature matches; The first causal variable of the Bayesian causal model is determined based on the statistical information.
6. The method for pathogen gene sequence detection and analysis based on nanopore sequencing according to claim 5, characterized in that, The steps for determining the preset mutation risk assessment rule include: A sliding time window sample set is constructed based on the mutation risk score, wherein the sliding time window sample set is all pathogen samples received within a preset time window, and the mutation risk score is calculated from the mutation feature vector; The mutation risk scores in the sliding time window sample set are statistically analyzed, and the high quantile value of the distribution is calculated as a dynamic threshold. The mutation risk score of the current sample is compared with the dynamic threshold to determine whether to update the preset mutation risk determination rule.
7. The method for pathogen gene sequence detection and analysis based on nanopore sequencing according to claim 4, characterized in that, The step of identifying the propagation potential of the potential propagation path scoring information to construct the second causal variable of the Bayesian causal model includes: Determine the threshold information for determining the propagation potential; Based on the propagation potential determination threshold information and the potential propagation path scoring information, obtain the propagation path quantity information, maximum propagation probability information, and structure centrality index; The second causal variable of the Bayesian causal model is determined based on the information on the number of propagation paths, the information on the maximum propagation probability, the structural centrality index, and the preset propagation potential determination rule.
8. The method for pathogen gene sequence detection and analysis based on nanopore sequencing according to claim 7, characterized in that, The step of determining the second causal variable of the Bayesian causal model based on the propagation path quantity information, the maximum propagation probability information, the structure centrality index, and the preset propagation potential determination rule includes: Based on the propagation path quantity information, the maximum propagation probability information, and the structure centrality index, determine the propagation probability driving term, the propagation depth adjustment term, the target node importance weighting term, and the spatial reachability adjustment term; The propagation potential scoring function is determined based on the propagation probability driving term, the propagation depth adjustment term, the target node importance weighting term, and the spatial reachability adjustment term; The second causal variable of the Bayesian causal model is determined based on the propagation potential scoring function and the preset propagation potential determination rule.
9. The method for pathogen gene sequence detection and analysis based on nanopore sequencing according to claim 1, characterized in that, The process of generating risk level labels based on the posterior probability and outputting pathogen analysis reports and early warning information includes: Obtain dynamic level threshold information; The risk level label is determined based on the posterior probability and the dynamic level threshold information; The pathogen analysis report and early warning information are output based on the risk level label and sample analysis results.
10. The method for pathogen gene sequence detection and analysis based on nanopore sequencing according to claim 9, characterized in that, The specific steps for determining the dynamic level threshold information include: Construct a historical pathogen sample set containing the posterior probability and the corresponding true risk label; Construct receiver operating characteristic curves for the historical pathogen sample set, using the posterior probability as the predictor variable and the risk label as the true classification criterion; Calculate the sensitivity and specificity indices corresponding to different posterior probability thresholds, and form the Youden index based on the sum of sensitivity and specificity minus one. The maximum posterior probability value of the Youden index is used as the dynamic judgment threshold for classifying the first-level risk level. By combining the macro-average AUC index and the classification accuracy evaluation index, the posterior probability value is optimized with multiple thresholds to form probability intervals corresponding to high, medium, high-medium and low risk levels.
Citation Information
Patent Citations
Rapid PM2.5 bacterial community composition source analysis and risk assessment method
CN108841942A
Pathogenic microorganism detection system and method based on nanopore sequencing
CN112967753A
Variant gene pathogenicity evaluation method and device based on Bayesian algorithm
CN116665773A
Microorganism tracing method, multi-strain variation map construction method and equipment
CN119207577A
Infectious disease early warning and monitoring method and system based on big data and deep learning
CN120221125A