Pathogenic microorganism detection system based on nanopore sequencing and intelligent interpretation

By acquiring biomolecular electrical signals, using multi-level differential coding and multi-dimensional parameter mapping, and combining it with a hypergraph neural network, the problems of insufficient data analysis and processing capabilities and insufficient fusion of multi-source information in nanopore sequencing have been solved, enabling efficient and accurate detection of pathogenic microorganisms.

CN120895093APending Publication Date: 2025-11-04ANIMAL & PLANT & FOOD INSPECTION CENT OF TIANJIN ENTRY EXIT INSPECTION & QUARANTINE BUREAU
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510909463.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing technologies lack the ability to analyze and process nanopore sequencing data, making it difficult to accurately capture key information about pathogenic microorganisms. Furthermore, their ability to fuse multi-source information and perform intelligent analysis is limited, resulting in low detection accuracy and reliability.

Method used

The system employs a biomolecular electrical signal acquisition module, a pathogen feature sequence extraction module, a hypergraph topology adaptive construction module, a multidimensional heterogeneous pathogen parameter fusion analysis module, and an optimized hypergraph neural network inference module. Combined with multilayer perceptron and hyperedge convolution operations, it achieves high-sensitivity signal acquisition, multi-level differential coding, multidimensional parameter mapping, and pathogen category prediction.

Benefits of technology

It significantly improves the accuracy and efficiency of pathogen detection, generating precise detection reports that include pathogen type, subtype, and risk level, meeting the needs for rapid and accurate detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895093A_ABST
    Figure CN120895093A_ABST
Patent Text Reader

Abstract

The invention discloses a pathogenic microorganism detection system based on nanopore sequencing and intelligent interpretation. The system comprises six modules including a biomolecule electric signal acquisition module and a pathogen characteristic sequence extraction module. Ion current change signals are collected through a nanopore chip array, a pathogen characteristic sequence is extracted through differential coding, a hypergraph topological structure self-adaption building module is used for building a pathogen relation model, a multi-dimensional heterogeneous pathogen parameter fusion analysis module integrates morphological parameters, metabonomics parameters and other multi-dimensional parameters, and the multi-dimensional heterogeneous pathogen parameter fusion analysis module is used for analyzing the multi-dimensional parameters. The optimized hypergraph neural network reasoning module realizes pathogen category prediction, and finally, the nanopore sequencing result intelligent interpretation output module generates a detection report containing pathogen categories, subtypes and risk levels. The system is combined with nanopore sequencing and an intelligent algorithm, pathogenic microorganism detection can be efficiently and accurately completed, the accuracy, reliability and efficiency of detection are greatly improved, and the requirement for rapid and accurate detection of pathogens in practical application is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pathogen detection, and more particularly to a pathogen detection system based on nanopore sequencing and intelligent interpretation. Background Technology

[0002] Pathogen detection is crucial in public health, clinical diagnostics, and biosafety. With the development of nanotechnology and biological sequencing technology, nanopore sequencing, with its advantages of long reads, single-molecule detection, and real-time analysis, has gradually become an important method for pathogen detection. Meanwhile, artificial intelligence technology has demonstrated powerful capabilities in data analysis and pattern recognition, providing new directions for the intelligent upgrading of pathogen detection. However, current nanopore sequencing-based pathogen detection still faces many challenges.

[0003] The primary problem with existing technologies lies in their insufficient ability to analyze and process nanopore sequencing data. Traditional data processing methods struggle to efficiently extract pathogenic feature sequences from the massive biomolecular electrical signals generated by nanopore sequencing. When processing weak signals under complex background noise interference, feature loss or misjudgment is prone to occur, making it impossible to accurately capture key information about pathogenic microorganisms and limiting the accuracy and reliability of detection.

[0004] On the other hand, existing pathogen detection systems have significant shortcomings in multi-source information fusion and intelligent analysis. Pathogen detection requires comprehensive consideration of multi-dimensional heterogeneous parameters such as genomics, proteomics, and metabolomics, but traditional systems often lack effective fusion mechanisms, failing to fully explore the potential correlations between different parameters. Furthermore, in the intelligent analysis process, traditional algorithm models have limited learning and reasoning capabilities regarding pathogen characteristics, making it difficult to adapt to the diversity and complexity of pathogens, resulting in low detection efficiency and diagnostic capabilities, and failing to meet the practical needs of rapid and accurate detection. Summary of the Invention

[0005] To overcome the shortcomings and deficiencies of existing technologies, this invention provides a pathogen detection system based on nanopore sequencing and intelligent interpretation.

[0006] The technical solution adopted in this invention is a pathogen detection system based on nanopore sequencing and intelligent interpretation, comprising: The biomolecular electrical signal acquisition module consists of a nanopore chip array, a high-sensitivity current amplifier, and a signal conditioning circuit. The nanopore chip array is used to form a nanoscale detection channel, the high-sensitivity current amplifier acquires the ion current change signal passing through the nanopore, and the signal conditioning circuit performs filtering, amplification, and analog-to-digital conversion processing on the acquired signal. The pathogen feature sequence extraction module receives the digital signal output by the biomolecular electrical signal acquisition module and extracts the pathogenic microorganism feature base sequence from the signal frame by frame by constructing a multi-level differential coding matrix. The hypergraph topology adaptive construction module uses the feature sequences output by the pathogen feature sequence extraction module as nodes, constructs the hypergraph topology based on the collinearity of pathogen microbial genomes and gene cluster distribution relationships, and optimizes the hypergraph structure through a dynamic edge weight update algorithm. The multidimensional heterogeneous pathogen parameter fusion analysis module collects morphological parameters, metabolomics parameters, proteomics parameters and genomic parameters of pathogenic microorganisms, and maps the multidimensional heterogeneous parameters to a unified feature space through tensor product operation; An optimized hypergraph neural network inference module takes the hypergraph generated by the hypergraph topology adaptive construction module as input, combines the feature vector of the multidimensional heterogeneous pathogen parameter fusion analysis module, and performs pathogen microorganism category prediction through multilayer perceptron and hyperedge convolution operation; The nanopore sequencing result intelligent interpretation and output module receives the prediction results from the optimized hypergraph neural network inference module, compares them with the pathogen database through confidence matrix analysis, and generates a detection report including pathogen type, subtype and potential risk level. The biomolecular electrical signal acquisition module, pathogen feature sequence extraction module, hypergraph topology adaptive construction module, multidimensional heterogeneous pathogen parameter fusion analysis module, optimized hypergraph neural network inference module and nanopore sequencing result intelligent interpretation and output module are connected in sequence.

[0007] Furthermore, the multidimensional heterogeneous pathogen parameter fusion analysis module adopts the following fusion model: ,in, The fused feature vector; For activation functions; The number of dimensions representing pathogenic microorganism parameters; For the first Weighting coefficients for dimensional parameters; For the first A mapping matrix of dimensional parameters; For the first The vector of pathogenic microorganism parameters, the first The parameters include cell diameter in morphological parameters, metabolite concentration in metabolomics parameters, expression level of characteristic proteins in proteomics parameters, and gene fragment length in genomics parameters.

[0008] Furthermore, the optimized hypergraph neural network inference module includes a hyperedge update model: ,in, For the first Layer super edge The update vector; ReLU is the linear rectified function; For the first Layer weight matrix; For the first Layer nodes eigenvectors; For the first The bias vector of the layer, the node feature vector is composed of the feature vectors of the hypergraph nodes generated by the hypergraph topology adaptive construction module and the feature vectors of the multidimensional heterogeneous pathogen parameter fusion analysis module.

[0009] Furthermore, the pathogen feature sequence extraction module is implemented through the following encoding model: ,in, for The differential coded sequence at time step; It is a difference operation function; and They are respectively Time and The original signal sequence at time ( ); This is an element-wise multiplication operation; This is the encoding mask matrix, which is initialized based on the known characteristic sequence patterns of pathogenic microorganisms to highlight the encoding differences of characteristic sequences.

[0010] Furthermore, the edge weight adjustment model of the hypergraph topology adaptive construction module is as follows: ,in, and Superedge before and after the update The weights; and This is the weighting adjustment factor; For nodes and nodes The similarity measure of feature vectors is calculated based on the sequence alignment score of the pathogenic microorganism genome feature fragments and the multidimensional parameter feature distance.

[0011] Furthermore, the confidence evaluation model of the intelligent interpretation and output module for nanopore sequencing results is: Conf Where Conf is the confidence level of the prediction result; To compare the number of similar pathogen records in the database; For the first Weighting factors for each record; For the first The matching score between each record and the prediction result is calculated by comprehensively considering genomic sequence similarity, phenotypic feature matching degree and epidemiological association degree.

[0012] Furthermore, the multidimensional heterogeneous pathogen parameter fusion analysis module includes: a morphological parameter acquisition unit, used to acquire cell morphology, size, and surface structure parameters of pathogenic microorganisms using electron microscopy; a metabolomics parameter acquisition unit, used to determine the types, concentrations, and metabolic pathway-related parameters of pathogenic microorganism metabolites using mass spectrometry and nuclear magnetic resonance technology; a proteomics parameter acquisition unit, used to analyze the expression level, modification state, and interaction parameters of pathogenic microorganism proteins using protein chip and liquid chromatography-mass spectrometry; and a genomics parameter acquisition unit, used to acquire the whole genome sequence, gene copy number, and mutation site parameters of pathogenic microorganisms using nanopore sequencing data and real-time PCR technology.

[0013] Furthermore, the optimized hypergraph neural network inference module includes: a hypergraph convolution operation unit, which performs inter-node information transfer and feature update through hyperedge aggregation operation; a multilayer perceptron classification unit, which maps the feature vector after hypergraph convolution to the pathogen category prediction space; an attention mechanism unit, which dynamically adjusts the weights of hypergraph nodes and hyperedges according to the importance of pathogen parameters; and a model parameter optimization unit, which uses a stochastic gradient descent algorithm combined with an adaptive learning rate adjustment strategy to iteratively update the hypergraph neural network parameters.

[0014] Furthermore, the intelligent interpretation and output module for nanopore sequencing results includes: a pathogen database management unit for storing and updating the genome, phenotype, and epidemiological data of known pathogenic microorganisms; a similarity comparison unit that uses a locality-sensitive hashing algorithm to quickly compare the predicted results with database records; a risk level assessment unit that calculates the potential risk level based on pathogen pathogenicity, transmissibility, and environmental adaptability parameters; and a test report generation unit that integrates pathogen type, subtype information, and risk level according to a preset template to generate a standardized test report.

[0015] Beneficial Effects: This invention proposes a pathogen detection system based on nanopore sequencing and intelligent interpretation. This system achieves high-sensitivity signal acquisition through a biomolecular electrical signal acquisition module, combined with a pathogen feature sequence extraction module, constructing a multi-level differential coding matrix to accurately extract characteristic base sequences of pathogenic microorganisms. This solves the problems of insufficient processing capability and easy feature loss in traditional methods for complex noise signals. In terms of multi-source information fusion and intelligent analysis, the multi-dimensional heterogeneous pathogen parameter fusion analysis module uses tensor product operations to map morphological, metabolomics, and other multi-dimensional parameters to a unified feature space. Combined with a hypergraph topology adaptive construction module, it constructs and optimizes the hypergraph structure based on pathogen genome relationships, providing an effective data foundation for subsequent analysis. The optimized hypergraph neural network inference module, based on the hypergraph and fused feature vectors, achieves accurate prediction of pathogen categories through multilayer perceptrons and hyperedge convolution operations. Compared with traditional algorithm models, it significantly improves the learning and inference capabilities for pathogen features. Ultimately, the intelligent interpretation and output module for nanopore sequencing results generates a test report that includes pathogen type, subtype, and risk level by comparing the confidence matrix analysis with the pathogen database. This enables comprehensive and efficient accurate detection of pathogenic microorganisms, greatly improving the accuracy, reliability, and efficiency of detection, and meeting the needs for rapid and accurate pathogen detection in practical applications. Attached Figure Description

[0016] Figure 1 This is a diagram showing the system module composition of the present invention; Figure 2 This is a flowchart of the system operation of the present invention. Detailed Implementation

[0017] It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. The following describes the application in further detail with reference to the accompanying drawings and specific embodiments.

[0018] like Figure 1 As shown, the pathogen detection system based on nanopore sequencing and intelligent interpretation includes: The biomolecular electrical signal acquisition module consists of a nanopore chip array, a high-sensitivity current amplifier, and a signal conditioning circuit. The nanopore chip array is used to form a nanoscale detection channel, the high-sensitivity current amplifier acquires the ion current change signal passing through the nanopore, and the signal conditioning circuit performs filtering, amplification, and analog-to-digital conversion processing on the acquired signal. Specifically, the module consists of a nanopore chip array, a high-sensitivity current amplifier, and a signal conditioning circuit. The nanopore chip array is made of materials such as silicon nitride or graphene, forming detection channels with a diameter of approximately 1-2 nanometers. This size is comparable to the diameter of a single biomolecule, ensuring that the passage of a biomolecule triggers a significant change in ion current. The high-sensitivity current amplifier achieves a current detection accuracy in the picoampere range (10⁻⁶ Ω·cm).-12 The sampling frequency, reaching megahertz levels, is high enough to capture in real time the weak current changes generated when biomolecules pass through nanopores. The signal conditioning circuit integrates a low-noise amplifier, bandpass filter, and analog-to-digital converter. It filters the acquired analog signal, eliminating 50 / 60Hz power frequency interference and high-frequency noise, amplifies the signal to a level suitable for digital processing, and then performs analog-to-digital conversion at a resolution of 16 bits or higher to generate a digitized sequence of biomolecular electrical signals.

[0019] The module is implemented as follows: A biological sample solution to be tested is introduced into the detection chamber of the nanopore chip array. A transmembrane voltage of 100-200 millivolts is applied across the nanopores, driving biomolecules such as DNA, RNA, or proteins through the nanopores. As the biomolecules pass through, they partially block the ion channels, causing characteristic changes in the ion current. A high-sensitivity current amplifier acquires these current change signals in real time. A signal conditioning circuit preprocesses the raw signal, removing baseline drift and random noise, enhancing the signal-to-noise ratio, and finally outputting a digitized biomolecular electrical signal, providing a high-quality data foundation for subsequent pathogen feature sequence extraction.

[0020] The pathogen feature sequence extraction module receives the digital signal output by the biomolecular electrical signal acquisition module and extracts the pathogenic microorganism feature base sequence from the signal frame by frame by constructing a multi-level differential coding matrix. Specifically, this module receives the digital signal output from the biomolecular electrical signal acquisition module and extracts the characteristic base sequences of pathogenic microorganisms by constructing a multi-level differential coding matrix. This module employs a multi-scale differential analysis method, first dividing the continuous electrical signal sequence into time windows of 20-50 milliseconds, each window corresponding to the passage time of approximately 5-15 bases. Within each time window, the rate of change of current between adjacent sampling points is calculated, forming a first-order differential sequence to highlight the local variation characteristics of the signal. Further, a second-order differential operation is performed on the first-order differential sequence to obtain a second-order differential sequence, enhancing the sensitivity to subtle feature changes. The multi-level differential coding matrix consists of multiple differential filters of different scales, each filter corresponding to a specific base feature pattern. Convolution operations are used to identify characteristic signal patterns related to pathogenic microorganisms.

[0021] The module is implemented as follows: First, the digitized electrical signal is synchronized in the time domain to eliminate time offsets caused by inconsistent biomolecule speeds. Then, a sliding window technique is applied to divide the electrical signal sequence into overlapping analysis windows, and multi-level differential operations are performed on the signal within each window. Each filter in the differential coding matrix performs matched filtering on the differential sequence, and when a signal pattern corresponding to a known pathogenic microorganism characteristic sequence is detected, a corresponding coded output is generated. Through a dynamic threshold adjustment mechanism, the threshold for signal pattern matching is adaptively determined to reduce false positives. Finally, the identified characteristic sequence fragments are spliced ​​and error-corrected to form a complete pathogenic microorganism characteristic base sequence, providing accurate sequence data for subsequent hypergraph construction.

[0022] The hypergraph topology adaptive construction module uses the feature sequences output by the pathogen feature sequence extraction module as nodes, constructs the hypergraph topology based on the collinearity of pathogen microbial genomes and gene cluster distribution relationships, and optimizes the hypergraph structure through a dynamic edge weight update algorithm. Specifically, this module uses the feature sequences output by the pathogen feature sequence extraction module as nodes and constructs a hypergraph topology based on the collinearity and gene cluster distribution relationships of the pathogenic microorganism genome. The module first performs cluster analysis on the feature sequences, classifying them into different categories based on sequence similarity. Each category corresponds to a potential pathogenic microorganism feature node. When constructing the hypergraph, in addition to considering the binary relationships between nodes, the concept of hyperedges is introduced. Each hyperedge connects multiple nodes, indicating that the feature sequences represented by these nodes are collinear in the genome or belong to the same gene cluster. The initial weights of the hyperedges are calculated and determined based on multiple dimensions such as the similarity score between node sequences, co-occurrence frequency, and evolutionary conservation.

[0023] The implementation of this module is as follows: A graph embedding-based method is used to map feature sequences to a low-dimensional vector space, where semantic similarity between nodes is calculated. Co-occurrence matrix analysis identifies sets of frequently occurring feature sequences, creating hyperedges for each set. A dynamic edge weight update algorithm, based on a reinforcement learning framework, adaptively adjusts the hyperedge weights according to the contribution of the hypergraph structure to subsequent classification tasks. During training, the classification performance of the optimized hypergraph neural network inference module under different hypergraph structures is compared, and the hyperedge weights are adjusted accordingly to ensure the hypergraph structure better reflects the intrinsic relationships between pathogenic microorganisms. After multiple iterative optimizations, a hypergraph topology that accurately represents the genomic feature relationships of pathogenic microorganisms is finally formed, providing an effective graph structure model for pathogen classification.

[0024] The multidimensional heterogeneous pathogen parameter fusion analysis module collects morphological parameters, metabolomics parameters, proteomics parameters and genomic parameters of pathogenic microorganisms, and maps the multidimensional heterogeneous parameters to a unified feature space through tensor product operation; Specifically, this module collects morphological, metabolomics, proteomics, and genomic parameters of pathogenic microorganisms, mapping these multidimensional heterogeneous parameters to a unified feature space through tensor product operations. The morphological parameter acquisition unit uses scanning electron microscopy or transmission electron microscopy, achieving nanometer-level resolution to obtain morphological characteristics such as cell diameter, aspect ratio, and surface protrusion density of pathogenic microorganisms. The metabolomics parameter acquisition unit utilizes liquid chromatography-mass spectrometry (LC-MS / MS) to detect hundreds of small molecule metabolites, determining the concentrations of specific metabolites such as ATP and pyruvate, as well as metabolic pathway activity parameters. The proteomics parameter acquisition unit employs protein chip technology, enabling simultaneous detection of thousands of proteins and analysis of characteristic protein expression levels, phosphorylation modification states, and protein-protein interaction network parameters. The genomic parameter acquisition unit combines nanopore sequencing data and quantitative real-time PCR to determine parameters such as the whole genome sequence of pathogenic microorganisms, copy numbers of specific genes, and SNP mutation sites.

[0025] The implementation of this module is as follows: First, the parameters of each dimension are standardized and preprocessed to eliminate dimensional differences caused by different detection technologies. Morphological parameters are represented as three-dimensional spatial vectors, metabolomics parameters as metabolite concentration matrices, proteomics parameters as adjacency matrices of protein-protein interaction graphs, and genomics parameters as gene sequence vectors. Through tensor product operations, these parameters with different dimensions and structures are mapped to a unified high-dimensional tensor space, in which the intrinsic structural information of each dimension parameter is preserved. An attention mechanism is adopted to dynamically allocate weights according to the contribution of different parameters to pathogen classification, enhancing the expressive power of key parameters. Finally, a unified feature vector comprehensively representing the multidimensional characteristics of pathogenic microorganisms is formed, providing a comprehensive and discriminative feature representation for subsequent neural network inference.

[0026] An optimized hypergraph neural network inference module takes the hypergraph generated by the hypergraph topology adaptive construction module as input, combines the feature vector of the multidimensional heterogeneous pathogen parameter fusion analysis module, and performs pathogen microorganism category prediction through multilayer perceptron and hyperedge convolution operation; Specifically, this module takes the hypergraph generated by the hypergraph topology adaptive construction module as input, combines it with the feature vectors from the multidimensional heterogeneous pathogen parameter fusion analysis module, and performs pathogen category prediction through multilayer perceptron and hyperedge convolution operations. The hypergraph convolution operation unit of this module employs a message passing mechanism, where each node receives information from its neighboring nodes through hyperedges. During information transmission, the weights and directions of the hyperedges are considered to achieve feature aggregation and updating between nodes. The multilayer perceptron classification unit consists of multiple fully connected layers, each containing 100-500 neurons. A ReLU activation function is used to introduce a nonlinear transformation, mapping the high-dimensional feature vectors after hypergraph convolution to a pathogen category prediction space. The dimension of this space corresponds to the number of known pathogen species.

[0027] The implementation of this module is as follows: First, the hypergraph nodes are initialized with features, and the unified feature vectors generated by the multidimensional heterogeneous parameter fusion analysis module are assigned to the corresponding hypergraph nodes. In the hyperedge convolution operation, the feature vector of each node is updated by weighted aggregation of the feature information of its neighboring nodes, with the weights determined by the weights of the hyperedges and the similarity between nodes. After multiple layers of hyperedge convolution, the node feature vectors gradually include global structural information. The attention mechanism unit dynamically adjusts the weights of the hypergraph nodes and hyperedges according to the importance of pathogen parameters, highlighting features that play a key role in classification. The model parameter optimization unit uses a stochastic gradient descent algorithm combined with an adaptive learning rate adjustment strategy to continuously adjust the network parameters during training, minimizing the cross-entropy loss function between the predicted results and the true pathogen categories. After multiple rounds of training, the network learns the mapping relationship between pathogen features and categories, and can accurately predict the pathogen category of unknown samples.

[0028] The nanopore sequencing result intelligent interpretation and output module receives the prediction results from the optimized hypergraph neural network inference module, compares them with the pathogen database through confidence matrix analysis, and generates a detection report including pathogen type, subtype and potential risk level. The biomolecular electrical signal acquisition module, pathogen feature sequence extraction module, hypergraph topology adaptive construction module, multidimensional heterogeneous pathogen parameter fusion analysis module, optimized hypergraph neural network inference module and nanopore sequencing result intelligent interpretation and output module are connected in sequence.

[0029] Specifically, this module receives the prediction results from the optimized hypergraph neural network inference module, compares them with the pathogen database through confidence matrix analysis, and generates a detection report including pathogen type, subtype, and potential risk level. The module's pathogen database management unit stores the genome sequences, phenotypic characteristics, and epidemiological data of over 100,000 known pathogens, and is regularly updated to include newly discovered pathogens. The similarity comparison unit uses a locality-sensitive hashing algorithm to quickly compare the prediction results with records in the database, calculating multi-dimensional matching scores such as genome sequence similarity, phenotypic feature matching degree, and epidemiological correlation. The confidence assessment model comprehensively considers these matching scores, combined with the classification probability output by the hypergraph neural network inference module, to calculate the confidence value of the prediction result.

[0030] The implementation of this module is as follows: First, the pathogen category prediction results output by the optimized hypergraph neural network inference module are converted into standard classification labels and compared with records in the pathogen database. The similarity comparison unit calculates a matching score for each candidate pathogen record from three dimensions: genome, phenotype, and epidemiology, generating a matching score matrix. The risk level assessment unit calculates the potential risk level using the analytic hierarchy process (AHP) based on parameters such as pathogen pathogenicity, transmissibility, and environmental adaptability. The test report generation unit integrates pathogen type, subtype information, confidence scores, and risk levels according to a preset template to generate a structured test report. The report content includes an overview of the testing method, pathogen identification results, confidence analysis, potential risk assessment, and prevention and control recommendations, presented in text, tables, and charts to provide users with a comprehensive and intuitive interpretation of the test results.

[0031] Preferably, the multidimensional heterogeneous pathogen parameter fusion analysis module adopts the following fusion model: ,in, The fused feature vector; For activation functions; The number of dimensions representing pathogenic microorganism parameters; For the first Weighting coefficients for dimensional parameters; For the first A mapping matrix of dimensional parameters; For the first The vector of pathogenic microorganism parameters, the first The parameters include cell diameter in morphological parameters, metabolite concentration in metabolomics parameters, expression level of characteristic proteins in proteomics parameters, and gene fragment length in genomics parameters.

[0032] Specifically, the multidimensional heterogeneous pathogen parameter fusion analysis module achieves the fusion of multidimensional parameters by constructing a weighted tensor product model. This module performs standardized preprocessing on the morphological, metabolomics, proteomics, and genomic parameters of pathogenic microorganisms, eliminating dimensional differences caused by different detection technologies. By assigning adaptive weight coefficients to each dimension parameter and combining them with a mapping matrix, parameters with different structures are mapped to a unified feature space. Activation functions enhance the nonlinear expressive power of the model, enabling the fused feature vector to more comprehensively represent the characteristics of pathogenic microorganisms. In implementation, the parameters of each dimension are first normalized, and then the weight coefficients are dynamically adjusted according to the contribution of each parameter to the classification task. A fused feature vector is generated through tensor product operations, providing a comprehensive and discriminative feature representation for subsequent pathogen classification.

[0033] Preferably, the optimized hypergraph neural network inference module includes a hyperedge update model: ,in, For the first Layer super edge The update vector; ReLU is the linear rectified function; For the first Layer weight matrix; For the first Layer nodes eigenvectors; For the first The bias vector of the layer, the node feature vector is composed of the feature vectors of the hypergraph nodes generated by the hypergraph topology adaptive construction module and the feature vectors of the multidimensional heterogeneous pathogen parameter fusion analysis module.

[0034] Specifically, the optimized hypergraph neural network inference module's hyperedge update model utilizes a message-passing mechanism to achieve information transfer and feature updates between nodes in the hypergraph structure. Based on the hypergraph topology, this model uses linear transformations and activation functions to non-linearly combine node features, generating updated hyperedge vectors. The weight matrix and bias vector are continuously optimized during training, enabling the model to adaptively learn the complex relationships between pathogen features. In implementation, hypergraph nodes are initialized with fused pathogen feature vectors. Multi-layer hyperedge convolution operations are used to abstract features layer by layer, ultimately mapping node features to the pathogen category space, achieving accurate classification of pathogens.

[0035] Preferably, the pathogen feature sequence extraction module is implemented using the following encoding model: ,in, for The differential coded sequence at time step; It is a difference operation function; and They are respectively Time and The original signal sequence at time ( ); This is an element-wise multiplication operation; This is the encoding mask matrix, which is initialized based on the known characteristic sequence patterns of pathogenic microorganisms to highlight the encoding differences of characteristic sequences.

[0036] Specifically, the encoding model of the pathogen feature sequence extraction module extracts features from nanopore sequencing electrical signals through differential operations and a mask matrix. This model performs differential operations on signal sequences within adjacent time windows to highlight local changes in the signal. The mask matrix is ​​initialized based on known pathogen characteristic patterns to enhance the ability to identify feature sequences. By processing the electrical signal sequence frame by frame, feature-encoded sequences of pathogenic microorganisms are generated. In implementation, the electrical signals are first synchronized in the time domain and noise filtered. Then, a sliding window technique is applied for differential encoding. False positives are reduced through dynamic threshold adjustment, ultimately generating accurate pathogen feature sequences.

[0037] Preferably, the edge weight adjustment model of the hypergraph topology adaptive construction module is as follows: ,in, and Superedge before and after the update The weights; and This is the weighting adjustment factor; For nodes and nodes The similarity measure of feature vectors is calculated based on the sequence alignment score of the pathogenic microorganism genome feature fragments and the multidimensional parameter feature distance.

[0038] Specifically, the edge weight adjustment model of the hypergraph topology adaptive construction module dynamically optimizes the hypergraph structure by combining historical weights and node similarity. This model uses reinforcement learning as a framework, adjusting edge weights based on the contribution of the hypergraph structure to the classification task, enabling the hypergraph to more accurately reflect the intrinsic relationships between pathogenic microorganisms. Node similarity is calculated based on genomic feature alignment scores and multidimensional parameter distance, ensuring that the adjustment of edge weights has biological significance. In implementation, an initial hypergraph structure is first constructed, and then the edge weights are adjusted based on classification performance feedback during training. After multiple iterations of optimization, a hypergraph structure that effectively represents pathogen relationships is formed.

[0039] Preferably, the confidence evaluation model of the intelligent interpretation and output module for nanopore sequencing results is: Conf Where Conf is the confidence level of the prediction result; To compare the number of similar pathogen records in the database; For the first Weighting factors for each record; For the first The matching score between each record and the prediction result is calculated by comprehensively considering genomic sequence similarity, phenotypic feature matching degree and epidemiological association degree.

[0040] Specifically, the confidence assessment model of the intelligent interpretation and output module for nanopore sequencing results calculates the reliability of the predicted results by integrating multi-dimensional matching scores. This model considers factors such as genomic sequence similarity, phenotypic feature matching degree, and epidemiological association, assigning weight factors to each alignment record and calculating the final confidence level through weighted averaging. The weight factors are dynamically adjusted according to the importance of different dimensional parameters to pathogen identification, ensuring the accuracy of the assessment results. During implementation, the predicted results are compared with a pathogen database to generate a multi-dimensional matching score matrix, and then the confidence value is calculated based on the weight factors, providing a reliable credibility assessment for the test report.

[0041] Preferably, the multidimensional heterogeneous pathogen parameter fusion analysis module includes: a morphological parameter acquisition unit for acquiring cell morphology, size, and surface structure parameters of pathogenic microorganisms using electron microscopy; a metabolomics parameter acquisition unit for determining the types, concentrations, and metabolic pathway-related parameters of pathogenic microorganism metabolites using mass spectrometry and nuclear magnetic resonance technology; a proteomics parameter acquisition unit for analyzing the expression levels, modification states, and interaction parameters of pathogenic microorganism proteins using protein chips and liquid chromatography-mass spectrometry; and a genomics parameter acquisition unit for acquiring the whole genome sequence, gene copy number, and mutation site parameters of pathogenic microorganisms using nanopore sequencing data and quantitative real-time PCR technology.

[0042] Specifically, the four units of the multidimensional heterogeneous pathogen parameter fusion analysis module are each responsible for collecting pathogen parameters from different dimensions. The morphological parameter acquisition unit obtains microscopic structural information of pathogenic microorganisms using electron microscopy; the metabolomics parameter acquisition unit analyzes metabolite characteristics using mass spectrometry and nuclear magnetic resonance (NMR) techniques; the proteomics parameter acquisition unit detects protein expression and interactions using protein chip and mass spectrometry combined techniques; and the genomics parameter acquisition unit combines nanopore sequencing and quantitative real-time PCR to determine genomic characteristics. Each unit operates independently but works collaboratively to provide comprehensive data support for subsequent parameter fusion. During implementation, each unit collects parameters according to standardized procedures, and performs quality control and preprocessing on the collected data to ensure accuracy and consistency.

[0043] Preferably, the optimized hypergraph neural network inference module includes: a hypergraph convolution operation unit, which performs inter-node information transfer and feature update through hyperedge aggregation operation; a multilayer perceptron classification unit, which maps the feature vector after hypergraph convolution to the pathogen category prediction space; an attention mechanism unit, which dynamically adjusts the weights of hypergraph nodes and hyperedges according to the importance of pathogen parameters; and a model parameter optimization unit, which uses a stochastic gradient descent algorithm combined with an adaptive learning rate adjustment strategy to iteratively update the hypergraph neural network parameters.

[0044] Specifically, the four units of the optimized hypergraph neural network inference module work together to achieve intelligent classification of pathogenic microorganisms. The hypergraph convolution operation unit realizes information transfer between nodes through hyperedge aggregation; the multilayer perceptron classification unit maps high-dimensional features to the category space; the attention mechanism unit dynamically adjusts the weights of nodes and hyperedges; and the model parameter optimization unit optimizes network parameters using an adaptive learning rate strategy. In implementation, pathogen features are first extracted through hypergraph convolution, then key features are enhanced using the attention mechanism, and after classification by the multilayer perceptron, model performance is continuously improved through parameter optimization, ultimately achieving accurate classification of pathogenic microorganisms. Preferably, the nanopore sequencing result intelligent interpretation and output module includes: a pathogen database management unit for storing and updating the genome, phenotype, and epidemiological data of known pathogenic microorganisms; a similarity comparison unit for quickly comparing the predicted results with database records using a locality-sensitive hashing algorithm; a risk level assessment unit for calculating the potential risk level based on pathogen pathogenicity, transmissibility, and environmental adaptability parameters; and a test report generation unit for generating a standardized test report by integrating pathogen type, subtype information, and risk level according to a preset template.

[0045] Specifically, the four units of the nanopore sequencing result intelligent interpretation and output module are responsible for analyzing the test results and generating reports. The pathogen database management unit maintains a comprehensive database including genomic, phenotypic, and epidemiological data; the similarity comparison unit uses a locality-sensitive hashing algorithm to quickly match candidate pathogens; the risk level assessment unit calculates potential risks based on pathogen characteristics; and the test report generation unit integrates the results according to a preset template. In practice, the predicted results are first compared with the database, then the risk level is assessed, and finally a structured report including pathogen type, subtype, and risk level is generated, providing users with a comprehensive interpretation of the test results.

[0046] like Figure 2 As shown, the pathogen detection system based on nanopore sequencing and intelligent interpretation includes the following steps in its operation: Step S1: The biomolecular electrical signal acquisition module acquires the biomolecular electrical signal generated by the change in ion current through the nanopore, and converts it into a digital signal through the signal conditioning circuit. Step S2: The pathogen feature sequence extraction module performs multi-level differential encoding on the digital signal to extract the characteristic base sequence of the pathogenic microorganism; Step S3: The hypergraph topology adaptive construction module constructs the hypergraph topology based on the feature base sequence and optimizes the hypergraph through a dynamic edge weight update algorithm. Step S4: The multidimensional heterogeneous pathogen parameter fusion analysis module collects morphological, metabolomics, proteomics and genomic parameters of pathogenic microorganisms and maps them to a unified feature space through tensor product operation; Step S5: The optimized hypergraph neural network inference module takes the hypergraph and fused feature vector as input, and performs pathogen category prediction through multilayer perceptron and hyperedge convolution operation; Step S6: The nanopore sequencing result intelligent interpretation and output module compares the results with the pathogen database through confidence matrix analysis to generate a test report including pathogen type, subtype and potential risk level.

[0047] This detection system innovates in multiple aspects, including data processing, information fusion and analysis, and result output, effectively overcoming the shortcomings of previous technologies and demonstrating significant advantages. In data processing, addressing the difficulties of traditional methods in extracting pathogen characteristic sequences from nanopore sequencing data and their susceptibility to noise interference, the system achieves high-sensitivity signal acquisition through a biomolecular electrical signal acquisition module. It then utilizes a pathogen characteristic sequence extraction module to construct a multi-level differential coding matrix, extracting the characteristic base sequences of pathogenic microorganisms frame by frame. This approach can accurately process weak signals under complex background noise, avoiding feature loss or misjudgment, and significantly improving the accuracy and reliability of data processing.

[0048] At the level of multi-source information fusion and intelligent analysis, traditional systems cannot effectively integrate multi-dimensional parameters and have limited intelligent analysis capabilities. However, the multi-dimensional heterogeneous pathogen parameter fusion analysis module of this system uses tensor product operations to map multi-dimensional heterogeneous parameters such as morphology, metabolomics, proteomics, and genomics of pathogenic microorganisms to a unified feature space, breaking down information silos. The hypergraph topology adaptive construction module constructs a hypergraph based on the collinearity of pathogen genomes and gene cluster distribution relationships, and optimizes the structure through a dynamic edge weight update algorithm, providing an efficient data structure for subsequent analysis. The optimized hypergraph neural network inference module combines the hypergraph with fused feature vectors, and through multilayer perceptron and hyperedge convolution operations, enhances the learning and inference capabilities of pathogen features, adapts to the diversity and complexity of pathogenic microorganisms, and significantly improves detection efficiency and diagnostic capabilities.

[0049] In the result output stage, the intelligent interpretation module for nanopore sequencing results compares the results with a pathogen database using a confidence matrix analysis. It comprehensively considers factors such as genomic sequence similarity, phenotypic feature matching, and epidemiological correlation to calculate the confidence level of the predicted results and generate a test report including pathogen type, subtype, and potential risk level. Compared to traditional systems, this system not only achieves accurate detection of pathogenic microorganisms but also provides more comprehensive and detailed detection information, meeting the urgent need for rapid and accurate detection of pathogenic microorganisms in practical applications. It has significant application value in public health, clinical diagnosis, and biosafety.

[0050] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0051] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A pathogen detection system based on nanopore sequencing and intelligent interpretation, characterized in that, include: The biomolecular electrical signal acquisition module consists of a nanopore chip array, a high-sensitivity current amplifier, and a signal conditioning circuit. The nanopore chip array is used to form a nanoscale detection channel, the high-sensitivity current amplifier acquires the ion current change signal passing through the nanopore, and the signal conditioning circuit performs filtering, amplification, and analog-to-digital conversion processing on the acquired signal. The pathogen feature sequence extraction module receives the digital signal output by the biomolecular electrical signal acquisition module and extracts the pathogenic microorganism feature base sequence from the signal frame by frame by constructing a multi-level differential coding matrix. The hypergraph topology adaptive construction module uses the feature sequences output by the pathogen feature sequence extraction module as nodes, constructs the hypergraph topology based on the collinearity of pathogen microbial genomes and gene cluster distribution relationships, and optimizes the hypergraph structure through a dynamic edge weight update algorithm. The multidimensional heterogeneous pathogen parameter fusion analysis module collects morphological parameters, metabolomics parameters, proteomics parameters and genomic parameters of pathogenic microorganisms, and maps the multidimensional heterogeneous parameters to a unified feature space through tensor product operation; An optimized hypergraph neural network inference module takes the hypergraph generated by the hypergraph topology adaptive construction module as input, combines the feature vector of the multidimensional heterogeneous pathogen parameter fusion analysis module, and performs pathogen microorganism category prediction through multilayer perceptron and hyperedge convolution operation; The nanopore sequencing result intelligent interpretation and output module receives the prediction results from the optimized hypergraph neural network inference module, compares them with the pathogen database through confidence matrix analysis, and generates a test report including pathogen type, subtype and potential risk level.

2. The pathogen detection system based on nanopore sequencing and intelligent interpretation according to claim 1, characterized in that, The multidimensional heterogeneous pathogen parameter fusion analysis module adopts the following fusion model: ,in, The fused feature vector; For activation functions; The number of dimensions representing pathogenic microorganism parameters; For the first Weighting coefficients for dimensional parameters; For the first A mapping matrix of dimensional parameters; For the first The vector of pathogenic microorganism parameters, the first The parameters include cell diameter in morphological parameters, metabolite concentration in metabolomics parameters, expression level of characteristic proteins in proteomics parameters, and gene fragment length in genomics parameters.

3. The pathogen detection system based on nanopore sequencing and intelligent interpretation according to claim 1, characterized in that, The optimized hypergraph neural network inference module includes a hyperedge update model: ,in, For the first Layer super edge The update vector; ReLU is the linear rectified function; For the first Layer weight matrix; For the first Layer nodes eigenvectors; For the first The bias vector of the layer, the node feature vector is composed of the feature vectors of the hypergraph nodes generated by the hypergraph topology adaptive construction module and the feature vectors of the multidimensional heterogeneous pathogen parameter fusion analysis module.

4. The pathogen detection system based on nanopore sequencing and intelligent interpretation according to claim 1, characterized in that, The pathogen feature sequence extraction module is implemented using the following encoding model: ,in, for The differential coded sequence at time step; It is a difference operation function; and They are respectively Time and The original signal sequence at time ( ); This is an element-wise multiplication operation; This is the encoding mask matrix, which is initialized based on the known characteristic sequence patterns of pathogenic microorganisms to highlight the encoding differences of characteristic sequences.

5. The pathogen detection system based on nanopore sequencing and intelligent interpretation according to claim 1, characterized in that, The edge weight adjustment model of the hypergraph topology adaptive construction module is as follows: ,in, and Superedge before and after the update The weights; and This is the weighting adjustment factor; For nodes and nodes The similarity measure of feature vectors is calculated based on the sequence alignment score of the pathogenic microorganism genome feature fragments and the multidimensional parameter feature distance.

6. The pathogen detection system based on nanopore sequencing and intelligent interpretation according to claim 1, characterized in that, The confidence evaluation model for the intelligent interpretation and output module of nanopore sequencing results is as follows: Conf Where Conf is the confidence level of the prediction result; To compare the number of similar pathogen records in the database; For the first Weighting factors for each record; For the first The matching score between each record and the prediction result is calculated by comprehensively considering genomic sequence similarity, phenotypic feature matching degree and epidemiological association degree.

7. The pathogen detection system based on nanopore sequencing and intelligent interpretation according to claim 1, characterized in that, The multidimensional heterogeneous pathogen parameter fusion analysis module includes: a morphological parameter acquisition unit, used to acquire cell morphology, size, and surface structure parameters of pathogenic microorganisms using electron microscopy; a metabolomics parameter acquisition unit, used to determine the types, concentrations, and metabolic pathway-related parameters of pathogenic microorganism metabolites using mass spectrometry and nuclear magnetic resonance technology; a proteomics parameter acquisition unit, used to analyze the expression level, modification state, and interaction parameters of pathogenic microorganism proteins using protein chip and liquid chromatography-mass spectrometry; and a genomics parameter acquisition unit, used to acquire the whole genome sequence, gene copy number, and mutation site parameters of pathogenic microorganisms using nanopore sequencing data and real-time PCR technology.

8. The pathogen detection system based on nanopore sequencing and intelligent interpretation according to claim 1, characterized in that, The optimized hypergraph neural network inference module includes: a hypergraph convolution operation unit, which performs inter-node information transfer and feature update through hyperedge aggregation operation; a multilayer perceptron classification unit, which maps the feature vector after hypergraph convolution to the pathogen category prediction space; an attention mechanism unit, which dynamically adjusts the weights of hypergraph nodes and hyperedges according to the importance of pathogen parameters; and a model parameter optimization unit, which uses a stochastic gradient descent algorithm combined with an adaptive learning rate adjustment strategy to iteratively update the hypergraph neural network parameters.

9. The pathogen detection system based on nanopore sequencing and intelligent interpretation according to claim 1, characterized in that, The intelligent interpretation and output module for nanopore sequencing results includes: a pathogen database management unit for storing and updating the genome, phenotype, and epidemiological data of known pathogenic microorganisms; a similarity comparison unit that uses a locality-sensitive hashing algorithm to quickly compare the predicted results with database records; a risk level assessment unit that calculates the potential risk level based on pathogen pathogenicity, transmissibility, and environmental adaptability parameters; and a test report generation unit that integrates pathogen type, subtype information, and risk level according to a preset template to generate a standardized test report.

Citation Information

Cited By

  • Multiplex immunofluorescence assay method for detecting multiple microbial pathogens

    CN121558698A