Targeted sequencing pathogen-drug resistance gene synchronous detection method and system

By constructing a composite knowledge graph and using the Kalman filter algorithm to correct the overlap between pathogens and drug-resistant genes, the problem of inaccurate detection caused by PCR inhibitors in targeted sequencing technology was solved, and higher-precision simultaneous detection of pathogens and drug-resistant genes was achieved.

CN121617467AActive Publication Date: 2026-03-06SHANGHAI HONGXU BIOTECHNOLOGY CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610141861.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-02
Publication Date
2026-03-06
Estimated Expiration
2046-02-02

AI Technical Summary

Technical Problem

Existing targeted sequencing technologies cannot accurately identify the "decoupling" phenomenon between pathogens and drug resistance genes when faced with PCR inhibitors, resulting in inaccurate test results.

Method used

A composite knowledge graph of pathogens and drug resistance genes was constructed. By introducing an inhibition intensity factor, the Kalman filter algorithm was used for iterative inverse estimation to correct the overlap rate between pathogen gene sequences and drug resistance gene sequences and eliminate the interference of PCR inhibitors.

Benefits of technology

It improves the accuracy of pathogen and drug resistance gene detection, effectively eliminates the interference of PCR inhibitors on sequencing data, and ensures the accuracy of test results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121617467A_ABST
    Figure CN121617467A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biomedical engineering, in particular to a targeted sequencing pathogen-drug-resistant gene synchronous detection method and system.The pathogen-drug-resistant gene synchronous detection method is characterized in that a composite knowledge graph about pathogens and drug-resistant genes is constructed, and the pathogen-drug-resistant gene synchronous detection result is obtained on the basis of the composite knowledge graph; the method comprises the following steps: determining an overlapping rate between a pathogen gene sequence and a drug-resistant gene sequence of a target pathogen in a to-be-detected clinical sample, analyzing sequence inhibition sensitivity of the pathogen gene sequence and the drug-resistant gene sequence of the target pathogen, and constructing a collaborative state evolution model in combination with a composite knowledge graph, and obtaining an optimal inhibition intensity factor of the drug-resistant gene sequence by using a Kalman filtering algorithm, correcting the overlapping ratio by using the optimal inhibition intensity factor, and finally determining the drug-resistant gene sequence in all pathogen gene sequences and the gene abundance thereof. According to the invention, the detection precision of the drug-resistant gene is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical engineering technology, specifically to a method and system for simultaneous detection of pathogen-drug resistance genes using targeted sequencing. Background Technology

[0002] Identifying drug-resistant pathogens is crucial in the diagnosis of infectious diseases. Traditional detection methods, such as bacterial culture and biochemical assays, typically take several days to obtain results, and their efficiency and accuracy are significantly reduced when dealing with multiple infections or difficult-to-culture pathogens. In recent years, targeted sequencing technology has been introduced into the field of pathogen detection due to its high throughput and high sensitivity, enabling the simultaneous identification of multiple pathogens and their drug resistance genes in a shorter time.

[0003] In current targeted sequencing technologies, pathogens (hosts) and drug resistance genes are typically analyzed quantitatively as independent events, and common correction methods (such as GC content-based bias correction) are performed independently for each individual sequence. However, clinical samples contain complex PCR inhibitors (such as heme, polysaccharides, and proteins), which interfere with different sequences to varying degrees (sequence specificity). This can lead to a physically linked DNA molecule (e.g., a drug resistance plasmid inside a bacterium) exhibiting inconsistent coverage depth in sequencing data (e.g., high pathogen detection abundance but low drug resistance gene detection abundance). Existing technologies cannot identify this "decoupling" phenomenon caused by inhibition differences, easily resulting in inaccurate pathogen and drug resistance gene detection results. Summary of the Invention

[0004] To address the aforementioned technical problem of inaccurate drug resistance gene detection results due to the influence of PCR inhibitors, the present invention aims to provide a method and system for simultaneous detection of pathogens and drug resistance genes via targeted sequencing. The specific technical solution adopted is as follows: In a first aspect, the present invention provides a method for simultaneous detection of pathogen-drug resistance genes via targeted sequencing, comprising the following steps: Based on the genomic sequence information and drug resistance gene sequence information of different pathogens, a composite knowledge graph about pathogens and drug resistance genes is constructed. Based on the genomic sequence information of the different pathogens, the target pathogen present in the clinical sample to be tested and the pathogen gene sequence of the target pathogen present in the clinical sample to be tested are identified. Based on the composite knowledge graph, a set of drug-resistant gene sequences that may exist in the target pathogen is determined, and the pathogen gene sequence is compared with the drug-resistant gene sequence in the set of drug-resistant gene sequences to obtain the overlap rate between the pathogen gene sequence and the drug-resistant gene sequence. The sequence inhibition sensitivity of the pathogen gene sequence and the drug resistance gene sequence of the target pathogen is analyzed, and a co-state evolution model is constructed in conjunction with the composite knowledge graph. The co-state evolution model introduces an inhibition intensity factor as a latent variable to characterize the degree of PCR amplification inhibition. The Kalman filter algorithm is used to estimate the state prior determined by the co-state evolution model and update the residuals by combining the state observation values. The optimal inhibition intensity factor of the drug resistance gene sequence is determined by iterative inverse estimation. The overlap rate is corrected using the optimal inhibition strength factor, and the drug resistance gene sequences and their abundances in all pathogen gene sequences are determined based on the overlap rate correction value.

[0005] In conjunction with the first aspect mentioned above, among some possible implementation methods, a composite knowledge graph about pathogens and drug resistance genes can be constructed, including: Each pathogen and each drug resistance gene sequence is treated as a node in the composite knowledge graph; Based on the genomic sequence information and drug resistance gene sequence information of different pathogens, when the drug resistance gene sequence exists in the genomic sequence of the pathogen, a directed edge is constructed between the nodes of the corresponding pathogen and the drug resistance gene sequence, and the edge weight of the directed edge is determined. The directed edge points from the drug resistance gene sequence to the pathogen, and the edge weight is used to reflect the tightness of the evolutionary association between the drug resistance gene sequence and the pathogen.

[0006] In conjunction with the first aspect above, in some possible implementations, determining the edge weight of the directed edge includes: For the drug-resistant gene sequence and pathogen corresponding to the directed edge, determine the copy number of the drug-resistant gene sequence in the genome sequence of the pathogen; For the drug resistance gene sequence and pathogen corresponding to the directed edge, determine the insertion position and detection frequency of the drug resistance gene sequence in the genome sequence of different strains of the pathogen; The edge weight corresponding to the directed edge is determined based on the number of copies, the stability of the insertion position, and the detection frequency.

[0007] In conjunction with the first aspect above, in some possible implementations, analyzing the sequence inhibition sensitivity of the pathogen gene sequence and the drug resistance gene sequence of the target pathogen includes: Identify the core gene sequence in the pathogen gene sequence of the target pathogen; For the core gene sequence and the drug resistance gene sequence, determine the GC base content and the minimum free energy of the secondary structure of the gene sequence; Based on the GC base content and minimum free energy of secondary structure of all the core gene sequences, the inhibition sensitivity index of the target pathogen is determined, and based on the GC base content and minimum free energy of secondary structure of the drug resistance gene sequence, the inhibition sensitivity index of the drug resistance gene sequence is determined. The inhibition sensitivity index is used to reflect the ease with which the target pathogen or the drug resistance gene sequence is inhibited by the inhibitor.

[0008] In conjunction with the first aspect mentioned above, among some possible implementation methods, a cooperative state evolution model can be constructed, including: Based on the inhibition sensitivity index, the baseline inhibition coefficient of the clinical sample to be tested is corrected to obtain the initial inhibition strength factor of the target pathogen and the drug resistance gene sequence; Based on the initial inhibition strength factor, a basic update equation for the inhibition strength factor corresponding to the target pathogen and the drug resistance gene sequence during PCR amplification is constructed. Based on the composite knowledge graph, each candidate pathogen in the clinical sample to be tested that is associated with the drug resistance gene sequence is identified, and based on the difference in inhibition intensity factors between the drug resistance gene sequence and each candidate pathogen during PCR amplification, the synergistic driving term in the PCR amplification process is determined. Based on the rate of change of the basic update equation of the inhibition intensity factor and the synergistic driving term, the rate of change of the inhibition intensity factor corresponding to the drug resistance gene sequence during PCR amplification is determined.

[0009] In conjunction with the first aspect above, among some possible implementations, the co-driving terms in the PCR amplification process are identified, including: The difference between the inhibition intensity factor corresponding to each candidate pathogen and the drug resistance gene sequence during PCR amplification is determined to obtain the inhibition intensity factor difference value; By using the edge weights of the directed edges between the candidate pathogens and the corresponding nodes of the drug resistance gene sequences in the composite knowledge graph, the difference values ​​of the inhibition intensity factor are weighted and accumulated to obtain the synergistic driving term in the PCR amplification process.

[0010] In conjunction with the first aspect mentioned above, in some possible implementations, the Kalman filter algorithm is used to determine the optimal inhibitory strength factor of the drug resistance gene sequence based on the prior state estimate determined by the cooperative state evolution model, combined with residual updates using state observations, and through iterative inverse estimation, including: The inhibitory factor and abundance values ​​of the target pathogen corresponding to each PCR amplification cycle in the Kalman filtering process are obtained and constituted as state observation values ​​corresponding to each PCR amplification cycle in the Kalman filtering process. Based on the initial inhibition strength factor, the basic update equation of the inhibition strength factor, and the rate of change of the inhibition strength factor, the predicted inhibition factor values ​​of the target pathogen and the drug resistance gene sequence corresponding to each PCR amplification cycle in the Kalman filtering process are determined. Based on the predicted value of the inhibitory factor of the target pathogen, the PCR exponential growth model of the target pathogen is modified to determine the predicted abundance value of the target pathogen for each PCR amplification cycle in the Kalman filtering process. Based on the predicted values ​​of the inhibitory factors and the abundance of the target pathogen, a priori state estimate corresponding to each PCR amplification cycle in the Kalman filtering process is determined. Based on the prior state estimate and the observed state values, the observation residuals are determined, and the predicted value of the inhibitory factor of the drug-resistant gene sequence is used as the state update object. Kalman filtering is performed iteratively with each PCR amplification cycle to update the residuals, and finally the optimal inhibitory strength factor corresponding to the drug-resistant gene sequence is obtained.

[0011] In conjunction with the first aspect above, in some possible implementations, the overlap rate is corrected using the optimal suppression strength factor, including: Determine the ratio of the overlap rate to the optimal suppression intensity factor; The product of the ratio and the probe capture efficiency coefficient is determined as the overlap rate correction value.

[0012] In conjunction with the first aspect above, in some possible implementations, based on overlap correction values, the drug resistance gene sequences and their abundance in all said pathogen gene sequences are determined, including: Pathogen gene sequences whose overlap correction value is less than or equal to the set overlap threshold are removed to obtain the remaining pathogen gene sequences after removal. The remaining pathogen gene sequences after the removal are subjected to duplicate sequence removal, and reads from the same original deoxyribonucleic acid template are merged to form a UMI family. The sequence consistency of pathogen gene sequences in the UMI family is calculated, and UMI families with sequence consistency greater than a set consistency threshold are retained. By removing pathogen gene sequences containing chimeric sequences from the retained UMI family, the drug resistance gene sequences and their abundances in all pathogen gene sequences are finally obtained.

[0013] Secondly, the present invention also provides a targeted sequencing-based pathogen-drug resistance gene simultaneous detection system, including a memory and a processor. The memory is used to store executable computer program code, and the processor is used to call and run the executable computer program code from the memory, causing the system to perform a targeted sequencing-based pathogen-drug resistance gene simultaneous detection method according to the first aspect or any possible implementation thereof.

[0014] This invention offers the following advantages: First, by constructing a composite knowledge graph of pathogens and drug-resistant genes to reflect the correlation between them, readable structured prior knowledge is obtained. Second, the presence of target pathogens and their original gene sequences in the clinical samples to be tested is identified to quickly pinpoint the core target. The pathogen gene sequence is then compared with drug-resistant gene sequences in the set of possible drug-resistant gene sequences of the target pathogen determined by the composite knowledge graph, yielding the overlap rate between the pathogen and drug-resistant gene sequences, thus providing initial input for subsequent fine-tuning. Next, the sequence inhibition sensitivity of the pathogen gene sequence and drug-resistant gene sequence of the target pathogen is analyzed to quantify the degree to which different genes are "naturally" susceptible to inhibition. This, combined with the composite knowledge graph, further enhances the ability to... By utilizing the biological principles of co-amplification of pathogens and drug-resistant genes, a co-state evolution model was constructed. This model introduces an inhibition strength factor as a latent variable to characterize the degree of PCR amplification suppression. Based on the state prior estimate determined by the co-state evolution model, and combined with the state observation values ​​from sequencing data, the Kalman filter algorithm was used for residual update. Through iterative inverse estimation, the optimal inhibition strength factor for the drug-resistant gene sequence was determined, thereby decoupling the interference term "inhibition factor" from the messy sequencing data. Finally, the optimal inhibition strength factor was used to correct the overlap rate between pathogen gene sequences and drug-resistant gene sequences, restoring the true biological signal, and thus determining the drug-resistant gene sequences and their abundance in all pathogen gene sequences, effectively improving the detection accuracy. Attached Figure Description

[0015] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart illustrating the steps of a targeted sequencing method for simultaneous detection of pathogen-drug resistance genes according to an embodiment of the present invention. Figure 2This is a flowchart illustrating the steps involved in constructing a composite knowledge graph about pathogens and drug resistance genes, as described in an embodiment of the present invention. Figure 3 This is a flowchart illustrating the steps of analyzing the sequence inhibition sensitivity of pathogen gene sequences and drug resistance gene sequences of a target pathogen according to an embodiment of the present invention. Figure 4 This is a flowchart illustrating the steps involved in constructing a cooperative state evolution model according to an embodiment of the present invention. Figure 5 This is a flowchart illustrating the steps for determining the optimal inhibitory strength factor of a drug resistance gene sequence according to an embodiment of the present invention. Figure 6 This is a flowchart illustrating the steps of determining drug resistance gene sequences and their abundance in all pathogen gene sequences according to an embodiment of the present invention. Figure 7 This is a schematic diagram of a targeted sequencing pathogen-drug resistance gene synchronous detection system according to an embodiment of the present invention. Detailed Implementation

[0017] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific embodiments and in conjunction with the accompanying drawings.

[0018] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.

[0019] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0020] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0021] It should be noted that the concepts of "first" and "second" mentioned in this invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0022] Although operations or steps are described in a specific order in the accompanying drawings in the embodiments of the present invention, this should not be construed as requiring these operations or steps to be performed in the specific order or serial order shown, or requiring all of the shown operations or steps to be performed to obtain the desired result. In the embodiments of the present invention, these operations or steps may be performed serially; they may be performed in parallel; or a portion of these operations or steps may be performed.

[0023] Meanwhile, it is understood that the data involved in the technical solutions of this invention (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions. Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by those skilled in the art to which this invention pertains. Furthermore, in all division and logarithmic operations involved in this invention, a protection mechanism is employed to prevent computational crashes or invalid values ​​due to a zero denominator or zero input. The implementation of this protection mechanism can be reasonably set according to the actual situation. For example, when the denominator term of a division operation or the argument term of a logarithmic function is zero, a protection parameter with the same dimension as or dimensionless as the denominator term or the argument term can be added. The value of this protection parameter can be a very small value greater than zero, thereby ensuring the robustness and feasibility of the algorithm under extreme conditions. In addition, the normalization function mentioned in this invention, unless otherwise specifically stated, uses maximum-minimum value normalization to normalize the normalization result to the [0, 1] interval or other continuous intervals. The maximum and minimum values ​​used in the maximum-minimum normalization can be obtained according to the actual situation. For example, when multiple values ​​can be obtained in the implementation process and it is necessary to compare the relationship between different values, multiple values ​​can be counted to obtain the maximum and minimum values. However, when only a single value can be obtained in the implementation process, the maximum and minimum values ​​can be obtained by counting based on a large amount of historical experimental data or prior data obtained in the early stage.

[0024] The following will provide a detailed description of a targeted sequencing method and system for simultaneous detection of pathogen-drug resistance genes provided by an embodiment of the present invention, with reference to the accompanying drawings.

[0025] Figure 1 This diagram illustrates the basic flowchart of a targeted sequencing-based method for simultaneous detection of pathogen-drug resistance genes provided by an embodiment of the present invention. Figure 1 As shown, the method specifically includes the following steps: Step S100: Based on the genomic sequence information and drug resistance gene sequence information of different pathogens, construct a composite knowledge graph about pathogens and drug resistance genes.

[0026] This study collects publicly available pathogen genome sequences and drug resistance gene sequences from public databases. By comparing the drug resistance gene sequences with the pathogen genome sequences, a composite knowledge graph of pathogens and drug resistance genes is constructed. This composite knowledge graph reflects the relationship between pathogens and drug resistance genes as a "drug resistance ecological map," showcasing which pathogens and drug resistance genes are carried by a particular pathogen, and the co-occurrence relationships between different drug resistance genes. This provides readable, structured prior knowledge for subsequent identification of the relationship between pathogens and drug resistance genes.

[0027] Step S200: Based on the genomic sequence information of different pathogens, identify the target pathogen present in the clinical sample to be tested and the pathogen gene sequence of the target pathogen present in the clinical sample to be tested.

[0028] The pathogen present in the clinical sample to be tested is identified as the target pathogen, and all DNA sequences of the sequencing data of the clinical sample to be tested are obtained by sequencing. Based on the publicly available genomic sequence information of the pathogen, the pathogen gene sequence of the target pathogen present in the clinical sample to be tested is determined.

[0029] Step S300: Based on the composite knowledge graph, determine the set of drug-resistant gene sequences that may exist in the target pathogen, and compare the pathogen gene sequence with the drug-resistant gene sequence in the set of drug-resistant gene sequences to obtain the overlap rate between the pathogen gene sequence and the drug-resistant gene sequence.

[0030] To identify the presence of drug-resistant genes in pathogen gene sequences, based on the aforementioned composite knowledge graph, a set of potentially present fixed-direction drug-resistant genes in the target pathogen is determined. Specifically, all drug-resistant gene sequence nodes with directed edges to the target pathogen node are identified in the composite knowledge graph, and the set of drug-resistant gene sequences corresponding to these nodes is taken as the drug-resistant gene sequence set. Using a sequence comparison tool, the pathogen gene sequence of the target pathogen is compared with any drug-resistant gene sequence in the aforementioned drug-resistant gene set, and the overlap rate between the gene sequences is calculated. This overlap rate can be used as a basis for verifying the presence of drug-resistant gene sequences in the pathogen gene sequence.

[0031] It should be understood that gene sequences are usually represented as a straight line of nucleotides. For example, ATGCGTAC is a simple gene fragment. The aforementioned overlap rate refers to the proportion of matching nucleotides to the length of the drug resistance gene sequence after sequence alignment (matching the same nucleotides) of two gene fragments. The core concept is to measure sequence "homology" (higher similarity likely indicates stronger sequence homology). For example, fragment 1: ATGCGTAC (length 8); fragment 2: ATGCGTAG (length 8); number of matching nucleotides = 7, total alignment length = 8, in this case, the overlap rate is 7 / 8. For fragments of unequal length, fragment X: ATGCGTAC (length 8); fragment Y: CGTAC (length 5); number of matches = 5, total length of the drug resistance gene sequence = 5, overlap rate (similarity) = 5 / 5 = 1.

[0032] Step S400: Analyze the sequence inhibition sensitivity of the pathogen gene sequence and drug resistance gene sequence of the target pathogen, and construct a co-state evolution model by combining a composite knowledge graph. The co-state evolution model introduces the inhibition intensity factor as a latent variable to characterize the degree of PCR amplification inhibition. Using the Kalman filter algorithm, the state prior determined by the co-state evolution model is estimated, and the residual is updated by combining the state observation values. The optimal inhibition intensity factor of the drug resistance gene sequence is determined by iterative inverse estimation.

[0033] Regarding the overlap rate between the identified pathogen gene sequences and the drug-resistant gene sequences in the drug-resistant gene sequence set, a high overlap rate verifies the presence of a drug-resistant gene sequence fragment within the pathogen gene sequence. The abundance of these verified drug-resistant genes can then be calculated to obtain the detection results. However, clinical samples (especially blood, sputum, and pus) may contain various PCR inhibitors such as heme, polysaccharides, and proteins. These inhibitors affect the amplification efficiency of different gene sequences differently, leading to systematic biases in the consistency of pathogen genome and drug-resistant gene detection across different samples. Specifically, if multiple samples contain the same pathogen and its drug-resistant genes, theoretically these genes should be detected in all samples (i.e., complete overlap). However, actual sequencing data shows inconsistencies, with some samples detecting and others missing, resulting in a lower-than-expected gene overlap rate. This is due to the combined effect of differential interference from inhibitors on amplification and the randomness of sequencing.

[0034] This embodiment considers using the state tracking mechanism of Kalman filtering to decouple and compensate for overlap rate errors, thereby eliminating suppressor interference. The core advantage of Kalman filtering lies in its ability to recursively estimate the optimal state value from noisy observations by utilizing the dynamic evolution of the system and the statistical characteristics of the observation data. However, because Kalman filtering assumes that state evolution is a Markov process—that is, the system state must "evolve independently" (e.g., changes in A do not directly affect B)—the pathogen-drug resistance gene system has a special constraint structure. The copy number of drug resistance genes does not evolve independently; it is entirely constrained by the number of pathogens carrying them. That is, more pathogens mean more drug resistance genes. This binding relationship cannot be handled by standard Kalman filtering; therefore, improvements and adjustments to Kalman filtering are necessary.

[0035] Therefore, this embodiment analyzes the sequence inhibition sensitivity of pathogen gene sequences and drug resistance gene sequences of the target pathogen, and combines a composite knowledge graph. Utilizing the biological characteristics of co-amplification of pathogens and drug resistance genes, it introduces an inhibition strength factor as a latent variable to characterize the degree of PCR amplification inhibition, constructing a co-state evolution model corresponding to the pathogen and drug resistance gene sequences. Then, using the Kalman filter algorithm, based on the state prior estimation determined by the co-state evolution model, and combining it with state observations for residual updates, it iterative inverse estimation determines the optimal inhibition strength factor for the drug resistance gene sequence. Subsequently, by using this optimal inhibition strength factor to correct the error in the overlap rate between pathogen gene sequences and drug resistance gene sequences, the accuracy of detecting drug resistance gene sequences in pathogen gene sequences can be effectively improved.

[0036] Step S500: Using the optimal inhibition strength factor, the overlap rate is corrected, and based on the overlap rate correction value, the drug resistance gene sequences and their gene abundance in all pathogen gene sequences are determined.

[0037] Using the optimal inhibition strength factor of the drug-resistant gene sequence output by the Kalman filter as the core, the overlap rate between pathogen gene sequences and drug-resistant gene sequences is corrected to eliminate measurement interference caused by inhibitors. Then, based on the overlap rate correction value, the drug-resistant gene sequences and their abundance in all pathogen gene sequences are accurately determined.

[0038] Based on the above technical solution, a composite knowledge graph of pathogens and drug-resistant genes is constructed to reflect the correlation between pathogens and drug-resistant genes, thereby obtaining readable structured prior knowledge. The core target is quickly identified by recognizing the presence of the target pathogen and its original gene sequence in the clinical sample to be tested. This pathogen gene sequence is then compared with drug-resistant gene sequences in the set of possible drug-resistant gene sequences of the target pathogen determined based on the composite knowledge graph, yielding the overlap rate between the pathogen gene sequence and the drug-resistant gene sequence. This provides raw input for subsequent fine-tuning. Furthermore, the sequence inhibition sensitivity of the pathogen gene sequence and drug-resistant gene sequence of the target pathogen is analyzed to quantify the degree to which different genes are "naturally" susceptible to inhibition. Combined with the composite knowledge graph, this allows for the utilization of pathogen... This study investigates the biological mechanisms of co-amplification of pathogen and drug-resistant genes, constructing a co-state evolution model. This model introduces an inhibition intensity factor as a latent variable to characterize the degree of PCR amplification suppression. Based on the state prior estimation determined by the co-state evolution model and combined with state observations from sequencing data, a Kalman filter algorithm is used for residual updates. Iterative inverse estimation is then employed to determine the optimal inhibition intensity factor for the drug-resistant gene sequence, thereby decoupling the interference term "inhibition factor" from the chaotic sequencing data. By utilizing the optimal inhibition intensity factor, the overlap rate between pathogen and drug-resistant gene sequences is corrected, restoring the true biological signal. This allows for the identification of drug-resistant gene sequences and their abundance in all pathogen gene sequences, effectively improving detection accuracy.

[0039] In one possible implementation, such as Figure 2 As shown, step S100 involves constructing a composite knowledge graph about pathogens and drug resistance genes, including: Step S101: Treat each pathogen and each drug resistance gene sequence as a node in the composite knowledge graph.

[0040] Step S102: Based on the genomic sequence information and drug resistance gene sequence information of different pathogens, when the drug resistance gene sequence exists in the genomic sequence of the pathogen, a directed edge is constructed between the nodes of the corresponding pathogen and the drug resistance gene sequence, and the edge weight of the directed edge is determined. The directed edge points from the drug resistance gene sequence to the pathogen.

[0041] Specifically, in the construction process of the aforementioned composite knowledge graph, each pathogen species and each drug-resistant gene sequence are treated as nodes. When a drug-resistant gene sequence is confirmed to exist in the genome of a pathogen, a directed edge is established between the corresponding pathogen node and the drug-resistant gene sequence node. The direction of the directed edge is "drug-resistant gene sequence node → pathogen node". This directed edge represents the subordinate relationship between the drug-resistant gene sequence and the pathogen, that is, the drug-resistant gene sequence belongs to one of the pathogen's genomes. Simultaneously, the edge weights of the directed edges are determined, and these weights are used to reflect the tightness of the evolutionary association between the drug-resistant gene and the pathogen.

[0042] Furthermore, in one possible implementation, the edge weights of directed edges can be determined as follows: First, for the drug-resistant gene sequence and pathogen corresponding to the directed edge, determine the copy number of the drug-resistant gene sequence in the pathogen's genome sequence. This copy number refers to how many copies of the drug-resistant gene are contained within a single pathogen (such as a bacterial cell).

[0043] Secondly, for the drug-resistant gene sequences and pathogens corresponding to directed edges, the insertion positions and detection frequencies of the drug-resistant gene sequences in the genome sequences of different strains of the pathogen are determined. Specifically, for each directed edge corresponding to a drug-resistant gene sequence node and a pathogen node, the genome sequences of multiple strains of the pathogen species are statistically analyzed, and the insertion positions of the drug-resistant gene sequence in different strains are compared (the insertion position can be characterized by the sequence number of the inserted gene sequence). A more stable insertion position indicates that the drug-resistant gene sequence is more conserved in the genome of the corresponding pathogen. Simultaneously, the proportion of occurrence of the drug-resistant gene sequence in the sequenced strains of the pathogen is calculated to obtain the detection frequency of the drug-resistant gene sequence in the pathogen strains. This detection frequency refers to the proportion of strains in the publicly available and sequenced genomes of all strains of the pathogen that have detected this drug-resistant gene sequence.

[0044] Finally, based on the copy number, the stability of the insertion position, and the detection frequency, the edge weights corresponding to the directed edges are determined. Specifically, for each directed edge corresponding to the drug-resistant gene sequence node and the pathogen node, the copy number is normalized to the range [0,1] using a minimum-maximum normalization function, resulting in a normalized copy number value, denoted as C. The standard deviation of the position coordinates of all insertion positions is calculated. This standard deviation reflects the degree of conservation of the drug-resistant gene position; a smaller standard deviation indicates greater conservation, while a larger standard deviation indicates that the drug-resistant gene may have been acquired through horizontal gene transfer, resulting in greater positional variability. This standard deviation is also normalized to the range [0,1] using a minimum-maximum normalization function, resulting in a normalized standard deviation value, denoted as D. The detection frequency is then normalized to the range [0,1] using a minimum-maximum normalization function, resulting in a normalized detection frequency value, denoted as P. Furthermore, based on the normalized copy number value C, the normalized standard deviation value D, and the normalized detection frequency value P, the edge weights corresponding to the directed edges are calculated. ;in, , and These represent the reference weights for copy number, standard deviation, and detection frequency, respectively. Their values ​​are determined based on the degree of influence of copy number, standard deviation, and detection frequency on the relationship between the drug-resistant gene sequence and the pathogen; the higher the degree of influence, the higher the reference weight. For example, in drug resistance surveillance scenarios, position conservation reflects whether a gene is fixed in the genome. The smaller the standard deviation (higher conservation), the greater the likelihood that the drug resistance gene is an inherent characteristic of the pathogen, rather than a temporary horizontal transfer. Increasing its weight helps to screen for more stable associations and reduce interference from plasmid exchange. This can be achieved by setting... , , The weight of a directed edge. The larger the value, the stronger and more stable the association between the drug resistance gene and the pathogen.

[0045] Furthermore, since co-occurrence relationships also exist among drug-resistant gene sequences, when multiple drug-resistant gene sequences frequently appear simultaneously on the same mobile genetic element, undirected edges are also established between the nodes corresponding to these drug-resistant gene sequences, and the co-occurrence frequency of these drug-resistant gene sequences in the known pathogen genome is used as the edge weight. Thus, a composite knowledge graph is constructed, containing pathogen nodes, drug-resistant gene sequence nodes, and descriptions of the relationships between them (directed graph) and the co-occurrence relationships among drug-resistant genes (undirected graph).

[0046] Based on the above technical solution, by transforming the biological pathogen-drug resistance gene attribution relationship into a quantified topological constraint weight, differentiated co-evolutionary stiffness is provided for subsequent algorithms. This effectively solves the ambiguity of unclear origin of drug resistance genes in mixed infections and prevents overcorrection of unstable associated genes, thereby significantly improving the accuracy of complex sample detection.

[0047] In one possible implementation, such as Figure 3 As shown, step S400 analyzes the sequence inhibition sensitivity of the pathogen gene sequence and drug resistance gene sequence of the target pathogen, including: Step S401: Determine the core gene sequence in the pathogen gene sequence of the target pathogen.

[0048] Specifically, for all pathogen gene sequences obtained from sequencing in the clinical samples to be tested (pathogens, as complete genomes, contain tens to thousands of gene sequences), gene sequences with a length of 500bp to 5000bp are retained (this is the mainstream length range for target genes in clinical PCR amplification; sequences that are too short lack representativeness, while sequences that are too long can lead to low amplification efficiency). For these retained gene sequences, sequences with sequencing coverage ≥90% are further retained to avoid deviations in GC content and free energy calculations due to incomplete sequences. Then, sequence similarity clustering (CD-HIT algorithm, similarity threshold 95%) is used to obtain several sequence clusters, with only one representative gene sequence retained in each cluster to avoid feature duplication caused by highly homologous genes. Finally, approximately 10 core gene sequences are selected from the representative gene sequences according to their copy number from high to low.

[0049] Step S402: For the core gene sequence and the drug resistance gene sequence, determine the GC base content and the minimum free energy of the secondary structure of the gene sequence.

[0050] Specifically, different DNA sequences have varying sensitivities to inhibitors. Sequences with high GC content and stable secondary structures are more likely to have their polymerase elongation blocked by inhibitors. Therefore, the GC content and minimum free energy of secondary structure are calculated for each core gene sequence, as well as for drug resistance gene sequences. Sequences with more GC bases are more susceptible to interference, and those with lower free energy are more stable and more easily targeted by inhibitors.

[0051] Step S403: Based on the GC base content and minimum free energy of secondary structure of all core gene sequences, determine the inhibition sensitivity index of the target pathogen, and based on the GC base content and minimum free energy of secondary structure of the drug resistance gene sequence, determine the inhibition sensitivity index of the drug resistance gene sequence. This inhibition sensitivity index is used to reflect the ease with which the target pathogen or drug resistance gene sequence is inhibited by the inhibitor.

[0052] Specifically, for each core gene sequence, the absolute value of the minimum free energy of the secondary structure and the GC base content corresponding to the sequence are normalized to the range of [0,1] using a maximum-minimum normalization function. The average of the two normalized values ​​is then used as the inhibition sensitivity index for each core gene sequence. The average inhibition sensitivity index of all core gene sequences is calculated and used as the inhibition sensitivity index for the corresponding target pathogen. Similarly, for drug-resistant gene sequences, the absolute value of the minimum free energy of the secondary structure and the GC base content corresponding to the sequence are also normalized to the range of [0,1] using a maximum-minimum normalization function. The average of the two normalized values ​​is then used as the inhibition sensitivity index for that drug-resistant gene sequence.

[0053] Based on the above technical solution, by determining the core gene sequence in the pathogen gene sequence of the target pathogen, and identifying the GC base content and minimum free energy of secondary structure in the core gene sequence and drug resistance gene sequence, the inhibitory sensitivity of the target pathogen and drug resistance gene sequence can be accurately quantified.

[0054] In one possible implementation, such as Figure 4 As shown, step S400 involves constructing a cooperative state evolution model, including: Step S411: Based on the inhibition sensitivity index, the baseline inhibition coefficient of the clinical sample to be tested is corrected to obtain the initial inhibition strength factor of the target pathogen and drug resistance gene sequence.

[0055] Specifically, sequences with high inhibition sensitivity (GC-rich and with stable secondary structures) should have a smaller initial value for their inhibition strength factor but a faster decay rate, while sequences with low inhibition sensitivity should have a larger initial value but a slower decay rate. Baseline inhibition coefficients should be obtained from different clinical samples (e.g., blood, sputum). The baseline suppression coefficient It can be obtained by calculating the average inhibition level of historical sample data, such as 0.8 for blood samples and 0.6 for sputum / pus.

[0056] The initial inhibition strength factor of the target pathogen is calculated based on the inhibition sensitivity index of the target pathogen and the baseline inhibition coefficient of the clinical sample to be tested. ;in: Indicates the loop duration; This represents the balance coefficient, used to balance the weights of baseline suppression and sequence-specific suppression, and is set empirically. ; This represents the inhibition sensitivity index of the target pathogen.

[0057] Similarly, based on the inhibition sensitivity index of the drug resistance gene sequence and the baseline inhibition coefficient of the clinical sample to be tested, the initial inhibition strength factor of the drug resistance gene sequence is calculated. ;in: An index representing the inhibition sensitivity of drug resistance gene sequences.

[0058] Step S412: Based on the initial inhibition strength factor, construct the basic update equation of the inhibition strength factor corresponding to the target pathogen and drug resistance gene sequence during PCR amplification.

[0059] The inhibition strength factor is a latent variable in the state space that changes dynamically with the PCR amplification cycle: it takes a small value when the inhibitor concentration is high in the early stage of amplification (less than 1 indicates inhibition, which will lead to the observed value being lower than the true value), and it tends to be close to 1 as the inhibitor is diluted during amplification (indicating no inhibition); each target pathogen and each drug resistance gene sequence has its own specific inhibition strength factor.

[0060] Specifically, based on the initial inhibition strength factors of the target pathogen and drug resistance gene sequences, the basic update equations for the corresponding inhibition strength factors during PCR amplification of the target pathogen and drug resistance gene sequences are constructed: In the formula: Indicates the target pathogen The updated value of the inhibition strength factor as PCR amplification proceeds to cycle number t; Indicating drug resistance gene sequences The updated value of the inhibition strength factor as PCR amplification proceeds to cycle number t; Indicates the target pathogen The initial inhibition strength factor; Indicating drug resistance gene sequences The initial inhibition strength factor; This represents the attenuation coefficient of the clinical sample being tested, and is set according to the type of inhibitor in the sample, such as the inhibitor corresponding to heme. The value is 0.05 per cycle. When multiple inhibitors coexist, the inhibitor that is the most difficult to eliminate (the one with the slowest decay) will determine the duration of the inhibition tail of the entire reaction. Therefore, the minimum value among the decay coefficients of all inhibitors is selected as the decay coefficient of the clinical sample to be tested.

[0061] In the above formula, As the loop progresses, it approaches 0, causing the first term to... or The initial inhibitory effect gradually disappears, and the second term... or (The compensation term for the unsuppressed state) corresponds to the gradual approximation of or This demonstrates the actual law that the inhibitory effect decays with each cycle.

[0062] Step S413: Based on the composite knowledge graph, identify each candidate pathogen in the clinical sample to be tested that is associated with the drug resistance gene sequence, and based on the difference in the inhibition intensity factor between the drug resistance gene sequence and each candidate pathogen during PCR amplification, identify the synergistic driving term in the PCR amplification process.

[0063] Directed edges between pathogens and drug-resistant genes are extracted from the composite knowledge graph to identify all pathogen types associated with drug-resistant gene sequences (i.e., directed edges exist between pathogen nodes and drug-resistant gene sequence nodes). Then, all pathogens present in the clinical sample to be tested are identified as candidate pathogens constituting the drug-resistant gene sequence. It should be understood that when the number of candidate pathogens is large, only the top K (e.g., Top 5) candidate pathogens in terms of abundance (Reads Count) can be selected for subsequent co-evolutionary calculations.

[0064] Standard Kalman filtering assumes independent state evolution, but pathogens and their drug-resistant genes (the same DNA molecule) should be subject to equal suppression; a large difference in suppression intensity factors between the two is illogical. Therefore, this embodiment analyzes the differences in suppression intensity factors between drug-resistant gene sequences and candidate pathogens during PCR amplification to define a co-evolutionary driving term for both. This term replaces the "independent state transition equation" of standard Kalman filtering, ensuring that the change in the suppression intensity factor of the drug-resistant gene is driven by both its own sequence characteristics and the associated pathogen factor, thus guaranteeing synergy between the two.

[0065] Furthermore, in one possible implementation, the co-driving term in the PCR amplification process is determined by: determining the difference in inhibition intensity factors between each candidate pathogen and the drug-resistant gene sequence during the PCR amplification process, and obtaining the inhibition intensity factor difference value; using the edge weights of the directed edges between the nodes corresponding to each candidate pathogen and the drug-resistant gene sequence in the composite knowledge graph, the inhibition intensity factor difference value is weighted and accumulated to obtain the co-driving term in the PCR amplification process.

[0066] Step S414: Based on the rate of change of the basic update equation of the inhibition intensity factor and the co-driving term, determine the rate of change of the inhibition intensity factor corresponding to the drug resistance gene sequence during PCR amplification.

[0067] Specifically, the rate of change of the inhibition intensity factor corresponding to the drug resistance gene sequence during PCR amplification is equal to the sum of the exponential decay term and the co-driving term of its own inhibition intensity factor. At this point, the rate of change of the inhibition intensity factor of the drug resistance gene sequence is: In the formula: Indicating drug resistance gene sequences The rate of change of the inhibition intensity factor; Indicating drug resistance gene sequences The updated value of the inhibition strength factor as PCR amplification proceeds to cycle number t; Indicating drug resistance gene sequences Nodes and candidate pathogens The edge weights of directed edges between nodes; Indicates candidate pathogens The update value of the inhibition strength factor as PCR amplification proceeds to cycle number t can be determined based on the basic update equation of the inhibition strength factor corresponding to the target pathogen during PCR amplification. Indicating drug resistance gene sequences The updated value of the inhibition strength factor as PCR amplification proceeds to cycle number t; This indicates the total number of candidate pathogens.

[0068] In the above formula, the first term is the derivative of the dynamic update equation of the inhibitory strength factor of the drug resistance gene sequence, which represents the basic update rate of the inhibitory strength of the drug resistance gene; the second term, the co-driving term, is the weighted sum of the differences between the inhibitory strength factors of all candidate pathogens and the inhibitory strength factor of the drug resistance gene sequence, i.e. ,in This constitutes negative feedback, when When the collaborative driving term is negative, it makes Decrease Growth slows or declines, and vice versa. Growth, ultimately achieving the goals of various candidate pathogens Inhibition strength factor With drug resistance gene sequence Inhibition strength factor synchronous.

[0069] Based on the above technical solution, a co-evolution mechanism is achieved by constructing a co-driving term in the PCR amplification process. This not only preserves the decay law of the inhibitory effect of the drug resistance gene sequence itself, but also transforms the topological constraints of the knowledge graph into dynamic coupling, so that the inhibitory intensity factors of pathogen-drug resistance gene pairs from the same DNA molecule automatically tend to be consistent.

[0070] Furthermore, using the Kalman filter algorithm, based on the state prior estimate determined by the cooperative state evolution model, and combined with the state observation values ​​for residual update, the optimal inhibition intensity factor of the drug resistance gene sequence is determined through iterative inverse estimation.

[0071] In one possible implementation, such as Figure 5 As shown, in step S400, the Kalman filter algorithm is used to estimate the state prior based on the cooperative state evolution model, and the residual is updated by combining the state observations. The optimal inhibitory strength factor for the drug-resistant gene sequence is determined through iterative inverse estimation, including: Step S421: Obtain the observed values ​​of the inhibitory factors and abundance of the target pathogen corresponding to each PCR amplification cycle in the Kalman filtering process, and construct the state observation values ​​corresponding to each PCR amplification cycle in the Kalman filtering process.

[0072] Specifically, in each PCR amplification cycle, two types of information need to be tracked simultaneously: The first category is the target detection quantity: the measurement abundance of the target pathogen, that is, the number of copies measured. The abundance of drug resistance genes is linked to the abundance of pathogens. The more pathogen genes there are, the more drug resistance genes there will be. In the case of uncertainty about drug resistance genes, the abundance of drug resistance genes can be indirectly observed by directly measuring the amplification abundance of pathogens.

[0073] Among them, the theoretical amplification abundance (copy number) of the target pathogen under experimental conditions, i.e., in an ideal environment without amplification inhibitors, can be obtained by traditional bacterial culture and biochemical detection methods (which take several days and are for preparing experimental data).

[0074] The second category is interference correction factor: the real-time inhibition intensity factor of the target pathogen, which is used to quantify the degree to which the target pathogen is inhibited during amplification.

[0075] The difference between the theoretical amplification abundance of the target pathogen and the measured abundance obtained above can be used to determine whether the amplification process is inhibited. The theoretical amplification abundance refers to the amount that should be reached under ideal conditions (efficiency 100%) without inhibition, after undergoing different PCR amplification cycles. At this time, the real-time inhibition intensity factor of the target pathogen = difference / theoretical amplification abundance. Completely uninhibited is recorded as zero, and completely unable to amplify is recorded as 1.

[0076] The amplification abundance of the target pathogen and the real-time inhibition intensity factor of the target pathogen obtained above during different PCR amplification cycles were used as the observed values ​​of the current cycle state corresponding to different PCR amplification cycles.

[0077] It should be understood that, since the amplification process cannot be interrupted, in one specific example, a limited number of discrete termination experiments (a common experimental method in this field: artificially terminating the PCR reaction at different cycle numbers t and sequencing to obtain the abundance of that cycle) can be performed beforehand to obtain a small number of discrete true observations. Then, interpolation fitting is used to generate the observations for each cycle. In another specific example, based on the abundance in the final sequencing results and the theoretical amplification efficiency E, the theoretical amplification efficiency E is obtained using the formula... t is the cycle number. Let be the abundance at cycle number t. Using the initial template quantity, the virtual observation sequence of abundance is obtained by reverse reconstruction, and then the observation value of each cycle is obtained.

[0078] Meanwhile, two types of basic noise are introduced: process noise (describing the random fluctuations of PCR amplification) and observation noise (reflecting the inherent errors of the sequencing platform), both of which conform to Gaussian distribution. The values ​​of these two types of noise can be determined by statistical analysis of historical experimental data. This is a basic component of Kalman filtering, which will not be elaborated here.

[0079] Step S422: Based on the initial inhibition strength factor, the basic update equation of the inhibition strength factor, and the rate of change of the inhibition strength factor, determine the predicted values ​​of the inhibition factors for the target pathogen and drug resistance gene sequences corresponding to each PCR amplification cycle during the Kalman filtering process.

[0080] Specifically, state prediction is based on co-evolution, which includes repressor prediction and abundance prediction. When predicting repressors, the initial repressor strength factor of the target pathogen in the constructed co-evolutionary state model, combined with the basic update equation of the repressor strength factor corresponding to the target pathogen during PCR amplification, can determine the predicted repressor value of the target pathogen for each PCR amplification cycle in the Kalman filter process. Simultaneously, using the initial repressor strength factor of the drug-resistant gene sequence, combined with the basic update equation and the rate of change of the repressor strength factor corresponding to the drug-resistant gene sequence during PCR amplification, the predicted repressor value of the drug-resistant gene sequence for each PCR amplification cycle in the Kalman filter process can be determined. This ensures that the repressor changes of drug-resistant genes are driven by both their own sequence characteristics and associated pathogens, guaranteeing that the repressive effects of sequences originating from the same DNA molecule tend to be consistent.

[0081] Step S423: Based on the predicted value of the inhibitory factor of the target pathogen, the PCR exponential growth model of the target pathogen is modified to determine the predicted abundance value of the target pathogen for each PCR amplification cycle in the Kalman filtering process.

[0082] Specifically, in the abundance prediction within state prediction, the existing PCR exponential growth model... ,in, The initial template amount is given, E is the amplification efficiency (generally assumed to be 1), and t is the cycle number. The abundance at cycle number t. In this existing PCR exponential growth model... The model incorporates the predicted value of the target pathogen's inhibitory factor, X2, to regulate amplification efficiency. Stronger inhibition results in slower abundance growth, leading to a modified PCR exponential growth model. Using this modified model, the predicted abundance of the target pathogen in different PCR amplification cycles can be determined.

[0083] Step S424: Based on the predicted values ​​of the inhibitory factors and abundance of the target pathogen, determine the state prior estimate corresponding to each PCR amplification cycle during the Kalman filtering process.

[0084] Specifically, the predicted values ​​of the inhibitory factors and abundance of the target pathogen and drug-resistant gene sequences are integrated and packaged into a vector, which is then used as the state prior estimate for the current loop.

[0085] Step S425: Determine the observation residuals based on the state prior estimate and state observations, and use the predicted value of the inhibitory factor of the drug resistance gene sequence as the state update object. Perform Kalman filter loop iteration with the PCR amplification cycle to update the residuals, and finally obtain the optimal inhibitory strength factor corresponding to the drug resistance gene sequence.

[0086] Specifically, based on the aforementioned prior state estimates and the observed and predicted values ​​of the target pathogen's inhibitory factor and abundance from the state observations, observational residuals can be obtained to indirectly assess the predicted values ​​of the inhibitory factor for the drug-resistant gene sequence. These predicted inhibitory factors are then used as the state value, i.e., the object of state update. Kalman gain is used to fuse the two values ​​to correct the residuals. If the prediction uncertainty is high, more observational data is accepted; if the observation noise is high, more prediction results are retained. Finally, the optimal value of the inhibitory factor for the drug-resistant gene sequence in the current PCR amplification cycle is obtained, allowing for the inference of its true abundance. This process completes a compensation cycle of "inhibition interference stripping - true abundance restoration," gradually improving the accuracy of state estimation as PCR amplification iterates. It should be understood that due to the binding relationship between the pathogen and the drug-resistant gene, the observational residuals are essentially an indirect projection of state error. They do not directly modify the state but rather complete the transmission from observational bias to state update through three steps: "bias decoding → weight allocation → incremental correction," thus driving the update of the drug-resistant gene inhibitory factor. PCR amplification typically involves 35-40 cycles, with filtering iterations ending as amplification terminates, ultimately outputting the optimal inhibition strength factor for the drug resistance gene sequence.

[0087] Based on the above technical solution, by utilizing the predicted values ​​of inhibitory factors and abundance of target pathogens and drug-resistant gene sequences, a priori estimates of the state of different PCR amplification cycles are constructed. Furthermore, by utilizing the observed values ​​of inhibitory factors and abundance of target pathogens and drug-resistant gene sequences, state observations of different PCR amplification cycles are constructed. Thus, Kalman filtering is used to realize the observation update process, ultimately achieving accurate state estimation and obtaining the optimal inhibitory strength factor of the drug-resistant gene sequence.

[0088] In one possible implementation, step S500 uses the optimal suppression intensity factor to correct the overlap rate, including: determining the ratio of the overlap rate to the optimal suppression intensity factor; and determining the product of the ratio and the probe capture efficiency coefficient as the overlap rate correction value.

[0089] Specifically, the overlap rate between pathogen gene sequences and drug-resistant gene sequences is corrected using the optimal inhibition strength factor of the drug resistance gene sequence, and the overlap rate correction value is determined by the following formula: In the formula: This represents the overlap correction value between pathogen gene sequences and drug resistance gene sequences; This indicates the overlap rate between pathogen gene sequences and drug resistance gene sequences; The optimal inhibitory strength factor representing the drug resistance gene sequence; This represents the probe capture efficiency coefficient, which can be determined through calibration using standard sequencing data (valued as the ratio of the known concentration of the standard to the concentration detected by sequencing). It should be understood that drug resistance genes can be inhibited to some extent by inhibitors. The value of is usually not 0, when When the value of is 0, it means that the drug resistance gene will not be inhibited by the inhibitor. In this case, it can be directly expressed by the formula. To obtain the overlap rate correction value.

[0090] Based on the above technical solution, by taking the optimal inhibition intensity factor of the drug resistance gene sequence as the core and combining it with the probe binding efficiency characteristics of targeted sequencing, a calibration model is constructed. This can eliminate the dual interference of inhibitor and probe bias, and effectively restore the true amplification level underestimated by the inhibitor.

[0091] In one possible implementation, such as Figure 6 As shown, step S500, based on the overlap correction value, determines the drug resistance gene sequences and their abundance in all pathogen gene sequences, including: Step S501: Remove pathogen gene sequences whose overlap correction value is less than or equal to the set overlap threshold to obtain the remaining pathogen gene sequences after removal.

[0092] Specifically, based on the sequence alignment results, reads that completely match the target gene region are retained, that is, pathogen gene sequences with an overlap correction value greater than 95% are retained, while interfering sequences that are homologous to the human genome or non-target pathogens are removed, thus obtaining the remaining pathogen gene sequences after removal.

[0093] Step S502: Perform duplicate sequence removal on the remaining pathogen gene sequences after deletion, and merge reads from the same original deoxyribonucleic acid template to form a UMI family. Calculate the sequence consistency of pathogen gene sequences in the UMI family, and retain UMI families with sequence consistency greater than the set consistency threshold.

[0094] Specifically, for the pathogen gene sequences remaining after deletion, duplicate sequences are removed using molecular tags (UMIs), and reads from the same original DNA template are merged to form UMI families. Then, the percentage of identical bases in all pathogen gene sequences within each UMI family is calculated to obtain the sequence identity of each UMI family, and families with a sequence identity greater than 95% are retained to exclude amplification errors.

[0095] Step S503: Remove pathogen gene sequences containing chimeric sequences from the retained UMI family to finally obtain the drug resistance gene sequences and their abundance in all pathogen gene sequences.

[0096] Specifically, within the retained UMI family, read alignment breakpoints are detected, reads containing chimeric sequences (cross-gene fragment fusion) are removed, and finally, the verified drug resistance genes are obtained. The abundance of these "verified" genes is recalculated, and the final detection results are output.

[0097] Based on the above technical solution, by sequentially performing noise reduction operations such as small overlap rate removal, repetitive sequence deduplication, small sequence consistency removal, and chimeric sequence removal on pathogen gene sequences, the drug resistance gene sequences and their abundance in all pathogen gene sequences can be accurately determined, effectively improving the detection accuracy.

[0098] Based on the same inventive concept, embodiments of the present invention also provide a targeted sequencing-based pathogen-drug resistance gene simultaneous detection system, such as... Figure 7 As shown, the system includes: a memory, a processor, and computer program code stored in the memory and running on the processor, wherein when the processor executes the computer program code, the system is able to perform any of the aforementioned targeted sequencing pathogen-drug resistance gene simultaneous detection methods.

[0099] In this embodiment of the invention, the system can be divided into functional modules according to the above method example. For example, each module can correspond to a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0100] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for simultaneous detection of pathogen-drug resistance genes by targeted sequencing, characterized in that, The method comprises the following steps: constructing a complex knowledge graph about pathogens and drug-resistant genes based on genomic sequence information and drug-resistant gene sequence information of different pathogens; identifying a target pathogen present in a to-be-detected clinical sample and a pathogen gene sequence of the target pathogen present in the to-be-detected clinical sample based on the genomic sequence information of the different pathogens; determining a set of drug-resistant gene sequences that the target pathogen may have based on the complex knowledge graph, and comparing the pathogen gene sequence with the drug-resistant gene sequences in the set to obtain an overlap rate between the pathogen gene sequence and the drug-resistant gene sequences; analyzing sequence inhibition sensitivity of the pathogen gene sequence and the drug-resistant gene sequence of the target pathogen, and constructing a collaborative state evolution model in combination with the complex knowledge graph, wherein the collaborative state evolution model characterizes a degree of PCR amplification inhibition by introducing an inhibition intensity factor as a hidden variable, and determines an optimal inhibition intensity factor of the drug-resistant gene sequence by iterative inverse estimation based on state prior estimation determined by the collaborative state evolution model and in combination with state observation values for residual update; correcting the overlap rate by using the optimal inhibition intensity factor, and determining drug-resistant gene sequences and gene abundances in all the pathogen gene sequences based on an overlap rate correction value.

2. The method according to claim 1, wherein, The method for constructing a complex knowledge graph about pathogens and drug-resistant genes comprises: regarding each pathogen and each drug-resistant gene sequence as a node in the complex knowledge graph; based on genomic sequence information and drug-resistant gene sequence information of different pathogens, when a drug-resistant gene sequence exists in a genomic sequence of a pathogen, constructing a directed edge between nodes corresponding to the pathogen and the drug-resistant gene sequence, and determining an edge weight of the directed edge, wherein the directed edge is directed from the drug-resistant gene sequence to the pathogen, and the edge weight is used to reflect closeness of evolutionary association between the drug-resistant gene sequence and the pathogen.

3. The method according to claim 2, wherein the pathogen-drug resistance gene is selected from the group consisting of the genes listed in Table 1. Determining the edge weight of the directed edge comprises: determining a copy number of the drug-resistant gene sequence in the genomic sequence of the pathogen for the drug-resistant gene sequence and the pathogen corresponding to the directed edge; determining an insertion position and a detection frequency of the drug-resistant gene sequence in different strain genomic sequences of the pathogen for the drug-resistant gene sequence and the pathogen corresponding to the directed edge; determining the edge weight corresponding to the directed edge based on the copy number, stability of the insertion position, and the detection frequency.

4. The method according to claim 1, wherein, Analyzing sequence inhibition sensitivity of the pathogen gene sequence and the drug-resistant gene sequence of the target pathogen comprises: determining a core gene sequence in the pathogen gene sequence of the target pathogen; determining GC base content and secondary structure minimum free energy of gene sequences for the core gene sequence and the drug-resistant gene sequence; Determine the inhibition susceptibility index of the target pathogen based on the GC base content and the minimum free energy of the secondary structure of all the core gene sequences, and determine the inhibition susceptibility index of the drug-resistant gene sequence based on the GC base content and the minimum free energy of the secondary structure of the drug-resistant gene sequence, which reflects the ease of inhibition of the target pathogen or the drug-resistant gene sequence by an inhibitor.

5. The method according to claim 4, wherein the pathogen-drug resistance gene is selected from the group consisting of the genes listed in Table 1. Construct a synergistic state evolution model, including: Based on the inhibition susceptibility index, correct the baseline inhibition coefficient of the to-be-detected clinical sample to obtain the initial inhibition intensity factor of the target pathogen and the drug-resistant gene sequence; Based on the initial inhibition intensity factor, construct the basic update equation of the corresponding inhibition intensity factor of the target pathogen and the drug-resistant gene sequence in the PCR amplification process; Based on the complex knowledge graph, determine each candidate pathogen associated with the drug-resistant gene sequence present in the to-be-detected clinical sample, and determine the synergistic driving term in the PCR amplification process based on the difference in inhibition intensity factor between the drug-resistant gene sequence and each candidate pathogen in the PCR amplification process. Determine the change rate of the inhibition intensity factor of the drug-resistant gene sequence in the PCR amplification process based on the change rate of the basic update equation of the inhibition intensity factor and the synergistic driving term.

6. The method according to claim 5, wherein the pathogen-drug resistance gene is selected from the group consisting of the genes listed in Table 1. Determine the synergistic driving term in the PCR amplification process, including: Determine the difference value of the inhibition intensity factor of the candidate pathogen and the drug-resistant gene sequence in the PCR amplification process to obtain the inhibition intensity factor difference value; Use the edge weight of the directed edge between the nodes corresponding to the candidate pathogen and the drug-resistant gene sequence in the complex knowledge graph to weight and accumulate the inhibition intensity factor difference value, thereby obtaining the synergistic driving term in the PCR amplification process.

7. The method according to claim 5, wherein the pathogen-drug resistance gene is selected from the group consisting of the genes listed in Table 1. Using Kalman filtering algorithm, based on the state priori estimation determined by the synergistic state evolution model, and combining with the state observation value to update the residual error, through iterative inverse estimation to determine the optimal inhibition intensity factor of the drug-resistant gene sequence, including: Obtain the inhibition factor observation value and abundance observation value of the target pathogen corresponding to each PCR amplification cycle in the Kalman filtering process, and form the state observation value corresponding to each PCR amplification cycle in the Kalman filtering process; Based on the initial inhibition intensity factor, the basic update equation of the inhibition intensity factor, and the change rate of the inhibition intensity factor, determine the inhibition factor prediction value of the target pathogen and the drug-resistant gene sequence corresponding to each PCR amplification cycle in the Kalman filtering process; Based on the inhibition factor prediction value of the target pathogen, correct the PCR exponential growth model of the target pathogen to determine the abundance prediction value of the target pathogen corresponding to each PCR amplification cycle in the Kalman filtering process; Based on the inhibition factor prediction value and the abundance prediction value of the target pathogen, determine the state priori estimation corresponding to each PCR amplification cycle in the Kalman filtering process; Based on the state prior estimation and the state observation value, an observation residual is determined, and a predicted value of an inhibition factor of the drug resistance gene sequence is taken as a state update object, and a Kalman filtering cycle iteration is performed along with a PCR amplification cycle to update the residual, and finally an optimal inhibition intensity factor corresponding to the drug resistance gene sequence is obtained.

8. The method according to claim 1, wherein, Using the optimal inhibition intensity factor, the overlap rate is corrected, including: determining a ratio of the overlap rate to the optimal inhibition intensity factor; determining a product of the ratio and a probe capture efficiency coefficient as an overlap rate correction value.

9. The method according to claim 1, wherein the method is characterized by, Based on the overlap rate correction value, the drug resistance gene sequence and its gene abundance in all the pathogen gene sequences are determined, including: eliminating the pathogen gene sequence with an overlap rate correction value less than or equal to a set overlap rate threshold value to obtain a remaining pathogen gene sequence after elimination; performing repeat sequence deduplication on the remaining pathogen gene sequence after elimination, and merging reads from the same original deoxyribonucleic acid template to form a UMI family, calculating sequence consistency of the pathogen gene sequence in the UMI family, and retaining the UMI family with a sequence consistency greater than a set consistency threshold value; eliminating the pathogen gene sequence containing a chimeric sequence in the retained UMI family, and finally obtaining the drug resistance gene sequence and its gene abundance in all the pathogen gene sequences.

10. A simultaneous detection system of pathogen-drug resistance gene targeted sequencing, characterized in that, The computer program product comprises a memory, a processor, and executable computer program code stored in the memory and executable on the processor, and the processor executes the computer program code to perform the pathogen-drug resistance gene synchronous detection method of targeted sequencing according to any one of claims 1-9.

Citation Information

Patent Citations

  • Non-therapeutic-purpose metatranscriptome-based pathogen detection method

    CN116949154A

  • Clinical microbial infection intelligent diagnosis system and method based on multi-omics data fusion

    CN120913824A

  • Thyroid cancer analysis method and device based on real-time fluorescent quantitative PCR

    CN121320499A