Pathogen drug-resistant single base mutation detection method, device, equipment and storage medium

By comparing the sequencing data of pathogen samples with the reference genome and numerically processing strongly correlated sites, combined with a single-base mutation detection model, the shortcomings of single-base mutation site identification in second-generation metagenomic sequencing were solved, and accurate detection of sites not covered by sequencing was achieved.

CN116959567BActive Publication Date: 2026-06-23ZHUHAI CARBON CHINA TESTING TECH CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHUHAI CARBON CHINA TESTING TECH CO LTD
Filing Date
2023-07-17
Publication Date
2026-06-23

Smart Images

  • Figure CN116959567B_ABST
    Figure CN116959567B_ABST
Patent Text Reader

Abstract

The application discloses a pathogen drug-resistant single-base mutation detection method, device, equipment and storage medium, and the method comprises the following steps: sequencing a to-be-detected pathogen sample, and obtaining sequencing data under a first preset sequencing depth; comparing the sequencing data with a reference genome corresponding to the to-be-detected pathogen sample, and confirming a to-be-detected site which is not covered by sequencing; confirming all first strongly-correlated sites corresponding to the to-be-detected site based on a pre-constructed strongly-correlated site corresponding relation table; numerically representing mutation information of all the first strongly-correlated sites based on a preset numerical representation mode, and obtaining a numerical result; inputting the numerical result into a pre-trained single-base mutation detection model for prediction, and confirming whether the to-be-detected site is a positive site or a negative site. Through the strong correlation between sites, the mutation information of the corresponding strongly-correlated sites is used to predict whether the to-be-detected site is mutated or not, and the detection accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of gene detection technology, and in particular to a method, apparatus, equipment and storage medium for detecting single-base mutations in drug-resistant pathogens. Background Technology

[0002] Antimicrobial resistance refers to the ability of pathogenic microorganisms to resist the effects of drugs in order to survive. Once resistance occurs, it means that the pathogen cannot be completely killed and its proliferation cannot be inhibited. Diseases caused by drug-resistant pathogens are difficult or even impossible to cure. Children, especially newborns, are more severely affected by drug-resistant pathogens due to their weaker immune systems. Researching, statistically analyzing, and predicting drug-resistant pathogens can help us better understand the mechanisms of drug resistance and thus guide us in better combating the harm caused by drug-resistant pathogens to humans.

[0003] Currently, the main methods for detecting bacterial drug resistance include: detection of bacterial resistance phenotypes (including resistance screening tests, breakpoint sensitivity tests, etc.), β-lactamase detection, detection of specific drug-resistant bacteria, and detection of drug resistance genes. The most intuitive method is in vitro drug detection, but its drawback is the long culture time. As research into drug resistance mechanisms deepens, molecular biology methods for detecting bacterial resistance are increasingly accepted clinically. Common methods for detecting bacterial drug resistance genes include: polymerase chain reaction (PCR), PCR-restriction fragment length polymorphism analysis (PCR-RFLP), PCR-single strand conformation polymorphism analysis (PCR-SSCP), microarray technology, and DNA sequencing. Traditional gene identification methods are PCR and DNA sequencing. Traditional PCR has a long detection cycle, typically requiring 3-4 hours; Sanger sequencing requires high primer specificity, is expensive, and has complex subsequent data analysis, making it unsuitable for widespread clinical use. Furthermore, with the development of sequencing technology, algorithms for pathogen drug resistance analysis based on second-generation and third-generation sequencing have also been developed. Third-generation sequencing is currently unsuitable for large-scale sample analysis due to its high cost. Second-generation metagenomic sequencing, on the other hand, is increasingly being used in pathogen detection due to its high efficiency and cost advantages. Various methods for analyzing drug resistance homologous genes based on second-generation metagenomic sequencing have also been developed.

[0004] Single-base mutation sites are also one of the reasons why pathogens develop drug resistance. Currently, most microbial SNP analysis methods based on next-generation sequencing (mNGS) rely on sequencing data from a single species. In metagenomic sequencing, data from individual species are often limited, leading to insufficient sequencing data coverage. This affects the accuracy of SNP detection for individual species based on mNGS data, resulting in the inability to effectively identify single-base mutation sites. Summary of the Invention

[0005] In view of this, this application provides a method, apparatus, device and storage medium for detecting single-base mutations in pathogens to solve the problem that existing second-generation metagenomic sequencing methods cannot effectively identify single-base mutation sites.

[0006] To address the aforementioned technical problems, this application provides a method for detecting single-base mutations in drug-resistant pathogens, comprising: sequencing a sample of the pathogen to be tested to obtain sequencing data at a first preset sequencing depth; comparing the sequencing data with a reference genome corresponding to the pathogen sample to be tested to identify unsequential sites; identifying all first strongly correlated sites corresponding to the unsequential sites based on a pre-constructed strongly correlated site correspondence table; numerically representing the mutation information of all first strongly correlated sites based on a preset numerical method to obtain numerical results; and inputting the numerical results into a pre-trained single-base mutation detection model for prediction to determine whether the site to be tested is a positive or negative site.

[0007] As a further improvement of this application, the mutation information of all first strongly correlated sites is numerically represented based on a preset numerical method to obtain numerical results, including: confirming whether each first strongly correlated site is covered by sequencing and whether a mutation has occurred based on the alignment results of sequencing data and the reference genome corresponding to the pathogen sample to be tested; constructing a first matrix as the numerical result, where the rows of the first matrix represent all strains of the pathogen sample to be tested, and the columns of the first matrix represent the first strongly correlated sites, wherein the first strongly correlated sites that are covered by sequencing and have mutated are assigned a value of 1, and the first strongly correlated sites that are not covered by sequencing or have not mutated are assigned a value of 0.

[0008] As a further improvement to this application, the first preset sequencing depth is a sequencing depth of 3 layers or more.

[0009] As a further improvement of this application, the single-base mutation detection model represents the mapping relationship between strongly correlated sites and related drug-resistant single-base mutations. The single-base mutation detection model is trained based on the numerical representation of the variation results of drug-resistant single-base sites and the strongly correlated sites corresponding to drug-resistant single-base sites.

[0010] As a further improvement to this application, the step of pre-training the single-base mutation detection model includes: obtaining genomic data of all bacterial strains at the species classification level of the target bacteria from the database, and performing simulated sequencing on each bacterial strain at a second preset sequencing depth to obtain simulated sequencing data; comparing the simulated sequencing data with the reference genome of the pathogen to confirm the variation results at all sites in the simulated sequencing results; quantifying the variation results, including: constructing a second matrix, where the rows of the second matrix represent all bacterial strains of the pathogen sample, and the columns of the second matrix represent amino acid sites, wherein mutated amino acid sites are assigned a value of 1, and non-mutated amino acid sites are assigned a value of 0; statistically analyzing the drug-resistant single-base sites of the target bacteria recorded in the database, and using... The second matrix calculates the correlation between each drug-resistant single-base site and any other site, defines a correlation threshold, and defines sites with a correlation less than or equal to the threshold as the second strongly correlated sites of the drug-resistant single-base site; constructs a mapping relationship between each drug-resistant single-base site and its corresponding second strongly correlated site to obtain the single-base mutation detection model to be trained; quantifies the mutation results of the second strongly correlated sites, including: constructing a third matrix, where the rows of the third matrix represent bacterial strains and the columns represent the second strongly correlated sites, where the second strongly correlated sites with mutations are assigned a value of 1, and the second strongly correlated sites without mutations are assigned a value of 0; and trains the single-base mutation detection model using the third matrix and the drug-resistant single-base sites.

[0011] As a further improvement of this application, the correlation between each drug-resistant single base site and any other site is calculated, and a correlation threshold is defined. Sites with a correlation less than or equal to the threshold are defined as the second strongly correlated sites of the drug-resistant site. This includes: calculating the Pearson coefficient between any two columns in the second matrix using the Pearson method, and obtaining the correlation threshold using a preset threshold acquisition method; when a strong correlation is confirmed between two columns based on the Pearson coefficient and the correlation threshold, confirming the correspondence between sites at corresponding positions in the two columns as strongly correlated sites; and confirming all second strongly correlated sites corresponding to the drug-resistant single base site based on the correspondence between the strongly correlated sites.

[0012] As a further improvement to this application, after confirming that the sites at corresponding positions between the two arrays are strongly correlated, the method further includes: recording the correspondence between the strongly correlated sites in a strongly correlated site correspondence table.

[0013] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a pathogen drug resistance single-base mutation detection device, comprising: a sequencing module for sequencing a pathogen sample to be tested and obtaining sequencing data at a first preset sequencing depth; an alignment module for aligning the sequencing data with a reference genome corresponding to the pathogen sample to be tested and identifying test sites not covered by sequencing; a confirmation module for confirming all first strongly correlated sites corresponding to the test sites not covered by sequencing based on a pre-constructed strongly correlated site correspondence table; a numericalization module for numerically representing the mutation information of all first strongly correlated sites based on a preset numericalization method and obtaining numerical results; and a prediction module for inputting the numerical results into a pre-trained single-base mutation detection model for prediction and confirming whether the test site is a positive or negative site.

[0014] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide a computer device, the computer device including a processor and a memory coupled to the processor, the memory storing program instructions, and when the program instructions are executed by the processor, causing the processor to perform the steps of the pathogen drug resistance single base mutation detection method as described above.

[0015] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a storage medium storing program instructions capable of implementing the pathogen drug resistance single base mutation detection method described above.

[0016] The beneficial effects of this application are as follows: The pathogen drug resistance single-base mutation detection method of this application obtains the sequencing data of the pathogen sample at a first preset sequencing depth, compares the sequencing data with the reference genome, identifies the test sites not covered by sequencing, and then identifies all strongly correlated sites corresponding to the test site through a pre-constructed strongly correlated site correspondence table. Based on the comparison results between the sequencing data and the reference genome, and combined with a preset numerical method, all strongly correlated sites are numerically quantified, and the numerical results are input into the single-base mutation detection model for detection to confirm whether the test site is a positive or negative site. It is based on the phenomenon of gene linkage disequilibrium that occurs during the evolution of pathogen genomes, utilizes the strong correlation between sites, and numerically represents the mutation information of all strongly correlated sites corresponding to the test site. Then, it uses the numerical representation results to predict whether the test site has mutated, thereby accurately predicting the mutation information of sites not covered by sequencing. Attached Figure Description

[0017] Figure 1 This is a schematic flowchart of a pathogen drug resistance single-base mutation detection method according to an embodiment of the present invention;

[0018] Figures 2-33 This is an ROC curve graph evaluating the accuracy of the single-base mutation detection model based on 32 drug-resistant single base sites of Staphylococcus aureus according to an embodiment of the present invention.

[0019] Figure 34 This is a schematic diagram showing the relationship between the accuracy of the single-base mutation detection model and the average sequencing depth in an embodiment of the present invention.

[0020] Figure 35 This is a schematic diagram of the functional modules of the pathogen drug resistance single-base mutation detection device according to an embodiment of the present invention;

[0021] Figure 36 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0022] Figure 37 This is a schematic diagram of the structure of the storage medium according to an embodiment of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0024] The terms "first," "second," and "third" in this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. All directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationships and movements between components in a specific orientation (as shown in the figures). If the specific orientation changes, the directional indications also change accordingly. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0025] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0026] Figure 1 This is a schematic flowchart of the pathogen drug resistance single-base mutation detection method according to an embodiment of the present invention. It should be noted that if substantially the same result is obtained, the method of the present invention is not necessarily identical. Figure 1 The illustrated process sequence is limited. For example... Figure 1 As shown, the method for detecting drug-resistant single-base mutations in this pathogen includes the following steps:

[0027] Step S101: Sequencing the pathogen sample to be tested to obtain sequencing data at the first preset sequencing depth.

[0028] Specifically, after obtaining the pathogen sample to be tested, second-generation metagenomic sequencing is performed on it to extract sequencing data at a first preset sequencing depth. This first preset sequencing depth refers to a pre-set average sequencing depth.

[0029] Step S102: Compare the sequencing data with the reference genome corresponding to the pathogen sample to be tested to identify the test sites that were not covered by sequencing.

[0030] It should be noted that when the amount of pathogen sample data obtained is small, the coverage of the sequencing data is often insufficient in second-generation metagenomic sequencing. Therefore, the sequencing data cannot cover all sites of the pathogen sample.

[0031] Specifically, the reference genome corresponding to the pathogen sample to be tested is downloaded from a gene database. This reference genome includes information on all known and detected sites in the pathogen sample. For example, for Staphylococcus aureus, there are currently 805 known strains, and the reference genome includes information on all sites of these 805 strains. After obtaining the sequencing data of the pathogen sample to be tested and the reference genome, the sequencing data and the reference genome are compared to identify the test sites not covered in the sequencing data. There may be multiple test sites not covered in the sequencing data, and the mutation information of each test site needs to be analyzed and confirmed individually. Each test site can be detected and confirmed using the single-base mutation detection method of this embodiment.

[0032] Step S103: Based on the pre-constructed strong correlation site correspondence table, identify all first strong correlation sites corresponding to the test sites that have not been covered by sequencing.

[0033] Specifically, for a test site not covered by sequencing, all the first strongly correlated sites of that test site are identified using a pre-constructed strong correlation mapping. It should be noted that the strong correlations between sites are constructed based on the phenomenon of gene linkage disequilibrium. Gene linkage disequilibrium refers to the phenomenon that, in a population, the probability of alleles belonging to two or more loci appearing simultaneously on a chromosome is higher than the frequency of random occurrence. Based on the phenomenon of gene linkage disequilibrium, it is known that when a site in a bacterial strain has a mutation, the probability of a mutation at sites strongly correlated with that site is extremely high.

[0034] Step S104: Based on the preset numerical method, the mutation information of all first strongly correlated sites is numerically represented to obtain the numerical result.

[0035] Specifically, after identifying all the first strongly correlated sites corresponding to the site to be tested, these first strongly correlated sites are located in the sequencing data. The alignment results between the sequencing data and the reference genome are used to confirm whether each first strongly correlated site is covered by sequencing, and whether any mutations have occurred at the covered first strongly correlated sites. It should be noted that the first strongly correlated sites corresponding to the site to be tested are also sites in the pathogen sample to be tested, and they may also be covered by the sequencing data.

[0036] Furthermore, step S104 specifically includes:

[0037] 1. Based on the alignment results of sequencing data with the reference genome corresponding to the pathogen sample to be tested, confirm whether each first strongly correlated site has been covered by sequencing and whether mutations have occurred.

[0038] 2. Construct a first matrix as the numerical result. The rows of the first matrix represent all strains of the pathogen sample to be tested, and the columns of the first matrix represent the first strongly correlated sites. The first strongly correlated sites that are covered by sequencing and have undergone mutation are assigned a value of 1, and the first strongly correlated sites that are not covered by sequencing or have not undergone mutation are assigned a value of 0.

[0039] The mutation information at a site can be numerically represented as follows:

[0040]

[0041] in, Indicates site The corresponding value, Indicates site The genotype.

[0042] Specifically, for all first strongly correlated loci, based on the alignment results of sequencing data with the reference genome, it is confirmed whether the first strongly correlated loci are covered by sequencing. If covered, it is further determined whether the first strongly correlated locus has mutated. If a mutation occurs, the first strongly correlated locus is assigned a value of 1; if the first strongly correlated locus is not covered by sequencing or is sequenced but has not mutated, the first strongly correlated locus is assigned a value of 0. For example, assuming that the pathogen sample to be tested corresponds to 4 bacterial strains, each bacterial strain has 4 loci, and the value of each locus of each bacterial strain is known, the first matrix constructed is as follows:

[0043] .

[0044] Step S105: Input the numerical results into the pre-trained single-base mutation detection model for prediction to confirm whether the site to be tested is a positive site or a negative site.

[0045] Specifically, after obtaining the numerical result of the first strongly correlated site, the numerical result is input into the pre-trained single-base mutation detection model for prediction. When the prediction result indicates that the test site has mutated, the test site is confirmed as a positive site; when the prediction result indicates that the test site has not mutated, the test site is confirmed as a negative site.

[0046] Furthermore, the single-base mutation detection model represents the mapping relationship between strongly correlated sites and related drug-resistant single-base mutations. This model is trained based on a numerical representation of the mutation results of drug-resistant single-base sites and their corresponding strongly correlated sites. The single-base mutation detection model is constructed based on machine learning algorithms, including any one of logistic regression, random forest, or support vector machine.

[0047] Further steps in pre-training the single-base mutation detection model include:

[0048] 1. Obtain the genomic data of all bacterial strains at the species classification level of the target bacteria from the database, and perform simulated sequencing on each bacterial strain at the second preset sequencing depth to obtain simulated sequencing data.

[0049] Specifically, DNA sequences of all bacterial strains corresponding to the pathogen sample are obtained from a gene database. Then, the art_illumina software is used to perform simulated sequencing for each bacterial strain, and simulated sequencing data at a second preset sequencing depth are obtained. This second preset sequencing depth is preset to be a high sequencing depth, and the sequencing data at this second preset sequencing depth covers as many sites as possible.

[0050] 2. Compare the simulated sequencing data with the pathogen's reference genome to confirm the variation results at all sites in the simulated sequencing results.

[0051] Specifically, the pathogen's reference genome is obtained from a gene database. The simulated sequencing data is compared with the reference genome to confirm the variation results at all sites in the simulated sequencing results.

[0052] 3. Numericalize the mutation results, including: constructing a second matrix, where the rows of the second matrix represent all strains of the pathogen sample, and the columns of the second matrix represent amino acid sites, where mutated amino acid sites are assigned a value of 1, and non-mutated amino acid sites are assigned a value of 0.

[0053] Specifically, a second matrix is ​​constructed by numerically representing all sites based on the variation results of all sites.

[0054] 4. Statistically analyze the drug-resistant single-base sites of the target bacteria recorded in the database, and use the second matrix to calculate the correlation between each drug-resistant single-base site and any other site. Define a correlation threshold, and define sites with a correlation less than or equal to the threshold as the second strongly correlated sites of the drug-resistant single-base sites.

[0055] Specifically, based on the phenomenon of gene linkage disequilibrium, the bacterial strain containing the drug-resistant single base site is sequentially subjected to strong correlation determination with other bacterial strains to confirm the existence of strongly correlated bacterial strains. Then, from the strongly correlated bacterial strains, the second strongly correlated site corresponding to the drug-resistant single base site is identified, thereby obtaining all the second strongly correlated sites corresponding to the drug-resistant single base site.

[0056] The steps involved in calculating the correlation between each drug-resistant single-base site and any other site, defining a correlation threshold, and defining sites with a correlation less than or equal to the threshold as the second most strongly correlated sites of the drug-resistant site. These steps specifically include:

[0057] 4.1. Calculate the Pearson coefficient between any two columns in the second matrix using the Pearson method, and obtain the correlation threshold using a preset threshold acquisition method.

[0058] Specifically, after obtaining the second matrix, any two columns in the second matrix are selected, each column corresponding to all sites of a bacterial strain. The Pearson coefficient between the two columns is calculated, and then the Bonferroni method is used to determine the p-value threshold, which is the correlation threshold. The above steps are repeated to obtain the Pearson coefficient between every two columns in the third matrix.

[0059] 4.2 When confirming a strong correlation between two columns of arrays based on the Pearson coefficient and the correlation threshold, confirm the correspondence between the corresponding positions in the two columns of arrays as strongly correlated positions.

[0060] Specifically, when the Pearson coefficient is less than or equal to the correlation threshold, a strong correlation is confirmed between the two arrays, and the corresponding positions in the two arrays are strongly correlated with each other.

[0061] 4.3. Based on the correspondence of strongly correlated sites, identify all the second strongly correlated sites corresponding to the drug-resistant single base sites.

[0062] Specifically, the correlation between the bacterial strain containing the drug-resistant single base site and other bacterial strains is calculated using the above method, thereby confirming all the second strongly correlated sites corresponding to the drug-resistant single base site.

[0063] Furthermore, after confirming that the corresponding positions in the two arrays are strongly correlated, the process also includes: recording the correspondence between the strongly correlated positions in the strongly correlated position correspondence table.

[0064] Specifically, in order to facilitate the rapid acquisition of strongly correlated sites corresponding to the test site in practical applications, when training the single base mutation detection model for the test site, the correspondence between strongly correlated sites obtained during the training process is recorded in a strongly correlated site correspondence table. Subsequently, the strongly correlated sites corresponding to the test site can be quickly obtained by looking up the table.

[0065] 5. Construct a mapping relationship between each drug-resistant single base site and the second strongly correlated site corresponding to that drug-resistant single base site to obtain the single base mutation detection model to be trained.

[0066] 6. Numericalize the variation results of the second strongly correlated site, including: constructing a third matrix, where the rows of the third matrix represent bacterial strains and the columns of the third matrix represent the second strongly correlated sites, wherein the second strongly correlated sites that have mutated are assigned a value of 1, and the second strongly correlated sites that have not mutated are assigned a value of 0.

[0067] Specifically, after identifying all the second strongly correlated sites corresponding to the drug-resistant single base sites, the alignment results of the simulated sequencing data with the pathogen's reference genome are used to determine whether each second strongly correlated site has mutated. Mutated second strongly correlated sites are assigned a value of 1, and unmutated second strongly correlated sites are assigned a value of 0, thus constructing a third matrix.

[0068] 7. Train a single-base mutation detection model using the third matrix and drug-resistant single-base sites.

[0069] Specifically, the drug-resistant single-base site is used as the output of the single-base mutation detection model, and the third matrix is ​​used as the input of the single-base mutation detection model to train the model until it reaches a preset accuracy. In this embodiment, hierarchical cross-validation is used to train the model.

[0070] It should be understood that in this embodiment, each drug-resistant single-base site corresponds to a single-base mutation detection model. After training a single-base mutation detection model for each drug-resistant single-base site through the above training method, in actual application, when a certain site is not covered by sequencing, the single-base mutation detection model corresponding to that site is used to perform true / false positive detection on that site.

[0071] Furthermore, the first preset sequencing depth is a sequencing depth of 3 layers or more.

[0072] It should be noted that, in order to study the accuracy of the single-base mutation detection model under different sequencing depths, this embodiment uses the two mutations shown in Table 1 below as positive sites to evaluate the accuracy of the single-base mutation detection model.

[0073] Table 1

[0074]

[0075] This invention utilizes an experimental analysis of ciprofloxacin resistance in Staphylococcus aureus. Sequencing sequences from Staphylococcus aureus samples were randomly selected at different sequencing depths (20 data points randomly selected from each sequencing depth). The randomly selected sequences were analyzed for mutations using bwa+samtools+varscan. For known drug-resistant single-base mutations not covered by sequencing data, strongly correlated mutations were first extracted and quantified. The aforementioned single-base mutation detection model was then used to identify related drug-resistant single-base mutations. Finally, all ciprofloxacin-related sites detected based on the randomly selected data were compared with positive sites. As shown in Table 2, 32 drug-resistant single-base mutations related to Staphylococcus aureus were recorded in the card database and showed strong correlations. The accuracy of the single-base mutation detection model was evaluated using these 32 drug-resistant single-base sites, and the ROC curves of the evaluation results are shown below. Figures 2 to 33 As shown in the figure, "Receiver operating characteristic example" represents the ROC curve, with the horizontal axis representing the false positive rate and the vertical axis representing the true positive rate.

[0076] Table 2

[0077]

[0078] Based on the ROC curves of the evaluation results using 32 drug-resistant single-base sites of Staphylococcus aureus, it can be seen that the AUC of most drug-resistant single-base sites is greater than 0.9, indicating that the single-base mutation detection model obtained through the above training method has high predictive accuracy. Figure 34 As shown in the figure (the horizontal axis represents the average sequencing depth, and the vertical axis represents the accuracy of the single base mutation detection model at that sequencing depth), when the average sequencing depth of Staphylococcus aureus reaches 3 layers or more, the consistency between the single base mutation detection model results and positive sites is higher than 98%. Therefore, in this embodiment, the first preset sequencing depth is set to 3 layers or more.

[0079] The pathogen drug resistance single-base mutation detection method of this invention acquires sequencing data of the pathogen sample at a first preset sequencing depth, compares the sequencing data with a reference genome to identify unsequential sites, and then uses a pre-constructed strong correlation site correspondence table to identify all strongly correlated sites corresponding to the test site. Based on the comparison results between the sequencing data and the reference genome, and combined with a preset numerical method, all strongly correlated sites are quantified. The numerical results are then input into a single-base mutation detection model for detection to determine whether the test site is a positive or negative site. This method is based on the phenomenon of gene linkage disequilibrium that occurs during the evolution of pathogen genomes. It utilizes the strong correlation between sites, quantifies the mutation information of all strongly correlated sites corresponding to the test site, and then uses the numerical representation results to predict whether the test site has mutated, thereby accurately predicting the mutation information of sites not covered by sequencing.

[0080] Figure 35 This is a schematic diagram of the functional modules of the pathogen drug resistance single-base mutation detection device according to an embodiment of the present invention. Figure 35 As shown, the pathogen drug resistance single base mutation detection device 20 includes a sequencing module 21, an alignment module 22, a confirmation module 23, a numericalization module 24, and a prediction module 25.

[0081] Sequencing module 21 is used to sequence the pathogen sample to be tested and obtain sequencing data at the first preset sequencing depth;

[0082] The comparison module 22 is used to compare the sequencing data with the reference genome corresponding to the pathogen sample to be tested, and to identify the test sites that have not been covered by sequencing.

[0083] The confirmation module 23 is used to confirm all first strongly correlated sites corresponding to the test sites that have not been covered by sequencing, based on a pre-constructed strongly correlated site correspondence table.

[0084] The numericalization module 24 is used to numerically represent the mutation information of all first strongly correlated sites based on a preset numericalization method, and obtain the numericalization result.

[0085] The prediction module 25 is used to input the numerical results into the pre-trained single-base mutation detection model for prediction, to confirm whether the site to be tested is a positive site or a negative site.

[0086] Optionally, the confirmation module 23 performs specific operations to numerically represent the mutation information of all first strongly correlated sites based on a preset numerical method to obtain numerical results, including: confirming whether each first strongly correlated site is covered by sequencing and whether a mutation has occurred based on the alignment results of sequencing data and the reference genome corresponding to the pathogen sample to be tested; constructing a first matrix as the numerical result, where the rows of the first matrix represent all strains of the pathogen sample to be tested, and the columns of the first matrix represent the first strongly correlated sites, wherein the first strongly correlated sites that are covered by sequencing and have mutated are assigned a value of 1, and the first strongly correlated sites that are not covered by sequencing or have not mutated are assigned a value of 0.

[0087] Optionally, the first preset sequencing depth includes a sequencing depth of 3 layers or more.

[0088] Optionally, the single-base mutation detection model represents the mapping relationship between strongly correlated sites and related drug-resistant single-base mutations. The single-base mutation detection model is trained based on the numerical representation of the variation results of drug-resistant single-base sites and the strongly correlated sites corresponding to drug-resistant single-base sites.

[0089] Optionally, the prediction module 25 is also used to perform operations of a pre-trained single-base mutation detection model, including: obtaining genomic data of all bacterial strains at the species classification level of the target bacteria from the database, and performing simulated sequencing on each bacterial strain at a second preset sequencing depth to obtain simulated sequencing data; comparing the simulated sequencing data with the reference genome of the pathogen to confirm the variation results at all sites in the simulated sequencing results; quantifying the variation results, including: constructing a second matrix, where the rows of the second matrix represent all bacterial strains of the pathogen sample, and the columns of the second matrix represent amino acid sites, wherein mutated amino acid sites are assigned a value of 1, and non-mutated amino acid sites are assigned a value of 0; and statistically analyzing the drug-resistant single-base sites of the target bacteria recorded in the database. The correlation between each drug-resistant single-base site and any other site is calculated using a second matrix. A correlation threshold is defined, and sites with a correlation less than or equal to the threshold are defined as the second strongly correlated sites of the drug-resistant single-base site. A mapping relationship is constructed between each drug-resistant single-base site and the corresponding second strongly correlated site to obtain the single-base mutation detection model to be trained. The mutation results of the second strongly correlated sites are numerically quantified, including: constructing a third matrix, where the rows of the third matrix represent bacterial strains and the columns of the third matrix represent the second strongly correlated sites, where the second strongly correlated sites with mutations are assigned a value of 1, and the second strongly correlated sites without mutations are assigned a value of 0; the single-base mutation detection model is trained using the third matrix and the drug-resistant single-base sites.

[0090] Optionally, the prediction module 25 performs the operation of calculating the correlation between each drug-resistant single base site and any other site, defining a correlation threshold, and defining sites with a correlation less than or equal to the threshold as the second strongly correlated sites of the drug-resistant site. Specifically, this includes: cyclically calculating the Pearson coefficient between any two columns in the second matrix based on the Pearson method, and obtaining the correlation threshold using a preset threshold acquisition method. When a strong correlation is confirmed between two columns based on the Pearson coefficient and the correlation threshold, the corresponding sites in the two columns are confirmed as strongly correlated sites. Based on the correspondence of strongly correlated sites, all second strongly correlated sites corresponding to the drug-resistant single base site are confirmed.

[0091] Optionally, after the prediction module 25 performs the operation of confirming that the corresponding positions between the two columns of arrays are strongly correlated, it is also used to: record the correspondence of the strongly correlated positions to the strongly correlated position correspondence table.

[0092] For further details regarding the implementation techniques of each module in the pathogen drug resistance single-base mutation detection device in the above embodiments, please refer to the description in the pathogen drug resistance single-base mutation detection method in the above embodiments, which will not be repeated here.

[0093] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0094] Please see Figure 36 , Figure 36 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Figure 36 As shown, the computer device 30 includes a processor 31 and a memory 32 coupled to the processor 31. The memory 32 stores program instructions. When the program instructions are executed by the processor 31, the processor 31 performs the steps of the pathogen drug resistance single base mutation detection method described in any of the above embodiments.

[0095] The processor 31 can also be referred to as a CPU (Central Processing Unit). The processor 31 may be an integrated circuit chip with signal processing capabilities. The processor 31 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor.

[0096] See Figure 37 , Figure 37 This is a schematic diagram of the structure of a storage medium according to an embodiment of the present invention. The storage medium of this embodiment stores program instructions 41 capable of implementing the above-described method for detecting drug-resistant single-base mutations in pathogens. These program instructions 41 can be stored in the storage medium in the form of a software product, including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or computer devices such as computers, servers, mobile phones, and tablets.

[0097] In the several embodiments provided in this application, it should be understood that the disclosed computer devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.

[0098] Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. The above are merely embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for detecting single-base mutations in drug-resistant pathogens, characterized in that, It includes: Sequencing is performed on the pathogen sample to be tested to obtain sequencing data at the first preset sequencing depth; The sequencing data is compared with the reference genome corresponding to the pathogen sample to be tested to identify the test sites that were not covered by sequencing. Based on the pre-constructed strong correlation site correspondence table, all first strong correlation sites corresponding to the test sites that were not covered by sequencing were identified; The mutation information of all first strongly correlated sites is numerically represented based on a preset numerical method to obtain numerical results; The numerical results are input into a pre-trained single-base mutation detection model for prediction to confirm whether the site to be tested is a positive or negative site. The single-base mutation detection model represents the mapping relationship between strongly correlated sites and related drug-resistant single-base mutations. The single-base mutation detection model is trained based on the numerical representation of the variation results of drug-resistant single-base sites and the strongly correlated sites corresponding to the drug-resistant single-base sites. The steps for pre-training a single-base mutation detection model include: The genome data of all bacterial strains at the species classification level of the target bacteria were obtained from the database, and each bacterial strain was simulated for sequencing at the second preset sequencing depth to obtain simulated sequencing data. The simulated sequencing data is compared with the reference genome of the pathogen to confirm the variation results at all sites in the simulated sequencing results; The mutation results are quantified, including: constructing a second matrix, where the rows of the second matrix represent all strains of the pathogen sample, and the columns of the second matrix represent amino acid sites, wherein mutated amino acid sites are assigned a value of 1, and non-mutated amino acid sites are assigned a value of 0; The drug-resistant single base sites of the target bacteria are recorded in the statistical database. The correlation between each drug-resistant single base site and any other site is calculated using the second matrix. A correlation threshold is defined. Sites with a correlation less than or equal to the threshold are defined as the second strongly correlated sites of the drug-resistant single base site. A mapping relationship is constructed between each drug-resistant single base site and the second strongly correlated site corresponding to that drug-resistant single base site to obtain the single base mutation detection model to be trained; The results of the mutations at the second strongly correlated sites are quantified, including: constructing a third matrix, wherein the rows of the third matrix represent bacterial strains and the columns of the third matrix represent the second strongly correlated sites, wherein the second strongly correlated sites that have mutated are assigned a value of 1 and the second strongly correlated sites that have not mutated are assigned a value of 0; The single-base mutation detection model is trained using the third matrix and the drug-resistant single-base site.

2. The method for detecting single-base mutations in drug-resistant pathogens according to claim 1, characterized in that, The mutation information of all first strongly correlated sites is numerically represented based on a preset numerical method to obtain numerical results, including: Based on the alignment results between the sequencing data and the reference genome corresponding to the pathogen sample to be tested, it is confirmed whether each first strongly correlated site has been covered by sequencing and whether a mutation has occurred. A first matrix is ​​constructed as the numerical result. The rows of the first matrix represent all strains of the pathogen sample to be tested, and the columns of the first matrix represent the first strongly correlated sites. The first strongly correlated sites that are covered by sequencing and have undergone mutation are assigned a value of 1, and the first strongly correlated sites that are not covered by sequencing or have not undergone mutation are assigned a value of 0.

3. The method for detecting single-base mutations in drug-resistant pathogens according to claim 1, characterized in that, The first preset sequencing depth is a sequencing depth of 3 layers or more.

4. The method for detecting single-base mutations in drug-resistant pathogens according to claim 1, characterized in that, The process involves calculating the correlation between each drug-resistant monobase site and any other site, defining a correlation threshold, and defining sites with a correlation less than or equal to the threshold as the second strongly correlated sites of the drug-resistant monobase site, including: The Pearson coefficient between any two columns in the second matrix is ​​calculated iteratively based on the Pearson method, and the correlation threshold is obtained using a preset threshold acquisition method. When a strong correlation is confirmed between two columns of arrays based on the Pearson coefficient and the correlation threshold, the corresponding positions in the two columns of arrays are confirmed to be strongly correlated. Based on the correspondence of the strongly correlated sites, all the second strongly correlated sites corresponding to the drug-resistant single base sites were identified.

5. The method for detecting single-base mutations in drug-resistant pathogens according to claim 4, characterized in that, After confirming that the corresponding positions in the two array columns are strongly correlated, the process further includes: The correspondence between strongly correlated sites is recorded in the strongly correlated site correspondence table.

6. A pathogen drug resistance single-base mutation detection device utilizing the pathogen drug resistance single-base mutation detection method of claim 1, characterized in that, It includes: The sequencing module is used to sequence the pathogen sample to be tested and obtain sequencing data at the first preset sequencing depth; The comparison module is used to compare the sequencing data with the reference genome corresponding to the pathogen sample to be tested, and to identify the test sites that have not been covered by sequencing. The confirmation module is used to confirm all first strongly correlated sites corresponding to the test sites that have not been covered by sequencing, based on a pre-constructed strongly correlated site correspondence table. The numericalization module is used to numerically represent the mutation information of all first strongly correlated sites based on a preset numericalization method, and obtain the numericalization result. The prediction module is used to input the numerical results into a pre-trained single-base mutation detection model for prediction, and to confirm whether the site to be tested is a positive site or a negative site.

7. A computer device, characterized in that, The computer device includes a processor and a memory coupled to the processor, the memory storing program instructions that, when executed by the processor, cause the processor to perform the steps of the pathogen drug resistance single-base mutation detection method as described in any one of claims 1-5.

8. A storage medium, characterized in that, It stores program instructions capable of implementing the pathogen drug resistance single-base mutation detection method as described in any one of claims 1-5.