Marker combination, method for predicting invasion risk of serotype 19F and application
By providing a combination of markers including Gene1 to Gene10 for detecting the invasiveness of Streptococcus pneumoniae serotype 19F, the shortcomings in identifying highly invasive serotypes in the prior art are solved, high sensitivity and specific detection are achieved, supporting early monitoring and clinical research, and providing strong support for clinical decision-making.
Patent Information
- Application Number
- CN202510022655.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has insufficient identification of highly invasive serotypes such as 19F in the detection and diagnosis of Streptococcus pneumoniae. The sensitivity and specificity of the detection method are limited, and sufficient molecular-level information cannot be provided to assist clinical decision-making.
A marker combination, including at least one nucleic acid sequence in Gene1 to Gene10, is provided for detecting the invasiveness of Streptococcus pneumoniae serotype 19F, is detected by a kit, a combination of diagnostic reagents, and a combination of marker in a detection device, and data analysis is performed in combination with a machine learning model to predict the risk of invasiveness of the 19F strain.
High sensitivity and specific detection of the invasiveness of Streptococcus pneumoniae serotype 19F is achieved, which can accurately identify highly invasive strains, meet the needs of early monitoring and clinical research, provide a more comprehensive disease risk assessment, and provide strong support for clinical decision-making.
Smart Images

Figure CN119932212A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pathogen detection and molecular diagnosis, and in particular to a marker combination, a method for predicting the invasion risk of serotype 19F and its application. Background Art
[0002] Pathogen detection and molecular diagnostic techniques are important means for identifying and evaluating pathogen infections in modern medicine. Streptococcus pneumoniae, a Gram-positive bacterium, can cause a variety of invasive diseases including pneumonia, meningitis, and bacteremia, and has become a global public health issue. The bacteria can be divided into more than 90 serotypes based on the different capsular polysaccharide antigens, among which serotype 19F poses a particular threat to children and people with weakened immunity due to its high invasiveness.
[0003] In the prior art, vaccination is the main strategy for preventing pneumococcal infection. However, there are significant differences in the responsiveness of different serotypes to vaccines, especially the 19F serotype, which not only shows a strong immune escape ability, but also has strong invasive properties. These characteristics are closely related to the molecular characteristics of the 19F serotype, but there is currently a lack of effective detection methods to evaluate its invasiveness, resulting in an inability to meet clinical needs for early monitoring and rapid diagnosis.
[0004] The drawback of the existing technology is that although some detection methods have been used to identify Streptococcus pneumoniae, these methods are often unable to accurately distinguish different serotypes, especially the identification of highly invasive serotypes. In addition, the existing detection methods are also limited in sensitivity and specificity, which makes it difficult to meet the needs of precision medicine. These methods cannot provide enough molecular information to help clinicians make more accurate diagnosis and treatment decisions.
[0005] In summary, existing technologies have multiple defects in the detection and diagnosis of Streptococcus pneumoniae, including insufficient recognition of highly invasive serotypes such as 19F, limited sensitivity and specificity of detection methods, and lack of molecular markers that can assist clinical decision-making. These defects limit the precise prevention and control of Streptococcus pneumoniae infection and highlight the urgent need to develop new detection markers.
[0006] In view of this, the present invention is proposed. Summary of the invention
[0007] The purpose of the present invention is to provide a marker combination, a method for predicting the invasion risk of serotype 19F and its application. The provided marker combination can be used for predicting the invasiveness of 19F and has good detection sensitivity and specificity.
[0008] In order to achieve the above-mentioned purpose of the present invention, the following technical solutions are particularly adopted:
[0009] In a first aspect, the present invention provides a marker combination, comprising at least one of Gene1, Gene2, Gene3, Gene4, Gene5, Gene6, Gene7, Gene8, Gene9, and Gene10;
[0010] The nucleic acid sequence of Gene1 is shown in SEQ ID NO.1;
[0011] The nucleic acid sequence of Gene2 is shown in SEQ ID NO.2;
[0012] The nucleic acid sequence of Gene3 is shown in SEQ ID NO.3;
[0013] The nucleic acid sequence of Gene4 is shown in SEQ ID NO.4;
[0014] The nucleic acid sequence of Gene5 is shown in SEQ ID NO.5;
[0015] The nucleic acid sequence of Gene6 is shown in SEQ ID NO.6;
[0016] The nucleic acid sequence of Gene7 is shown in SEQ ID NO.7;
[0017] The nucleic acid sequence of Gene8 is shown in SEQ ID NO.8;
[0018] The nucleic acid sequence of Gene9 is shown in SEQ ID NO.9;
[0019] The nucleic acid sequence of Gene10 is shown in SEQ ID NO.10.
[0020] In an optional embodiment, the marker combination is composed of a first marker group and a second marker group;
[0021] The first marker combination is at least one of Gene4, Gene6, Gene8 and Gene9;
[0022] The second marker group is at least one of Gene1, Gene2, Gene3, Gene5, Gene7 and Gene10.
[0023] In an optional embodiment, the marker combination is selected from any one of the following combinations:
[0024] Gene4 and Gene1, Gene4 and Gene2, Gene4 and Gene3, Gene4 and Gene5, Gene4 and Gene7, Gene4 and Gene10, Gene6 and Gene1, Gene6 and Gene2, Gene6 and Gene3, Gene6 and Gene5, Gene6 and Gene7, Gene6 and Gene10, Gene8 and Gene1, Gene8 and Gene2, Gene8 and Gene3, Gene8 and Gene5, Gene8 and Gene7, Gene8 and Gene10, Gene9 and Gene1, Gene9 and Gene2, Gene9 and Gene3, Gene9 and Gene5, Gene9 and Gene7, Gene9 and Gene10.
[0025] In an optional embodiment, the marker combination is selected from any one of the following combinations:
[0026] A: Gene4, Gene6, Gene8, Gene9 and Gene1;
[0027] B: Gene4, Gene6, Gene8, Gene9 and Gene2;
[0028] C: Gene4, Gene6, Gene8, Gene9 and Gene3;
[0029] D: Gene4, Gene6, Gene8, Gene9 and Gene5;
[0030] E: Gene4, Gene6, Gene8, Gene9 and Gene7;
[0031] F: Gene4, Gene6, Gene8, Gene9 and Gene10.
[0032] In a second aspect, the present invention provides a kit comprising the marker combination as described in any one of the aforementioned embodiments.
[0033] In an optional embodiment, the kit includes: a primer pair and a probe for detecting the marker combination, and a gene chip containing the marker combination.
[0034] In a third aspect, the present invention provides a use of a marker combination as described in any one of the aforementioned embodiments in the preparation of a product for detecting the invasiveness of Streptococcus pneumoniae serotype 19F;
[0035] The product for detecting the invasiveness of Streptococcus pneumoniae serotype 19F includes at least one of a kit, a diagnostic reagent combination and a detection device.
[0036] In a fourth aspect, the present invention provides a method for predicting the risk of serotype 19F invasion, comprising:
[0037] Obtaining a test result of a target sample obtained after testing with the kit according to any one of the aforementioned embodiments; and determining whether the marker combination in the kit is present in the test result;
[0038] If yes, the prediction result of the target sample is determined to be positive for serotype 19F invasion risk;
[0039] If not, the predicted result of the target sample is determined to be negative for the risk of serotype 19F invasion.
[0040] In an optional embodiment, the determining whether the detection result contains the marker combination comprises:
[0041] The test results are used as input, and data analysis is performed through the trained prediction model to obtain the invasive risk prediction result of the 19F strain corresponding to the target sample.
[0042] In an optional embodiment, the method for constructing the trained prediction model includes:
[0043] According to the training samples in the sample pool, a training set and a validation set are established; wherein each of the training samples includes marker detection data and invasion detection data;
[0044] constructing the prediction model;
[0045] The prediction model is trained using the training set; including: extracting features from each training sample in the training set to obtain the presence or absence results of each marker corresponding to the marker detection data corresponding to the training sample and the labeling results corresponding to the invasion detection data; and training the prediction model based on the marker presence or absence results and the labeling results using a machine learning algorithm;
[0046] Using a test machine to evaluate the prediction model to obtain an evaluation result; the evaluation index includes at least one of accuracy, recall, precision and AUC value;
[0047] The prediction model is repeatedly trained and adjusted according to the evaluation results to obtain the trained prediction model.
[0048] The present invention provides a marker combination, including at least one of Gene1, Gene2, Gene3, Gene4, Gene5, Gene6, Gene7, Gene8, Gene9, and Gene10; the nucleic acid sequence of Gene1 is shown in SEQ ID NO.1; the nucleic acid sequence of Gene2 is shown in SEQ ID NO.2; the nucleic acid sequence of Gene3 is shown in SEQ ID NO.3; the nucleic acid sequence of Gene4 is shown in SEQ ID NO.4; the nucleic acid sequence of Gene5 is shown in SEQ ID NO.5; the nucleic acid sequence of Gene6 is shown in SEQ ID NO.6; the nucleic acid sequence of Gene7 is shown in SEQ ID NO.7; the nucleic acid sequence of Gene8 is shown in SEQ ID NO.8; the nucleic acid sequence of Gene9 is shown in SEQ ID NO.9; the nucleic acid sequence of Gene10 is shown in SEQ ID NO.10. The marker combination provided by the present invention can be used for 19F invasive prediction to determine whether there are differences in the genes of patients with invasive infection of Streptococcus pneumoniae 19F and strains isolated from respiratory tract colonization population, and has good detection sensitivity and specificity. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0050] Figure 1 This is a pie-shaped diagram of the pan-genomic gene distribution of the 19F isolate in Example 1 of the present application;
[0051] Figure 2 This is a schematic diagram of the pan-genome accumulation curve in Example 1 of the present application;
[0052] Figure 3 This is a schematic diagram of the phylogenetic tree in Example 1 of the present application;
[0053] Figure 4 This is a box diagram showing the comparison of CFU values of respiratory isolates and invasive isolates after co-culture of lung epithelial cells in the embodiments of the present application;
[0054] Figure 5 This is a Venn diagram of the significant genes analyzed by Pyseer GWAS and Scoary in the examples of this application;
[0055] Figure 6This is a schematic diagram of the sensitivity density of adhesion and invasion related genes in the examples of this application;
[0056] Figure 7 It is a regression diagram of the Pyseer beta value and the Scoary Odds Ratio result in the embodiment of the present application;
[0057] Figure 8 This is a schematic diagram of the importance ranking of features of the random forest model used to distinguish invasive isolates in the embodiments of the present application;
[0058] Fig. 9 Schematic diagram of ROC curve analysis in the embodiment of this application. DETAILED DESCRIPTION
[0059] The embodiments of the present invention will be described in detail below in conjunction with the examples, but it will be appreciated by those skilled in the art that the following examples are only used to illustrate the present invention and should not be considered as limiting the scope of the present invention. If no specific conditions are specified in the examples, the conditions are carried out according to normal conditions or the conditions recommended by the manufacturer. If the manufacturer of the reagents or instruments used is not specified, they are all conventional products that can be obtained commercially.
[0060] In the present application, a marker combination is provided, including at least one of Gene1, Gene2, Gene3, Gene4, Gene5, Gene6, Gene7, Gene8, Gene9, and Gene10; as shown in Table 1:
[0061] Table 1. Markers in the marker combination and their corresponding nucleic acid sequences
[0062] SEQ ID NO. Name of landmark SEQ ID NO. Name of landmark 1 Gene1 6 Gene6 2 Gene2 7 Gene7 3 Gene3 8 Gene8 4 Gene4 9 Gene9 5 Gene5 10 Gene10
[0063] This embodiment relates to a marker combination for detecting the invasiveness of Streptococcus pneumoniae serotype 19F. This combination includes at least one specific nucleic acid sequence, namely any one or more of Gene1 to Gene10. Each Gene sequence has its corresponding nucleic acid sequence, which is identified by SEQ ID NO.1 to SEQ ID NO.10 respectively.
[0064] The above combination of markers can predict the invasiveness of pneumococcal serotype 19F strains. This prediction is helpful for early monitoring and clinical research for non-diagnostic purposes, thereby assisting clinical decision-making through the test results and other related test results.
[0065] This marker combination can provide highly sensitive and specific detection, helping to accurately identify highly invasive strains; it can meet the needs of early monitoring and clinical research on the invasiveness of Streptococcus pneumoniae serotype 19F for non-diagnostic purposes; and by predicting the invasiveness of the strain, it supports technicians in developing more comprehensive research plans.
[0066] In summary, the marker combination provided in this embodiment can provide a new molecular detection tool for the invasiveness detection of Streptococcus pneumoniae serotype 19F through specific gene sequences and detection methods, which helps to improve the accuracy and efficiency of detection.
[0067] It should be noted that the marker combination provided in the present application for detecting the invasiveness of Streptococcus pneumoniae serotype 19F is not directly used for the diagnosis of the disease, but is used as a research tool for identifying and studying molecular markers associated with the invasiveness of Streptococcus pneumoniae serotype 19F. Specifically, the marker combination of the present invention includes at least one nucleic acid sequence from Gene1 to Gene10, and the nucleic acid sequence of each gene is clearly provided (SEQ ID NO.1 to SEQ ID NO.10). The discovery of these gene sequences is based on in-depth molecular biological research and a large amount of experimental data, and they play an important role in the invasive behavior of Streptococcus pneumoniae.
[0068] By using these marker combinations, the invasion mechanism of Streptococcus pneumoniae can be better understood, and a scientific basis can be provided for the development of new prevention and treatment strategies. The marker combinations in the embodiments of the present application do not involve methods for the purpose of analyzing samples and directly obtaining disease diagnosis results. Therefore, they do not constitute a diagnostic method in the traditional sense. On the contrary, they provide a new tool for biomedical research related to Streptococcus pneumoniae, which helps to promote scientific progress in related fields.
[0069] In some embodiments, the marker combination consists of a first marker group and a second marker group.
[0070] In this embodiment, the marker combination divides the 10 markers into two groups, the first group is the first marker group, and the second group is the second marker group. The first marker combination is at least one of Gene4, Gene6, Gene8 and Gene9; the second marker group is at least one of Gene1, Gene2, Gene3, Gene5, Gene7 and Gene10.
[0071] Therefore, according to the above definition in this embodiment, any marker selected in the first marker group and any marker selected in the second marker group can be combined to form a marker combination, and this grouping is for more accurately detecting the invasiveness of Streptococcus pneumoniae serotype 19F.
[0072] Through the above grouping and combination, a more accurate prediction of the invasiveness of Streptococcus pneumoniae serotype 19F can be achieved. This prediction is helpful for early monitoring and rapid diagnosis, thereby assisting clinical decision-making.
[0073] By grouping and combining different markers, the accuracy of predicting the aggressiveness of S. pneumoniae serotype 19F can be improved. This grouping approach allows different marker combinations to be selected as needed, increasing the flexibility and scalability of the detection method. Through targeted testing, medical resources can be allocated more effectively, focusing on high-risk strains.
[0074] In some embodiments, the marker combination is selected from any one of the 24 combinations in Table 2:
[0075] Table 2. Markers in each marker combination (two markers in each combination)
[0076] NO. The first marker group Second marker panel NO. The first marker group Second marker panel 1 Gene4 Gene1 13 Gene8 Gene1 2 Gene4 Gene2 14 Gene8 Gene2 3 Gene4 Gene3 15 Gene8 Gene3 4 Gene4 Gene5 16 Gene8 Gene5 5 Gene4 Gene7 17 Gene8 Gene7 6 Gene4 Gene10 18 Gene8 Gene10 7 Gene6 Gene1 19 Gene9 Gene1 8 Gene6 Gene2 20 Gene9 Gene2 9 Gene6 Gene3 21 Gene9 Gene3 10 Gene6 Gene5 22 Gene9 Gene5 11 Gene6 Gene7 23 Gene9 Gene7 12 Gene6 Gene10 24 Gene9 Gene10
[0077] It should be noted that a single marker as a marker combination may have the risk of false positives or false negatives, while a pairwise combination can provide more molecular information and increase the accuracy of detection. The phenotype of invasive Streptococcus pneumoniae may involve the interaction of multiple genes, and a pairwise combination can better reflect these complex biological processes. In addition, when constructing a predictive model, the combination of multiple markers can improve the robustness of the model, allowing the model to maintain stable predictive performance when facing different samples and conditions.
[0078] In summary, the pairwise combination of markers provides higher accuracy, sensitivity, and robustness in detecting the invasiveness of S. pneumoniae serotype 19F compared with a single marker, which helps to improve the quality of diagnosis and the reliability of clinical decision-making.
[0079] In some embodiments, the marker combination is selected from any one of the following combinations:
[0080] Table 3. Markers in each marker combination (two markers in each combination)
[0081] Group Markers Group Markers A Gene4, Gene6, Gene8, Gene9 and Gene1 D Gene4, Gene6, Gene8, Gene9 and Gene5 B Gene4, Gene6, Gene8, Gene9 and Gene2 E Gene4, Gene6, Gene8, Gene9 and Gene7 C Gene4, Gene6, Gene8, Gene9 and Gene3 F Gene4, Gene6, Gene8, Gene9 and Gene10
[0082] A kit is provided in an embodiment of the present application, comprising a marker combination as described in any one of the aforementioned embodiments.
[0083] The above kit may include the above different marker combinations, which involve specific genes (Gene1 to Gene10) and are used to detect the invasiveness of Streptococcus pneumoniae serotype 19F.
[0084] In some embodiments, the kit includes: a primer pair and a probe for detecting the marker combination, and a gene chip containing the marker combination.
[0085] The kit comprises a primer pair and a probe for detecting the marker combination, and a gene chip containing the marker combination.
[0086] The design of primer pairs and probes can target specific gene sequences to improve the sensitivity and specificity of detection. Gene chip technology allows the simultaneous detection of multiple gene markers, improving detection efficiency. Gene chip technology provides intuitive results that are easy to analyze and interpret. The standardization of the kit reduces human errors in experimental operations. This kit has the advantages of easy operation, accurate detection, and wide application, which can be achieved through molecular biology technology and gene chip technology.
[0087] In an example of the present application, a marker combination as described in any one of the aforementioned embodiments is provided for preparing a product for detecting the invasiveness of Streptococcus pneumoniae serotype 19F.
[0088] The product for detecting the invasiveness of Streptococcus pneumoniae serotype 19F includes at least one of a kit, a diagnostic reagent combination and a detection device.
[0089] The present application provides a method for predicting the risk of serotype 19F invasion, comprising:
[0090] Step S1, obtaining the test result of the target sample obtained after testing with the kit as described in any one of the aforementioned embodiments; and determining whether the marker combination in the kit is present in the test result.
[0091] Step S2: If yes, the prediction result of the target sample is determined to be positive for serotype 19F invasion risk.
[0092] Step S3: If not, the prediction result of the target sample is determined to be negative for the risk of serotype 19F invasion.
[0093] As mentioned above, the target sample may refer to a biological sample collected from a subject (who may be suspected of being infected with Streptococcus pneumoniae serotype 19F, or an arbitrary or random test subject), which is used to detect the invasive risk of bacteria. Common sample types include, but are not limited to, blood samples (whole blood or plasma), respiratory samples (sputum, nasopharyngeal swabs or nasopharyngeal washes), cerebrospinal fluid samples (for patients suspected of meningitis), urine samples, tissue samples (such as lung tissue biopsy), pleural effusion samples, middle ear fluid samples, traumatic exudate, and bone marrow samples. The type of sample depends on the specific symptoms, infection site, and specific experimental requirements of the subject to ensure the representativeness of the sample and the accuracy of the test results. In addition, the target sample may also include environmental samples containing the above-mentioned biological samples, as well as artificially prepared standard samples, and the like.
[0094] The prediction method provided in this embodiment is not for the direct purpose of diagnosing or treating a disease, but is used to study the correlation between markers and the invasiveness of infection with Streptococcus pneumoniae serotype 19F. There are many situations where the direct purpose is not to diagnose or treat a disease, for example, when the sample to be tested is an environmental sample containing a biological sample or an artificially prepared standard, the direct purpose of the test is not to diagnose or treat a disease.
[0095] In the prediction method, the target sample is first tested using a kit. This usually involves extracting nucleic acid (DNA or RNA) from the sample, and then using the primer pairs and probes or gene chips in the kit to perform specific molecular biological reactions, such as PCR amplification or hybridization, to obtain preliminary data on whether a specific marker combination is present in the target sample.
[0096] Then, the data obtained in the analysis is used to determine whether the sample contains the marker combination in the kit. It is determined whether the sample contains the marker combination associated with the invasiveness of S. pneumoniae serotype 19F. Based on the analysis results of step 2, the sample is classified as positive for the risk of invasiveness of serotype 19F if the marker combination is present; if not, it is determined to be negative.
[0097] As mentioned above, the prediction result is the result corresponding to the target sample obtained after judgment, which may include a positive risk of serotype 19F invasion and a negative risk of serotype 19F invasion. A positive risk of serotype 19F invasion may be a prediction that the sample has a high risk state of serotype 19F invasion, and a negative risk of serotype 19F invasion may be a prediction that the sample has a low risk state of serotype 19F invasion.
[0098] The above prediction method achieves a rapid and accurate prediction of the invasive risk of pneumococcal serotype 19F through a series of molecular biology techniques and data analysis steps. This method not only improves the efficiency and accuracy of detection, but also provides a more comprehensive disease risk assessment by integrating information from multiple markers, providing strong support for clinical decision-making. In addition, the application of this method is not limited to clinical samples, but can also be extended to the detection of environmental samples and standards, and has broad application prospects.
[0099] In an optional embodiment, in step S1, determining whether the detection result contains the marker combination includes:
[0100] Step S11, taking the detection result as input, performing data analysis through the trained prediction model, and obtaining the invasive risk prediction result of the 19F strain corresponding to the target sample.
[0101] In the above, the obtained test results are used as input data, and a pre-trained prediction model (which can be a machine learning model) is used to analyze these data, so as to obtain the prediction results of the invasive risk of the 19F strain corresponding to the target sample. Specifically, the target sample can be classified according to the analysis results of the prediction model: if the test results show the presence of the marker combination, the prediction result is positive for the invasive risk of serotype 19F; if the marker combination is not detected, the prediction result is negative.
[0102] In summary, in this embodiment, by combining molecular detection technology and advanced data analysis, an efficient and accurate method for assessing the invasive risk of pneumococcal serotype 19F is provided. This method not only improves the efficiency and accuracy of detection, but also enhances the ability to predict the invasive risk of samples through the application of prediction models, providing strong support for non-diagnostic purposes for the study of data correlation. In addition, the automation and integration of this method further improves the efficiency of laboratory testing and the reliability of results.
[0103] In some embodiments, the method for constructing the trained prediction model comprises:
[0104] Step S111, establishing a training set and a validation set according to the training samples in the sample pool; wherein each of the training samples includes marker detection data and invasion detection data.
[0105] As mentioned above, a part of the samples are selected from the sample pool as the training set for model training; the other part is selected as the validation set for evaluating the performance of the model, thus forming a sample set for training and validation.
[0106] The sample size of the training set may be greater than or equal to 40. For example, the sample size may be 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, and the like.
[0107] For example, all the data in the sample pool are divided into 5 subsets, one of which is used as a validation set, and the remaining 4 subsets are used as training sets. Cross-validation is performed in sequence to optimize the model parameters to ensure that the model has stronger generalization and robustness.
[0108] Step S112, constructing the prediction model.
[0109] As mentioned above, the prediction model can be a machine learning model, and the machine learning model can be selected from: any one or a combination of support vector machine, decision tree, random forest, logistic regression, Bayesian, K nearest neighbor, K means, Markov and regression ridge algorithm.
[0110] For example, using the random forest algorithm, multiple decision trees are constructed to improve the accuracy and robustness of predictions. Random forest models generally have good prediction performance and are robust to outliers.
[0111] Step S113, training the prediction model using the training set; including: extracting features from each training sample in the training set to obtain the presence or absence results of each marker corresponding to the marker detection data corresponding to the training sample and the labeling results corresponding to the invasion detection data; training the prediction model based on the marker presence or absence results and the labeling results using a machine learning algorithm;
[0112] Marker detection data is extracted from the training samples, and the presence or absence of the marker is converted into numerical data (such as 0 and 1).
[0113] For example, in the process of training the prediction model using the training set, after feature extraction of the markers, the detection result of each marker (presence or absence) is converted into 0 and 1 that can be used for model training, and the annotation result (invasive or non-invasive) is also converted into 0 and 1; model training is performed based on the extracted markers and their corresponding label data.
[0114] In addition, according to the correlation of the markers, feature selection methods (such as Boruta, L1 regularization, correlation coefficient analysis, information gain, etc.) can be used to further optimize the input features.
[0115] Step S114: using a test machine to evaluate the prediction model to obtain an evaluation result. The evaluation index includes at least one of accuracy, recall, precision and AUC value.
[0116] As mentioned above, the trained model is evaluated using the test set, including accuracy, recall, precision, AUC value, etc.; based on the evaluation results, the model's hyperparameters are adjusted, or different algorithms are tried to improve the prediction accuracy; the model is repeatedly trained and adjusted until it meets the predetermined performance requirements.
[0117] Step S115, repeatedly training and adjusting the prediction model according to the evaluation result to obtain the trained prediction model.
[0118] As described above, after obtaining the trained prediction model, the presence or absence of the marker of the target sample can be detected. Specifically, the result can be input into the pre-built trained prediction model to calculate the prediction result of the sample, that is, the risk probability of invasion by serotype 19F strain.
[0119] This prediction method can determine the risk of 19F strain invasion in samples and provide a reference for clinical decision-making. At the same time, the actual application effect of the prediction model is verified in different clinical environments, and feedback data is collected; according to the performance of the model in actual use, further optimization and update are carried out to ensure that the prediction performance of the model can remain stable and efficient in different environments.
[0120] Example 1: Initial screening of markers
[0121] In this example, a preliminary screening of markers was performed.
[0122] Experimental methods:
[0123] 1. Experimental subjects:
[0124] Select the following objects as target samples:
[0125] (1) 42 invasive 19F strains, including 3 strains isolated from cerebrospinal fluid and 34 strains isolated from blood.
[0126] (2) 153 non-invasive 19F strains were isolated only from the respiratory tract, including 135 from sputum and 18 from bronchoalveolar lavage fluid.
[0127] All strains were from confirmed patients at Shenzhen Children's Hospital.
[0128] 2. Phenotypic detection of co-culture of strains and epithelium:
[0129] Co-culture experiments were performed by co-culturing S. pneumoniae 19F strain with human lung epithelial cells.
[0130] First, A549 cells were cultured in T75 flasks and passaged at a ratio of 1:3 to maintain optimal growth.
[0131] At the same time, Streptococcus pneumoniae was revived from a glycerol stock and cultured overnight on blood agar plates at 37°C and 5% CO2. Selected colonies were transferred to THB medium to ensure optimal bacterial growth.
[0132] Next, the bacterial culture was added to a 96-well plate, with 200 μL of culture medium per well, and cultured and monitored for optical density at 600 nm until the optical density reached approximately 0.1, which is approximately 5 × 10 7 CFU / mL, which is suitable for subsequent experiments. For co-culture experiments, these S. pneumoniae cultures were added to A549 cells without antibiotics, with a cell density of approximately 5×10 per bottle. 6 The infection index (MOI) was set at 10.
[0133] This setup allowed bacteria to adhere to the cell surface within 1 hour.
[0134] After incubation, non-adherent and loosely attached bacteria were washed away with PBS.
[0135] Subsequently, the epithelial cells were treated with trypsin to detach from the flask surface and the trypsin was removed by centrifugation.
[0136] The resulting cell pellet was resuspended in PBS containing 3% FBS to stabilize the cells. The suspension was serially diluted and plated on blood agar plates for the growth of bacterial colonies. The number of attached bacteria was quantitatively assessed by counting these colonies.
[0137] 3. Determination of differentially expressed genes:
[0138] Genomic DNA was extracted from bacterial cultures, and impurities were removed using cell lysis solution, RNase solution, and protease solution. Subsequently, whole genome sequencing (WGS) was performed using the Illumina Novaseq6000 platform provided by Novogene Co., Ltd.
[0139] The raw reads were filtered using Trimmomatic v0.36, and the filtered reads were assembled de novo using SPAdesv3.11.2 software, and contigs smaller than 500 bp were removed.
[0140] The core genome and additional genome of S. pneumoniae isolates were determined by pan-genome analysis using Roary v3.11.2. Pan-genome association analysis (pan-GWAS) based on gene presence / absence tables was performed using Scoary software, while correcting for covariates such as age and sex.
[0141] To control false positive results, the Bonferroni correction method was used for multiple hypothesis testing. Pyseer software was used to perform association analysis between the presence or absence of genes and the co-culture results, and covariates such as age and gender were also adjusted.
[0142] The screening criteria for significant genes were: the Bonferroni_p value in the Scoary result was less than 0.01, and the lrt-p value in the Pyseer result was less than 0.001.
[0143] 4. Initial screening of gene markers:
[0144] The Boruta algorithm was used to select markers that distinguish invasive strains from respiratory strains. The algorithm iteratively compares the importance of pan-genome association analysis (pan-GWAS) markers with simulated markers through statistical tests. Those that are consistently significantly higher than the simulated markers are considered relevant, while those that are not significant are excluded. This process allows Boruta to capture markers that range from weakly correlated to strongly correlated in high-dimensional data sets, retaining many markers that may have subtle but important relationships with the target variable.
[0145] 5. Selection of markers:
[0146] After the Boruta algorithm screened out irrelevant markers, the remaining markers were input into the RandomForestClassifier model to calculate the importance score of each marker. Subsequently, the features were sorted according to these scores, and the top features with a cumulative importance of 80% were selected for model construction.
[0147] 6. ROC analysis:
[0148] The ROC analysis was constructed using the python package sklearn. The prediction accuracy of the RandomForest classifier was evaluated by stratified 5-fold cross-validation, and the performance of the model was further verified by receiver operating characteristic (ROC) curve analysis.
[0149] Experimental results:
[0150] Pan-genome analysis of the 19F strain revealed that there were 3864 genes in the 19F strain, such as Figure 1The figure shows the distribution of genes in the 19F isolate, which are classified into core genes, soft core genes, shell genes and cloud genes according to their presence in different strains; among them, 1649 genes were identified as core genes (appearing in 99% to 100% of the strains), 202 genes were soft core genes (appearing in 95% to 99% of the strains), 223 genes were shell genes (appearing in 15% to 95% of the strains), and 1790 genes were cloud genes (appearing in less than 15% of the strains).
[0151] As the number of genomes increases, the number of conserved genes tends to stabilize, while the total number of genes continues to rise, e.g. Figure 2 The pan-gene organization cumulative curve shown in Figure 1 shows the changes in the total number of genes and the number of conserved genes as the number of genomes increases. The curve shows that with the increase in the number of genomes, the total number of genes continues to increase, while the number of conserved genes tends to stabilize. This trend reflects the high diversity and openness of the 19F strain on the additional genome. At the same time, the phylogenetic tree also highlights the genetic complexity and variability between non-invasive and invasive 19F strains, such as Figure 3 The phylogenetic tree diagram shown in the figure shows the phylogenetic tree of 3864 gene clusters and the Roary matrix of gene presence / absence. The phylogenetic tree on the left shows the evolutionary relationship between the 19F isolates, and the heat map on the right visualizes the distribution of gene clusters in different isolates. Dark blue indicates gene presence and white indicates gene absence. The matrix highlights the diversity of gene content between individual isolates, thus emphasizing that the role of both the core genome and the additional genome must be considered when understanding their pathogenic potential.
[0152] Scoary analysis was further used to investigate the presence of specific genes in 19F strains associated with non-invasive and invasive infections. This analysis statistically evaluated the distribution patterns of genes in the two groups of strains in an effort to identify unique potential genetic markers.
[0153] The analysis results showed that 37 genes were significantly different between non-invasive and invasive 19F strains (Tables 4 and 5). These genes include a variety of hypothetical proteins and transposases, such as genes from the ISL3, IS5, and IS630 families, suggesting that they promote the adaptability and immune escape of 19F in invasion through gene recombination and mutation. The top-ranked Gene_6 is annotated as the MarR gene family, which includes transcription factors related to virulence factors such as SlyA and SarX, and has functions such as promoting pathogen biofilm formation and invasion. These results highlight specific genetic elements that may play an important role in the pathogenicity and characteristic differences between invasive and non-invasive 19F strains.
[0154] Table 4. Differential genes and annotation results between invasive 19F and non-invasive 19F (1-18)
[0155] Gene ID Notes Odds_ratio Bonferroni_p Gene_1 IS5 family transposase IS1381 0.006643357 4.63E-22 Gene_2 IS5 family transposase IS1381 0.005862069 4.75E-23 Gene_3 IS5 family transposase ISSpn7 0.006643357 4.63E-22 Gene_4 transposase 177.1 1.44E-13 Gene_5 IS630 family transposase ISSpn2 0.031968032 6.63E-12 Gene_6 MarR family 96.84210526 1.86E-13 Gene_7 hypothetical protein 18.4 2.39E-08 Gene_8 hypothetical protein 0.011363636 1.54E-12 Gene_9 IS5 family transposase ISSpn7 0 6.08E-14 Gene_10 IS630 family transposase ISSpn2 0.044829343 2.35E-10 Gene_11 hypothetical protein 0.059808612 1.44E-07 Gene_12 hypothetical protein 0.020913594 8.77E-12 Gene_13 hypothetical protein 0.04516129 3.15E-08 Gene_14 Lactococcin A secretion protein LcnD 0.037307153 7.34E-10 Gene_15 hypothetical protein 0.092764378 1.16E-05 Gene_16 IS5 family transposase ISSpn7 0.092764378 1.16E-05 Gene_17 hypothetical protein 0.150425114 0.001900761 Gene_18 hypothetical protein 0.054635762 1.07E-08
[0156] Table 5. Differential genes and annotation results between invasive 19F and non-invasive 19F (19-38)
[0157] Gene ID Annotation Odds_ratio Bonferroni_p Gene_19 ISL3 family transposase ISSpn14 0.033755274 1.37E-07 Gene_20 IS5 family transposase ISSpn7 0.084294587 2.27E-06 Gene_21 hypothetical protein 0.072727273 9.69E-07 Gene_22 Lactococcin A secretion protein LcnD 23.30263158 2.55E-09 Gene_23 IS630 family transposase ISSpn2 0.079724138 3.73E-07 Gene_24 hypothetical protein 0.054347826 2.39E-08 Gene_25 IS5 family transposase ISSpn7 0.117857143 6.01E-05 Gene_26 IS630 family transposase ISSpn2 0.10326087 2.09E-05 Gene_27 hypothetical protein 0.106583072 4.95E-05 Gene_28 IS630 family transposase ISSpn2 0.04859335 8.11E-09 Gene_29 hypothetical protein 0.157423971 0.001314531 Gene_30 hypothetical protein 0.182153846 0.007789225 Gene_31 ISL3 family transposase ISSpn14 109.48 3.73E-09 Gene_32 ISL3 family transposase IS1193 0.10326087 2.09E-05 Gene_33 hypothetical protein 0.054347826 2.39E-08 Gene_34 ISL3 family transposase IS1193 9.8 5.55E-05 Gene_35 hypothetical protein 0.117241379 0.000223071 Gene_36 transposase 109.48 3.73E-09 Gene_37 ISL3 family transposase IS1167 109.48 3.73E-09
[0158] To further analyze whether the genes associated with invasiveness have biological functions, the co-culture experiment of invasive and non-invasive strains with epithelial cells found that the adhesion ability of invasive 19F strains was significantly higher than that of non-invasive 19F strains. Figure 4 The results showed that the invasive strains showed significantly higher bacterial CFU counts during the exponential growth phase (median 3200 vs 400, p value < 0.001, Mann-Whitney U test). These findings suggest that the invasive 19F strain is able to colonize and invade epithelial cells more efficiently, which may be related to its pathogenicity. By analyzing the results of Scoary and Pyseer, Figure 5 Significant overlap was shown between genes associated with invasion and genes associated with co-culture experiments. Figure 6 This indicates that genes identified by only Scoary or Pyseer have lower sensitivity, while genes detected by both show higher sensitivity. Figure 7 This showed that there was significant collinearity between Scoary and Pyseer results. This result suggests that genetic variation between groups leads to differences in adhesion in the lung epithelial microenvironment and may be a significant trigger of invasiveness. This provides an explanation for the biological functions of some gene features constructed in subsequent prediction models. For example, genes encoding the HTH-type transcriptional regulator sarX were enriched in the invasive group and significantly associated with increased adhesion indicators, highlighting the importance of these genes in bacterial pathogenicity.
[0159] Example 2: Determination of marker combinations
[0160] Table 6. AUC values of different marker combinations (single marker)
[0161] Gene combination AUC value Gene combination AUC value Gene_1 0.9176 Gene_6 0.7660 Gene_2 0.9238 Gene_7 0.7428 Gene_3 0.9176 Gene_8 0.7536 Gene_4 0.7566 Gene_9 0.7486 Gene_5 0.7837 Gene_10 0.7744
[0162] Table 7. AUC values of different marker combinations (two markers)
[0163] Gene combination AUC value Gene combination AUC value Gene_4 + Gene_1 0.9688 Gene_8 + Gene_1 0.9612 Gene_4 + Gene_2 0.9714 Gene_8 + Gene_2 0.9638 Gene_4 + Gene_3 0.9688 Gene_8 + Gene_3 0.9612 Gene_4 + Gene_5 0.7941 Gene_8 + Gene_5 0.7919 Gene_4 + Gene_7 0.8024 Gene_8 + Gene_7 0.7993 Gene_4 + Gene_10 0.7908 Gene_8 + Gene_10 0.7885 Gene_6 + Gene_1 0.9651 Gene_9 + Gene_1 0.9469 Gene_6 + Gene_2 0.9681 Gene_9 + Gene_2 0.9495 Gene_6 + Gene_3 0.9651 Gene_9 + Gene_3 0.9469 Gene_6 + Gene_5 0.7953 Gene_9 + Gene_5 0.7961 Gene_6 + Gene_7 0.8022 Gene_9 + Gene_7 0.8051 Gene_6 + Gene_10 0.7455 Gene_9 + Gene_10 0.7927
[0164] Table 8. AUC values of different marker combinations (five markers)
[0165] Group Name Marker AUC Value A Gene4, Gene6, Gene8, Gene9 and Gene1 0.9711 B Gene4, Gene6, Gene8, Gene9 and Gene2 0.9733 C Gene4, Gene6, Gene8, Gene9 and Gene3 0.9711 D Gene4, Gene6, Gene8, Gene9 and Gene5 0.7431 E Gene4, Gene6, Gene8, Gene9 and Gene7 0.7990 F Gene4, Gene6, Gene8, Gene9 and Gene10 0.7417
[0166] Through Tables 6 to 8 and Figure 8 The experimental results in show that the cumulative importance of the first 10 genes (Gene1-10) reaches 80%. Figure 9 The AUC (area under the curve) was 0.96, showing excellent discrimination ability.
[0167] In addition to genes of unknown function, the most important genes in the model are those encoding transposases, such as the IS630 family transposase ISSpn2 and the ISL3 family transposase ISSpn14; as well as virulence factor-related transcription factor MarR family genes, which are significantly correlated with the bacterial adhesion ability.
[0168] Even the AUC values of a single gene ranged from 0.74 to 0.92, and the AUC value of two genes reached 0.97 (Gene_4 and Gene_2). Overall, the prediction model was able to effectively identify invasive strains.
[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A marker combination, characterized in that: including at least one of Gene1, Gene2, Gene3, Gene4, Gene5, Gene6, Gene7, Gene8, Gene9, and Gene10; The nucleic acid sequence of Gene1 is shown in SEQ ID NO.1; The nucleic acid sequence of Gene2 is shown in SEQ ID NO.2; The nucleic acid sequence of Gene3 is shown in SEQ ID NO.3; The nucleic acid sequence of Gene4 is shown in SEQ ID NO.4; The nucleic acid sequence of Gene5 is shown in SEQ ID NO.5; The nucleic acid sequence of Gene6 is shown in SEQ ID NO.6; The nucleic acid sequence of Gene7 is shown in SEQ ID NO.7; The nucleic acid sequence of Gene8 is shown in SEQ ID NO.8; The nucleic acid sequence of Gene9 is shown in SEQ ID NO.9; The nucleic acid sequence of Gene10 is shown in SEQ ID NO.
10.
2. The marker combination according to claim 1, characterized in that The marker combination is composed of a first marker group and a second marker group; The first marker combination is at least one of Gene4, Gene6, Gene8 and Gene9; The second marker group is at least one of Gene1, Gene2, Gene3, Gene5, Gene7 and Gene10.
3. The marker combination according to claim 1, characterized in that: The marker combination is selected from any one of the following combinations: Gene4 and Gene1, Gene4 and Gene2, Gene4 and Gene3, Gene4 and Gene5, Gene4 and Gene7, Gene4 and Gene10, Gene6 and Gene1, Gene6 and Gene2, Gene6 and Gene3, Gene6 and Gene5, Gene6 and Gene7, Gene6 and Gene10, Gene8 and Gene1, Gene8 and Gene2, Gene8 and Gene3, Gene8 and Gene5, Gene8 and Gene7, Gene8 and Gene10, Gene9 and Gene1, Gene9 and Gene2, Gene9 and Gene3, Gene9 and Gene5, Gene9 and Gene7, Gene9 and Gene10.
4. The marker combination according to claim 1, characterized in that The marker combination is selected from any one of the following combinations: A: Gene4, Gene6, Gene8, Gene9 and Gene1; B: Gene4, Gene6, Gene8, Gene9 and Gene2; C: Gene4, Gene6, Gene8, Gene9 and Gene3; D: Gene4, Gene6, Gene8, Gene9 and Gene5; E: Gene4, Gene6, Gene8, Gene9 and Gene7; F: Gene4, Gene6, Gene8, Gene9 and Gene10.
5. A kit, characterized in that: Comprising the marker combination as described in any one of claims 1-4.
6. The kit according to claim 5, characterized in that include: A primer pair and a probe for detecting the marker combination, and a gene chip containing the marker combination.
7. Use of the marker combination according to any one of claims 1 to 4 in the preparation of a product for detecting the invasiveness of Streptococcus pneumoniae serotype 19F; The product for detecting the invasiveness of Streptococcus pneumoniae serotype 19F includes at least one of a kit, a diagnostic reagent combination and a detection device.
8. A method for predicting the risk of serotype 19F invasion, characterized in that: include: Obtaining a test result of a target sample obtained after testing with the kit according to any one of claims 5 to 6; and determining whether the marker combination in the kit exists in the test result; If yes, the prediction result of the target sample is determined to be positive for serotype 19F invasion risk; If not, the predicted result of the target sample is determined to be negative for the risk of serotype 19F invasion.
9. The method for predicting the risk of serotype 19F invasion according to claim 8, characterized in that: The determining whether the detection result contains the marker combination comprises: The test results are used as input, and data analysis is performed through the trained prediction model to obtain the invasive risk prediction result of the 19F strain corresponding to the target sample.
10. The method for predicting the risk of serotype 19F invasion according to claim 9, characterized in that: The method for constructing the trained prediction model includes: According to the training samples in the sample pool, a training set and a validation set are established; wherein each of the training samples includes marker detection data and invasion detection data; constructing the prediction model; The prediction model is trained using the training set; including: extracting features from each training sample in the training set to obtain the presence or absence results of each marker corresponding to the marker detection data corresponding to the training sample and the labeling results corresponding to the invasion detection data; and training the prediction model based on the marker presence or absence results and the labeling results using a machine learning algorithm; Using a test machine to evaluate the prediction model to obtain an evaluation result; the evaluation index includes at least one of accuracy, recall, precision and AUC value; The prediction model is repeatedly trained and adjusted according to the evaluation results to obtain the trained prediction model.
Citation Information
Cited By
Clinical microbial infection intelligent diagnosis system and method based on multi-omics data fusion
CN120913824A
Intelligent diagnosis system and method for clinical microbial infection based on multi-omics data fusion
CN120913824B