Method for predicting invasiveness risk of 19F streptococcus pneumoniae
By evaluating the gene expression levels of serotype 19F Streptococcus pneumoniae using a combination of biomarkers and incorporating a risk prediction model, this approach addresses the limitation of existing technologies in directly determining invasiveness. It achieves highly sensitive and specific invasiveness assessment, making it suitable for clinical diagnosis and treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN CHILDRENS HOSPITAL
- Filing Date
- 2025-04-18
- Publication Date
- 2026-05-19
AI Technical Summary
Existing pneumococcal typing technology cannot directly determine whether the same serotype is invasive, and it fails to make full use of gene expression information for invasiveness assessment, making it difficult to achieve early warning and precise intervention.
An assessment of invasiveness risk was conducted using a combination of evaluation biomarkers, including translational genomes (rluB and rsmE), cellular synthetic genomes (mscL and murI), and ligand transport metabolism genomes (czcD, sitB, mtsC, and pstS), by detecting the expression levels of these genes and combining them with a risk prediction model.
It achieves high sensitivity and specificity in assessing the invasiveness of serotype 19F Streptococcus pneumoniae, supports high-throughput detection and automated analysis, provides a reliable basis for clinical diagnosis and treatment, and promotes the understanding of pathogenic mechanisms.
Smart Images

Figure CN122060879A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biology technology, and more specifically, to a method for predicting the invasive risk of 19F Streptococcus pneumoniae. Background Technology
[0002] Streptococcus pneumoniae is an important pathogenic bacterium. Among its many serotypes, serotype 19F is considered one of the most invasive, causing serious invasive infections, especially in children, such as meningitis and bacteremia. The pathogenicity and invasiveness of Streptococcus pneumoniae are closely related to its surface capsular polysaccharides, cell wall components, and various virulence factors. In clinical practice, predicting the invasive risk of Streptococcus pneumoniae in advance is crucial for preventing infection and improving treatment outcomes. However, current detection technologies still face many challenges in this area.
[0003] Current Streptococcus pneumoniae typing techniques primarily rely on methods such as mass spectrometry (MALDI-TOF) and nuclear magnetic resonance (NMR) detection of capsular polysaccharides, and PCR detection. These techniques can effectively distinguish different serotypes of Streptococcus pneumoniae, with more than ninety serotypes identified. However, these traditional methods have significant limitations in assessing the invasiveness of strains. They mainly focus on the phenotypic characteristics of the strain, such as the structure and composition of capsular polysaccharides, and cannot directly determine whether the same serotype (such as 19F) will cause invasive infection. In addition, these methods usually require complex experimental procedures and long detection times, making it difficult to meet the needs of rapid clinical diagnosis.
[0004] Invasive phenotypes are closely related to pathogen adaptability, adhesion, immune evasion, and stress resistance. Studies have shown significant differences in gene expression levels between invasive and non-invasive 19F strains, with these differentially expressed genes involved in multiple functions, including transcription and translation, biofilm formation, metabolic regulation, and ion pumps. However, existing detection technologies fail to fully utilize this gene expression information for invasiveness assessment. Traditional molecular typing methods cannot provide direct evidence of a strain's invasive potential, making early warning and precise intervention for invasive infections difficult in clinical practice. Furthermore, current technologies lack a comprehensive detection method capable of assessing the expression levels of multiple functional genes, thus failing to fully reflect the strain's biological characteristics and invasive risk.
[0005] In summary, existing technologies have several shortcomings in detecting the invasiveness of Streptococcus pneumoniae. On the one hand, traditional typing methods cannot directly determine whether a single serotype is invasive; on the other hand, current technologies fail to fully utilize gene expression information for invasiveness assessment, lacking a rapid, accurate, and comprehensive detection method. These problems limit the early diagnosis and precise treatment of invasive Streptococcus pneumoniae infections in clinical practice, necessitating the development of a new technology capable of assessing invasiveness risk based on gene expression levels to meet the needs of clinical and public health fields.
[0006] In view of this, the present invention is hereby proposed. Summary of the Invention
[0007] The purpose of this invention is to provide a method for predicting the invasive risk of serotype 19F Streptococcus pneumoniae. The proposed combination of assessment biomarkers can accurately assess the invasive risk of serotype 19F Streptococcus pneumoniae, enabling molecular typing and risk prediction. It possesses high sensitivity and specificity, is suitable for high-throughput detection, and contributes to clinical diagnosis and pathogenic mechanism research.
[0008] In order to achieve the above-mentioned objectives of the present invention, the following technical solution is adopted:
[0009] In a first aspect, the present invention provides an evaluation biomarker combination, comprising: a translational genome, a cellular synthetic genome, and a ligand transport metabolism genome;
[0010] The translational genome includes rluB and rsmE;
[0011] The cellular synthetic genome includes mscL and murI;
[0012] The ligand transport metabolism genome includes czcD, sitB, mtsC, and pstS.
[0013] Secondly, the present invention provides a detection reagent comprising a combination of evaluation markers as described in the foregoing embodiments.
[0014] In an optional implementation, the detection reagent may be any one or more of primer pairs, probes, and chips.
[0015] In an optional implementation, the primer pair includes:
[0016] Primer pairs targeting the rluB gene: the amino acid sequence of the forward primer is shown in SEQ ID NO.1; the amino acid sequence of the reverse primer is shown in SEQ ID NO.2;
[0017] Primer pairs targeting the rsmE gene: the nucleotide sequence of the forward primer is shown in SEQ ID NO.3; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.4;
[0018] Primer pairs targeting the mscL gene: the nucleotide sequence of the forward primer is shown in SEQ ID NO.5; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.6;
[0019] Primer pairs targeting the murI gene: the nucleotide sequence of the forward primer is shown in SEQ ID NO.7; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.8;
[0020] Primer pairs for the czcD gene: the nucleotide sequence of the forward primer is shown in SEQ ID NO.9; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.10;
[0021] Primer pairs targeting the gene sitB: the nucleotide sequence of the forward primer is shown in SEQ ID NO.11; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.12;
[0022] Primer pairs for the mtsC gene: the nucleotide sequence of the forward primer is shown in SEQ ID NO.13; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.14;
[0023] Primer pairs for the pstS gene: the nucleotide sequence of the forward primer is shown in SEQ ID NO.15; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.16.
[0024] Thirdly, the present invention provides a kit comprising the detection reagents as described in any of the foregoing embodiments.
[0025] Fourthly, the present invention provides a method for detecting the expression level of a biomarker, comprising:
[0026] Using the evaluation biomarker combination described in the foregoing embodiments as biomarkers, the expression level of the evaluation biomarker combination of the target sample is detected to obtain the expression level detection result corresponding to the target sample.
[0027] Fifthly, the present invention provides a method for predicting the invasive risk of Streptococcus pneumoniae 19F, comprising:
[0028] Obtain the expression level detection results as described in the foregoing embodiments;
[0029] Using the expression level detection result as input, a trained risk prediction model is used to make a prediction, and the prediction result corresponding to the target sample is obtained.
[0030] In a sixth aspect, the present invention provides the application of the evaluation biomarker combination as described in the foregoing embodiments in serotype 19F molecular typing and / or invasiveness assessment.
[0031] In an optional implementation, the invasiveness assessment includes an assessment of the invasiveness risk of serotype 19F after colonization of the respiratory tract.
[0032] In an optional embodiment, the typing includes: molecular typing of invasive and non-invasive serotype 19F strain.
[0033] In an optional embodiment, the serotype 19F molecule is the pneumococcal serotype 19F molecule.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] This invention provides a method for predicting the invasive risk of serotype 19F Streptococcus pneumoniae. The assessment biomarker set, through comprehensive analysis of the expression levels of specific genes in the translational genome (including rluB and rsmE), cellular synthetic genome (including mscL and murI), and ligand transport metabolic genome (including czcD, sitB, mtsC, and pstS), can more accurately assess the invasive risk of serotype 19F Streptococcus pneumoniae. Compared with traditional detection methods, it not only improves the sensitivity and specificity of detection but also combines molecular typing with invasiveness prediction, providing a more reliable basis for clinical diagnosis and treatment. Furthermore, this gene set supports high-throughput detection and automated analysis, making it suitable for large-scale clinical applications and public health surveillance. It also contributes to a deeper understanding of the pathogenic mechanisms of Streptococcus pneumoniae, providing a theoretical basis for developing new therapeutic targets and prevention strategies. Attached Figure Description
[0036] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0037] Figure 1 This is a heatmap of overall transcriptome expression levels in the embodiments of this application;
[0038] Figure 2 This is a schematic diagram of the COG with the highest enrichment level in the embodiments of this application;
[0039] Figure 3 This is a box plot of gene scores for each gene in two groups of samples in the embodiments of this application;
[0040] Figure 4 This is a graph showing the qPCR ratios of each gene between the invasive and non-invasive groups in the embodiments of this application.
[0041] Figure 5 This is a graph showing the AUC results of gene combinations predicting invasiveness in the embodiments of this application. Detailed Implementation
[0042] The embodiments of the present invention will be described in detail below with reference to examples. However, those skilled in the art will understand that the following examples are for illustrative purposes only and should not be considered as limiting the scope of the invention. Unless otherwise specified in the examples, conventional conditions or conditions recommended by the manufacturer are followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products.
[0043] In this embodiment of the application, an evaluation biomarker combination is provided, including: translational genome, cellular synthetic genome, and ligand transport metabolism genome. See Table 1 for details.
[0044] Table 1. Combination of evaluation markers
[0045] NO. Genome classification Gene name 1 Translational genome rluB、rsmE 2 Cellular Synthesis Genome mscL、murI 3 Ligand transport metabolism genome czcD、sitB、mtsC、pstS
[0046] In this embodiment, by comprehensively analyzing the expression levels of specific genes in the translational genome, cellular synthetic genome, and ligand transport metabolic genome, the invasive risk of serotype 19F Streptococcus pneumoniae can be assessed more accurately. Compared with traditional phenotype-based detection methods, this gene set assessment method can directly reflect the biological characteristics and potential pathogenicity of the strain, thus providing a more reliable basis for clinical diagnosis.
[0047] Based on a combination of assessment biomarkers, molecular typing and invasiveness prediction can be combined. This gene set can not only distinguish between invasive and non-invasive strains, but also perform molecular typing of serotype 19F Streptococcus pneumoniae. This combination allows clinicians to predict the invasiveness risk of strains before infection occurs, thereby taking targeted preventive measures to reduce the occurrence of invasive infections.
[0048] Furthermore, by detecting the expression levels of these genes, the sensitivity and specificity of detection can be improved, enabling highly sensitive and specific assessment of the invasive risk of serotype 19F Streptococcus pneumoniae. For example, qPCR validation results showed that these genes exhibited significant expression differences between invasive and non-invasive strains, effectively distinguishing the two types of strains.
[0049] Supporting high-throughput detection and automated analysis, the gene set assessment method can be combined with high-throughput detection technologies such as microarrays and probes, enabling rapid detection and automated analysis of large numbers of samples. This not only improves detection efficiency but also reduces human error, making it suitable for large-scale clinical applications and public health surveillance.
[0050] Providing auxiliary guidance for clinical treatment: By assessing the risk of invasiveness in advance, clinicians can more rationally choose treatment options, such as whether antibiotics are needed or whether enhanced monitoring is required. This gene expression-based assessment method provides more precise auxiliary guidance for clinical treatment, helping to improve treatment outcomes.
[0051] The analysis of this gene set contributes to a deeper understanding of the pathogenic mechanisms of serotype 19F Streptococcus pneumoniae. By studying the function and expression regulation of these genes, we can reveal the mechanisms by which the strain adapts to the host environment and evades the immune system, providing a theoretical basis for developing new therapeutic targets and preventive strategies.
[0052] In summary, this combination of assessment biomarkers provides a more precise, efficient, and comprehensive tool for assessing the invasiveness and molecular typing of serotype 19F Streptococcus pneumoniae by comprehensively analyzing the expression levels of multiple functional genomes, and has significant clinical and research value.
[0053] In this embodiment, a detection reagent is also provided, comprising the combination of evaluation markers as described in the foregoing embodiments.
[0054] In some embodiments, the detection reagents include any one or more of primer pairs, probes, and chips.
[0055] Furthermore, the primer pair includes:
[0056] Table 2. Primer pairs
[0057]
[0058] In this application embodiment, a reagent kit is provided, including the detection reagents as described in any of the foregoing embodiments.
[0059] It should be noted that the kit may include, but is not limited to, the following components:
[0060] (1) Primer pairs, used for specific amplification of target genes (such as rluB, rsmE, mscL, murI, czcD, sitB, mtsC, and pstS) in the biomarker combination. The composition may include: each primer pair includes a forward primer and a reverse primer, designed separately for the above genes.
[0061] (2) Probes, used to detect the expression level of target genes, are usually combined with fluorescent labels for quantitative PCR (qPCR) or other molecular hybridization techniques. Their composition may include fluorescent probes designed for the genes mentioned above.
[0062] (3) A chip for high-throughput detection of the expression levels of multiple genes, suitable for simultaneous detection of multiple samples. Its components may include probes for the aforementioned genes immobilized on the chip, enabling simultaneous detection of gene expression in the translational genome, cellular synthetic genome, and ligand transport metabolic genome.
[0063] (4) Other auxiliary reagents, such as nucleic acid extraction reagents, used to extract total RNA or DNA from samples; reverse transcription reagents, used to reverse transcribe RNA into cDNA for qPCR detection; qPCR reaction system, including SYBR Green or TaqMan fluorescent probes, dNTPs, MgCl2, reaction buffer, etc.; standards, used to calibrate and verify the accuracy of the detection system; negative controls, used to exclude non-specific reactions; positive controls, used to verify the effectiveness of the detection system.
[0064] (5) It may also include instruction manuals, consumables, etc.
[0065] In summary, this kit can contain a variety of detection reagents, including primer pairs, probes, and chips, comprehensively covering the detection needs of translational genomics, cellular synthetic genomics, and ligand transport metabolism genomics. It is suitable not only for high-throughput laboratory detection but also for rapid clinical diagnosis, providing a comprehensive solution for assessing the invasiveness of serotype 19F Streptococcus pneumoniae.
[0066] In this application embodiment, a method for detecting the expression level of a biomarker for non-diagnostic purposes is provided, comprising:
[0067] Using the evaluation biomarker combination described in the foregoing embodiments as biomarkers, the expression level of the evaluation biomarker combination of the target sample is detected to obtain the expression level detection result corresponding to the target sample.
[0068] It should be noted that the biomarker expression level detection method provided in this embodiment is explicitly limited to non-diagnostic purposes. This method assesses the expression levels of a combination of biomarkers and can only obtain expression data for specific biomarkers in the target sample. This data is only used as research information to illustrate the human condition and its relevance. This technology cannot directly draw diagnostic conclusions related to human health or disease. Actual diagnosis requires a comprehensive evaluation using other detection methods, relevant evidence, and the experience and judgment of professionals. The detection results of this technology can serve as one of the auxiliary reference data to support a comprehensive assessment.
[0069] The detection method uses a combination of biomarkers as markers to detect their expression levels, and the target sample is the biochemical sample to be tested. This sample can be obtained from the subject, such as sputum, blood, or cerebrospinal fluid. The specific detection method may include the following procedure:
[0070] (1) Sample preparation:
[0071] Samples are processed according to standard operating procedures to extract total RNA or DNA. For example, RNA is extracted using a commercially available nucleic acid extraction kit and converted to cDNA using reverse transcriptase for subsequent detection.
[0072] (2) Detection of gene expression levels:
[0073] Choose the appropriate detection technology based on the experimental requirements, such as qPCR or microarray hybridization.
[0074] qPCR detection: Reaction system preparation: Add cDNA template, specific primer pairs, fluorescent probe (if used), and qPCR reaction mixture (such as SYBR Green or TaqMan reagent) to the reaction system. Reaction conditions: Perform the amplification reaction on a qPCR instrument, which typically includes initial denaturation (95℃), cycling reaction (95℃ denaturation, 55-65℃ annealing, 72℃ extension), etc.
[0075] Data analysis: The relative expression level of each gene was calculated based on the Ct value (cycle threshold).
[0076] Microarray hybridization detection: Sample labeling: cDNA samples are labeled with fluorescent dyes. Hybridization reaction: The labeled samples are hybridized with probes immobilized on the microarray. Signal detection: The fluorescence signal is detected using a microarray scanner, and gene expression levels are analyzed.
[0077] In this embodiment of the application, a method for predicting the invasive risk of 19F Streptococcus pneumoniae is also provided, including:
[0078] Obtain the expression level detection results as described in the foregoing embodiments;
[0079] Using the expression level detection result as input, a trained risk prediction model is used to make a prediction, and the prediction result corresponding to the target sample is obtained.
[0080] The aforementioned risk prediction model can obtain prediction results by inputting the expression levels of biomarkers from samples in the test set into a pre-constructed risk prediction model; and then iteratively optimize the model using a validation set to obtain a trained risk prediction model. The predictive ability of different gene combinations can be evaluated using AUC values.
[0081] In some implementations, the number of training samples in the test and validation sets is greater than or equal to 10. Specifically, this number can be any one or a range between any two of the following: 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, and 1000.
[0082] In some implementations, the risk prediction model may be selected from any one or a combination of several of the following algorithms: support vector machine, decision tree, random forest, logistic regression, Bayesian algorithm, K-nearest neighbors, K-means algorithm, Markov algorithm, and regression ridge algorithm.
[0083] In this application embodiment, an application of the evaluation biomarker combination as described in the foregoing embodiments is provided in serotype 19F molecular typing and / or invasiveness assessment.
[0084] In an optional implementation, the invasiveness assessment includes an assessment of the invasiveness risk of serotype 19F after colonization of the respiratory tract.
[0085] In an optional embodiment, the typing includes: molecular typing of invasive and non-invasive serotype 19F strain.
[0086] In an optional embodiment, the serotype 19F molecule is the pneumococcal serotype 19F molecule.
[0087] Example 1
[0088] In this embodiment, the gene expression levels of the target sample biomarkers were detected.
[0089] Experimental methods:
[0090] (1) Selection of implementation targets:
[0091] All 19F strains were isolated from inpatients in the Department of Respiratory Medicine at Shenzhen Children's Hospital.
[0092] A. Three invasive Streptococcus pneumoniae 19F target samples: including two blood-derived 19F strains and one cerebrospinal fluid-derived 19F strain;
[0093] B. Three non-invasive Streptococcus pneumoniae 19F target samples: including three sputum-derived 19F strains and one bronchoalveolar lavage fluid-derived 19F strain.
[0094] (2) Identification of differentially expressed genes:
[0095] To screen for differentially expressed genes, Bowtie2 was used to align RNASeq data, and then Deseq2 (R package) was used to perform differential analysis on gene expression matrix data of invasive and non-invasive 19F strains. FoldChange was set to 2, and p-value to 0.05 was used as the significance threshold.
[0096] (3) Perform functional enrichment analysis on differentially expressed genes.
[0097] The selected differentially expressed gene sequences were compared with the COG (Cluster of Orthologous Groups) database, and each differentially expressed gene was classified and annotated using BLAST or Diamond tools.
[0098] Then, based on the COG annotation results, the differentially expressed genes were assigned to different functional categories, and enrichment analysis was performed on different COG functional categories using Fisher's exact test or hypergeometric distribution test to compare which COG categories in the differentially expressed gene set showed a significant bias towards the random distribution of the background gene set.
[0099] (4) Calculation of the gene set module score for each sample
[0100] For each gene involved in each functional classification, a z-score transformation was performed on the log2(CPM+1) normalized matrix, and the gene score of each sample was calculated as the z-score value of all involved genes. A t-test was used to compare module scores from different samples, with a p-value of 0.05 as the threshold for statistical significance.
[0101] (5) qPCR detection
[0102] Twenty invasive and twenty non-invasive bacterial strains were selected for qPCR validation of each differentially expressed gene. Total RNA was extracted using a bacterial total RNA extraction kit, and the OD260 / 280 and OD260 / 230 ratios were assessed using Nanodrop.
[0103] cDNA was synthesized using a reverse transcription kit, followed by qPCR detection. The qPCR reaction system consisted of SYBR Green or TaqMan fluorescent probes, specific primers, template cDNA, and nuclease-free water to make up the reaction volume.
[0104] qPCR reaction program: initial denaturation at 95℃ for 2 minutes, followed by 40 cycles of reaction (95℃ for 10 seconds, 55-65℃ for 20 seconds, and 72℃ for 20 seconds).
[0105] For each gene, qPCR detection was performed in triplicate to obtain the Ct value for each gene and calculate... ΔΔCt value (comparing invasive and non-invasive strains), expressed as 2^ (-ΔΔCt) This indicates the relative expression level of each detected gene in the invasive group.
[0106] Example 2
[0107] In this embodiment, the risk prediction model was constructed, trained, and used for prediction.
[0108] Experimental methods:
[0109] Step 1: Data Preparation
[0110] (1) Loading data:
[0111] Load the DataFrame qpcr_data containing qPCR data, where each row represents a sample, each column represents the expression level of a gene (2-ΔCt value), and includes a label column Group indicating whether the sample is invasive or non-invasive.
[0112] (2) Data preprocessing: Separating features and labels from the data:
[0113] X: Contains data on all gene expression levels, i.e., all columns in qpcr_data except for the Group column.
[0114] y: The label of the sample, mapping the text labels (Invasive and Non-invasive) in the Group column to numerical values (1 and 0).
[0115] Step 2: Data Segmentation
[0116] (3) Split the dataset: Use the train_test_split function to split the dataset into a training set and a test set, with the test set accounting for 30% of the total data. Set the random seed random_state=10 to ensure that the results are reproducible.
[0117] The segmented dataset includes: X_train: training set features; X_test: test set features; y_train: training set labels; y_test: test set labels.
[0118] Step 3: Model Training
[0119] (4) Model selection: Logistic Regression is selected as the classification model, which is suitable for binary classification problems.
[0120] (5) Training the model: The model is trained using the training set data X_train and the corresponding label y_train.
[0121] Step 4: Model Evaluation
[0122] (6) Prediction results: Predict the training set and the test set to obtain y_train_pred and y_test_pred respectively.
[0123] (7) Calculate accuracy: Calculate the accuracy of the model on the training set and the test set, and print the results.
[0124] (8) Generate confusion matrix: Generate confusion matrix of test set to evaluate model performance.
[0125] Step 5: Plot the ROC curve
[0126] (9) Calculate the ROC curve parameters:
[0127] Use the `predict_proba` method to obtain the probability that each sample in the test set belongs to class 1. Use the `roc_curve` function to calculate the false positive rate (FPR) and the true positive rate (TPR), as well as the corresponding thresholds. Use the `roc_auc_score` function to calculate the AUC (Area Under Curve).
[0128] (10) Plot the ROC curve: Use matplotlib to plot the ROC curve, label the AUC value, and add a 45-degree dashed line as a reference.
[0129] Step 6: Gene Combination Assessment
[0130] (11) Gene combination evaluation: Iterate through all gene pairwise combinations to construct a dataset containing two genes. Use cross-validation (cross_val_score) to calculate the AUC value of each gene combination and store the results in auc_scores.
[0131] Experimental results:
[0132] (1) The results showed that there were significant differences in the transcriptome characteristics of the three invasive strains of strain 19F. Clustering and heatmap analysis of all genes revealed ( Figure 1 The expression patterns of invasive and non-invasive bacterial strains in the transcriptome expression profiles are significantly differentiated. These differentially expressed genes have a wide range of functions, including ribosome structure and biosynthesis; replication, recombination and repair; cell wall / membrane / enveloping biosynthesis; ligand transport and metabolism, etc., which are closely related to bacterial environmental adaptation, metabolic homeostasis, and bacterial virulence.
[0133] (2) Functional enrichment of differentially expressed genes and screening of biomarkers
[0134] Through functional enrichment, differentially expressed genes were classified into 26 COG classes in this embodiment. Statistical analysis revealed that the three most enriched genes in invasive strains were M: cell wall / membrane / enveloping biogenesis, J: translation, ribosome structure and biosynthesis, and P: inorganic ion transport. Figure 2 ).
[0135] The genes upregulated in each COG category include those shown in Table 3: rluB, rsmE, mscL, murI, czcD, sitB, mtsC, and pstS. The selection criteria for this gene series were: a fold change in expression level greater than 2-fold, a p-value less than 0.05, and a clearly defined gene function.
[0136] Table 3. List of Selected Differential Genes
[0137] Gene ID logFC p-value COG Genetic tags L3475_08565 1.1563 0.0085 J rluB L3475_08135 1.0795 0.0085 J rsmE L3475_05665 1.0270 0.0219 M mscL L3475_08600 1.0140 0.0178 M murI L3475_08485 1.8542 0.0000 P czcD L3475_07580 1.4183 0.0266 P sitB L3475_07585 1.3970 0.0283 P mtsC L3475_10025 1.1509 0.0006 P pstS
[0138] Table 2 shows that these genes play crucial roles in bacterial physiology, with the following specific functions: rluB belongs to the RsuA family of pseudouracil synthases and participates in the modification of ribosomal RNA, which is essential for the stability of ribosome structure and function. rsmE specifically methylates the N3 position of the uracil ring in uracil 1498 of 16S rRNA, forming m3U1498. This modification occurs on the fully assembled 30S ribosomal subunit and has a significant impact on ribosome biosynthesis and function. mscL encodes a mechanosensitive channel protein that opens in response to stretching forces in the lipid bilayer and may participate in the regulation of intracellular osmotic pressure changes, helping bacteria adapt to different environmental conditions. murI provides (R)-glutamate for cell wall biosynthesis and is a key enzyme in the cell wall synthesis process, playing an important role in maintaining the structure and integrity of the cell wall. czcD belongs to the cation efflux family and participates in the transport and expulsion of intracellular cations, helping to maintain intracellular ion balance and is of great significance for bacterial growth and adaptation. sitB encodes an ABC transporter involved in the transport of specific substances, playing a crucial role in bacterial nutrient acquisition and maintaining intracellular homeostasis. mtsC belongs to the ABC3 transporter family and participates in substance transport processes, influencing bacterial metabolism and adaptation. pstS is part of the ABC transporter complex PstSACB, which is involved in phosphate import and is responsible for phosphate uptake, essential for bacterial metabolism and growth.
[0139] III. Using qPCR to obtain the relative expression levels of each gene in the two strains
[0140] Universal primers were designed for each strain and gene (Table 4), and then qPCR was performed on 20 invasive and 20 non-invasive strains of strain 19F. The results showed that each gene was significantly overexpressed in both groups of samples, consistent with the transcriptome results. Figure 3 , Figure 4 ).
[0141] Table 4. Primer pairs for each gene
[0142]
[0143] IV. Model Construction Using qPCR Results
[0144] The relative expression levels of rluB, rsmE, mscL, murI, czcD, sitB, mtsC, and pstS in 20 samples were measured by qPCR (2 –ΔCt Using these as input values to predict strain invasiveness, a mathematical model is constructed with known invasiveness tags. The model is then built using various gene combinations, and its overall predictive performance is evaluated using ROC. Figure 5 The study found that the combination of murI and sitB predicted an AUC of 0.80, while the combination of pstS and each gene predicted an AUC of over 0.70.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A combination of evaluation markers, characterized in that, include: Translational genome, cellular synthetic genome, and ligand transport metabolic genome; The translational genome includes rluB and rsmE; The cellular synthetic genome includes mscL and murI; The ligand transport metabolism genome includes czcD, sitB, mtsC, and pstS.
2. A detection reagent, characterized in that, The detection reagent includes the combination of evaluation markers as described in claim 1.
3. The detection reagent as described in claim 2, characterized in that, The detection reagents include any one or more of primer pairs, probes, and chips; Preferably, the primer pair comprises: Primer pairs targeting the rluB gene: the amino acid sequence of the forward primer is shown in SEQ ID NO.1; the amino acid sequence of the reverse primer is shown in SEQ ID NO.2; Primer pairs targeting the rsmE gene: the nucleotide sequence of the forward primer is shown in SEQ ID NO.3; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.4; Primer pairs targeting the mscL gene: the nucleotide sequence of the forward primer is shown in SEQ ID NO.5; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.6; Primer pairs targeting the murI gene: the nucleotide sequence of the forward primer is shown in SEQ ID NO.7; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.8; Primer pairs for the czcD gene: the nucleotide sequence of the forward primer is shown in SEQ ID NO.9; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.10; Primer pairs targeting the gene sitB: the nucleotide sequence of the forward primer is shown in SEQ ID NO.11; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.12; Primer pairs for the mtsC gene: the nucleotide sequence of the forward primer is shown in SEQ ID NO.13; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.14; Primer pairs for the pstS gene: the nucleotide sequence of the forward primer is shown in SEQ ID NO.15; the nucleotide sequence of the reverse primer is shown in SEQ ID NO.
16.
4. A reagent kit, characterized in that, Includes the detection reagents as described in any one of claims 2-3.
5. A method for detecting the expression level of a biomarker, characterized in that, include: Using the evaluation biomarker combination described in claim 1 as a biomarker, the expression level of the evaluation biomarker combination of the target sample is detected to obtain the expression level detection result corresponding to the target sample.
6. A method for predicting the invasive risk of 19F Streptococcus pneumoniae, characterized in that, include: Obtain the expression level detection results as described in claim 5; Using the expression level detection result as input, a trained risk prediction model is used to make a prediction, and the prediction result corresponding to the target sample is obtained.
7. The use of the assessment biomarker combination as described in claim 1 in serotype 19F molecular typing and / or invasiveness assessment.
8. The application as described in claim 7, characterized in that, The invasiveness assessment includes an assessment of the invasiveness risk of serotype 19F after colonization of the respiratory tract.
9. The application as described in claim 7, characterized in that, The typing includes molecular typing of invasive and non-invasive serotype 19F strains.
10. The application as described in claim 7, characterized in that, The serotype 19F molecule is the serotype 19F molecule of Streptococcus pneumoniae.