Biomarker for individual recognition and application thereof

Through high-resolution analysis of skin microbial markers, using nine microorganisms such as Propionibacterium acnes and a random forest model, the problem of individual identification in forensic medicine was solved, and individual identification with high accuracy and sensitivity was achieved.

CN120683284APending Publication Date: 2025-09-23SOUTHERN MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510787619.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In traditional forensic medicine, biological samples left at the scene may not contain human DNA or the DNA may be degraded and difficult to detect, making individual identification difficult, and existing technologies cannot meet the needs.

Method used

Skin microbial markers, especially nine microorganisms such as Propionibacterium acnes, were used for amplification and sequencing using a 16S rDNA hypervariable region primer set, combined with a random forest model for individual identification, and high-resolution analysis was performed using species-level information of skin bacteria.

Benefits of technology

It achieved high accuracy and sensitivity in individual identification, with an accuracy rate of 88.12%, and provided rich forensic evidence for individual identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005447471570000042
    Figure BDA0005447471570000042
  • Figure HDA0005447471580000011
    Figure HDA0005447471580000011
  • Figure HDA0005447471580000012
    Figure HDA0005447471580000012
Patent Text Reader

Abstract

The invention relates to a biomarker for individual recognition and application thereof, and relates to the technical field of biology. The biomarker comprises at least one of the following microorganisms: propionibacterium acnes, staphylococcus epidermidis, marine paracoccus, streptococcus, rhodococcus fence celebrating, staphylococcus, pseudomonas, streptococcus sanguineus, achromobacter xylosoxidans, acinetobacter calcoaceticus, acinetobacter johnsonii, staphylococcus lentus, pseudomonas toralaris and brevundimonas vesicularis. Or enterococcus faecium. The biomarker can provide analysis with higher resolution in the aspect of differential strain information, has a better distinguishing effect, higher accuracy and sensitivity on individual recognition, and can be used as a new technical means for supplementing forensic individual recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, and in particular to a biomarker for individual identification and application thereof. Background Art

[0002] Individual identification is a crucial task in forensic medicine, and determining the source of crime scene biological specimens is a key goal. Traditional genetic markers for individual identification, including STRs, SNPs, and InDels, are molecular markers derived from human cell DNA. However, in actual forensic work, biological specimens left at the scene may not contain human DNA or may be degraded, making it difficult to detect. This limits the application of human DNA and makes it inadequate for individual identification. Therefore, it is necessary to explore non-human DNA biomarkers, such as microbial markers, to meet the needs of forensic applications.

[0003] Skin contact samples are one of the most common biological samples found at the scene. Compared with traditional detection methods, methods for individual identification based on skin microbial markers have the following advantages: (1) The number of skin microorganisms is large, they are easy to fall off, and they are highly resistant to the environment. (2) Sequencing requires less sample and has high sensitivity. (3) Bacterial DNA is more resistant to degradation than human nuclear DNA. (4) Third-generation sequencing technology (single-molecule sequencing technology) can efficiently obtain the full length of bacterial 16S rDNA and reveal the composition of microbial communities at the species level. (5) Skin microbial combinations are individual-specific. Therefore, a method for individual identification based on the analysis of microorganisms carried by the skin as biomarkers is developed to supplement forensic human DNA analysis, provide more clues and information for the case, and thus achieve the purpose of individual identification in forensic identification. Summary of the Invention

[0004] In response to the above problems, the present invention provides a biomarker for individual identification. The biomarker can provide higher-resolution analysis of differential bacterial strain information, has better resolution, higher accuracy and sensitivity for individual identification, and can serve as a new technical means to supplement forensic individual identification.

[0005] To achieve the above objectives, the present invention provides a biomarker for individual identification, comprising at least one of the following microorganisms: Propionibacterium acnes, Staphylococcus epidermidis, Paracoccus marineus, Streptococcus, Rhodococcus fanqingsheng, Staphylococcus, Pseudomonas, Streptococcus sanguinis, Achromobacter xylosoxidans, Acinetobacter calcoaceticus, Acinetobacter johnsonii, Staphylococcus lentus, Pseudomonas toraea, Brevundimonas vesicularis, or Enterococcus faecium.

[0006] In one embodiment, the biomarkers include the following microorganisms: Propionibacterium acnes, Staphylococcus epidermidis, Paracoccus marineus, Streptococcus, Rhodococcus fanqingsheng, Staphylococcus, Pseudomonas, Streptococcus sanguinis, Achromobacter xylosoxidans, Acinetobacter calcoaceticus, Acinetobacter johnsonii, Staphylococcus lentus, Pseudomonas toraea, Brevundimonas vesicularis, and Enterococcus faecium.

[0007] The present invention also provides the use of the biomarker in preparing products for individual identification.

[0008] In one embodiment, the individual identification includes: distinguishing different individuals and identifying the same individual.

[0009] The present invention also provides a primer set for detecting the biomarker, which comprises primers for amplifying the V1-V9 hypervariable region of the biomarker 16S rDNA.

[0010] In one embodiment, the primer set comprises:

[0011] Primer 27F: 5′-AGRGTTYGATYMTGGCTCAG-3′ (SEQ NO. 1);

[0012] Primer 1492R: 5′-RGYTACCTTGTTACGACTT-3′ (SEQ NO. 2).

[0013] The present invention also provides a product for individual identification, which comprises: a reagent for detecting the expression level of the biomarker in a biological sample, or the primer set.

[0014] In one embodiment, the biological sample is derived from skin.

[0015] The present invention also provides a detection method for individual identification, comprising the following steps: extracting DNA from a sample to be tested, using the product for detection, using a species-level taxonomic level analysis method to obtain the relative abundance of the biomarker, and inputting the species type and relative abundance of the biomarker species level as input features into a random forest model to obtain the individual identification result of the sample to be tested.

[0016] In one embodiment, the sample to be tested is derived from skin.

[0017] The above-mentioned products and detection methods for individual identification are mainly based on skin samples. Skin bacteria can easily fall off and transfer when touching or contacting surfaces and objects at the scene during forensic examinations. They are a common and abundant forensic evidence. Therefore, the above-mentioned products and detection methods have broad application prospects in individual identification.

[0018] The present invention also provides a method for screening biomarkers for individual identification, comprising the following steps: extracting total DNA from the microbiome of a biological sample, constructing a gene library of the V1-V9 hypervariable regions of microbial 16S rDNA to obtain a gene library, sequencing the gene library to obtain sequencing data, and analyzing and processing the sequencing data to obtain biomarkers for individual identification.

[0019] In one embodiment, the steps of constructing the V1-V9 hypervariable region gene library of the microbial 16S rDNA include: PCR amplification using specific primers with barcodes, PCR product purification and quantification; library acquisition; library purification; and library quality inspection.

[0020] In one embodiment, the sequencing step is 16S amplicon third-generation sequencing.

[0021] In one embodiment, the analysis and processing steps include: filtering and correcting the sequencing data, performing OUT cluster analysis, species taxonomy annotation, analyzing bacterial species composition and abundance information at the species level, performing ANOSIM analysis, and machine learning analysis.

[0022] In one embodiment, the ANOSIM analysis step includes: analyzing the distribution of intra-individual and inter-individual differences, and a p value less than 0.05 indicates that the inter-group difference is significantly greater than the intra-group difference;

[0023] The machine learning analysis step includes: using the Random Forest algorithm to construct an individual classification prediction model based on species-level skin microbial markers, and evaluating the model performance through feature selection, model training and cross-validation to ensure the accuracy and generalization ability of the model.

[0024] In one embodiment, the step of obtaining biomarkers for individual identification includes: using the species composition and abundance information of species-level microbial markers as input features, constructing multiple decision trees and forming multiple weak classifiers to construct an individual identification prediction model, and obtaining the biomarkers for individual identification through screening through a random forest classification model.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] The present invention provides a biomarker for individual identification and its application. The biomarker can be used to obtain microbial community information of the individual skin microbiome at the species level through third-generation sequencing technology, and higher-resolution analysis can be performed by using differential bacterial species information. The accuracy of the biomarker was verified by a random forest classification model, reaching 88.12%, proving that it has high accuracy and sensitivity. The product and detection method for individual identification of the present invention are mainly based on skin samples. Skin bacteria are easily shed and transferred when touching or contacting surfaces and objects at the scene in forensic examinations. They are a common and abundant forensic evidence. Therefore, the above-mentioned product and detection method have a broad application prospect in individual identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a bar graph of the microbial community composition of different skin sites of different individuals at different time points in Example 1;

[0028] Figure 2 The weighted UniFrac distance distribution diagram between samples of different individuals and the same individual in Example 1;

[0029] Figure 3 This is the importance ranking diagram of microbial markers at the species level of random forest in Example 1;

[0030] Figure 4 Graph showing the ROC curve test results of the random forest prediction classification model in Example 1;

[0031] Figure 5 The confusion matrix results of the correct prediction of the random forest model classification in Example 1 are shown in Figure 1, where (A) is based on the top 10 microorganisms, (B) is based on the top 15 microorganisms, and (C) is based on the top 20 microorganisms;

[0032] Figure 6 This is the confusion matrix result diagram of the correct prediction of the random forest model classification on the new evaluation sample set in Example 2. DETAILED DESCRIPTION

[0033] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. Preferred embodiments of the present invention are shown in the accompanying drawings. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present disclosure.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0035] source:

[0036] Unless otherwise specified, the reagents, materials, and equipment used in this example are all commercially available; and the experimental methods, unless otherwise specified, are all conventional experimental methods in the art.

[0037] Example 1

[0038] Screening and validation of skin microbial markers for individual identification.

[0039] 1. Skin sample collection.

[0040] A total of 160 skin samples were collected using cotton swabs from the fingers, palms, arms, and foreheads of ten volunteers on days 1, 3, 5, and 7. The sample naming convention was: individual number - skin site number - time number, with individual numbers S01-S10. The skin sites were Fi, P, A, and Fo, representing samples from the fingers, palms, arms, and forehead, respectively. Time numbers were 1d, 3d, 5d, and 7d, representing samples collected on days 1, 3, 5, and 7, respectively.

[0041] 2. Sample DNA extraction and quality inspection.

[0042] use Soil DNA kit (Omega Bio-tek, Norcross, GA, US) was used to extract total microbial DNA, and the quality of DNA extraction was tested by 1% agarose gel electrophoresis. NanoDrop 2000 (ThermoScientific, USA) was used to determine DNA concentration and purity.

[0043] 3. Sequencing library construction, purification and quantification.

[0044] The extracted DNA was used as a template to amplify the V1-V9 region of the 16S rRNA gene using the barcoded primers 27F (SEQ NO. 1: 5'-AGRGTTYGATYMTGGCTCAG-3') and 1492R (SEQ NO. 2: 5'-RGYTACCTTGTTACGACTT-3').

[0045] The PCR reaction system is shown in Table 1. Three replicates were performed for each sample. The amplification procedure is shown in Table 2 (PCR instrument: T100 ThermalCycler PCR, USA). After 2% agarose gel electrophoresis, the product was purified by magnetic beads, and the purified product was quantified using Qubit 4.0 (ThermoFisher Scientific, USA). The appropriate proportions of the mixture were then mixed according to the sequencing requirements for each sample.

[0046] Library construction was performed using the SMRTbell prep kit 3.0: ① DNA damage repair; ② end repair; and ③ adapter ligation. Sequencing was performed using the Pacbio Sequel IIe System. HiFi reads were generated from sequencing subreads using the CCS mode of SMRT-Link v11.0 for subsequent data analysis.

[0047] Table 1 PCR reaction system

[0048] Reagent name Dosage 5×FastPfu buffer 4 μL 2.5mM dNTPs 2μL Upstream primer (5 μM) 0.8μL Downstream primer (5 μM) 0.8μL FastPfu polymerase 0.4μL BSA 0.2μL Template DNA 10ng <![CDATA[Add ddH2O]]> Make up to 20 μL

[0049] Table 2 Amplification procedures

[0050]

[0051] 4. Bioinformatics analysis.

[0052] Data from each sample was distinguished based on barcode sequence, and length filtering and orientation correction were performed, retaining sequences between 1000-1800 bp (bacterial). Sequences were clustered into OTUs at 97% similarity using UPARSE 7.1 (http: / / drive5.com / uparse / , version 7.1), and chimeras were removed. Sequences annotated to chloroplasts and mitochondria were removed from all samples. After the number of sequences from all samples was equalized, OTU species taxonomy was annotated using the RDP classifier (http: / / rdp.cme.msu.edu / , version 2.11) against the Silva 16S rRNA gene database (v138) with a 70% confidence threshold. Community composition for each sample was calculated at the kingdom, phylum, class, order, family, genus, and species levels. Sequencing of the V1-V9 regions of the 16S rDNA was performed, and species annotation to the species level was more accurate, so this study was performed at the bacterial species level.

[0053] Based on the analysis results at different taxonomic levels, a species distribution histogram based on relative abundance at the species level was drawn to visually display the composition of the microbial community. The species distribution histogram showed that the composition of the skin microbial communities of the fingers, palms, and arms of the same individual on the same day was highly similar, while the microbial composition of the forehead skin was dominated by Propionibacterium acnes, and the forehead skin microbial composition remained relatively stable over time; the composition of the skin microbial community varied between individuals ( Figure 1 ); similarity analysis (ANOSIM) was used to evaluate the similarities and differences between groups or samples. ANOSIM analysis of the skin microbiome between and within individuals showed that there were significant differences between individuals based on the species-level skin microbiome, which were greater than the intra-individual differences. Samples from the same individual shared their skin microbial communities to a certain extent ( Figure 2 All analyses were performed using the R software package (v4.3.1).

[0054] 5. Random forest model construction.

[0055] According to the random forest method (RF) in the R package RandomForest (v 4.7-1.1), the species and relative abundance data of the skin microbiome at the species level of different individuals were input as input features. 112 skin samples were used as the training set for training to build the model, and 160 samples were used as the test set for validation.

[0056] 6. Screening of differential bacterial marker combinations.

[0057] Based on the random forest classification prediction model constructed above, the importance() function of the randomForest package was used to analyze the importance of each feature in the random forest model. Mean Decrease Gini measures the degree to which a feature reduces impurity when used to split nodes, which helps identify microbial features that contribute to the model. A ten-fold cross-validation was performed. After all ten iterations, the importance scores of each feature were averaged to determine the relative importance of each feature in the entire dataset. The cross-validation curve represents the relationship between model error and the number of features used for fitting, which can be used to screen the number of features required to minimize model error. According to the principle of parsimony, the top 15 species-level microbial marker combinations with the highest contribution were selected as the features with the greatest explanatory or predictive power for distinguishing individuals, including Cutibacterium macnes (Propionibacterium acnes), Staphylococcus epidermidis (Staphylococcus epidermidis), Paracoccus marinus (Paracoccus marinus), unclassified_g__Streptococcus (unclassified Streptococcus), Rhodococcus qingshengii (Rhodococcus qingshengii), Staphylococcus hominis (Staphylococcus hominis), unclassified_g__Pseudomonas (unclassified Pseudomonas), Streptococcus sanguinis (Streptococcus sanguinis), Achromobacter xylosoxidans (Achromobacter xylosoxidans), Acinetobacter calcoaceticus (Acinetobacter johnsonii), Mammaliicoccus lentus (Staphylococcus lentus), Pseudomonas stolaasi (Pseudomonas stolaasii), Brevundimonas vesicularis (Brevundimonas vesicularis), Enterococcus faecium (Enterococcus faecium) ( Figure 3 In addition, to verify that the top 15 species-level microbial marker combinations were the optimal combination, the top 10 and top 20 species-level microbial markers were also included for comparative analysis to comprehensively evaluate the effectiveness and stability of different marker combinations.

[0058] 7. Classification model verification.

[0059] The model was rebuilt using the selected top 15 features, and the classification model was validated using the test set skin samples. The receiver operating characteristic curve (ROC) was used to evaluate the constructed model, and the area under the ROC curve (AUC) was used to evaluate the ROC effect. The results showed that the random forest classification model constructed based on the species-level skin microbiome had an accuracy rate of 88.12% for the correct classification of samples, and the AUC value was 0.791-0.977 ( Figure 4 ), the AUC value reflects the overall performance evaluation indicator of the model in distinguishing different categories. In the field of machine learning and statistics, a model with an AUC value between 0.8-0.9 is considered to be well-performing, while a model close to 1 is considered to be excellent. In addition, confusion matrix analysis was performed on the models built based on the selected Top10, Top15, and Top20 features. The results showed that the accuracy of these models in sample classification was 86.25%, 88.12%, and 85%, respectively. Figure 5 ).

[0060] Example 2

[0061] To evaluate the effectiveness of screened microbial markers for individual identification

[0062] 1. Skin sample collection.

[0063] A total of 60 skin swabs were collected from the heads and faces of 10 new volunteers on days 1, 2, and 3. The sample naming convention was: individual number - skin site number - time number, with individual numbers S01-S10, skin sites "Arm" and "Cheek" representing hand and face samples, respectively, and time numbers "1d", "2d", and "3d" representing skin samples collected on days 1, 2, and 3, respectively.

[0064] 2. DNA extraction, amplification, detection analysis and bioinformatics analysis were carried out according to the detection conditions in Example 1.

[0065] 3. Construction and verification of random forest classification model.

[0066] According to the results of the taxonomic level analysis at the species level, the relative abundance distribution of the 15 microbial markers screened in Example 1 was obtained, including Cutibacterium macnes (Propionibacterium acnes), Staphylococcus epidermidis (Staphylococcus epidermidis), Paracoccus marinus (Marine Paracoccus), unclassified_g__Streptococcus (unclassified Streptococcus), Rhodococcus qingshengii (Fan Qingsheng Rhodococcus), Staphylococcus hominis (human Staphylococcus), unclassified_g__Pseudomonas (unclassified Pseudomonas), Streptococcus sanguinis (Streptococcus sanguinis), Achromobacter xylosoxidans (Xylose oxidizing Colorless Bacterium), Acinetobacter calcoaceticus (Calcium Acinetobacter johnsonii (Johnsonii), Mammaliicoccus lentus (Staphylococcus lentus), Pseudomonas stolaasi (Pseudomonas stolaasi), Brevundimonas vesicularis (Brevundimonas vesicularis), Enterococcus faecium (Enterococcus faecium). According to the random forest method (RF) in the R package RandomForest (v 4.7-1.1), the species and relative abundance data of the skin microbial species level of different individuals were input as input features. 42 skin samples were used as training sets to train and build the model, and 60 samples were used as test sets for validation. The confusion matrix results showed that the random forest classification model constructed based on the species-level skin microbiome had an accuracy rate of 95.00% ( Figure 6 ).

[0067] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0068] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.

Claims

1. A biomarker for individual identification, characterized in that: The method includes at least one of the following microorganisms: Propionibacterium acnes, Staphylococcus epidermidis, Paracoccus marineus, Streptococcus, Rhodococcus fanqingsheng, Staphylococcus, Pseudomonas, Streptococcus sanguinis, Achromobacter xylosoxidans, Acinetobacter calcoaceticus, Acinetobacter johnsonii, Staphylococcus lentus, Pseudomonas toraea, Brevundimonas vesicularis, or Enterococcus faecium.

2. The biomarker according to claim 1, characterized in that The biomarkers include the following microorganisms: Propionibacterium acnes, Staphylococcus epidermidis, Paracoccus marineus, Streptococcus, Rhodococcus fanqingsheng, Staphylococcus, Pseudomonas, Streptococcus sanguinis, Achromobacter xylosoxidans, Acinetobacter calcoaceticus, Acinetobacter johnsonii, Staphylococcus lentus, Pseudomonas toraea, Brevundimonas vesicularis, and Enterococcus faecium.

3. Use of the biomarker according to any one of claims 1 to 2 in the preparation of a product for individual identification.

4. The use according to claim 3, characterized in that The individual identification includes: distinguishing different individuals and identifying the same individual.

5. A primer set for detecting the biomarker according to any one of claims 1-2, characterized in that: The primer set includes primers for amplifying the V1-V9 hypervariable regions of the biomarker 16S rDNA.

6. The primer set according to claim 5, characterized in that The primer set includes: Primer 27F: 5′-AGRGTTYGATYMTGGCTCAG-3′ (SEQ NO. 1); Primer 1492R: 5′-RGYTACCTTGTTACGACTT-3′ (SEQ NO. 2).

7. A product for individual identification, characterized in that: The product comprises: a reagent for detecting the expression level of the biomarker according to any one of claims 1-2 in a biological sample, or a primer set according to any one of claims 5-6.

8. The product according to claim 7, characterized in that The biological sample is derived from skin.

9. A method for detecting individual identification, characterized in that: The following steps are involved: Extract DNA from the sample to be tested, and use the product described in any one of claims 7-8 for detection. Use the species-level taxonomic level analysis method to obtain the relative abundance of the biomarker described in any one of claims 1-2. The species type and relative abundance of the biomarker species level are input as input features into the random forest model to obtain the individual identification result of the sample to be tested.

10. The detection method according to claim 9, characterized in that: The sample to be tested is derived from skin.

Citation Information

Patent Citations

  • Application of bacteria as nasopharynx cancer detection, treatment or prognosis marker

    CN117535416A

  • Methods and systems for analyzing microbiota

    US20210057046A1