Consensus matrix-based female vagina flora structure typing method, device and equipment and storage medium
Through the consensus matrix-based method, multiple sampling and multi-clustering of female vaginal flora data is solved, and the problem of sensitivity to outliers and noise in the prior art is improved, and the robustness and accuracy of the typing results are improved.
Patent Information
- Application Number
- CN202510098070.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is sensitive to outliers and noise in the data when typing female vaginal microbiome structure, resulting in the typing results being not robust and accurate.
A consensus matrix-based method is used to generate different subsets of female vaginal microbial data by sampling multiple times without retrieval, and each data subset is clustered using multiple clustering methods. Finally, the consensus matrix is classified through the hierarchical clustering method to obtain the final CST typing result.
It reduces the impact of data noise on the classification results, reduces the algorithm preference, and improves the robustness and accuracy of female vaginal microbial structure classification.
Smart Images

Figure CN119993279A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of biological detection technology, and in particular to a method, device, equipment and storage medium for typing the structure of female vaginal flora based on a consensus matrix. Background Art
[0002] Community-state types (CSTs) are defined by gene sequencing technology based on the composition and abundance of microorganisms in the female vagina. The vaginal microbial community can be divided into at least five types, of which four are dominated by lactic acid-producing Lactobacillus, and the other is mainly composed of mixed Lactobacillus and a large number of non-lactobacillus anaerobic bacteria. Microbiome studies have shown that the composition and changes of female vaginal flora are closely related to HPV invasion, persistent infection, and the occurrence and development of cervical cancer. Therefore, accurate typing of the female vaginal flora structure can provide a more reliable basis for related medical research and development, scientific research, etc.
[0003] In related technologies, hierarchical clustering methods or K-means clustering methods are usually used to classify female vaginal microbial composition and abundance to determine the type of community state to which it belongs. However, in the process of obtaining data on female vaginal microbial composition and abundance, there are usually some outliers and noise, and these methods are sensitive to outliers and noise in the data, so the results are not robust and have low accuracy.
[0004] Therefore, there is an urgent need to redesign a consensus matrix-based method for classifying the female vaginal flora structure and overcome the above-mentioned defects. Summary of the invention
[0005] The embodiments of the present application provide a method, device, electronic device and storage medium for typing the structure of female vaginal flora based on a consensus matrix, so as to improve the robustness and accuracy of typing the structure of female vaginal flora.
[0006] In a first aspect, the present application provides a method for typing the structure of female vaginal flora based on a consensus matrix, the method comprising:
[0007] DNA was extracted from vaginal secretion swabs from multiple women;
[0008] Using the polymerase chain reaction method, 16s rRNA was amplified for each obtained DNA to obtain amplified DNA;
[0009] Each amplified DNA was subjected to NGS high-throughput sequencing, and the sequencing data of the corresponding amplified DNA was stored in fastq format, and clustered using the DADA2 algorithm of qiime2 software to generate ASV representative sequences.
[0010] The relative abundance matrix data of female vaginal flora is obtained by classifying and annotating each ASV representative sequence through the greengene2 database, wherein the relative abundance matrix data includes the relative abundance of flora of multiple female vaginal flora DNA samples;
[0011] By sampling without replacement, the samples and strains of the relative abundance matrix data are repeatedly sampled according to the set ratio to generate multiple female vaginal flora data subsets. By multiple clustering methods, each female vaginal flora data subset is clustered to obtain multiple clustering results; the consensus matrix method is used to obtain the co-occurrence rate of any two female vaginal flora DNA samples in the relative abundance matrix data in the multiple clustering results; by the hierarchical clustering method, all the co-occurrence rates obtained are clustered to obtain the flora CST typing results, which are used to characterize the microbial community state type to which each female vaginal flora DNA sample belongs after clustering.
[0012] In a second aspect, an embodiment of the present application provides a device for typing the structure of female vaginal flora based on a consensus matrix, the device comprising:
[0013] A sample acquisition unit is used to extract DNA from multiple female vaginal secretion swabs; using the polymerase chain reaction method, 16s rRNA amplification is performed on each obtained DNA to obtain amplified DNA; each amplified DNA is subjected to NGS high-throughput sequencing, and the sequencing data of the corresponding amplified DNA obtained is stored in a fastq format, and clustered by the DADA2 algorithm of qiime2 software to generate ASV representative sequences; each ASV representative sequence is classified and annotated by the greengene2 database to obtain relative abundance matrix data of female vaginal flora, wherein the relative abundance matrix data contains the relative abundance of flora of multiple female vaginal flora DNA samples;
[0014] The typing unit is used to repeatedly sample the samples and strains of the relative abundance matrix data according to a set ratio by a sampling method without replacement to generate multiple female vaginal flora data subsets; cluster each female vaginal flora data subset respectively by multiple clustering methods to obtain multiple clustering results; adopt a consensus matrix method to obtain the co-occurrence rate of any two female vaginal flora DNA samples in the relative abundance matrix data in the multiple clustering results; cluster all the obtained co-occurrence rates by a hierarchical clustering method to obtain a flora CST typing result, and the flora CST typing result is used to characterize the type of microbial community state to which each female vaginal flora DNA sample belongs after clustering processing.
[0015] Optionally, the typing unit is specifically used for:
[0016] Using a sample sampling method, performing n1 samplings without replacement on the relative abundance matrix data to obtain n1 female vaginal flora sample data subsets, wherein the sample sampling method is to randomly extract female vaginal flora DNA sample data from the relative abundance matrix data;
[0017] The relative abundance matrix data were sampled n2 times without replacement to obtain n2 female vaginal flora species data subsets. The species sampling method was to randomly extract the relative abundance data of a certain species in the relative abundance matrix data in all female vaginal flora DNA samples, and n1 and n2 were positive integers.
[0018] Optionally, the number of rows of the relative abundance matrix data is m, and the number of columns is n, indicating that the relative abundance matrix data contains m bacterial species and n sample data, and the sampling ratio method is set to, in each sampling, the sampling ratio of the sample sampling method is 80%, and the sampling ratio of the bacterial species sampling method is 80%, and the typing unit is specifically used for:
[0019] The sample function of R 4.3 software was used, and the parameter replace of the function was set to False. The sample index of random sampling was obtained by setting the parameter size. The sample sampling method without replacement was adopted each time, and the sample ratio was n*80% to obtain the corresponding female vaginal flora sample data subset;
[0020] The typing unit is specifically used for:
[0021] The sample function of R 4.3 software was used, and the function parameter replace was set to False. The parameter size was set to obtain the randomly sampled bacterial species index. The sample sampling method without replacement was adopted each time. According to the strategy of bacterial species ratio of m*80%, the female vaginal flora species data subset was obtained.
[0022] Optionally, the multiple clustering methods include a hierarchical clustering method, a KMeans clustering method and a non-negative matrix decomposition method, the multiple female vaginal flora data subsets include a female vaginal flora sample data subset and a female vaginal flora species data subset, and the multiple clustering methods cluster each female vaginal flora data subset respectively to obtain multiple clustering results, including:
[0023] For each female vaginal flora sample data subset, the flora Pearson correlation distance between any two female vaginal flora sample data in each female vaginal flora sample data subset is calculated by the cor function, and for each female vaginal flora species data subset, the flora Pearson correlation distance between any two relative abundance data of species in each female vaginal flora species data subset is calculated by the cor function;
[0024] For each subset of female vaginal flora sample data, the distances between samples were hierarchically clustered using the hclust function, and the distances between samples were centrally clustered using the kmeans function to obtain two corresponding clustering results;
[0025] And for each subset of female vaginal flora species data, the distance between the relative abundance data of the species is hierarchically clustered using the hclust function, and the distance between the relative abundance data of the species is centrally clustered using the kmeans function to obtain two corresponding clustering results;
[0026] For each subset of female vaginal flora data, matrix decomposition was performed using the NMF R package to obtain the corresponding feature matrix W, and the hclust function was used to cluster the W matrix to obtain the corresponding clustering results.
[0027] Optionally, the typing unit is specifically used for:
[0028] A consensus matrix method is adopted to count the frequencies of any two female vaginal flora sample data in the relative abundance matrix data being assigned to the same cluster in the multiple clustering results, and an n×n consensus matrix C is initialized. The initial value of the consensus matrix C is zero. For any two female vaginal flora sample data (i, j), if i and j are assigned to the same cluster in a certain clustering result, the value of C[i][j] is added by 1 to obtain a first frequency matrix C. The clustering result contains at least one cluster, and the first frequency matrix C contains the co-occurrence frequencies between samples in all clusters.
[0029] Count the frequencies of any two female vaginal flora sample data in the relative abundance matrix data appearing in the same clustering result in the multiple clustering results, initialize an n×n consensus matrix E, the initial value of the consensus matrix E is zero, for any two female vaginal flora sample data (i, j), if i and j are in a certain clustering result, then the value of E[i][j] is added by 1, and a second frequency matrix E is obtained, wherein the second frequency matrix E contains the co-occurrence frequencies between samples in all clustering results;
[0030] Initialize an n×n consensus matrix H, H[i][j] = C / E, H[i][j] represents the proportion of i and j being assigned to the same cluster in all clustering methods.
[0031] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes any one of the methods for classifying the structure of female vaginal flora based on a consensus matrix in the above-mentioned first aspect.
[0032] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program is run on a computer device, the computer program is used to enable the computer device to execute any one of the methods for classifying the structure of female vaginal flora based on a consensus matrix in the above-mentioned first aspect.
[0033] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program, and the computer program is stored in a computer-readable storage medium; when a processor of a computer device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the computer device executes any one of the methods for classifying the structure of female vaginal flora based on a consensus matrix in the above-mentioned first aspect.
[0034] The beneficial effects of this application are as follows:
[0035] The female vaginal flora structure typing method, device, computer equipment and storage medium provided in the embodiments of the present application are for the female vaginal flora types and the relative abundance matrix data of the relative abundance of each species used to characterize the DNA of multiple female vaginal secretion swabs, and multiple clustering methods are used to perform CST typing and analysis of the female vaginal flora structure on the relative abundance matrix data, which can reduce the algorithm preference and improve the robustness and accuracy of the CST typing of the female vaginal flora structure. In addition, in the process of CST typing and analysis of the female vaginal flora structure, multiple samplings without replacement are performed on the relative abundance matrix data to generate different subsets of female vaginal flora data, and the obtained different subsets of female vaginal flora data are subjected to CST typing and analysis of the female vaginal flora structure, which can reduce the influence of outlier data and characteristic outliers in the relative abundance matrix data of the female vaginal flora on the typing results, and improve the accuracy of the CST typing results. In addition, based on the co-occurrence rate of any two female vaginal flora DNA sample data in the relative abundance matrix data in all clustering results, the female vaginal flora DNA sample data are clustered to obtain the CST typing results, which can reduce the impact of data noise on the CST typing results and further improve the accuracy of the CST typing results.
[0036] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or be understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0038] Figure 1 A schematic diagram of a method for typing the structure of female vaginal flora provided in an embodiment of the present application;
[0039] Figure 2 A schematic diagram of a consensus matrix clustering result provided in an embodiment of the present application;
[0040] Figure 3 A schematic diagram of classification results using a single hierarchical clustering method in a related technology provided in an embodiment of the present application;
[0041] Figure 4 A schematic diagram of the typing results of a female vaginal flora structure typing method provided in an embodiment of the present application;
[0042] Figure 5 A schematic diagram of a female vaginal flora structure typing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other arbitrarily. In addition, although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that here.
[0044] It is understandable that in the following specific implementations of the present application, related data such as female vaginal flora DNA sample data, relative abundance matrix data containing relative abundance of flora of multiple female vaginal flora DNA samples, etc., when the various embodiments of the present application are applied to specific products or technologies, relevant licenses or consents need to be obtained, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, when it is necessary to obtain relevant data, relevant volunteers can be recruited and relevant agreements for volunteer authorization data can be signed, and then the data of these volunteers can be used for implementation; or, by implementing within the scope of the authorized organization, the following implementation methods can be implemented by using the data of internal members of the organization to make relevant recommendations to internal members; or, the relevant data used in the specific implementation are all simulated data, such as simulated data generated in a virtual scene.
[0045] To facilitate understanding of the technical solutions provided in the embodiments of the present application, some key terms used in the embodiments of the present application are explained here:
[0046] 16SrRNA is a subunit of ribosomal RNA, and 16SrDNA is the gene encoding this subunit. Bacterial rRNA (ribosomal RNA) is divided into three types according to the sedimentation coefficient, namely 5S, 16S and 23SrRNA. 16S rDNA is the DNA sequence corresponding to the rRNA encoded on the bacterial chromosome, and exists in all bacterial chromosomal genes. 16SrDNA is the most useful and most commonly used molecular clock in the systematic classification of bacteria. It has few species and a large content (accounting for about 80% of the bacterial DNA content). The molecular size is moderate and it exists in all organisms. Its evolution has good clock properties and is highly conservative in structure and function. It is known as the "bacterial fossil".
[0047] Relative abundance, also known as isotopic abundance ratio, refers to the ratio of the abundance of the light component C in the gas to the sum of the abundances of the remaining components. In ecology, relative abundance refers to the number of species in a community. The species abundance in different communities is different. From the equator to the North and South Poles, the species abundance of the community gradually decreases. The greater the species richness, the more complex its structure and the greater its resistance and stability.
[0048] Amplicon sequence variant (ASV) refers to the determination and analysis of DNA sequences in microbial communities by high-throughput sequencing technology, thereby obtaining the ASV number of each microorganism. Compared with the traditional OTU (operational taxonomic unit method), ASV can more accurately identify microbial species, distinguish microorganisms at the level of closely related species and subspecies, and increase the accuracy and sensitivity of microbial diversity analysis. The working principle of ASV is to cluster similar DNA sequences into the same ASV by calculating the differences and similarities of DNA sequences. Unlike the OTU method, ASV does not rely on a set threshold, but accurately clusters according to the differences of each DNA sequence. The ASV method can reduce clustering bias, improve the resolution of microbial communities, and better study the diversity and dynamic changes of microbial communities. In microbial diversity analysis, ASV can be used to describe the composition and structure of microbial communities. By classifying and annotating ASV species, the functions and ecological functions of microorganisms can be inferred. At the same time, ASV can also be used to compare the differences in microbial communities between different samples, helping to find the correlation between key microbial species and community structures and array health status or environmental factors.
[0049] Polymerase chain reaction (PCR) is a molecular biology technique used to amplify specific DNA fragments. It can be regarded as a special DNA replication outside the body. The biggest feature of PCR is that it can greatly increase trace amounts of DNA. PCR uses the fact that DNA denatures into single strands at a high temperature of 95°C in vitro. At low temperatures (usually around 60°C), primers and single strands combine according to the principle of complementary base pairing. The temperature is then adjusted to the optimal reaction temperature of DNA polymerase (around 72°C). DNA polymerase synthesizes complementary chains along the direction from phosphate to pentose (5'-3'). The PCR instrument manufactured based on polymerase is actually a temperature control device that can well control the denaturation temperature, annealing temperature, and extension temperature.
[0050] DADA2 (Divisive Amplicon Denoising Algorithm 2), an algorithm for modeling and correcting Illumina amplicon sequencing errors, can accurately infer sample sequences and resolve differences of only 1 nucleotide. In some simulated communities, DADA2 identified more true variants and output fewer pseudo sequences than other methods. For paired-end Illumina Miseq data: Raw amplicon sequencing data is processed into a table of accurate amplicon sequence variants (ASVs) present in each sample.
[0051] Next-Generation Sequencing (NGS) is a high-throughput sequencing technology that is widely used in research fields such as genomics, transcriptomics, and epigenetics. This article will introduce the principles and experimental methods of NGS sequencing.
[0052] The Greengenes database contains 1,262,986 16S RNA sequences. The Greengenes2 database allows the direct integration of 16S rRNA and metagenomic datasets and unifies them in a reference tree. The analysis results show that the 16S rRNA and metagenomic data generated by the same samples are consistent in principal coordinate space, taxonomy, and phenotypic effect size.
[0053] The Silva database provides more than 1.7 million SSU (ribosomal small subunit, referred to as "SSU") sequences and more than 90,000 LSU (ribosomal large subunit, referred to as "LSU") sequences. It provides comprehensive, high-quality, comparable small and large subunit rRNA sequences for bacterial, archaea, and fungal analysis.
[0054] It provides ribosome-related data and services (Ribosomal Database Project, RDP), including online data analysis, alignment, and annotation of 16sRNA sequences. It contains more than 3.2 million 16sRNA sequences and more than 100,000 fungal 28sRNA sequences. It is one of the most widely used 16sRNA sequence databases and can provide convenient data processing functions such as 16sRNA sequence alignment, classification, evolutionary tree construction, species classification heatmap, functional gene analysis, etc.
[0055] The consensus matrix is a method for evaluating the stability and effectiveness of clustering. It determines the optimal number of clusters (K) by performing multiple clustering and analyzing the consistency. The elements of the consensus matrix represent the probability that two data points are in the same category, ranging from 0 to 1.
[0056] Hierarchical Clustering (HC) first calculates the distance between samples, and each time merges the closest points into the same class, then calculates the distance between classes, merges the closest classes into a large class, and merges different classes until a class is synthesized. Among them, the distance calculation methods between classes include: shortest distance method, longest distance method, middle distance method, class average method, etc. For example, the shortest distance method defines the distance between classes as the shortest distance between samples between classes. The order of hierarchical decomposition is: bottom-up, top-down, that is, agglomerative hierarchical clustering algorithm and classified hierarchical clustering algorithm.
[0057] The k-means clustering algorithm is an iterative clustering analysis algorithm, referred to as K-means. Its steps are: pre-divide the data into K groups, then randomly select K objects as the initial cluster centers, and then calculate the distance between each object and each seed cluster center, and assign each object to the cluster center closest to it. The cluster centers and the objects assigned to them represent a cluster. Each time a sample is assigned, the cluster center of the cluster will be recalculated based on the existing objects in the cluster. This process will be repeated until a certain termination condition is met. The termination condition can be that no (or the minimum number) of objects are reassigned to different clusters, no (or the minimum number) of cluster centers change again, and the sum of squared errors is locally minimized.
[0058] Partitioning Around Medoid (PAM) is a clustering method in the partitioning method of cluster analysis algorithms and is one of the earliest proposed k-medoid algorithms. The basic idea is: select the object with the most central position in the cluster and try to give k partitions for n objects; the representative object is also called the medoid, and other objects are called non-representative objects; initially, k objects are randomly selected as medoids, and the algorithm repeatedly replaces the representative objects with non-representative objects to try to find better medoids to improve the quality of clustering; in each iteration, all possible pairs of objects are analyzed, and one object in each pair is the medoid, while the other is a non-representative object. For various possible combinations, the quality of the clustering results is estimated; an object O i It can be replaced by the object that reduces the maximum square error value; the best set of objects generated in one iteration becomes the center point of the next iteration. The advantages of the PAM algorithm are: the PAM algorithm is more robust than the K-means algorithm, is not sensitive to "noise" and isolated point data, it can handle different types of data points, and is very effective for small data sets.
[0059] The basic idea of non-negative matrix factorization (NMF) can be described as: for any given non-negative matrix V, the NMF algorithm can find a non-negative matrix W and a non-negative matrix H, so that V = W*H holds, thereby decomposing a non-negative matrix into the product of two non-negative matrices on the left and right. NMF is essentially a matrix decomposition method, which is characterized by decomposing a large non-negative matrix into two small non-negative matrices. Because the decomposed matrix is also non-negative, it can be further decomposed. The application of NMF includes but is not limited to feature extraction, rapid recognition, gene and speech detection, etc.
[0060] The following is a brief introduction to the design concept of the embodiment of the present application:
[0061] The pH value in the vagina of healthy women is usually between 3.8 and 4.5. This acidic environment is mainly maintained by lactobacilli colonizing in the vagina, which can effectively resist the invasion of pathogenic bacteria. In the normal vaginal flora, Lactobacillus spp. is dominant, including more than 20 species such as Lactobacillus crispatus, Lactobacillus gasseri, Lactobacillus jensenii and Lactobacillus iners. In addition, there are a small number of other bacterial species in the vagina, such as Gardnerella vaginalis, Enterococcus spp., Staphylococcus spp., Enterobacteriaceae and BVAB1 bacteria associated with bacterial vaginosis (BV).
[0062] With the development of high-throughput sequencing technology, human microecology research has entered a new era. Using 16S rRNA sequencing technology, researchers have divided the vaginal flora state (community state type, CST) into five types: CST I (with Lactobacillus crispatus as the dominant bacteria), CST II (with Lactobacillus gasseri as the dominant bacteria), CST III (with Lactobacillus iners as the dominant bacteria), CST IV (with anaerobic bacteria and other bacteria as the dominant bacteria) and CSTV (with Lactobacillus jensenii as the dominant bacteria). Microbiome studies have shown that the composition and changes of female vaginal flora are closely related to HPV invasion, persistent infection and the occurrence and development of cervical cancer. Therefore, accurate typing of female vaginal flora structure is of great significance for early detection of potential flora imbalance and providing a scientific basis for the diagnosis and intervention of related diseases.
[0063] In related technologies, DNA extracted from fecal samples is usually sequenced using 16s rRNA sequencing technology. The sequencing results are used to generate ASV representative sequences using the dada2 clustering method of qiime2 software, and the species are classified and annotated using databases such as Silva and Greengene2, ultimately obtaining the relative abundance of each species in the corresponding sample. The most commonly used method for obtaining CST typing of samples is to perform unsupervised learning clustering based on the relative abundance of each species in the sample. Clustering methods include traditional clustering methods such as HC, K-means, and PAM. However, there are many factors that generate data noise during the 16s rRNA high-throughput sequencing experiment. For example, in PCR amplification, the amplification efficiency of different DNA fragments may be different, resulting in some sequences being over-amplified and other sequences being underestimated. Errors in the sequencing technology itself (such as mismatches, insertions, and deletions) will introduce noise that affects the accuracy of the sequence. Cross-contamination during sample processing and the experiment may introduce exogenous DNA. As a result, the above-mentioned traditional clustering methods are sensitive to outliers and noise in the sample, which can easily lead to typing results that are not robust and have low accuracy. In addition, the preferences of different clustering algorithms for input sample data will also affect the reliability of the results. For example, on a specific data set, different clustering algorithms will show different effects and selection tendencies. For example, K-means works better on data with relatively uniform distribution, while HC works better on data with uneven distribution.
[0064] In view of this, the embodiment of the present application provides a method, device, computer equipment and storage medium for typing the structure of female vaginal flora, using repeated sampling without replacement to generate different subsets of female vaginal flora data, reducing the influence of outlier sample data and characteristic outliers on the typing results, and using different clustering methods to cluster each subset of female vaginal flora data to obtain multiple clustering results, reducing the influence of different algorithms on the results, and finally using a consensus matrix to record the co-occurrence frequency between female vaginal flora DNA sample data in the relative abundance matrix data in multiple clustering results, and using a hierarchical clustering method to classify the consensus matrix to obtain the final CST typing result. In this way, the consensus matrix method is used to reduce the impact of data noise, reduce algorithm preference, and improve the accuracy of female vaginal flora structure typing.
[0065] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. In addition, the embodiments and features in the embodiments of the present application may be combined with each other if there is no conflict.
[0066] See also Figure 1As shown, it is a schematic diagram of a process of a method for typing the structure of female vaginal flora based on a consensus matrix provided in an embodiment of the present application. The specific implementation process of the method is as follows:
[0067] Step 101, extracting DNA from multiple female vaginal secretion swabs;
[0068] Step 102, using a polymerase chain reaction method to amplify each obtained DNA using a 16s rRNA sequencing method to obtain an amplified DNA;
[0069] Step 103, performing NGS high-throughput sequencing on each amplified DNA, storing the sequencing data of the corresponding amplified DNA in fastq format, and clustering the ASV representative sequence using the DADA2 algorithm of qiime2 software;
[0070] Step 104: Classify and annotate each ASV representative sequence through the greengene2 database to obtain relative abundance matrix data of female vaginal flora, wherein the relative abundance matrix data includes relative abundances of flora of multiple female vaginal flora DNA samples;
[0071] Step 105: Repeat sampling of samples and bacterial species of the relative abundance matrix data according to a set ratio by a sampling method without replacement to generate multiple female vaginal flora data subsets;
[0072] Step 106: clustering each female vaginal flora data subset respectively by multiple clustering methods to obtain multiple clustering results;
[0073] Step 107, using a consensus matrix method to obtain the co-occurrence rate of any two female vaginal flora DNA sample data in the relative abundance matrix data in multiple clustering results;
[0074] Step 108: Cluster all the co-occurrence rates obtained by hierarchical clustering method to obtain the flora CST typing results. The flora CST typing results are used to characterize the microbial community state type to which each female vaginal flora DNA sample belongs after clustering processing.
[0075] In one embodiment, the relative abundance matrix data of the DNA of multiple female vaginal secretion swabs can be obtained by 16s rRNA sequencing technology, specifically: First, the DNA of the female vaginal secretion swab is extracted, and the bacterial 16srRNA gene PCR amplification is performed. Then the amplified sequence is subjected to NGS high-throughput sequencing. The sequencing data is stored in fastq format, and the ASV representative sequence is generated by clustering the DADA2 algorithm of qiime2 software. Subsequently, the species are classified and annotated through the greengene2 database, and the relative abundance matrix table of the flora (relative abundance matrix data) is finally obtained.
[0076] In one embodiment, there are 1000 female vaginal flora DNA sample data (relative abundance of female vaginal flora DNA samples) in the relative abundance matrix data, and 700 female vaginal flora DNA sample data can be extracted from them each time without replacement sampling.
[0077] In one embodiment, the multiple clustering methods may include: one or more of decision tree, naive Bayes, K-nearest neighbor (KNN), HC, k-means, PAM, and NMF, or there may be multiple clustering methods, such as, the multiple clustering methods include two decision trees. There is no specific restriction on the type, quantity, and combination of clustering methods, which can be set as needed.
[0078] In one embodiment, the relative abundance matrix data includes 100 female vaginal flora DNA sample data and 4950 pairs of female vaginal flora DNA sample data, and the co-occurrence rate of each pair of female vaginal flora samples in each clustering result is determined respectively.
[0079] In order to reduce data noise and the influence of data preference of different clustering algorithms on the typing results of female vaginal flora, the above method uses multiple sampling methods without replacement to process the acquired relative abundance matrix data, generates different female vaginal flora data subsets, reduces the influence of data noise such as outlier samples and characteristic outliers on the typing results, and integrates different clustering algorithms to calculate each data subset to generate multiple clustering results, reduces the influence of data preference of different algorithms on the classification results, and finally uses the consensus matrix to count the co-occurrence frequency between different female vaginal flora DNA sample data in different clustering results, and uses the hierarchical clustering method to classify the consensus matrix to obtain the final CST typing results. In this way, the stability and accuracy of the typing results of female vaginal flora structure can be improved.
[0080] Based on the above Figure 1 In the method flow, the embodiment of the present application provides a sampling method. In step 104, the samples and strains of the relative abundance matrix data are repeatedly sampled according to a set ratio by a sampling method without replacement to generate multiple female vaginal flora data subsets, including:
[0081] Step 201: adopt a sample sampling method to perform n1 samplings without replacement on the relative abundance matrix data to obtain n1 female vaginal flora sample data subsets. The sample sampling method is to randomly extract female vaginal flora DNA sample data from the relative abundance matrix data.
[0082] Step 202: adopt a bacterial species sampling method, perform n2 samplings without replacement on the relative abundance matrix data, and obtain n2 female vaginal flora bacterial species data subsets. The bacterial species sampling method is to randomly extract the relative abundance data of a certain bacterial species in the relative abundance matrix data in all female vaginal flora DNA samples, and n1 and n2 are positive integers.
[0083] It should be noted here that there is no restriction on the execution order of step 201 and step 202 , and step 201 may be performed before step 202 , or after step 202 , or may be performed simultaneously with step 202 .
[0084] In one embodiment, the number of columns of the relative abundance matrix data represents the number of female vaginal flora DNA sample data x, the number of rows in the matrix represents the number of bacterial species y, and the proportion of female vaginal flora DNA sample data extracted without replacement each time is set to 60%, and the species proportion is 60%. For example, there are 100 female vaginal flora DNA sample data in the relative abundance matrix data, and each female vaginal flora DNA sample data corresponds to the relative abundance of 10 bacterial species. Then, the sample sampling method is adopted, with the column as a reference, and the number of female vaginal flora DNA sample data extracted without replacement each time is 60. The bacterial species sampling method is adopted, with the row as a reference, and the relative abundance data of 6 bacterial species of the species are extracted.
[0085] Based on the above Figure 2 In the method flow, the number of rows of the relative abundance matrix data is m, and the number of columns is n, indicating that the relative abundance matrix data contains m bacterial species and n sample data. The sampling ratio method is set to, in each sampling, the sampling ratio of the sample sampling method is 80%, and the sampling ratio of the bacterial species sampling method is 80%;
[0086] In step 201, a sample sampling method is adopted to perform n1 samplings without replacement on the relative abundance matrix data to obtain n1 female vaginal flora sample data subsets, including:
[0087] The sample function of R 4.3 software was used, and the parameter replace of the function was set to False. The sample index of random sampling was obtained by setting the parameter size. The sample sampling method without replacement was adopted each time, and the sample ratio was n*80% to obtain the corresponding female vaginal flora sample data subset;
[0088] In step 202, a bacterial species sampling method is used to perform n2 samplings without replacement on the relative abundance matrix data to obtain n2 female vaginal bacterial species data subsets, including:
[0089] The sample function of R 4.3 software was used, and the function parameter replace was set to False. The parameter size was set to obtain the randomly sampled bacterial species index. The sample sampling method without replacement was adopted each time. According to the strategy of bacterial species ratio of m*80%, the female vaginal flora species data subset was obtained.
[0090] Finally, the obtained female vaginal flora data subset contains two subsets, namely, n1 sample subsets and n2 feature subsets.
[0091] Based on the above method flows, the embodiment of the present application provides a clustering method, wherein the multiple clustering methods include a hierarchical clustering method, a KMeans clustering method, and a non-negative matrix decomposition method. In step 105, the multiple clustering methods cluster each female vaginal flora data subset respectively to obtain multiple clustering results, including:
[0092] Step 301: for each female vaginal flora sample data subset, the flora Pearson correlation distance of any two female vaginal flora sample data in each female vaginal flora sample data subset is calculated by the cor function, and for each female vaginal flora species data subset, the flora Pearson correlation distance of the relative abundance data of any two species in each female vaginal flora species data subset is calculated by the cor function;
[0093] Step 302: for each subset of female vaginal flora sample data, hierarchical clustering is performed on the distances between samples by using the hclust function, and central clustering is performed on the distances between samples by using the kmeans function, to obtain two corresponding clustering results; and for each subset of female vaginal flora strain data, hierarchical clustering is performed on the distances between strain relative abundance data by using the hclust function, and central clustering is performed on the distances between strain relative abundance data by using the kmeans function, to obtain two corresponding clustering results;
[0094] Step 303: for each female vaginal flora data subset, matrix decomposition is performed using the NMF R package to obtain a corresponding feature matrix W, and the hclust function is used to cluster the W matrix to obtain a corresponding clustering result.
[0095] It should be noted here that there is no restriction on the execution order between step 301, step 302 and step 303, for example, step 301, step 302 and step 303 can be executed simultaneously, or step 301 can be executed before or after any step between step 302 and step 303, and step 302 can be executed before or after any step between step 301 and step 303.
[0096] In one embodiment, first, each female vaginal flora data subset uses R4.3 software to calculate the flora Pearson correlation distance between female vaginal flora samples in the female vaginal flora data subset through the cor function. Then, the distance between female vaginal flora DNA sample data is hierarchically clustered by the hclust function, and the distance between sample data is centrally clustered by the kmeans function to obtain two clustering results for each female vaginal flora data subset. Finally, each female vaginal flora data subset is matrix decomposed by the NMF R package to obtain a feature matrix W, and the hclust function is used to cluster the W matrix to obtain a new clustering result. Finally, each female vaginal flora data subset will generate three clustering results.
[0097] Based on the above Figure 1 In the method flow, the embodiment of the present application provides a method for typing the female vaginal flora structure based on a consensus matrix. In step 106, the consensus matrix method is used to obtain the co-occurrence rate of any two female vaginal flora DNA sample data in the relative abundance matrix data in multiple clustering results, including:
[0098] Step 401: using a consensus matrix method, the frequency of any two female vaginal flora DNA sample data in the relative abundance matrix data being assigned to the same cluster in multiple clustering results is counted, and an n×n consensus matrix C is initialized. The initial value of the consensus matrix C is zero. For any two female vaginal flora DNA sample data (i, j), if i and j are assigned to the same cluster in a certain clustering result, the value of C[i][j] is added by 1 to obtain a first frequency matrix C. The clustering result contains at least one cluster, and the first frequency matrix C contains the co-occurrence frequency between samples in all clusters.
[0099] Step 402: Count the frequencies of any two female vaginal flora DNA sample data in the relative abundance matrix data appearing in the same clustering result in multiple clustering results, initialize an n×n consensus matrix E, the initial value of the consensus matrix E is zero, for any two female vaginal flora DNA sample data (i, j), if i and j are in a certain clustering result, then the value of E[i][j] is added by 1, and a second frequency matrix E is obtained, and the second frequency matrix E contains the co-occurrence frequencies between samples in all clustering results;
[0100] Step 403: Initialize an n×n consensus matrix H, where H[i][j]=C / E, and H[i][j] represents the proportion of i and j being assigned to the same cluster in all clustering methods.
[0101] That is to say, the consensus matrix method is used to count the co-occurrence rates of female vaginal flora DNA sample data in different clustering results of relative abundance matrix data, and hierarchical clustering is performed to obtain the flora CST typing results of the samples.
[0102] In one embodiment, n is the total number of female vaginal flora DNA sample data included in the relative abundance matrix data, and m is the number of all bacterial species covered in the DNA of multiple female vaginal secretion swabs in the relative abundance matrix data.
[0103] Based on the above method processes and their corresponding embodiments, the present application embodiment provides a schematic diagram of a consensus matrix clustering result, such as Figure 2 The present application embodiment also provides a schematic diagram of the typing results of a related technology using a single hierarchical clustering method, as shown in FIG. Figure 3 As shown in FIG. 1 , and a schematic diagram of the typing results using the female vaginal flora structure typing method of the present application, as shown in FIG. Figure 4 As shown,
[0104] like Figure 3 As shown in the figure, from left to right, cluster 1 is dominated by Lactobacillus iners, cluster 2 is dominated by Lactobacillus crispatus, and cluster 3 is unclear, and some samples are dominated by Bifidobacterium, accounting for nearly 50% of the total number of samples. In other words, nearly half of the female vaginal flora DNA samples cannot obtain clear clustering results through the single method of hierarchical clustering.
[0105] In contrast, Figure 4 As shown, the female vaginal flora structure typing method of the present application shows a higher resolution. From left to right, cluster 1 is mainly composed of Lactobacillus iners, cluster 2 is mainly composed of Lactobacillus iners and Lactobacillus crispatus, and cluster 3 also belongs to an unclear typing, but the number of samples in this cluster only accounts for nearly 15% of the total number of samples. The samples involved in this calculation are all healthy women. The results show that compared with the single clustering method, the results calculated by the female vaginal flora structure typing method of the present application are more consistent with the clinical results and more accurate.
[0106] Based on the same concept, Figure 5 As shown, an apparatus 500 for typing the structure of female vaginal flora based on a consensus matrix is provided in an embodiment of the present application, and the apparatus comprises:
[0107] The sample acquisition unit 501 is used to extract DNA from multiple female vaginal secretion swabs; amplify each obtained DNA by a 16s rRNA sequencing method using a polymerase chain reaction method to obtain amplified DNA; perform NGS high-throughput sequencing on each amplified DNA, store the obtained sequencing data of the corresponding amplified DNA in a fastq format, and cluster the ASV representative sequence using the DADA2 algorithm of the qiime2 software; classify and annotate each ASV representative sequence using the greengene2 database to obtain relative abundance matrix data of female vaginal flora, wherein the relative abundance matrix data contains the relative abundance of flora of multiple female vaginal flora DNA samples;
[0108] The typing unit 502 is used to repeatedly sample the samples and strains of the relative abundance matrix data according to a set ratio by a sampling method without replacement to generate multiple female vaginal flora data subsets; cluster each female vaginal flora data subset by multiple clustering methods to obtain multiple clustering results; use a consensus matrix method to obtain the co-occurrence rate of any two female vaginal flora DNA sample data in the relative abundance matrix data in the multiple clustering results; cluster all the obtained co-occurrence rates by a hierarchical clustering method to obtain flora CST typing results, which are used to characterize the microbial community state type to which each female vaginal flora DNA sample belongs after clustering. .
[0109] Optionally, the typing unit 502 is specifically used for:
[0110] Using a sample sampling method, performing n1 samplings without replacement on the relative abundance matrix data to obtain n1 female vaginal flora sample data subsets, wherein the sample sampling method is to randomly extract female vaginal flora DNA sample data from the relative abundance matrix data;
[0111] The relative abundance matrix data were sampled n2 times without replacement to obtain n2 female vaginal flora species data subsets. The species sampling method was to randomly extract the relative abundance data of a certain species in the relative abundance matrix data in all female vaginal flora DNA samples, and n1 and n2 were positive integers.
[0112] Optionally, the number of rows of the relative abundance matrix data is m, and the number of columns is n, indicating that the relative abundance matrix data contains m bacterial species and n sample data, and the sampling ratio method is set to, in each sampling, the sampling ratio of the sample sampling method is 80%, and the sampling ratio of the bacterial species sampling method is 80%;
[0113] The typing unit 502 is specifically used for:
[0114] The sample function of R 4.3 software was used, and the parameter replace of the function was set to False. The sample index of random sampling was obtained by setting the parameter size. The sample sampling method without replacement was adopted each time, and the sample ratio was n*80% to obtain the corresponding female vaginal flora sample data subset;
[0115] The typing unit 502 is specifically used for:
[0116] The sample function of R 4.3 software was used, and the function parameter replace was set to False. The parameter size was set to obtain the randomly sampled bacterial species index. The sample sampling method without replacement was adopted each time. According to the strategy of bacterial species ratio of m*80%, the female vaginal flora species data subset was obtained.
[0117] Optionally, the multiple clustering methods include a hierarchical clustering method, a KMeans clustering method and a non-negative matrix decomposition method, and the typing unit 502 is specifically used to:
[0118] For each female vaginal flora sample data subset, the flora Pearson correlation distance between any two female vaginal flora sample data in each female vaginal flora sample data subset is calculated by the cor function, and for each female vaginal flora species data subset, the flora Pearson correlation distance between any two relative abundance data of species in each female vaginal flora species data subset is calculated by the cor function;
[0119] For each subset of female vaginal flora sample data, the distances between samples were hierarchically clustered using the hclust function, and the distances between samples were centrally clustered using the kmeans function to obtain two corresponding clustering results;
[0120] And for each subset of female vaginal flora species data, the distance between the relative abundance data of the species is hierarchically clustered using the hclust function, and the distance between the relative abundance data of the species is centrally clustered using the kmeans function to obtain two corresponding clustering results;
[0121] For each subset of female vaginal flora data, matrix decomposition was performed using the NMF R package to obtain the corresponding feature matrix W, and the hclust function was used to cluster the W matrix to obtain the corresponding clustering results.
[0122] Optionally, the typing unit 502 is specifically used for:
[0123] A consensus matrix method is adopted to count the frequencies of any two female vaginal flora DNA sample data in the relative abundance matrix data being assigned to the same cluster in the multiple clustering results, and an n×n consensus matrix C is initialized, wherein the initial value of the consensus matrix C is zero. For any two female vaginal flora DNA sample data (i, j), if i and j are assigned to the same cluster in a certain clustering result, the value of C[i][j] is added by 1 to obtain a first frequency matrix C, wherein the clustering result contains at least one cluster, and the first frequency matrix C contains the co-occurrence frequencies between samples in all clusters;
[0124] Count the frequencies of any two female vaginal flora DNA sample data in the relative abundance matrix data appearing in the same clustering result in the multiple clustering results, initialize an n×n consensus matrix E, the initial value of the consensus matrix E is zero, for any two female vaginal flora DNA sample data (i, j), if i and j are in a certain clustering result, then the value of E[i][j] is added by 1, and a second frequency matrix E is obtained, wherein the second frequency matrix E contains the co-occurrence frequencies between samples in all clustering results;
[0125] Initialize an n×n consensus matrix H, H[i][j] = C / E, H[i][j] represents the proportion of i and j being assigned to the same cluster in all clustering methods.
[0126] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0127] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable female vaginal flora structure typing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable female vaginal flora structure typing device generate instructions for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0128] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable female vaginal flora structure typing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0129] These computer program instructions can also be loaded onto a computer or other programmable female vaginal flora structure typing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions executed on the computer or other programmable device for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0130] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A method for typing the structure of female vaginal flora based on a consensus matrix, characterized in that: The method comprises: Extract DNA from vaginal secretion swabs from multiple women; Using the polymerase chain reaction method, each obtained DNA was amplified by PCR using 16s rRNA primers to obtain amplified DNA; Each amplified DNA was subjected to NGS high-throughput sequencing, and the sequencing data of the corresponding amplified DNA was stored in fastq format, and clustered using the DADA2 algorithm of qiime2 software to generate ASV representative sequences; Classify and annotate each ASV representative sequence through the greengene2 database to obtain relative abundance matrix data of female vaginal flora, wherein the relative abundance matrix data includes relative abundance of flora of multiple female vaginal flora DNA samples; By using a sampling method without replacement, the samples and bacterial species of the relative abundance matrix data are repeatedly sampled according to a set ratio to generate multiple female vaginal flora data subsets; by using multiple clustering methods, each female vaginal flora data subset is clustered to obtain multiple clustering results; A consensus matrix method is used to obtain the co-occurrence rate of any two female vaginal flora DNA sample data in the relative abundance matrix data in the multiple clustering results; a hierarchical clustering method is used to cluster all the obtained co-occurrence rates to obtain the flora CST typing results, which are used to characterize the type of microbial community state to which each female vaginal flora DNA sample belongs after clustering processing.
2. The method according to claim 1, characterized in that The method of sampling without replacement is used to repeatedly sample the samples and bacterial species of the relative abundance matrix data according to a set ratio to generate multiple female vaginal flora data subsets, including: Using a sample sampling method, performing n1 samplings without replacement on the relative abundance matrix data to obtain n1 female vaginal flora sample data subsets, wherein the sample sampling method is to randomly extract female vaginal flora DNA sample data from the relative abundance matrix data; The relative abundance matrix data were sampled n2 times without replacement to obtain n2 female vaginal flora species data subsets. The species sampling method was to randomly extract the relative abundance data of a certain species in the relative abundance matrix data in all female vaginal flora DNA samples, and n1 and n2 were positive integers.
3. The method according to claim 2, characterized in that The number of rows and columns of the relative abundance matrix data is m, and the number of columns is n, indicating that the relative abundance matrix data contains m bacterial species and n sample data, and the sampling ratio method is set to, in each sampling, the sampling ratio of the sample sampling method is 80%, and the sampling ratio of the bacterial species sampling method is 80%; The sample sampling method is adopted to perform n1 samplings without replacement on the relative abundance matrix data to obtain n1 female vaginal flora sample data subsets, including: The sample function of R 4.3 software was used, and the parameter replace of the function was set to False. The sample index of random sampling was obtained by setting the parameter size. The sample sampling method without replacement was adopted each time, and the sample ratio was n*80% to obtain the corresponding female vaginal flora sample data subset; The bacterial species sampling method is adopted to perform n2 samplings without replacement on the relative abundance matrix data to obtain n2 female vaginal flora bacterial species data subsets, including: The sample function of R 4.3 software was used, and the function parameter replace was set to False. The parameter size was set to obtain the randomly sampled bacterial species index. The sample sampling method without replacement was adopted each time. According to the strategy of bacterial species ratio of m*80%, the female vaginal flora species data subset was obtained.
4. The method according to claim 1, characterized in that The multiple clustering methods include a hierarchical clustering method, a KMeans clustering method and a non-negative matrix decomposition method. The multiple female vaginal flora data subsets include a female vaginal flora sample data subset and a female vaginal flora strain data subset. The multiple clustering methods cluster each female vaginal flora data subset respectively to obtain multiple clustering results, including: For each female vaginal flora sample data subset, the flora Pearson correlation distance between any two female vaginal flora sample data in each female vaginal flora sample data subset is calculated by the cor function, and for each female vaginal flora species data subset, the flora Pearson correlation distance between any two relative abundance data of species in each female vaginal flora species data subset is calculated by the cor function; For each subset of female vaginal flora sample data, the distances between samples were hierarchically clustered using the hclust function, and the distances between samples were centrally clustered using the kmeans function to obtain two corresponding clustering results; And for each subset of female vaginal flora species data, the distance between the relative abundance data of the species is hierarchically clustered using the hclust function, and the distance between the relative abundance data of the species is centrally clustered using the kmeans function to obtain two corresponding clustering results; For each subset of female vaginal flora data, matrix decomposition was performed using the NMF R package to obtain the corresponding feature matrix W, and the hclust function was used to cluster the W matrix to obtain the corresponding clustering results.
5. The method according to claim 1, characterized in that The consensus matrix method is used to obtain the co-occurrence rate of any two female vaginal flora DNA samples in the relative abundance matrix data in the multiple clustering results, including: A consensus matrix method is adopted to count the frequencies of any two female vaginal flora DNA samples in the relative abundance matrix data being assigned to the same cluster in the multiple clustering results, and an n×n consensus matrix C is initialized. The initial value of the consensus matrix C is zero. For any two female vaginal flora DNA samples (i, j), if i and j are assigned to the same cluster in a certain clustering result, the value of C[i][j] is added by 1 to obtain a first frequency matrix C. The clustering result contains at least one cluster, and the first frequency matrix C contains the co-occurrence frequencies between samples in all clusters. Count the frequencies of any two female vaginal flora DNA sample data in the relative abundance matrix data appearing in the same clustering result in the multiple clustering results, initialize an n×n consensus matrix E, the initial value of the consensus matrix E is zero, for any two female vaginal flora DNA sample data (i, j), if i and j are in a certain clustering result, then the value of E[i][j] is added by 1, and a second frequency matrix E is obtained, wherein the second frequency matrix E contains the co-occurrence frequencies between samples in all clustering results; Initialize an n×n consensus matrix H, H[i][j] = C / E, H[i][j] represents the proportion of i and j being assigned to the same cluster in all clustering methods.
6. A device for typing the structure of female vaginal flora based on a consensus matrix, characterized in that: The device comprises: The sample acquisition unit is used to extract DNA from multiple female vaginal secretion swabs; the polymerase chain reaction method is used to perform 16s rRNA amplification on each obtained DNA to obtain amplified DNA; each amplified DNA is subjected to NGS high-throughput sequencing, and the sequencing data of the corresponding amplified DNA is stored in fastq format, and clustered by the DADA2 algorithm of qiime2 software to generate ASV representative sequences. The relative abundance matrix data of female vaginal flora is obtained by classifying and annotating each ASV representative sequence through the greengene2 database, and the relative abundance matrix data contains the relative abundance of flora of multiple female vaginal flora DNA samples; The typing unit is used to repeatedly sample the samples and strains of the relative abundance matrix data according to a set ratio by a sampling method without replacement to generate multiple female vaginal flora data subsets; cluster each female vaginal flora data subset respectively by multiple clustering methods to obtain multiple clustering results; adopt a consensus matrix method to obtain the co-occurrence rate of any two female vaginal flora DNA samples in the relative abundance matrix data in the multiple clustering results; cluster all the obtained co-occurrence rates by a hierarchical clustering method to obtain a flora CST typing result, and the flora CST typing result is used to characterize the type of microbial community state to which each female vaginal flora DNA sample belongs after clustering processing.
7. The method according to claim 1, characterized in that The typing unit is specifically used for: Using a sample sampling method, performing n1 samplings without replacement on the relative abundance matrix data to obtain n1 female vaginal flora sample data subsets, wherein the sample sampling method is to randomly extract female vaginal flora DNA sample data from the relative abundance matrix data; The relative abundance matrix data were sampled n2 times without replacement to obtain n2 female vaginal flora species data subsets. The species sampling method was to randomly extract the relative abundance data of a certain species in the relative abundance matrix data in all female vaginal flora DNA samples, and n1 and n2 were positive integers.
8. A computer-readable non-volatile storage medium, characterized in that: The computer-readable non-volatile storage medium stores a program, and when the program is run on a computer, the computer is enabled to implement the method according to any one of claims 1 to 5.
9. A computer device, characterized in that: include: Memory for storing computer programs; A processor, configured to call a computer program stored in the memory, and execute the method according to any one of claims 1 to 5 according to the obtained program.
10. A computer program product, characterized in that The method comprises a computer program, wherein the computer program is stored in a computer-readable storage medium; when a processor of a computer device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the computer device executes the method according to any one of claims 1 to 5.