Method for detecting microbial components based on high-throughput sequencing data

Through high-throughput sequencing technology and relative frequency quantitative calculation methods, a microbial component detection model was constructed, which solved the problems of inaccurate detection of microbial abundance and low accuracy of community structure quantification in traditional methods, and achieved more efficient microbial component detection.

CN119339810BActive Publication Date: 2025-08-05NINGXIA HONGGUO BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411421574.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-12
Publication Date
2025-08-05
Estimated Expiration
2044-10-12

AI Technical Summary

Technical Problem

Traditional microbial component detection methods based on high-throughput sequencing data have problems such as inaccurate detection of microbial abundance and low accuracy in quantification of microbial community structure.

Method used

Deep sequencing is carried out through high-throughput sequencing technology, combined with relative frequency quantitative calculation, agglomeration behavior recognition, functional gene recognition and metabolic pathway network construction, a microbial component detection model is constructed and sent to a cloud platform for detection.

Benefits of technology

It improves the accuracy of microbial abundance detection and the accuracy of microbial community structure quantification, and improves research efficiency and data processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339810B_ABST
    Figure CN119339810B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of microbial component detection technology, and in particular to a microbial component detection method based on high-throughput sequencing data. The method comprises the following steps: performing deep sequencing on a microbial test sample based on high-throughput sequencing technology to obtain microbial sample sequencing data; performing relative frequency quantitative calculation on the microbial sample sequencing data, and identifying clustering behavior between different categories to obtain microbial category clustering behavior data; constructing a metabolic pathway network between different microbial categories based on the microbial category clustering behavior data to obtain a microbial metabolic pathway network; constructing a microbial component detection model based on the microbial metabolic pathway network to obtain a microbial component detection model. The present invention makes the microbial component detection technology more perfect by optimizing the microbial component detection technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of microbial component detection, and in particular to a microbial component detection method based on high-throughput sequencing data. Background Art

[0002] High-throughput sequencing (HTS) technology has a significant application in microbial composition analysis. Microbial detection methods such as culture and microscopy are limited by culture conditions and microbial morphology, making it difficult to comprehensively and accurately identify microbial populations in complex environments. In contrast, high-throughput sequencing technology offers ultra-high sensitivity and parallel processing capabilities, enabling the sequencing of millions of DNA or RNA molecules in a single experiment. This technology offers significant advantages in microbial composition analysis. By sequencing the entire genome or specific regions (such as the 16S rRNA gene) of microbial DNA or RNA, scientists can comprehensively analyze the diversity, abundance, and function of microbial communities associated with the environment, the host, or disease. HTS technology not only detects uncultivable microbial species but also analyzes the dynamics of microbial populations, thereby providing researchers with a deeper understanding of the relationship between microbial communities and environmental or host health. However, traditional microbial composition analysis methods based on high-throughput sequencing data suffer from inaccurate microbial abundance measurements and low precision in quantifying microbial community structure. Summary of the Invention

[0003] Based on this, it is necessary to provide a method for detecting microbial components based on high-throughput sequencing data to solve at least one of the above technical problems.

[0004] To achieve the above objectives, a method for detecting microbial components based on high-throughput sequencing data is provided, the method comprising the following steps:

[0005] Step S1: obtaining a microbial test sample; performing deep sequencing on the microbial test sample based on high-throughput sequencing technology to obtain microbial sample sequencing data;

[0006] Step S2: performing relative frequency quantitative calculation on the microbial sample sequencing data to obtain microbial category relative frequency quantitative data; performing clustering behavior identification between different categories on the microbial category relative frequency quantitative data to obtain microbial category clustering behavior data;

[0007] Step S3: Identifying functional genes between different microbial categories based on the microbial category clustering behavior data to obtain microbial functional gene data; constructing a metabolic pathway network between different microbial categories based on the microbial functional gene data to obtain a microbial metabolic pathway network;

[0008] Step S4: construct a microbial component detection model based on the microbial metabolic pathway network to obtain a microbial component detection model; send the microbial component detection model to the cloud platform to execute the microbial component detection method.

[0009] The present invention first requires obtaining a microbial test sample. The sample can come from a variety of environments, such as soil, water, or organisms. After obtaining the sample, deep sequencing is performed using high-throughput sequencing technology. This technology offers high sensitivity and throughput, generating large amounts of sequence data in a short period of time. The results of deep sequencing can reveal the diversity and relative abundance of the various microorganisms present in the microbial sample, laying the foundation for subsequent analysis. Analysis of the resulting sequencing data using bioinformatics software accurately determines the microbial composition of the sample and provides the necessary information for subsequent relative frequency quantification. Successful implementation of this stage is fundamental to the entire research, ensuring the scientific and effective nature of subsequent analysis. After obtaining sequencing data from the microbial sample, relative frequency quantification is next required. This step primarily involves normalizing the sequence data of different microorganisms to determine the proportion of each microorganism in the total sample. This relative frequency quantification data helps understand the microbial community structure and abundance distribution. Subsequently, statistical analysis of the relative frequency data can identify clustering behaviors among different microbial classes. This can be performed using cluster analysis or other multidimensional scaling methods, aiming to reveal inter-microbial relationships and ecological functions. Identifying clustering behavior helps understand the complexity of microbial communities and their roles in ecosystems, providing crucial information for subsequent functional gene identification. The clustering data obtained from the previous steps allows for further identification of functional genes across different microbial categories. This process aims to identify genes associated with specific functions, such as those involved in carbohydrate metabolism, nitrogen cycling, or antibiotic resistance. By comparing these identified functional genes against known gene databases, we provide key insights into microbial ecological functions. Next, based on the identified functional gene data, a microbial metabolic pathway network is constructed. Metabolic pathway networks reveal the metabolic connections between different microorganisms and how they interact in the environment. This network not only helps to elucidate the metabolic mechanisms of microbial communities but also provides a theoretical foundation for applied microbial research, advancing the field of microbial ecology. After constructing the microbial metabolic pathway network, the next step is to develop a microbial component detection model based on this network information. This model involves using statistical and machine learning methods to integrate microbial functional genes and their roles in metabolic pathways, thereby enabling accurate detection of microbial community composition. The model design should ensure its applicability and reliability across diverse samples to ensure its practical application. Finally, the constructed microbial composition detection model is sent to a cloud platform for real-time application of the microbial composition detection method. The use of a cloud platform makes data processing and analysis more efficient, allowing researchers to access and analyze sample data at any time. This mechanism not only improves research efficiency but also provides a broader platform for microbial ecological monitoring and application.Therefore, the present invention is an optimization treatment of a traditional microbial component detection method based on high-throughput sequencing data, which solves the problems of inaccurate microbial abundance detection and low precision in quantification of microbial community structure in the traditional microbial component detection method based on high-throughput sequencing data, improves the accuracy of microbial abundance detection, and enhances the precision of quantification of microbial community structure.

[0010] Preferably, step S1 includes the following steps:

[0011] Step S11: obtaining a microbial test sample;

[0012] Step S12: filtering and separating the microbial test sample to obtain a microbial filtration test sample;

[0013] Step S13: centrifuging the microbial filtration test sample to obtain a microbial filtration centrifuge test sample;

[0014] Step S14: performing deep sequencing on the microbial filtration centrifugation test sample based on high-throughput sequencing technology to obtain microbial sample sequencing data.

[0015] Obtaining microbial test samples is the primary step in the present invention's research. Samples can be obtained from a variety of sources, such as natural water bodies, soil, plants, or animals. When obtaining samples, representativeness should be considered to ensure that they reflect the true structure and diversity of the microbial community. The sample collection process should be as free of contamination as possible and is generally conducted under sterile conditions and sealed in appropriate containers. Effective sample acquisition provides a reliable data foundation for subsequent experiments, providing important information for studying the species, abundance, and potential ecological functions of microorganisms. In addition, sample collection and storage conditions also have a significant impact on the results of subsequent analysis. Therefore, the sampling environment and process should be strictly controlled at this stage to ensure sample integrity and usability. After obtaining the microbial test sample, the next step is filtration separation. The purpose of this process is to separate the microorganisms in the sample from other impurities through filtration, thereby improving subsequent analysis. Filter membranes of different pore sizes can be used for filtration, selected based on the size of the microorganisms in the sample to ensure effective capture of the target microorganisms. During this process, the sterility of the filtration equipment and environment must be ensured to prevent contamination of the sample during processing. Filtration-separated samples enrich for microbial species and abundance, reduce background noise, and improve the sensitivity and accuracy of subsequent experiments. This process not only provides a purer sample for subsequent centrifugation and sequencing, but also lays the foundation for in-depth analysis of the diversity and ecological functions of microbial communities. Following filtration separation, the microbial filtration test sample undergoes centrifugation. This step further concentrates the microorganisms in the sample and removes any liquid components and other unwanted impurities. Centrifugation, through the centrifugal force generated by high-speed rotation, causes the microbial cells to settle at the bottom of the tube, separating them from the supernatant. This process effectively increases the concentration of microorganisms in the sample, facilitating subsequent deep sequencing and analysis. Furthermore, the centrifuged sample is purer, reducing background signals that can affect sequencing results and enhancing data reliability. Furthermore, the centrifugation process requires strict temperature and time control to avoid damage to the microorganisms, ensuring high-quality microbial samples for subsequent deep sequencing. Following centrifugation, the final step involves deep sequencing of the microbial filtration centrifugation test sample using high-throughput sequencing technology. High-throughput sequencing technology is highly efficient and rapid, generating large amounts of sequence data in a short period of time, making it particularly important for studying microbial communities. Deep sequencing of microbial samples can yield rich genetic information, enabling the identification of microbial species and their relative abundance. This step not only reveals microbial diversity but also provides foundational data for subsequent functional analyses and ecological studies. Furthermore, deep sequencing can identify low-abundance microorganisms, which is crucial for understanding the structure and function of microbial communities.Ultimately, the results of deep sequencing will provide important data support for fields such as microbial ecology research, environmental monitoring, and biotechnology applications.

[0016] Preferably, step S2 includes the following steps:

[0017] Step S21: recombining DNA fragments of the microbial sample sequencing data to obtain recombined microbial DNA fragment data;

[0018] Step S22: using a preset microbial DNA identification database to perform sequence comparison on the recombinant microbial DNA fragment data to obtain microbial DNA fragment classification data;

[0019] Step S23: performing relative frequency quantitative calculation on the microbial sample sequencing data according to the microbial DNA fragment category data to obtain microbial category relative frequency quantitative data;

[0020] Step S24: performing clustering behavior identification between different categories on the quantitative data of relative frequency of microorganism categories to obtain clustering behavior data of microorganism categories.

[0021] After obtaining sequencing data from microbial samples, the present invention first requires that these data be subjected to DNA fragment recombination. DNA fragment recombination involves splicing short fragments (reads) obtained through sequencing into longer sequences to more accurately identify the genomic information of the microorganism. This step can fill gaps in the sequencing by aligning and splicing overlapping fragments, thereby improving the continuity and integrity of the sequence. Fragment recombination not only helps to compensate for gaps in short sequence sequencing, but also improves data quality, reduces errors, and ensures the accuracy of subsequent comparisons and analyses. Through this process, the genetic information of the microorganism is effectively integrated, facilitating the classification and functional analysis of microbial communities in subsequent steps, and providing more accurate genomic data support for understanding the diversity and ecological role of microorganisms. The recombined microbial DNA fragments need to be sequenced with a preset microbial DNA identification database. This step is a key step in microbial classification and identification. Through this comparison, the microbial category to which each DNA fragment belongs can be identified, and the classification information of the fragment can be obtained. The preset database typically contains a large amount of known microbial genomic information, and similar sequences can be found through alignment algorithms (such as BLAST) to determine the type of microorganism in the sample. The accuracy of this step directly impacts the effectiveness of subsequent microbial classification and ecological studies. The resulting microbial DNA fragment classification data provides a foundation for microbial community diversity analysis, helping researchers identify the microbial species and their relative abundance within a sample. This provides crucial genetic information for functional analysis of microorganisms, ecological studies, and environmental monitoring. After obtaining the microbial DNA fragment classification data, relative frequency quantification is performed. By comparing the number of fragments representing each microbial class with the total number of fragments, the relative abundance of each microbial class within the sample can be calculated. This calculation reveals the distribution of different microorganisms within the sample and demonstrates their relative importance. Relative frequency quantification data facilitates the study of microbial community structure, identifying dominant and minority microbial populations. This data analysis provides a clearer understanding of the diversity and community characteristics of microorganisms within a sample, supporting subsequent microbial functional and ecological niche analyses. This step also lays the foundation for subsequent statistical and bioinformatics analyses, providing quantitative insights into microbial community succession and environmental adaptability. After obtaining quantitative data on the relative frequencies of microbial species, the next step is to identify clustering behaviors among different microbial species. Clustering is primarily accomplished through statistical methods such as cluster analysis, correlation analysis, or network analysis. These methods can identify interactions between microbial species and reveal whether symbiosis, competition, or other forms of ecological interaction exist within the microbial community. Analyzing clustering behavior helps understand the structure and function of microbial communities and assess the interdependencies and ecological roles of different microbial species.Through this process, researchers can better understand the complexity of microbial communities and predict how microorganisms behave and adapt under different environmental conditions. This phase of analysis provides strong data support for fields such as microbial community ecology research, environmental monitoring, and microbial engineering applications.

[0022] Preferably, step S23 includes the following steps:

[0023] Step S231: performing relative abundance calculations between different categories on the microbial sample sequencing data based on the microbial DNA fragment category data to obtain microbial category sequence relative abundance data;

[0024] Step S232: performing an abundance distribution skewness analysis among different categories on the relative abundance data of the microbial category sequences to obtain abundance distribution skewness data;

[0025] Step S233: performing a balanced quantitative analysis on the relative abundance data of the microbial category sequence according to the skewed abundance distribution data to obtain balanced quantitative data of the category sequence;

[0026] Step S234: performing relative frequency quantitative calculation on the microbial sample sequencing data based on the abundance distribution skewness data and the category sequence balance quantitative data to obtain microbial category relative frequency quantitative data.

[0027] The present invention calculates the relative abundance of different microbial classes in microbial sample sequencing data based on the class data of microbial DNA fragments. This process involves comparing the number of DNA fragments for each microorganism with the total number of DNA fragments to calculate the relative abundance of each microorganism in the sample. This analysis can effectively reveal the structural characteristics of microbial communities and help researchers understand the distribution of different microbial classes. For example, certain microorganisms dominate in a particular environment, while others are minor components. Relative abundance data not only provides a foundation for ecological research but also lays a data foundation for subsequent functional analyses and community comparisons. This step can provide important clues for analyzing the functional diversity and ecological adaptability of microbial communities. The relative abundance data of microbial class sequences is analyzed for abundance distribution skewness. This analysis aims to evaluate the characteristics of the abundance distribution of different microbial classes and identify skewness in the abundance distribution. Statistical analysis of the abundance data can determine which microbial classes dominate the community and which are at low abundance. Skewed abundance distribution can reveal the ecological characteristics and functional diversity of microbial communities. For example, positive skewness indicates that certain microbial groups are abundant, while negative skewness indicates that a small number of microorganisms dominate the resources. This analysis allows researchers to gain a deeper understanding of microbial niches, interactions, and adaptive mechanisms in the environment, providing important insights for subsequent ecological research and environmental monitoring. Based on the abundance distribution skewness data, a balanced quantitative analysis of the relative abundance data of microbial group sequences is performed. This process aims to eliminate bias within the sample, thereby ensuring a more balanced and accurate statistical analysis of the abundance data of different groups. Balanced quantitative analysis adjusts the abundance data of each microbial group within a sample to conform to a predetermined standard or distribution, making the results more comparable. This is crucial for assessing the true structure of microbial communities, particularly in the context of environmental change or experimental interventions, helping to identify microbial response mechanisms and adaptive capacities. This analysis enables researchers to obtain more accurate quantitative data on the balance of microbial group sequences, laying the foundation for further ecological model development and community dynamics analysis. By combining the abundance distribution skewness data with the quantitative data on the balance of group sequences, relative frequencies of microbial sample sequencing data were quantitatively calculated. This calculation process integrates the data obtained in the previous steps to ultimately obtain relative frequency data for each microbial category in the sample. These relative frequency data can reflect the actual presence of different microbial categories in the sample, helping researchers identify their roles and importance in the ecosystem. Quantitative relative frequency data not only helps understand the composition and dynamics of microbial communities but also provides more comprehensive and reliable basic data for ecological research. Ultimately, these data will support the analysis of microbial ecological functions, environmental adaptability research, and the development and application of related biotechnologies, providing a scientific basis for theoretical research and practical application of microbial ecology.

[0028] Preferably, performing balanced quantitative analysis on the relative abundance data of microbial class sequences based on the skewed abundance distribution data comprises the following steps:

[0029] The abundance distribution skewed data are processed by abundance stratification using the preset abundance stratification interval to obtain the abundance distribution skewed stratified data;

[0030] The extreme value fluctuation data of skewed abundance distribution were analyzed in different layers to obtain the extreme value fluctuation data of skewed layers.

[0031] The median difference of different strata is calculated for the skewed stratified extreme value fluctuation data to obtain the skewed stratified median difference data;

[0032] The skewed stratified extreme value fluctuation data are calculated by balanced interpolation according to the skewed stratified median difference data to obtain the skewed stratified balanced interpolation data;

[0033] Based on the skewed stratified balanced interpolation data, the relative abundance data of microbial category sequences were subjected to balanced quantitative analysis to obtain balanced quantitative data of category sequences.

[0034] The present invention uses preset abundance stratification intervals to perform abundance stratification processing on skewed abundance distribution data to obtain skewed abundance distribution stratified data. The core of this process is to divide the microbial abundance data into different levels so as to more clearly reveal the distribution characteristics of microorganisms within different abundance ranges. By stratifying the data, researchers can analyze the diversity and distribution patterns of microorganisms within each abundance interval. This stratification process can help identify which microbial categories dominate within a specific abundance interval and which microorganisms represent rare populations. By clearly defining the abundance intervals, we can better understand the ecological characteristics of the microbial community and its response to environmental changes, providing a good foundation for subsequent analysis. The skewed abundance distribution stratified data is subjected to extreme value fluctuation analysis within different strata to obtain skewed stratified extreme value fluctuation data. The purpose of extreme value fluctuation analysis is to identify extreme values and their changes within each abundance stratum. These extreme values usually refer to very high or very low observed values in the abundance distribution. By analyzing the fluctuations of these extreme values, researchers can reveal the survival strategies and ecological adaptability of microorganisms under specific environmental conditions. For example, some microorganisms exhibit extremely high abundances when resources are abundant, but their numbers plummet when resources are scarce. Extreme fluctuation data provide important information for studying the stability, resilience, and ecological interactions of microbial communities, supporting assessments of ecosystem health. Based on skewed stratified extreme fluctuation data, median differences within different strata are calculated to generate skewed stratified median difference data. The calculation of median differences aims to better understand the trends and concentration of microbial abundance within each abundance stratum. By comparing the medians within each abundance stratum, the significance and trends of abundance changes can be identified, revealing the biological significance of microbial species within each abundance stratum. For example, a significant change in the median within a stratum indicates changes in environmental conditions or altered competitive relationships between microorganisms. Median difference data facilitate the study of microbial community dynamics, revealing their ecological adaptability and succession processes, and providing a quantitative basis for ecological and microbiological research. Based on skewed stratified median difference data, balanced interpolation calculations are performed on the skewed stratified extreme fluctuation data to generate skewed stratified balanced interpolation data. The purpose of balanced interpolation is to smooth data between different abundance levels to eliminate existing biases and outliers, ensuring the accuracy and comparability of analytical results. Interpolation calculations can achieve reasonable estimates of data within each abundance stratum, making the abundance distribution more uniform and reflecting the true abundance of microorganisms in the sample. This step helps eliminate the influence of extreme values on the results and provides a more reliable data foundation for obtaining scientific and accurate conclusions in subsequent analyses. Based on the skewed stratified balanced interpolation data, a balanced quantitative analysis of the relative abundance data of the microbial class series was performed to obtain balanced quantitative data for the class series.The key to this step is to integrate the data processed in the previous steps to form a more stable and accurate abundance data set. Balanced quantitative analysis will provide quantitative support for the true distribution of microbial categories in the sample, which will help to study the overall structure and diversity of the microbial community. At the same time, this analysis result can also provide strong data support for ecological research, environmental monitoring and microbial engineering applications. Through this comprehensive analysis, researchers can have a deeper understanding of the role and function of microorganisms in the ecosystem, and provide a solid theoretical foundation and data support for future related research.

[0035] Preferably, performing balanced interpolation calculation on the relative abundance data of microbial class sequences based on the skewed stratified median difference data comprises the following steps:

[0036] Perform high-order difference accumulation on the skewed stratified median difference data to obtain the skewed stratified high-order difference accumulation data;

[0037] According to the skewness stratification high-order difference cumulative data and the skewness stratification median difference data, the skewness stratification extreme value fluctuation data is evaluated for data point continuity and smoothness, and the skewness stratification fluctuation continuous and smooth data are obtained;

[0038] The skewed stratified balanced interpolation data are obtained by performing balanced interpolation calculation on the skewed stratified continuous and smooth fluctuation data.

[0039] The present invention performs high-order difference accumulation on the skewed stratified median difference data to obtain skewed stratified high-order difference accumulation data. High-order difference accumulation is to gradually and deeply analyze the trend of microbial abundance changes by calculating multiple differences of median differences. This method can effectively capture subtle changes in abundance data, especially nonlinear trends that appear in certain abundance levels. Through the calculation of high-order differences, researchers can reveal the complex relationship between each abundance level, thereby more comprehensively understanding the dynamic changes and ecological adaptability of microbial communities. This step lays the foundation for subsequent interpolation calculations, so that data analysis can more accurately reflect the actual abundance status and distribution characteristics of microbial categories. Based on the skewed stratified high-order difference accumulation data and the skewed stratified median difference data, the skewed stratified extreme value fluctuation data is continuously and gently evaluated for data points, thereby obtaining skewed stratified fluctuation continuous and gentle data. Through continuous and gentle evaluation, researchers can identify and eliminate the instability caused by extreme value fluctuations, so that the data presents a smoother and more coherent trend. This process is particularly important because microbial abundance data are often influenced by environmental and other external factors, resulting in sudden fluctuations and irregular changes. By eliminating these interferences, the evaluated data can more realistically reflect the dynamic behavior and abundance trends of microbial communities, providing a more reliable data foundation for subsequent balanced interpolation calculations. Balanced interpolation is performed on the skewed, stratified, and continuously smoothed data to produce skewed, stratified, and balanced interpolated data. The purpose of balanced interpolation is to further process the smoothed data to achieve a more balanced and continuous data distribution across different abundance levels. This step uses interpolation methods to establish connections between data points, fill in gaps or missing data, and ensure the integrity and consistency of the final dataset. Balanced interpolation eliminates noise and outliers in the data, making the presentation of relative abundance data of microbial categories more accurate and reliable. This balanced data allows researchers to better understand the distribution characteristics and ecological functions of microorganisms in ecosystems, providing a scientific basis for ecological monitoring, microbial community management, and conservation.

[0040] Preferably, step S3 includes the following steps:

[0041] Step S31: Calculating the spatial dispersion between different microbial categories based on the microbial category relative frequency quantitative data on the microbial category aggregation behavior data to obtain microbial aggregation spatial dispersion data;

[0042] Step S32: performing functional gene identification between different microbial categories on the microbial sample sequencing data based on the microbial cluster spatial dispersion data and the microbial category cluster behavior data to obtain microbial functional gene data;

[0043] Step S33: constructing a metabolic pathway network among different microbial categories based on the microbial functional gene data according to the relative frequency quantitative data of the microbial categories to obtain a microbial metabolic pathway network.

[0044] Based on quantitative data on the relative frequency of microbial species, the present invention calculates the spatial dispersion of microbial clustering behavior to obtain spatial dispersion data. The core of this process is to quantify the spatial distribution patterns and dispersion of different microbial species. By calculating dispersion, researchers can identify the clustering characteristics of microbial communities in an environment and the interactions between different microorganisms. For example, a higher spatial dispersion indicates the clustering of certain microbial species within a specific area, while a lower dispersion indicates a more uniform distribution. This analysis not only helps understand the niche allocation of microbial communities but also reveals how environmental factors influence the spatial distribution of microorganisms, providing an important basis for microbial ecology research. Based on the spatial dispersion data and the clustering behavior data of microbial species, functional genes between different microbial species are identified in microbial sample sequencing data to obtain microbial functional gene data. Functional gene identification is a key step in analyzing the functional potential of microbial communities. By comparing and analyzing the genomic information in the samples, researchers can identify genes associated with specific biological functions. This process not only helps reveal the metabolic capacity and ecological functions of microorganisms but also further understands their adaptability to specific environmental conditions. By combining spatial dispersion with functional gene data, the roles and importance of different microbial classes in community function can be explored, providing a new perspective for assessing the ecological functions of microbial communities. Based on quantitative data on the relative frequencies of microbial classes, metabolic pathway networks between different microbial classes were constructed using microbial functional gene data, resulting in a microbial metabolic pathway network. This step aims to integrate microbial functional gene information to construct a network reflecting the metabolic interactions between different microbial classes. The construction of a metabolic pathway network can clearly demonstrate the interactions and dependencies between microorganisms during metabolic processes. For example, some microorganisms cooperate to participate in carbon or nitrogen cycling. By analyzing metabolic pathway networks, researchers can gain a deeper understanding of how microbial communities function in ecosystems and how metabolic networks vary under different environmental conditions. This information has important guiding significance for fields such as ecological restoration, agricultural microbial applications, and environmental protection.

[0045] Preferably, step S33 includes the following steps:

[0046] Step S331: performing gene interaction frequency relationships between different microbial categories on the microbial functional gene data based on the relative frequency quantitative data of the microbial categories to obtain gene interaction frequency relationship data;

[0047] Step S332: evaluating the functional impact of gene fragments on the microbial functional gene data based on the gene interaction frequency relationship data to obtain gene fragment functional impact data;

[0048] Step S333: performing effective association analysis of metabolic pathways in different functional regions on the gene fragment functional impact data to obtain effective association data of metabolic pathways;

[0049] Step S334: constructing a metabolic pathway network among different microbial categories based on the gene fragment functional impact data and the metabolic pathway effective association data to obtain a microbial metabolic pathway network.

[0050] The present invention analyzes microbial functional gene data using quantitative data on the relative frequency of microbial categories to calculate the frequency relationship of gene interactions between different microbial categories. The core of this process is to determine which genes interact with each other and the frequency of such interactions. In a microbial ecosystem, genes from different species affect each other's growth and metabolic characteristics through complex interactions, so understanding these interactions is crucial. By constructing gene interaction frequency relationship data, researchers can identify the contribution of genes with high frequency interactions to community function, thereby providing a quantitative basis for the synergistic effects between microorganisms. This step lays an important foundation for subsequent functional evaluation and network construction. Based on the gene interaction frequency relationship data, the functional impact of gene fragments is evaluated on the microbial functional gene data. The purpose of this evaluation is to identify the importance and influence of specific gene fragments in the function of the microbial community. By analyzing gene fragments with high interaction frequency, researchers can determine which genes play a key role in the metabolic activity and ecological adaptability of microorganisms. For example, certain gene fragments enhance the survival ability of microorganisms or promote their metabolic efficiency under specific environmental conditions. The resulting gene fragment functional impact data provides detailed information for subsequent metabolic pathway effective association analysis, enabling researchers to focus on gene fragments that significantly influence community function and thus deepen their understanding of microbial ecological function. Gene fragment functional impact data are then analyzed for metabolic pathway effective association across different functional regions to generate metabolic pathway effective association data. This step aims to identify effective associations within microbial metabolic pathways by analyzing the functional impact of gene fragments. This analysis reveals which gene fragments play key roles in specific metabolic pathways and how these genes interact to maintain microbial function and ecological balance. For example, certain functional regions contain multiple genes with high interaction frequencies, which together promote specific biotransformations. Metabolic pathway effective association data provides a foundation for constructing more complex metabolic networks, enabling researchers to gain a deeper understanding of how microorganisms interact within ecosystems through metabolic pathways. Based on the gene fragment functional impact data and metabolic pathway effective association data, metabolic pathway networks across different microbial classes are constructed, resulting in a microbial metabolic pathway network. This construction process is a key step in synthesizing the results of the previous analyses, integrating the functional associations of all relevant genes and metabolic pathways to form a comprehensive metabolic pathway network model. This network not only visually displays the metabolic interactions between different microorganisms, but also reveals the functions of key nodes and genes in metabolic pathways. The construction of this network is crucial for understanding the role of microbial communities in ecosystems, helping researchers predict how microorganisms respond to environmental changes and how their metabolic activities influence overall ecological functions.In this way, researchers can gain important insights into microbial ecology and provide a scientific basis for future ecological management and conservation strategies.

[0051] Preferably, step S4 includes the following steps:

[0052] Step S41: screening the key metabolic node functional genes of the microbial metabolic pathway network to obtain key metabolic node functional gene data;

[0053] Step S42: performing expression detection on the functional gene data of key metabolic nodes to obtain the node functional gene expression data;

[0054] Step S43: constructing a microbial component detection model for the node functional gene expression data according to the microbial metabolic pathway network to obtain a microbial component detection model;

[0055] Step S44: Send the microbial component detection model to the cloud platform to execute the microbial component detection method.

[0056] The present invention performs functional gene screening of key metabolic nodes in the microbial metabolic pathway network, aiming to identify key genes that play an important role in the metabolic process. This step determines which genes play a core role in regulating metabolic flow and metabolite production by analyzing the connectivity and functional importance in the metabolic pathway network. Key metabolic nodes are usually bottlenecks for microbial growth and reproduction. Understanding the functions of these genes can reveal the adaptation mechanism of microorganisms in specific environments and their metabolic characteristics. This screening result provides basic data for subsequent expression detection and functional analysis, helping researchers focus on the most potential genes for in-depth research. The expression level of the screened key metabolic node functional genes is detected to obtain the expression level data of the node functional genes. This step measures the transcription level of these genes under specific conditions by experimental methods (such as qPCR, RNA-seq, etc.), reflecting their actual activity in the physiological process of the microorganism. Changes in gene expression can provide important information about how microorganisms regulate their metabolic activities under environmental pressure. For example, certain genes are highly expressed when nutrients are abundant, but are downregulated when resources are scarce. By systematically documenting these changes, researchers can uncover metabolic regulation mechanisms and their relationships with environmental changes, providing an empirical basis for understanding microbial ecology and bioengineering applications. By analyzing node-level functional gene expression data based on a microbial metabolic pathway network, a microbial composition detection model was constructed. This model aims to predict the relative abundance of various components within a microbial community based on the expression characteristics of node genes. Using machine learning algorithms and statistical models, researchers were able to develop a model that comprehensively considers multiple factors, such as environmental factors, metabolic pathways, and gene expression. This model provides a new approach for quantitative analysis of microbial communities, enabling researchers to accurately predict and monitor changes in microbial communities and their role in ecosystems. Furthermore, the reliability and accuracy of the model directly impact the results and applications of microbiome research. The constructed microbial composition detection model is then transmitted to a cloud platform for execution of the microbial composition detection method. This step leverages the powerful computing and storage capabilities of cloud computing to process and analyze large amounts of microbial data, improving the efficiency and accuracy of data analysis. Through the cloud platform, researchers can not only obtain real-time detection results but also perform more complex computational and visual analyses to further optimize detection methods. Furthermore, the use of cloud platforms has made data sharing and collaboration between different research institutions more convenient, promoting the openness and transparency of microbiome research. The implementation of this approach not only accelerates the detection process of microbial components but also provides a solid data foundation for subsequent ecological research and application development.

[0057] Preferably, step S43 includes the following steps:

[0058] Step S431: performing metabolic flux balance calculation on the node functional gene expression data according to the microbial metabolic pathway network to obtain metabolic flux balance data;

[0059] Step S432: performing cross-influence analysis on the metabolic flux balance data to obtain metabolic flux cross-influence data;

[0060] Step S433: constructing a microbial component detection model for the node functional gene expression data based on the metabolic flow cross-influence data to obtain a microbial component detection model.

[0061] The present invention performs metabolic flux balance calculations based on the expression data of functional genes at nodes in a microbial metabolic pathway network. This process analyzes the flux of each metabolite in a specific metabolic pathway to determine the equilibrium between metabolite production and consumption under given conditions. This balance calculation can identify the main metabolic flux pathways and bottleneck reactions, thereby helping to understand how microorganisms utilize resources and convert energy under different environments. Metabolic flux balance data not only reflects the metabolic capacity of a microbial community but also provides fundamental information for subsequent cross-influence analysis, enabling researchers to gain a more comprehensive understanding of the role and function of microorganisms in an ecosystem. Cross-influence analysis is performed on metabolic flux balance data to obtain metabolic flux cross-influence data. The main purpose of this step is to identify the interactions and influence relationships between different metabolic pathways. For example, some metabolic pathways influence each other through the same metabolic precursors or common regulatory factors. Through cross-influence analysis, researchers can reveal how these interactions affect overall metabolic efficiency and the adaptability of microorganisms. In addition, this analysis can help identify key metabolic nodes and potential regulatory points for metabolic flux, providing data support for further metabolic engineering or ecological research. This analysis provides an important basis for establishing more complex metabolic models and optimizing microbial metabolic networks. Based on the metabolic flux cross-influence data and node functional gene expression data, the researchers constructed a microbial component detection model. The model aims to predict the composition and relative abundance of microbial communities by quantifying metabolic flux and its cross-influence. By integrating information on metabolic flux balance and cross-influence, the model can reflect the importance of different microbial species in metabolic processes and the dynamic characteristics of their interactions. The establishment of this model not only enhances the understanding of the complexity of microbial communities, but also provides feasible methodological support for environmental monitoring, microbiome research and bioengineering applications. In addition, the constructed model can also be applied to future experimental design and data analysis, providing a reliable tool for microbial ecology research.

[0062] To achieve the beneficial effects of the present invention, microbial test samples must first be obtained. These samples can come from various environments, such as soil, water, or organisms. After obtaining the samples, deep sequencing is performed using high-throughput sequencing technology. This technology offers high sensitivity and throughput, generating large amounts of sequence data in a short period of time. The results of deep sequencing can reveal the diversity and relative abundance of the various microorganisms present in the microbial sample, laying the foundation for subsequent analysis. Analysis of the resulting sequencing data using bioinformatics software can accurately determine the microbial composition of the sample and provide the necessary information for subsequent relative frequency quantification. Successful implementation of this stage is fundamental to the entire research, ensuring the scientific and effective nature of subsequent analysis. After obtaining sequencing data from the microbial sample, relative frequency quantification is next required. This step primarily involves normalizing the sequence data of different microorganisms to determine the proportion of each microorganism in the total sample. This relative frequency quantification data helps understand the microbial community structure and abundance distribution. Subsequently, statistical analysis of the relative frequency data can identify clustering behaviors among different microbial classes. This can be performed using cluster analysis or other multidimensional scaling methods, aiming to reveal inter-microbial relationships and ecological functions. Identifying clustering behavior helps understand the complexity of microbial communities and their roles in ecosystems, providing crucial information for subsequent functional gene identification. The clustering data obtained from the previous steps allows for further identification of functional genes across different microbial categories. This process aims to identify genes associated with specific functions, such as those involved in carbohydrate metabolism, nitrogen cycling, or antibiotic resistance. By comparing these identified functional genes against known gene databases, we provide key insights into microbial ecological functions. Next, based on the identified functional gene data, a microbial metabolic pathway network is constructed. Metabolic pathway networks reveal the metabolic connections between different microorganisms and how they interact in the environment. This network not only helps to elucidate the metabolic mechanisms of microbial communities but also provides a theoretical foundation for applied microbial research, advancing the field of microbial ecology. After constructing the microbial metabolic pathway network, the next step is to develop a microbial component detection model based on this network information. This model involves using statistical and machine learning methods to integrate microbial functional genes and their roles in metabolic pathways, thereby enabling accurate detection of microbial community composition. The model design should ensure its applicability and reliability across diverse samples to ensure its practical application. Finally, the constructed microbial composition detection model is sent to a cloud platform for real-time application of the microbial composition detection method. The use of a cloud platform makes data processing and analysis more efficient, allowing researchers to access and analyze sample data at any time. This mechanism not only improves research efficiency but also provides a broader platform for microbial ecological monitoring and application.Therefore, the present invention is an optimization treatment of a traditional microbial component detection method based on high-throughput sequencing data, which solves the problems of inaccurate microbial abundance detection and low precision in quantification of microbial community structure in the traditional microbial component detection method based on high-throughput sequencing data, improves the accuracy of microbial abundance detection, and enhances the precision of quantification of microbial community structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 A schematic diagram of the steps of a method for detecting microbial components based on high-throughput sequencing data;

[0064] Figure 2 for Figure 1 Detailed implementation steps of step S2 in FIG.

[0065] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0066] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. It is obvious that the embodiments described are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.

[0067] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller approaches.

[0068] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.

[0069] To achieve this, please refer to Figures 1 to 2 , a method for detecting microbial components based on high-throughput sequencing data, the method comprising the following steps:

[0070] Step S1: obtaining a microbial test sample; performing deep sequencing on the microbial test sample based on high-throughput sequencing technology to obtain microbial sample sequencing data;

[0071] Step S2: performing relative frequency quantitative calculation on the microbial sample sequencing data to obtain microbial category relative frequency quantitative data; performing clustering behavior identification between different categories on the microbial category relative frequency quantitative data to obtain microbial category clustering behavior data;

[0072] Step S3: Identifying functional genes between different microbial categories based on the microbial category clustering behavior data to obtain microbial functional gene data; constructing a metabolic pathway network between different microbial categories based on the microbial functional gene data to obtain a microbial metabolic pathway network;

[0073] Step S4: construct a microbial component detection model based on the microbial metabolic pathway network to obtain a microbial component detection model; send the microbial component detection model to the cloud platform to execute the microbial component detection method.

[0074] In the embodiment of the present invention, reference Figure 1 The above is a schematic flow chart of the steps of a method for detecting microbial components based on high-throughput sequencing data of the present invention. In this example, the method for detecting microbial components based on high-throughput sequencing data includes the following steps:

[0075] Step S1: obtaining a microbial test sample; performing deep sequencing on the microbial test sample based on high-throughput sequencing technology to obtain microbial sample sequencing data;

[0076] In the embodiment of the present invention, the acquisition of microbial test samples requires strict aseptic operation to avoid interference from external contamination sources. Using precise microbial collection tools, such as sterile sampling tubes and filter membranes, samples are collected from the target environment (such as water, soil or human body surface, etc.) and immediately stored at 4°C to ensure that the activity and structural integrity of the microorganisms in the sample are not damaged. Subsequently, the samples are sent to the laboratory for processing and deep sequencing is performed using high-throughput sequencing technology. High-throughput sequencing uses next-generation sequencing platforms, such as Illumina HiSeq or Ion Torrent, to first extract DNA from the sample. The extraction step requires the use of commercially available DNA extraction kits, such as the Qiagen DNeasy kit, combined with centrifugation and buffer washing steps to ensure the high purity and integrity of the DNA in the sample. After extraction is completed, specific primers are used to amplify the 16S rRNA gene (applicable to bacteria) or the 18S rRNA gene (applicable to eukaryotic organisms such as fungi), and library construction is performed to prepare for sequencing. The amount of data from deep sequencing is often as high as hundreds of millions of sequences, which can cover the genetic information of a large number of microbial populations in the sample. After sequencing is completed, initial sequence data is generated, i.e., microbial sample sequencing data. This data is usually saved in FASTQ format and contains the sequence itself and sequencing quality scores.

[0077] Step S2: performing relative frequency quantitative calculation on the microbial sample sequencing data to obtain microbial category relative frequency quantitative data; performing clustering behavior identification between different categories on the microbial category relative frequency quantitative data to obtain microbial category clustering behavior data;

[0078] In an embodiment of the present invention, the relative frequency quantitative calculation is performed on the obtained microbial sample sequencing data. First, the sequencing data needs to be quality controlled. Commonly used tools such as Trimmomatic or FASTQC software packages can remove low-quality sequences, connector contamination and other unreliable sequencing fragments. The data after quality control will be de-redundant to delete repeated sequencing fragments. Then, sequence classification is performed using platforms such as QIIME2 or Mothur. These tools assign microbial sequences to different taxonomic units (Operational Taxonomic Units, OTUs) by comparing databases such as Greengenes or SILVA, or divide them into different taxonomic groups based on sequence similarity. After the OTU or taxonomic unit is clear, its relative abundance in the sample is calculated based on the sequencing depth of each microorganism in the taxonomic group. The relative abundance is equal to the number of sequencing sequences of each microorganism divided by the total number of valid sequences, and the quantitative data of the relative frequency of each microbial category is obtained. This data characterizes the proportion of different microorganisms in the entire sample community to ensure the accuracy of subsequent analysis. Next, based on the quantitative data on the relative frequencies of these microbial categories, statistical cluster analysis methods, such as hierarchical clustering or K-means clustering, were used to identify clustering behaviors among different microbial categories. This clustering behavior indicates that certain microorganisms are more likely to interact or coexist under certain conditions, reflecting the ecological relationships or symbiotic mechanisms between microorganisms. Cluster analysis generates data on the clustering behavior of microbial categories, which can be used for further ecological analysis of microbial community structure and function.

[0079] Step S3: Identifying functional genes between different microbial categories based on the microbial category clustering behavior data to obtain microbial functional gene data; constructing a metabolic pathway network between different microbial categories based on the microbial functional gene data to obtain a microbial metabolic pathway network;

[0080] In an embodiment of the present invention, based on the clustering behavior data of microbial categories, a known microbial functional gene database (such as KEGG or COG database) is used to identify the functional genes between each microbial category. First, by comparing the known gene sequences in these functional gene libraries, the gene sequence of each microbial category is matched with the functional genes in the database. The specific operation includes utilizing a gene annotation tool (such as Prokka) to perform functional annotation on the microbial genome and identify the functional genes contained in each microbial species. In this process, it is necessary to strictly follow the sequence alignment algorithm, such as BLAST or HMMER, to ensure that the identification result of the functional gene has a high degree of confidence. The functional gene data of each microbial category will be output separately to form the microbial functional gene data. After the functional gene identification is completed, the microbial functional gene data is further analyzed to construct the metabolic pathway network between different microbial categories. The specific steps include, according to the functional gene information of each microorganism, in combination with a known metabolic pathway database (such as MetaCyc or KEGG PATHWAY), a metabolic network reconstruction tool (such as Pathway Tools or MetaboAnalyst) is used to reconstruct the metabolic network. This step involves the precise calculation of metabolic reactions, specifically the transfer relationships between different metabolites among microorganisms, the generation of metabolic intermediates, and their consumption pathways. The final output is a microbial metabolic pathway network, which illustrates the complex relationships between microorganisms through metabolite interactions and lays the foundation for subsequent microbial community function prediction and composition detection.

[0081] Step S4: construct a microbial component detection model based on the microbial metabolic pathway network to obtain a microbial component detection model; send the microbial component detection model to the cloud platform to execute the microbial component detection method.

[0082] In an embodiment of the present invention, after obtaining a microbial metabolic pathway network, a microbial component detection model is constructed based on the network. The construction of this model relies on the laws of microbial interactions in metabolic pathways. In order to construct the model, the nodes and edges in the metabolic network are first parsed by linear algebra and topological analysis methods to identify the most critical metabolic pathways and microbial components in the microbial community. The selection of key metabolic nodes depends on parameters such as degree distribution and path centrality. Through mathematical tools such as graph theory analysis methods (for example, using Cytoscape for graph structure analysis), the key nodes in the metabolic network are mapped to microbial components, and the contribution rate of each microorganism in the community is finally determined. After the model is constructed, the microbial component detection model needs to be sent to a cloud platform to facilitate remote execution of the microbial component detection method. This step first requires the model to be exported in a standard format (such as JSON or XML format) to ensure the integrity and compatibility of the model data. The microbial component detection model is uploaded to a predetermined cloud server via a secure transmission protocol (such as SFTP or HTTPS). After the cloud platform receives the model, it can call the detection parameters in the model and apply it to subsequent large-scale microbial sample component detection processes. The cloud platform uses the model to perform component analysis on newly input microbial samples, and quickly and accurately quantifies the microbial components by calling the metabolic network and microbial contribution rate information in the model.

[0083] Preferably, step S1 includes the following steps:

[0084] Step S11: obtaining a microbial test sample;

[0085] Step S12: filtering and separating the microbial test sample to obtain a microbial filtration test sample;

[0086] Step S13: centrifuging the microbial filtration test sample to obtain a microbial filtration centrifuge test sample;

[0087] Step S14: performing deep sequencing on the microbial filtration centrifugation test sample based on high-throughput sequencing technology to obtain microbial sample sequencing data.

[0088] In an embodiment of the present invention, the acquisition of microbial test samples involves collecting microbial samples to be analyzed from the target environment. The specific operation includes using a sterile collection device (such as a sterile sampling tube or filter) to collect microorganisms from different sources such as water, soil, food or human samples. Strictly follow the aseptic operation procedures during collection to ensure that interference from environmental pollution is avoided, especially for the collection of pathogenic microorganisms, which must be carried out in a sterile laboratory or a clean environment. After the sample is collected, it is immediately placed in a cold storage condition at a low temperature (4°C) to prevent the microorganisms in the sample from denaturing or prematurely decomposing. After sampling, the sample needs to be sent to the laboratory for subsequent processing in time to avoid degradation or contamination of the microbial sample by external factors. After the microbial test sample arrives at the laboratory, the sample is first filtered and separated to obtain the target microbial population. This step uses filter membranes with different pore sizes (such as 0.22μm or 0.45μm filter membranes) and selects a suitable filter according to the size of the microorganisms in the sample. The sample is passed through a vacuum filtration system or pressure filtration device. When the liquid sample passes through the filter membrane, particles larger than a specific pore size (such as impurities or non-microbial components) are trapped, effectively retaining the microorganisms on the membrane. This process requires controlling the pressure and filtration time to ensure that the microorganisms are completely retained without being damaged or lost. After filtration, the resulting filter membrane is the microbial filtration test sample, which will be used for further processing in subsequent steps. After filtration separation, the microbial filtration test sample is centrifuged to remove residual liquid and further concentrate the microbial population. First, the microbial filtration sample is gently scraped or eluted from the filter membrane into a sterile centrifuge tube. Next, the sample is centrifuged in a high-speed centrifuge (e.g., at a speed of 12,000 rpm or above) for 10-15 minutes to ensure that the microorganisms are sedimented to the bottom of the centrifuge tube due to the centrifugal force. The centrifugal force and time must be carefully monitored to ensure that the microbial sample is not excessively damaged or decomposed. After centrifugation, the supernatant is discarded, and the sediment at the bottom of the centrifuge tube is the concentrated microbial filtration centrifuge test sample. After obtaining the microbial filtration centrifugation test sample, high-throughput sequencing technology is used to perform deep sequencing analysis on it. First, high-quality DNA needs to be extracted from the microbial sample using a commercially available DNA extraction kit (such as Qiagen DNeasy). After DNA extraction, PCR amplification is required. Normally, specific primers are used to amplify the 16S rRNA gene fragment (for bacterial analysis) or the 18S rRNA gene (for fungal or other eukaryotic analysis). After the amplified product is confirmed by electrophoresis to confirm the fragment size, a high-throughput sequencing library is constructed. After the library construction is completed, the sample is loaded onto a next-generation sequencing platform (such as Illumina HiSeq or Ion Torrent) for sequencing. The sequencing platform generates a large amount of microbial sequence data, covering the complete genetic information of the microbial sample.Finally, the output microbial sample sequencing data is saved in FASTQ format, containing each sequence and its sequencing quality information for subsequent analysis.

[0089] Preferably, step S2 includes the following steps:

[0090] Step S21: recombining DNA fragments of the microbial sample sequencing data to obtain recombined microbial DNA fragment data;

[0091] Step S22: using a preset microbial DNA identification database to perform sequence comparison on the recombinant microbial DNA fragment data to obtain microbial DNA fragment classification data;

[0092] Step S23: performing relative frequency quantitative calculation on the microbial sample sequencing data according to the microbial DNA fragment category data to obtain microbial category relative frequency quantitative data;

[0093] Step S24: performing clustering behavior identification between different categories on the quantitative data of relative frequency of microorganism categories to obtain clustering behavior data of microorganism categories.

[0094] As an example of the present invention, refer to Figure 2 As shown, in this example, step S2 includes:

[0095] Step S21: recombining DNA fragments of the microbial sample sequencing data to obtain recombined microbial DNA fragment data;

[0096] In an embodiment of the present invention, first, the sequencing data of the microbial sample contains a large number of short DNA fragments, which are usually directly output by a high-throughput sequencing platform, but the order is disordered. In order to facilitate subsequent analysis, the DNA fragments must be recombined to ensure that these fragments can correctly reflect the complete genome information of the microorganism. Using a DNA fragment recombination algorithm, such as the Overlapping Consensus Assembly algorithm, short fragments with overlapping sequences are spliced into longer continuous sequences to form a complete microbial genome or a larger part of the genome sequence. This step requires precise pairing calculations to ensure that the overlapping regions between the fragments can be fully matched, thereby generating seamless recombinant DNA fragments. Through this process, microbial DNA fragment recombinant data is obtained, and the recombined data can be used as the basis for further analysis.

[0097] Step S22: using a preset microbial DNA identification database to perform sequence comparison on the recombinant microbial DNA fragment data to obtain microbial DNA fragment classification data;

[0098] In an embodiment of the present invention, after obtaining the recombinant microbial DNA fragment data, a preset microbial DNA identification database is used to perform sequence alignment on it. This step uses a sequence alignment algorithm, such as BLAST (Basic Local Alignment Search Tool) or BWA (Burrows-Wheeler Aligner), to align each fragment in the microbial DNA fragment recombinant data with the known sequences in the microbial DNA identification database. During the alignment process, the best matching microbial category is screened out by calculating the similarity between the fragment sequence and the known microbial gene sequence in the database, and each DNA fragment is classified. The accuracy of the alignment is achieved by setting a strict threshold. For example, when the sequence similarity is greater than 95%, it is considered to belong to the corresponding microbial category. The microbial DNA fragment category data finally outputted records the specific microbial category to which each fragment belongs, providing a clear category classification basis for subsequent analysis.

[0099] Step S23: performing relative frequency quantitative calculation on the microbial sample sequencing data according to the microbial DNA fragment category data to obtain microbial category relative frequency quantitative data;

[0100] In an embodiment of the present invention, a relative frequency quantitative calculation of the sequencing data of the microbial sample is performed based on the microbial DNA fragment category data. This step first counts the relative abundance of each microbial category in the sample based on the number of DNA fragments belonging to each microbial category. The relative frequency calculation is performed using the following formula: the ratio of the number of fragments of a certain microbial category to the total number of sequenced fragments. By calculating the relative frequencies of all microbial categories, the relative abundance information of different microbial categories in the sample is obtained. The relative frequency quantitatively reflects the proportion of each microorganism in the sample, and the relative frequency quantitative data of microbial categories are used for subsequent aggregation behavior identification and ecological function analysis.

[0101] Step S24: performing clustering behavior identification between different categories on the quantitative data of relative frequency of microorganism categories to obtain clustering behavior data of microorganism categories.

[0102] In an embodiment of the present invention, the clustering behavior between different microbial categories is identified based on the quantitative data of the relative frequency of microbial categories. By using a clustering analysis algorithm (such as Moran's I index or Ripley's K function), the clustering characteristics of microbial categories in the spatial or ecological dimension are analyzed to explore the distribution pattern of microorganisms in the sample. Clustering behavior identification determines which microbial populations have similar abundance distribution characteristics and act together in a certain functional ecological niche by performing cluster analysis on the quantitative data of relative frequency. For example, it is analyzed whether there are certain microorganisms that tend to co-occur at high frequencies under specific environmental conditions, or whether they exhibit ecological cooperation phenomena. The identified microbial clustering behavior data will provide new insights into microbial ecology and functional research.

[0103] Preferably, step S23 includes the following steps:

[0104] Step S231: performing relative abundance calculations between different categories on the microbial sample sequencing data based on the microbial DNA fragment category data to obtain microbial category sequence relative abundance data;

[0105] Step S232: performing an abundance distribution skewness analysis among different categories on the relative abundance data of the microbial category sequences to obtain abundance distribution skewness data;

[0106] Step S233: performing a balanced quantitative analysis on the relative abundance data of the microbial category sequence according to the skewed abundance distribution data to obtain balanced quantitative data of the category sequence;

[0107] Step S234: performing relative frequency quantitative calculation on the microbial sample sequencing data based on the abundance distribution skewness data and the category sequence balance quantitative data to obtain microbial category relative frequency quantitative data.

[0108] In an embodiment of the present invention, the relative abundance of different microbial categories is calculated based on the microbial DNA fragment category data. First, the DNA fragments of each microbial category need to be counted to determine the number of fragments in each category. Then, these fragment numbers are compared with the total number of fragments in the microbial sample to calculate the relative abundance of each category in the sample. This process uses the relative abundance formula: relative abundance = number of fragments of a certain microbial category / total number of fragments. The relative abundance of each microbial category reflects its proportion in the entire microbial population. Through this calculation process, the relative abundance data of the microbial category sequence is obtained, providing an accurate basis for the subsequent abundance distribution skewness analysis. After obtaining the relative abundance data of the microbial category sequence, it is necessary to perform a skewness analysis on the abundance distribution of different microbial categories. The abundance distribution skewness analysis is to use statistical methods to evaluate whether the abundance data of each category presents a skewed distribution, that is, whether the abundance of some microbial categories is much higher or lower than that of other categories. A commonly used method is to calculate the skewness coefficient. Based on the abundance distribution skewness data, a balanced quantitative analysis of the relative abundance data of the microbial category sequence is performed. Balanced quantitative analysis aims to investigate whether the distribution of various microbial classes in a sample is balanced, or whether certain classes dominate. The degree of population balance can be quantitatively assessed using the Shannon diversity index or the Simpson diversity index. Based on the aforementioned abundance distribution skewness data and class sequence balance quantitative data, a final relative frequency quantitative calculation is performed on the microbial sample sequencing data. This step integrates the skewness and balance between different microbial classes, ensuring that the relative frequency calculation more accurately reflects the microbial composition of the sample. During this process, the abundance values of each class are reweighted, so that classes exhibiting skewness or extreme abundance are appropriately corrected to reduce calculation errors. The relative frequency quantitative formula is adjusted based on the aforementioned steps, ultimately resulting in quantitative data on the relative frequency of microbial classes, providing a stable quantitative basis for microbial functional analysis and ecological niche research.

[0109] Preferably, performing balanced quantitative analysis on the relative abundance data of microbial class sequences based on the skewed abundance distribution data comprises the following steps:

[0110] The abundance distribution skewed data are processed by abundance stratification using the preset abundance stratification interval to obtain the abundance distribution skewed stratified data;

[0111] The extreme value fluctuation data of skewed abundance distribution were analyzed in different layers to obtain the extreme value fluctuation data of skewed layers.

[0112] The median difference of different strata is calculated for the skewed stratified extreme value fluctuation data to obtain the skewed stratified median difference data;

[0113] The skewed stratified extreme value fluctuation data are calculated by balanced interpolation according to the skewed stratified median difference data to obtain the skewed stratified balanced interpolation data;

[0114] Based on the skewed stratified balanced interpolation data, the relative abundance data of microbial category sequences were subjected to balanced quantitative analysis to obtain balanced quantitative data of category sequences.

[0115] In an embodiment of the present invention, abundance stratification processing is performed on the skewed abundance distribution data according to preset abundance stratification intervals. First, an abundance stratification standard is set. For example, the abundance can be divided into four levels: low abundance (0%-25%), medium abundance (26%-50%), high abundance (51%-75%), and extremely high abundance (76%-100%). Next, the skewed abundance distribution data is traversed, and the relative abundance values of each microbial category are mapped to the above-mentioned levels. For each microbial category, its abundance value is recorded in the corresponding hierarchical data structure, ultimately forming a data set after abundance stratification processing, namely, the skewed abundance distribution stratified data. This process ensures that the distribution characteristics of microorganisms in different abundance levels can be more clearly identified and compared in subsequent analysis. The skewed abundance distribution stratified data is analyzed for extreme value fluctuations within different levels. Specifically, it is first necessary to calculate the maximum and minimum values within each abundance level. On this basis, the extreme value fluctuations of each level are calculated, and for each abundance level, the extreme value fluctuation amplitude within that level is recorded. This analysis identifies the stability of microbial abundance within each abundance stratum. Greater fluctuations indicate more dramatic changes in abundance, while smaller fluctuations indicate more stable abundance. Ultimately, this generates extreme fluctuation data for the skewed stratification, providing data support for the subsequent calculation of median differences. After obtaining the extreme fluctuation data for the skewed stratification, the median differences within each stratum are calculated. For each abundance stratum, the median abundance of all microbial categories within that stratum is first calculated. This median is calculated using a sorting method, where the abundance values are arranged from smallest to largest. For even-numbered abundances, the average of the two middle values is taken, while for odd-numbered abundances, the middle value is taken. Once the median is obtained, the median differences between adjacent abundance strata are calculated. The median differences between all abundance strata are recorded to form skewed stratified median difference data, reflecting the differences in abundance distribution characteristics across the different abundance strata. Based on the skewed stratified median difference data, a balanced interpolation calculation is performed on the extreme fluctuation data for the skewed stratification. This step aims to reduce the deviation of the abundance distribution between different abundance layers through interpolation, so that the data is more balanced. The linear interpolation method is used to infer the new abundance value based on the extreme value fluctuation data and median difference of adjacent abundance layers. In the specific calculation process, the extreme value fluctuations of adjacent layers are weighted averaged to obtain the interpolated abundance data. After the interpolation calculation, the skewed stratified balanced interpolation data is finally generated, which lays the foundation for the subsequent analysis of the relative abundance data of microbial category sequences. After obtaining the skewed stratified balanced interpolation data, a balanced quantitative analysis is performed on the relative abundance data of microbial category sequences. This step uses the balanced abundance data to adjust the original relative abundance to eliminate the deviation caused by uneven abundance. The specific implementation method is to combine the skewed stratified balanced interpolation data with the relative abundance data of microbial category sequences, and recalculate the abundance of each microbial category by weighted averaging or other appropriate quantitative methods.This process ensures that the final category sequence balanced quantitative data truly reflects the actual distribution of microorganisms in the sample and provides reliable data support for microbial community analysis.

[0116] Preferably, performing balanced interpolation calculation on the relative abundance data of microbial class sequences based on the skewed stratified median difference data comprises the following steps:

[0117] Perform high-order difference accumulation on the skewed stratified median difference data to obtain the skewed stratified high-order difference accumulation data;

[0118] According to the skewness stratification high-order difference cumulative data and the skewness stratification median difference data, the skewness stratification extreme value fluctuation data is evaluated for data point continuity and smoothness, and the skewness stratification fluctuation continuous and smooth data are obtained;

[0119] The skewed stratified balanced interpolation data are obtained by performing balanced interpolation calculation on the skewed stratified continuous and smooth fluctuation data.

[0120] In an embodiment of the present invention, first, the skewed stratified median difference data is solved by high-order difference accumulation. This step aims to capture the changing trend of the abundance data through the high-order difference method. The specific implementation process is: first determine the order of the high-order difference to be calculated, usually taking the 1st and 2nd order differences. On this basis, calculate the median difference of each layer of abundance and accumulate it. The skewed stratified high-order difference accumulation data finally formed is the high-order difference and its accumulation, which can reflect the changing rate and trend of the microbial abundance data. The acquisition of this data will lay the foundation for the subsequent continuous and gentle assessment of fluctuations. Based on the obtained skewed stratified high-order difference accumulation data and skewed stratified median difference data, the skewed stratified extreme value fluctuation data is evaluated for data point continuity and gentleness. This step mainly focuses on how to smooth the data between different abundance layers in order to better reflect the overall trend of microbial abundance. First, the skewed stratified extreme value fluctuation data is divided into multiple data point intervals. In each interval, the average value of the continuous data points is calculated by the weighted average method, and it is combined with the median difference data. After obtaining continuous, smooth data with skewed stratified fluctuations, balanced interpolation calculations are performed. The purpose of this step is to accurately interpolate the relative abundance data of the microbial category series based on the continuous, smooth data, thereby ensuring data balance. The specific implementation process is: setting an interpolation model, for example, using linear interpolation or spline interpolation methods, and applying the smooth data to the abundance series of the microbial category. During the interpolation process, the interpolation nodes are first determined. These nodes are selected based on the changing trend of the smooth data. For each pair of interpolation nodes, the following formula is used for interpolation calculation: I(x) = (1-t)Y1 + tY2, where I(x) is the interpolation result, Y1 and Y2 are the abundance values of the adjacent nodes, and t is the number of interpolation points. Using the above formula, the abundance value of the new data point is calculated and incorporated into the final relative abundance data of the microbial category series. Interpolation calculations not only eliminate data discontinuities but also enhance the overall stability of the data. The resulting skewed, stratified, balanced interpolation data will provide a more accurate and reliable basis for microbial component detection.

[0121] Preferably, step S3 includes the following steps:

[0122] Step S31: Calculating the spatial dispersion between different microbial categories based on the microbial category relative frequency quantitative data on the microbial category aggregation behavior data to obtain microbial aggregation spatial dispersion data;

[0123] Step S32: performing functional gene identification between different microbial categories on the microbial sample sequencing data based on the microbial cluster spatial dispersion data and the microbial category cluster behavior data to obtain microbial functional gene data;

[0124] Step S33: constructing a metabolic pathway network among different microbial categories based on the microbial functional gene data according to the relative frequency quantitative data of the microbial categories to obtain a microbial metabolic pathway network.

[0125] In an embodiment of the present invention, when calculating the spatial dispersion of microbial clustering behavior based on quantitative data on the relative frequency of microbial classes, spatial distribution data of microbial samples is first collected. For each microbial class, its position distribution in space is defined, and the coordinates of each sample are represented in a two-dimensional or three-dimensional coordinate system. Next, a spatial dispersion calculation method is used, typically a variance-based dispersion index. In specific implementation, the mean position of each microbial class in its spatial location is first calculated, and then the dispersion is calculated based on the difference between the position of each sample and the mean position. After completing the calculation of the spatial dispersion data of microbial clustering, functional genes in the microbial sample sequencing data are identified based on this data and the microbial class clustering behavior data. This step first requires the establishment of a functional gene database containing genomic information and functional annotations of known microorganisms. The microbial sample sequence data obtained by high-throughput sequencing is compared with the functional gene database. Here, an alignment algorithm, such as BLAST (Basic Local Alignment Search Tool), is used to perform a sequence similarity search to identify sequences in the sequencing data that correspond to functional genes in the database. During the alignment process, an appropriate similarity threshold is set to ensure the accuracy of the identified functional genes. Each successfully identified gene records its corresponding microbial class and its clustering behavior characteristics. The resulting microbial functional gene data includes the name of the functional gene, functional annotations, and its abundance information within each microbial class, providing support for further analysis. After obtaining the microbial functional gene data, a microbial metabolic pathway network is constructed based on quantitative data on the relative frequencies of microbial classes. This step first requires integrating information on the roles of microbial functional genes. Leveraging existing metabolic pathway databases (such as the KEGG database), each functional gene is mapped to its corresponding metabolic pathway. The specific implementation process involves integrating the microbial functional gene data with its corresponding metabolic pathway information to form a basic dataset for metabolic pathway construction. A network analysis algorithm is then used to construct the metabolic pathway network. Based on the relationships between functional genes and their metabolic pathways, a graph structure is constructed, representing nodes (functional genes) and edges (interactions between genes). Each node represents a functional gene, and each edge represents the gene's role in the metabolic pathway. The constructed metabolic pathway network is optimized and analyzed using graph theory methods such as the shortest path algorithm and connected component analysis to identify key genes and metabolic pathways that play important roles across different microbial classes. The resulting microbial metabolic pathway network contains functional genes of different microbial categories and their interactions in the metabolic process, providing a basis for in-depth research on the biological functions of microbial components and their ecological significance.

[0126] Preferably, step S33 includes the following steps:

[0127] Step S331: performing gene interaction frequency relationships between different microbial categories on the microbial functional gene data based on the relative frequency quantitative data of the microbial categories to obtain gene interaction frequency relationship data;

[0128] Step S332: evaluating the functional impact of gene fragments on the microbial functional gene data based on the gene interaction frequency relationship data to obtain gene fragment functional impact data;

[0129] Step S333: performing effective association analysis of metabolic pathways in different functional regions on the gene fragment functional impact data to obtain effective association data of metabolic pathways;

[0130] Step S334: constructing a metabolic pathway network among different microbial categories based on the gene fragment functional impact data and the metabolic pathway effective association data to obtain a microbial metabolic pathway network.

[0131] In an embodiment of the present invention, when calculating the gene interaction frequency relationship of microbial functional gene data based on the quantitative data of the relative frequency of microbial categories, it is first necessary to extract the functional gene information under each microbial category from the previously obtained microbial functional gene data. Subsequently, an interaction matrix is constructed, in which the rows and columns represent different microbial categories and functional genes, respectively. Each element of the matrix represents the interaction frequency between functional genes under a specific microbial category. The calculation formula is as follows: "The interaction frequency of functional gene j in microbial category i is equal to the number of times functional gene j co-occurs in category i, divided by the total number of times functional gene j appears in all categories." Through the construction of this interaction matrix, gene interaction frequency relationship data is obtained, which provides a basis for the subsequent functional impact evaluation of gene fragments. After obtaining the gene interaction frequency relationship data, the functional impact evaluation of the gene fragments is performed. First, the functional genes are mapped to their corresponding gene fragments, and a mapping relationship between the gene fragments and the functional genes is constructed. For each gene fragment, the interaction frequency relationship data is used to evaluate its functional impact in different microbial categories. The specific implementation process includes the following steps: Functional impact indicators, such as interaction strength, are defined. The influence of each gene segment is calculated by weighted summation of interaction frequency data. Based on the calculated functional impact indicators, the gene segments are classified into different functional impact levels, and specific thresholds are set for classification. Based on the gene segment functional impact data, effective metabolic pathway association analysis is performed for different functional regions. First, the previously constructed functional gene database and metabolic pathway database are integrated to facilitate subsequent analysis. Functional genes are mapped to metabolic pathways to form a functional gene distribution table for metabolic pathways. This table records the functional genes included in each metabolic pathway and their impact levels in the gene segment functional impact data. Statistical analysis methods, such as chi-square tests or correlation analysis, are used to evaluate the functional impact relationship between each metabolic pathway and its associated gene segments. Appropriate statistical significance levels are set to determine effective associations between metabolic pathways and gene segments with high functional impact. The analysis results are organized into effective metabolic pathway association data, clearly defining the functional genes included in each metabolic pathway and their functional impact levels. Metabolic pathways are considered nodes in a network, with functional genes as connecting edges. The weight of each edge can be set based on the functional impact data of gene fragments, reflecting the importance of genes in metabolic pathways. Graph theory methods are used to construct metabolic pathway networks, generating a directed graph in which nodes represent metabolic pathways and edges represent the roles of functional genes. Shortest path algorithms can be used to analyze the shortest connections between nodes in the network. Network analysis techniques, such as cluster analysis and centrality measures, can be used to identify metabolic pathways and functional genes that play key roles across different microbial groups. These analytical results can help reveal the metabolic characteristics of microorganisms and their ecological functions.Ultimately, the microbial metabolic pathway network data is output, reflecting the metabolic pathway relationships between different microbial categories and their interactions at the functional gene level, providing an important foundation for the research of microbial ecology and bioengineering.

[0132] Preferably, step S4 includes the following steps:

[0133] Step S41: screening the key metabolic node functional genes of the microbial metabolic pathway network to obtain key metabolic node functional gene data;

[0134] Step S42: performing expression detection on the functional gene data of key metabolic nodes to obtain the node functional gene expression data;

[0135] Step S43: constructing a microbial component detection model for the node functional gene expression data according to the microbial metabolic pathway network to obtain a microbial component detection model;

[0136] Step S44: Send the microbial component detection model to the cloud platform to execute the microbial component detection method.

[0137] In an embodiment of the present invention, after the microbial metabolic pathway network is constructed, the key metabolic nodes in the network need to be screened to identify functional genes therein. The screening process first needs to analyze the importance of each node in the network through a graph theory algorithm (such as PageRank or betweenness centrality) to calculate the weight distribution of each metabolic node. Then, through these weight data, the nodes that play a key metabolic role in the metabolic pathway are screened, particularly those with high connectivity or located on important metabolic pathways in the metabolic pathway network. The metabolic nodes screened out have significant functional genes in different microbial categories, and the expression of these genes directly affects the efficiency of metabolic function. The output obtained by this step is key metabolic node functional gene data, which specifically includes each metabolic node and its corresponding functional gene set. After the key metabolic node functional gene data is screened, the next step is to detect the expression of these functional genes. Gene expression detection needs to be realized by high-throughput sequencing data, and the specific process includes extracting the mRNA sequence in the microbial sample related to the key metabolic node. Then, quantitative PCR or transcriptome sequencing technology (RNA-Seq) is used to amplify and sequence mRNA to obtain the expression value of each gene. These values reflect the activity levels of functional genes at key metabolic nodes in different microbial categories. This data is structured and stored as node functional gene expression data for subsequent use in building a microbial component detection model. After obtaining the node functional gene expression data, it is necessary to use this data to construct a microbial component detection model. The model construction process requires the use of statistical methods (such as principal component analysis (PCA) or linear discriminant analysis (LDA)) based on the topological structure of the metabolic pathway network and gene expression data to extract features and identify key gene expression patterns that can distinguish different microbial categories. The specific steps include: first, using principal component analysis to extract several gene sets with the greatest differences from the large amount of gene expression data; then, combining the node weights of each functional gene in the metabolic pathway network, the extracted gene sets are weighted to construct a linear equation or classifier for predicting the composition of unknown microbial samples. The model output is a microbial component detection model that can be used to detect and classify the components of future microbial samples. Once the microbial component detection model is constructed, it must be uploaded to a cloud platform to execute the microbial component detection task. First, the model data is formatted into a portable file (such as JSON or XML format) through a stable network connection and transferred to a dedicated cloud storage server. Next, the cloud platform will call the uploaded model and perform microbial component detection in combination with real-time sequencing data. The process includes inputting the newly acquired high-throughput sequencing data into the model, calculating the component detection results of the microbial sample through the model, and returning the test results to the data terminal for display. The entire process realizes the automation of microbial component detection and improves detection efficiency and accuracy.

[0138] Preferably, step S43 includes the following steps:

[0139] Step S431: performing metabolic flux balance calculation on the node functional gene expression data according to the microbial metabolic pathway network to obtain metabolic flux balance data;

[0140] Step S432: performing cross-influence analysis on the metabolic flux balance data to obtain metabolic flux cross-influence data;

[0141] Step S433: constructing a microbial component detection model for the node functional gene expression data based on the metabolic flow cross-influence data to obtain a microbial component detection model.

[0142] In the embodiment of the present invention, the purpose of metabolic flux balance calculation is to carry out dynamic analysis to the node functional gene expression amount data by the microbial metabolic pathway network, thereby calculate the reaction flow that each key gene participates in the metabolic process.In this process, need adopt steady-state metabolic flux analysis method (FBA, Flux Balance Analysis) to solve the material balance of each metabolic reaction in the network.Specific implementation step comprises: first extract the reaction equation relevant to metabolism from the microbial metabolic pathway network, define the stoichiometric relationship of these reactions.Then, utilize the expression amount of node functional gene as constraint condition, solve how metabolic flux is balanced and distributed in the network by linear programming algorithm (such as Simplex algorithm).Finally, obtain metabolic flux balance data, this data specifically reflects the material flow direction and relative flow rate of each metabolic reaction in the metabolic network.After obtaining metabolic flux balance data, need further analyze the cross influence between each metabolic flux, to explore the interaction between different metabolic reactions.Cross influence analysis is based on network topology and metabolic flux data, adopts quantitative impact factor analysis method (such as random walk algorithm) to calculate the influence of each metabolic reaction in the network. The specific process involves first associating metabolic flux balance data with each reaction node in the metabolic pathway network and constructing an influence matrix to quantify the interactions between metabolic fluxes. Next, weighted nodes and edges in the network are used to determine the transmission pathways and influence strengths between metabolic fluxes. A comprehensive assessment of these interactions generates metabolic flux cross-influence data, which demonstrates how different metabolic reactions influence each other within the network and regulate the balance of the entire metabolic system. After completing the metabolic flux cross-influence analysis, a microbial composition detection model is constructed based on this metabolic flux cross-influence data and node functional gene expression data. During model construction, key correlation patterns between metabolic fluxes and gene expression are identified through methods such as regression analysis and feature selection. The specific implementation steps are as follows: First, based on the metabolic flux cross-influence data, the metabolic reactions and corresponding functional genes with the greatest impact on microbial composition are extracted. Next, a detection model for predicting microbial composition is constructed using methods such as multivariate linear regression or logistic regression. The model construction is cross-validated to ensure its robustness and accuracy. Ultimately, the output microbial composition detection model accurately predicts the relative proportions of different microbial components in a sample, providing a foundation for subsequent microbial detection.

[0143] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.

[0144] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.

Claims

1. A method for detecting microbial components based on high-throughput sequencing data, characterized in that: The following steps are involved: Step S1: obtaining a microbial test sample; performing deep sequencing on the microbial test sample based on high-throughput sequencing technology to obtain microbial sample sequencing data; Step S2: performing relative frequency quantitative calculation on the microbial sample sequencing data to obtain microbial category relative frequency quantitative data; performing clustering behavior identification between different categories on the microbial category relative frequency quantitative data to obtain microbial category clustering behavior data; Step S3: Identifying functional genes between different microbial categories based on the microbial category clustering behavior data to obtain microbial functional gene data; constructing a metabolic pathway network between different microbial categories based on the microbial functional gene data to obtain a microbial metabolic pathway network; Step S4: construct a microbial component detection model based on the microbial metabolic pathway network to obtain a microbial component detection model; send the microbial component detection model to the cloud platform to execute the microbial component detection method.

2. The method for detecting microbial components based on high-throughput sequencing data according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: obtaining a microbial test sample; Step S12: filtering and separating the microbial test sample to obtain a microbial filtration test sample; Step S13: centrifuging the microbial filtration test sample to obtain a microbial filtration centrifuge test sample; Step S14: performing deep sequencing on the microbial filtration centrifugation test sample based on high-throughput sequencing technology to obtain microbial sample sequencing data.

3. The method for detecting microbial components based on high-throughput sequencing data according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: recombining DNA fragments of the microbial sample sequencing data to obtain recombined microbial DNA fragment data; Step S22: using a preset microbial DNA identification database to perform sequence comparison on the recombinant microbial DNA fragment data to obtain microbial DNA fragment classification data; Step S23: performing relative frequency quantitative calculation on the microbial sample sequencing data according to the microbial DNA fragment category data to obtain microbial category relative frequency quantitative data; Step S24: performing clustering behavior identification between different categories on the quantitative data of relative frequency of microorganism categories to obtain clustering behavior data of microorganism categories.

4. The method for detecting microbial components based on high-throughput sequencing data according to claim 3, characterized in that: Step S23 includes the following steps: Step S231: performing relative abundance calculations between different categories on the microbial sample sequencing data based on the microbial DNA fragment category data to obtain microbial category sequence relative abundance data; Step S232: performing an abundance distribution skewness analysis among different categories on the relative abundance data of the microbial category sequences to obtain abundance distribution skewness data; Step S233: performing a balanced quantitative analysis on the relative abundance data of the microbial category sequence according to the skewed abundance distribution data to obtain balanced quantitative data of the category sequence; Step S234: performing relative frequency quantitative calculation on the microbial sample sequencing data based on the abundance distribution skewness data and the category sequence balance quantitative data to obtain microbial category relative frequency quantitative data.

5. The method for detecting microbial components based on high-throughput sequencing data according to claim 4, characterized in that: The balanced quantitative analysis of the relative abundance data of microbial class sequences based on the skewed abundance distribution data includes the following steps: The abundance distribution skewed data are processed by abundance stratification using the preset abundance stratification interval to obtain the abundance distribution skewed stratified data; The extreme value fluctuation data of skewed abundance distribution were analyzed in different layers to obtain the extreme value fluctuation data of skewed layers. The median difference of different strata is calculated for the skewed stratified extreme value fluctuation data to obtain the skewed stratified median difference data; The skewed stratified extreme value fluctuation data are calculated by balanced interpolation according to the skewed stratified median difference data to obtain the skewed stratified balanced interpolation data; Based on the skewed stratified balanced interpolation data, the relative abundance data of microbial category sequences were subjected to balanced quantitative analysis to obtain balanced quantitative data of category sequences.

6. The method for detecting microbial components based on high-throughput sequencing data according to claim 5, characterized in that: The balanced interpolation calculation of skewed stratified extreme value fluctuation data based on skewed stratified median difference data includes the following steps: Perform high-order difference accumulation on the skewed stratified median difference data to obtain the skewed stratified high-order difference accumulation data; According to the skewness stratification high-order difference cumulative data and the skewness stratification median difference data, the skewness stratification extreme value fluctuation data is evaluated for data point continuity and smoothness, and the skewness stratification fluctuation continuous and smooth data are obtained; The skewed stratified balanced interpolation data are obtained by performing balanced interpolation calculation on the skewed stratified continuous and smooth fluctuation data.

7. The method for detecting microbial components based on high-throughput sequencing data according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: Calculating the spatial dispersion between different microbial categories based on the microbial category relative frequency quantitative data on the microbial category aggregation behavior data to obtain microbial aggregation spatial dispersion data; Step S32: performing functional gene identification between different microbial categories on the microbial sample sequencing data based on the microbial cluster spatial dispersion data and the microbial category cluster behavior data to obtain microbial functional gene data; Step S33: constructing a metabolic pathway network among different microbial categories based on the microbial functional gene data according to the relative frequency quantitative data of the microbial categories to obtain a microbial metabolic pathway network.

8. The method for detecting microbial components based on high-throughput sequencing data according to claim 7, characterized in that: Step S33 includes the following steps: Step S331: performing gene interaction frequency relationships between different microbial categories on the microbial functional gene data based on the relative frequency quantitative data of the microbial categories to obtain gene interaction frequency relationship data; Step S332: evaluating the functional impact of gene fragments on the microbial functional gene data based on the gene interaction frequency relationship data to obtain gene fragment functional impact data; Step S333: performing effective association analysis of metabolic pathways in different functional regions on the gene fragment functional impact data to obtain effective association data of metabolic pathways; Step S334: constructing a metabolic pathway network among different microbial categories based on the gene fragment functional impact data and the metabolic pathway effective association data to obtain a microbial metabolic pathway network.

9. The method for detecting microbial components based on high-throughput sequencing data according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: screening the key metabolic node functional genes of the microbial metabolic pathway network to obtain key metabolic node functional gene data; Step S42: performing expression detection on the functional gene data of key metabolic nodes to obtain the node functional gene expression data; Step S43: constructing a microbial component detection model for the node functional gene expression data according to the microbial metabolic pathway network to obtain a microbial component detection model; Step S44: Send the microbial component detection model to the cloud platform to execute the microbial component detection method.

10. The method for detecting microbial components based on high-throughput sequencing data according to claim 9, characterized in that: Step S43 includes the following steps: Step S431: performing metabolic flux balance calculation on the node functional gene expression data according to the microbial metabolic pathway network to obtain metabolic flux balance data; Step S432: performing cross-influence analysis on the metabolic flux balance data to obtain metabolic flux cross-influence data; Step S433: constructing a microbial component detection model for the node functional gene expression data based on the metabolic flow cross-influence data to obtain a microbial component detection model.

Citation Information

Patent Citations

  • Method for constructing, optimizing and visualizing genome metabolism model based on high-throughput sequencing technology

    CN113035269A

  • Microbiological analysis device and method based on 16S sequencing data

    CN117133364A