Group health evaluation method and system based on intestinal flora
By using high-throughput sequencing technology and gut microbiota analysis, the standardization problem of population health assessment has been solved, realizing a complete technical system from sample collection to health strategy formulation, supporting the effective application of public health management and health services.
Patent Information
- Application Number
- CN202511523934.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies are insufficient for early assessment and trend prediction of health status at the population level. They lack standardized operating procedures and unified evaluation standards, which fails to meet the population health management needs of government agencies and enterprises. Furthermore, the results of gut microbiota analysis are difficult to apply in public health management and health services.
Gut microbiota sample data are obtained through high-throughput sequencing technology. The composition and relative abundance of microbiota are analyzed, microbiota diversity and core microbiota characteristics are extracted, individual characteristics are compared with health reference characteristics, health status scores are calculated and classified, and a group health evaluation report is generated, including visualization charts and text analysis.
A complete technical system has been established, from sample collection to health strategy development, enabling systematic assessment of population health status and supporting the effective application of public health management and health services.
Smart Images

Figure CN121506473A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of health assessment technology, and in particular to a method and system for population-based health assessment of gut microbiota. Background Technology
[0002] Traditional health assessment methods primarily rely on individual physiological indicator testing and clinical symptom observation. These methods often focus on diagnosis after disease onset, making it difficult to achieve early assessment and trend prediction of health status at the population level. Although existing technologies have recognized the close relationship between gut microbiota and human health, related research has mostly focused on the correlation analysis between specific diseases and microbiota composition, lacking a systematic method to transform microbiota characteristics into a population health assessment system.
[0003] Currently, most gut microbiota-based detection technologies remain in the research and exploration stage, and their analysis results are often presented in the form of professional data, making them difficult to apply directly to public health management and social health services. Existing methods lack standardized operating procedures and unified evaluation criteria when processing population samples, making it difficult to compare and integrate results from different studies, thus limiting their practical application value in large-scale health management.
[0004] Furthermore, current technologies for analyzing gut microbiota are mostly limited to individual-level diagnostic recommendations, failing to establish a complete technical pathway from microbiota data to a population health profile. This limitation makes existing methods unable to meet the needs of organizations such as government agencies and enterprises for conducting population health management, and also difficult to support the formulation and effectiveness evaluation of regional public health policies. Summary of the Invention
[0005] According to a first aspect of the present invention, the present invention claims protection for a method for assessing gut microbiota-based population health, comprising the following steps: Step S1: Obtain gut microbiota sample data of the target population, wherein the gut microbiota sample data is obtained through high-throughput sequencing technology and contains gut microbial sequence information of each individual in the target population; Step S2: Analyze the gut microbial sequence information to identify the gut microbial composition of each individual and calculate the relative abundance of each species based on sequence count; Step S3: Based on the gut microbiota composition and relative abundance, extract the microbiota characteristics of each individual. The microbiota characteristics include microbiota diversity characteristics and core microbiota characteristics. Microbiota diversity characteristics are obtained by counting the number of unique microbiota and evaluating the uniformity of the relative abundance distribution of microbiota. Core microbiota characteristics are obtained by checking the existence and overall relative abundance of predefined core microbiota. Step S4: Compare the microbial characteristics of each individual with the pre-stored health reference characteristics, which are derived from the healthy population and include the average values of microbial diversity characteristics and core microbial characteristics. Calculate the health status score of each individual based on the comparison results. Step S5: Based on the health status score, divide each individual into multiple health levels, count the number of individuals in each health level in the target group, and calculate the percentage of individuals in each health level to obtain the health level distribution. Step S6: Based on the health level distribution, generate a group health assessment report, including a health level distribution chart and a text summary. The report describes the overall health status and potential risk trends of the group.
[0006] Furthermore, step S1 also includes: Fecal samples were collected from each individual in the target population using sterile sampling tools. The fecal samples were immediately frozen and stored in an ultra-low temperature environment. Intestinal microbial DNA was extracted from the fecal samples using a DNA extraction kit. The extracted intestinal microbial DNA was amplified by PCR targeting the 16S rRNA gene region and then sequenced using the Illumina sequencing platform to obtain intestinal microbial sequence information. The sampling process employs standardized sampling protocols to ensure sample consistency and comparability, including the use of uniform sampling containers, fixed sampling times, and storage conditions.
[0007] Furthermore, step S2 also includes: The gut microbial sequence information is compared with a pre-constructed gut microbiota reference database, which contains complete sequence information of a variety of known gut microbiota species. Sequence alignment software was used for similarity matching. Based on the matching results, the species identity corresponding to each sequence was determined, and the number of sequences for each species was counted. When calculating the relative abundance of each species, the number of sequences for each species is divided by the total number of sequences to obtain a relative abundance value in percentage form. In particular, setting a minimum sequence length threshold and quality filtering conditions during sequence alignment ensures the accuracy of identification.
[0008] Furthermore, step S3 also includes: The total number of unique bacterial species in each individual is counted as a richness index. The uniformity of the relative abundance distribution of bacterial species is assessed by analyzing the distribution of relative abundance values, where the uniformity of distribution is indirectly represented by calculating the coefficient of variation of relative abundance values. Extracting core microbial community features includes identifying the core microbial species present in each individual from a pre-stored list of core microbial species, which is defined based on species that are common in healthy individuals and positively correlated with health, and calculating the sum of the relative abundance of all core microbial species as the overall relative abundance; The absence of the core bacterial species is recorded as an auxiliary feature.
[0009] Furthermore, step S4 also includes: A large number of gut microbiota samples were collected from representative healthy individuals. After sequence analysis and feature extraction, the mean, median, and range of microbiota diversity characteristics and core microbiota characteristics were calculated, and these statistics were stored in a database. When comparing microbial community characteristics, for each microbial community characteristic of each individual, the absolute deviation of its value from the corresponding average value in the health reference characteristics is calculated, and the weighted sum of the deviation values of all characteristics is used to obtain the health status score. The weights are based on the predefined importance of the characteristics.
[0010] Furthermore, step S5 also includes: Set thresholds for health status scores. Individuals with health status scores above the first threshold are classified as excellent, those with scores between the first and second thresholds are classified as good, those with scores between the second and third thresholds are classified as average, and those with scores below the third threshold are classified as poor. The score thresholds are predefined based on the score distribution of healthy individuals and validated using historical data. After classification, a descriptive label is attached to each health level, including the level of health risk and directions for improvement.
[0011] Furthermore, step S6 also includes: Calculate the percentage of individuals at each health level relative to the total number of individuals in the target population, and calculate the cumulative percentage; Analyze the distribution pattern of the health levels, including identifying the main health levels and abnormal levels, and calculating distribution statistics such as mode and standard deviation; Compare the distribution of health grades among different subgroups, where subgroups are divided based on age, sex, or geographic factors.
[0012] Furthermore, step S6 also includes: Use data visualization tools to create pie charts and bar charts showing the distribution of health levels, displaying the percentage and number of individuals at each health level; The written summary includes a description of the overall health status of the group, identification of major health problems, risk level assessment, and trend analysis. Report output formats include electronic documents and interactive dashboards, allowing users to drill down into detailed data; The report generation process also includes quality control steps to ensure data consistency and report accuracy.
[0013] Furthermore, the method also includes step S7: Verify the reliability of the population health assessment results, including internal consistency checks and comparison with external benchmarks. Internal consistency checks assess the stability of the health grade distribution through resampling methods, while external benchmark comparisons compare the results with independent health survey data. After verification, a verification report is generated, including consistency indicators and deviation analysis, to optimize the subsequent evaluation process.
[0014] Furthermore, the method also includes step S8: Based on the population health assessment report, a population health management strategy is developed, which includes identifying the health levels that require priority intervention and designing targeted health intervention programs. The strategy formulation process takes into account group characteristics and resource constraints, and outputs strategy documents and action plans; Set up a tracking mechanism and update the evaluation regularly to monitor changes.
[0015] According to a second aspect of the present invention, the present invention claims protection for a gut microbiota-based population health assessment system, comprising: One or more processors; A memory storing one or more programs, which, when executed by one or more processors, enable the processors to implement the gut microbiota-based population health assessment method.
[0016] This invention relates to the field of health assessment technology, and particularly to a method and system for population-based health assessment based on gut microbiota. The method involves collecting gut microbiota samples from a target population through a standardized process, obtaining microbial sequence information using high-throughput sequencing technology, extracting microbial diversity and core microbiota characteristics through species identification and relative abundance analysis, comparing individual microbiota characteristics with a reference database of healthy individuals, calculating health status scores and classifying health levels, and then statistically analyzing the distribution of population health levels to generate a population health assessment report including visual charts and textual analysis. This invention establishes a complete technical system from sample collection to health strategy formulation, realizing a systematic assessment of population health status from the perspective of gut microbiota, overcoming the limitations of traditional methods in early population health assessment. This invention is applicable to health monitoring of populations of different sizes, providing effective technical support for public health management and health services. Attached Figure Description
[0017] Figure 1This is a flowchart illustrating a gut microbiota-based population health assessment method for which protection is sought in this invention. Figure 2 This is a second flowchart of a gut microbiota-based population health assessment method for which protection is claimed in this invention. Figure 3 This is a third flowchart of a gut microbiota-based population health assessment method for which protection is claimed in this invention. Figure 4 The fourth flowchart of a gut microbiota-based population health assessment method for which protection is sought in this invention is shown. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0019] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0020] According to the first embodiment of the present invention, referring to Figure 1 This invention claims protection for a gut microbiota-based population health assessment method, comprising the following steps: Step S1: Obtain gut microbiota sample data of the target population, wherein the gut microbiota sample data is obtained through high-throughput sequencing technology and contains gut microbial sequence information of each individual in the target population; Step S2: Analyze the gut microbial sequence information to identify the gut microbial composition of each individual and calculate the relative abundance of each species based on sequence count; Step S3: Based on the gut microbiota composition and relative abundance, extract the microbiota characteristics of each individual. The microbiota characteristics include microbiota diversity characteristics and core microbiota characteristics. Microbiota diversity characteristics are obtained by counting the number of unique microbiota and evaluating the uniformity of the relative abundance distribution of microbiota. Core microbiota characteristics are obtained by checking the existence and overall relative abundance of predefined core microbiota. Step S4: Compare the microbial characteristics of each individual with the pre-stored health reference characteristics, which are derived from the healthy population and include the average values of microbial diversity characteristics and core microbial characteristics. Calculate the health status score of each individual based on the comparison results. Step S5: Based on the health status score, each individual is divided into multiple health levels. The number of individuals in each health level in the target group is counted, and the percentage of individuals in each health level is calculated to obtain the health level distribution. Step S6: Based on the health level distribution, generate a group health assessment report, including a health level distribution chart and a text summary. The report describes the overall health status and potential risk trends of the group.
[0021] In this embodiment, each step further includes: Gut microbiota sample data of the target population were obtained through high-throughput sequencing technology. Specifically, this involved collecting fecal samples from each individual using sterile sampling tools, immediately freezing the fecal samples in an ultra-low temperature environment, extracting gut microbial DNA from the fecal samples using a DNA extraction kit, performing PCR amplification on the extracted DNA targeting the 16S rRNA gene region, and generating gut microbial sequence information through a high-throughput sequencing platform. The sampling process adopted standardized protocols to ensure sample consistency and comparability, including uniform sampling containers, fixed sampling time points, and storage conditions.
[0022] The gut microbiome sequence information was analyzed to identify the gut microbiota composition of each individual. Specifically, the sequence information was compared with a pre-constructed gut microbiota reference database, and similarity matching was performed using sequence alignment software. Based on the matching results, the species identity corresponding to each sequence was determined, and the number of sequences for each species was counted. The relative abundance of each species was calculated based on the sequence count, and the number of sequences for each species was divided by the total number of sequences to obtain a percentage relative abundance value. During the sequence alignment process, a minimum sequence length threshold and quality filtering conditions were set to exclude low-quality sequences.
[0023] Based on the composition and relative abundance of gut microbiota, the microbiota characteristics of each individual are extracted. These characteristics include microbiota diversity and core microbiota characteristics. Microbiota diversity is assessed by counting the number of unique species as a richness index, and the evenness of the relative abundance distribution is evaluated by analyzing the distribution of relative abundance values. Core microbiota characteristics are assessed by checking the presence of a predefined set of core species, which are defined based on species common in healthy individuals and positively correlated with health. The sum of the relative abundance of all core species is calculated as the overall relative abundance. In addition, the absence of core species is recorded as an auxiliary feature.
[0024] Each individual's microbial community characteristics are compared with pre-stored health reference characteristics, which are derived from a healthy population and include the mean, median, and range of microbial diversity characteristics and core microbial community characteristics. During the comparison, for each microbial community characteristic of each individual, the absolute deviation of its value from the corresponding mean in the health reference characteristics is calculated, and the deviation values of all characteristics are weighted and summed to obtain the health status score. The weights are based on the predefined importance of the features, and the importance is determined by the strength of the association between the features and health status in historical data.
[0025] Based on health status scores, each individual is divided into multiple health levels, including excellent, good, average, and poor. Thresholds are set for health status scores during the classification: individuals with scores above the first threshold are classified as excellent; those with scores between the first and second thresholds are classified as good; those with scores between the second and third thresholds are classified as average; and those with scores below the third threshold are classified as poor. These thresholds are predefined based on the score distribution of the healthy population and validated through statistical methods.
[0026] The number of individuals at each health level in the target group is counted, and the percentage of individuals at each health level is calculated to obtain the health level distribution. At the same time, the cumulative percentage is calculated, and the distribution pattern of health levels is analyzed, including identifying the major health levels and abnormal levels, and calculating distribution statistics such as mode and standard deviation.
[0027] Based on the health level distribution, a population health assessment report is generated, including health level distribution charts and text summaries; pie charts and bar charts are created using data visualization tools to display the percentage and number of individuals at each health level; the text summary includes a description of the overall health status of the population, identification of major health problems, risk level assessment, and trend analysis; the report output formats include electronic documents and interactive dashboards, and include quality control steps to ensure data consistency.
[0028] Furthermore, step S1 also includes: Fecal samples were collected from each individual in the target population using sterile sampling tools. The fecal samples were immediately frozen and stored in an ultra-low temperature environment. Intestinal microbial DNA was extracted from the fecal samples using a DNA extraction kit. The extracted intestinal microbial DNA was amplified by PCR targeting the 16S rRNA gene region and then sequenced using the Illumina sequencing platform to obtain intestinal microbial sequence information. The sampling process employs standardized sampling protocols to ensure sample consistency and comparability, including the use of uniform sampling containers, fixed sampling times, and storage conditions.
[0029] In this embodiment, a sterile fecal collection container was used to collect fecal samples from each individual in the target population. Immediately after sampling, the samples were stored in an ultra-low temperature freezer at -80°C or below. Dry ice was used to maintain the low temperature during sample transport. Intestinal microbial DNA was extracted from the fecal samples using a commercial DNA extraction kit. The extraction process included cell lysis, DNA binding, washing, and elution steps, strictly following the kit instructions. The extracted DNA was amplified by PCR, targeting the V3-V4 region of the 16S rRNA gene. Universal primers were used for amplification, and the PCR reaction conditions included initial denaturation, cyclic denaturation, annealing, and extension steps. The number of cycles was determined according to a standard protocol. After purification, the amplified product was sequenced using the Illumina sequencing platform to generate paired-end sequence reads. The sequencing process included library construction and cluster generation steps to ensure the quality and length of the sequence reads met requirements. The entire sampling and sequencing process adopted standardized protocols, including using uniformly sized sampling containers, fixing sampling time points to avoid dietary interference, and recording sample storage timestamps to ensure sample consistency and comparability.
[0030] Furthermore, referring to Figure 2 Step S2 also includes: The gut microbial sequence information is compared with a pre-constructed gut microbiota reference database, which contains complete sequence information of a variety of known gut microbiota species. Sequence alignment software was used for similarity matching. Based on the matching results, the species identity corresponding to each sequence was determined, and the number of sequences for each species was counted. When calculating the relative abundance of each species, the number of sequences for each species is divided by the total number of sequences to obtain a relative abundance value in percentage form. In particular, setting a minimum sequence length threshold and quality filtering conditions during sequence alignment ensures the accuracy of identification.
[0031] In this embodiment, gut microbial sequence information is compared with a pre-constructed gut microbiota reference database. The reference database contains complete 16S rRNA sequence information of various known gut microbiota species and is regularly updated to cover newly discovered species. Sequence alignment software is used for similarity matching. During matching, a minimum sequence similarity threshold of 97% is set to define the species level, and a minimum sequence length threshold is set to filter out sequences shorter than 200 bp. After alignment, each sequence is assigned to a specific species based on the matching results, and the number of sequences for each species is counted. When calculating relative abundance, the number of sequences for each species is divided by the total number of sequences to obtain a percentage relative abundance value. At the same time, quality control is performed to exclude sequences with low alignment quality to ensure identification accuracy. The entire analysis process is performed in a computing environment, including sequence preprocessing such as noise reduction and chimera removal.
[0032] Furthermore, referring to Figure 3 Step S3 also includes: The total number of unique bacterial species in each individual is counted as a richness index. The uniformity of the relative abundance distribution of bacterial species is assessed by analyzing the distribution of relative abundance values, where the uniformity of distribution is indirectly represented by calculating the coefficient of variation of relative abundance values. Extracting core microbial community features includes identifying the core microbial species present in each individual from a pre-stored list of core microbial species, which is defined based on species that are common in healthy individuals and positively correlated with health, and calculating the sum of the relative abundance of all core microbial species as the overall relative abundance; The absence of the core bacterial species is recorded as an auxiliary feature.
[0033] In this embodiment, the total number of unique bacterial species in each individual is counted as a richness index, and the richness calculation is based on the count of unique bacterial species identifiers in the sequence alignment results. Then, the uniformity of the relative abundance distribution of bacterial species is evaluated by analyzing the distribution of relative abundance values, specifically including calculating the coefficient of variation of relative abundance values or observing the distribution range of abundance values. The uniformity assessment also involves visually inspecting by plotting relative abundance distribution curves. Extracting core microbial community features includes: identifying the core bacterial species present in each individual from a pre-stored list of core bacterial species, the core bacterial species list being based on the definition of bacteria common in healthy individuals and positively correlated with health, including key species in Bacteroidetes and Firmicutes; calculating the sum of the relative abundance of all core bacterial species as the overall relative abundance; simultaneously, recording the absence of core bacterial species, and counting the number and type of missing species; the feature extraction process also includes verifying the consistency of features, ensuring the reproducibility of results by repeatedly analyzing a subset of samples.
[0034] Furthermore, step S4 also includes: A large number of gut microbiota samples were collected from representative healthy individuals. After sequence analysis and feature extraction, the mean, median, and range of microbiota diversity characteristics and core microbiota characteristics were calculated, and these statistics were stored in a database. When comparing microbial community characteristics, for each microbial community characteristic of each individual, the absolute deviation of its value from the corresponding average value in the health reference characteristics is calculated, and the weighted sum of the deviation values of all characteristics is used to obtain the health status score. The weights are based on the predefined importance of the characteristics.
[0035] In this embodiment, a large number of gut microbiota samples are collected from a representative healthy population. The inclusion criteria for healthy individuals include the absence of chronic diseases and no recent history of antibiotic use. Sequence analysis and feature extraction are performed on the healthy population samples to calculate the mean, median, and range of microbiota diversity characteristics and core microbiota characteristics, and these statistics are stored in a database. The database is updated regularly to reflect population changes. When comparing microbiota characteristics, for each microbiota characteristic of each individual, the absolute deviation from the corresponding mean in the health reference characteristics is calculated. Then, the deviation values of all characteristics are weighted and summed to obtain a health status score. The weights are based on the predefined importance of the characteristics in health. The importance is determined by the correlation strength between the characteristics and health status in historical data, and the correlation strength is obtained through statistical analysis such as correlation calculation. The score calculation process also includes adjusting the weights based on the stability of the characteristics to ensure the reliability of the evaluation.
[0036] Furthermore, step S5 also includes: Set thresholds for health status scores. Individuals with health status scores above the first threshold are classified as excellent, those with scores between the first and second thresholds are classified as good, those with scores between the second and third thresholds are classified as average, and those with scores below the third threshold are classified as poor. The score thresholds are predefined based on the score distribution of healthy individuals and validated using historical data. After classification, a descriptive label is attached to each health level, including the level of health risk and directions for improvement.
[0037] In this embodiment, classifying individuals into multiple health levels includes: setting threshold values for health status scores, where the threshold values are predefined based on the score distribution of the healthy population and determined using statistical methods such as percentiles. The first threshold corresponds to the upper quartile of the healthy population's scores, the second threshold corresponds to the median, and the third threshold corresponds to the lower quartile. During classification, individuals with health status scores above the first threshold are classified as excellent, those with scores between the first and second thresholds are classified as good, those with scores between the second and third thresholds are classified as average, and those with scores below the third threshold are classified as poor. After classification, a descriptive label is attached to each health level, including the level of health risk and suggested improvement directions. The classification process also includes verifying the rationality of the thresholds by testing their effectiveness using independent datasets through cross-validation to ensure the accuracy of the classification.
[0038] Furthermore, referring to Figure 4 Step S6 also includes: Calculate the percentage of individuals at each health level relative to the total number of individuals in the target population, and calculate the cumulative percentage; Analyze the distribution pattern of the health levels, including identifying the main health levels and abnormal levels, and calculating distribution statistics such as mode and standard deviation; Compare the distribution of health grades among different subgroups, where subgroups are divided based on age, sex, or geographic factors.
[0039] In this embodiment, the percentage of individuals in each health level relative to the total number of individuals in the target group is calculated, and the cumulative percentage is also calculated. Simultaneously, the distribution pattern of health levels is analyzed, including identifying major and abnormal health levels, and calculating distribution statistics such as mode and standard deviation. The distribution analysis also includes comparing the health level distributions among different subgroups, where subgroups are defined based on age, gender, or geographic factors. The statistical process also includes generating a distribution summary table displaying the count, percentage, and cumulative percentage for each level. Furthermore, a distribution consistency check is performed, and the stability of the distribution is assessed using a resampling method to ensure the reliability of the results.
[0040] Furthermore, step S6 also includes: Use data visualization tools to create pie charts and bar charts showing the distribution of health levels, displaying the percentage and number of individuals at each health level; The written summary includes a description of the overall health status of the group, identification of major health problems, risk level assessment, and trend analysis. Report output formats include electronic documents and interactive dashboards, allowing users to drill down into detailed data; The report generation process also includes quality control steps to ensure data consistency and report accuracy.
[0041] In this embodiment, data visualization tools are used to create pie charts and bar charts showing the distribution of health levels. The pie charts display the percentage of each health level, and the bar charts display the number of individuals at each level. The text summary includes a description of the overall health status of the group, identification of major health problems, risk level assessment, and trend analysis. The report output formats include PDF documents and interactive dashboards. The interactive dashboards allow users to drill down into detailed data by clicking on charts, such as viewing a list of individuals at a specific health level. The report generation process includes quality control steps to ensure data consistency and report accuracy. Quality control includes verifying data sources, validating the consistency between charts and data, and checking the logic of the text descriptions. The report is ultimately generated automatically by the system and supports manual editing and annotation.
[0042] Furthermore, the method also includes step S7: Verify the reliability of the population health assessment results, including internal consistency checks and comparison with external benchmarks. Internal consistency checks assess the stability of the health grade distribution through resampling methods, while external benchmark comparisons compare the results with independent health survey data. After verification, a verification report is generated, including consistency indicators and deviation analysis, to optimize the subsequent evaluation process.
[0043] In this embodiment, the reliability of the population health assessment results is verified. The verification process includes internal consistency checks and external benchmark comparisons. The internal consistency check assesses the stability of the health level distribution through a resampling method, specifically by randomly selecting subsamples from the target population and repeating the evaluation process multiple times, calculating the standard deviation of the distribution change. The external benchmark comparison compares the results with independent health survey data, using only health indicators from the survey data as a reference. After verification, a verification report is generated, including consistency indicators and bias analysis. The consistency indicators include distribution similarity measures, and the bias analysis includes identifying systematic biases. The verification report is used to optimize subsequent evaluation processes, such as adjusting thresholds or feature weights.
[0044] Furthermore, the method also includes step S8: Based on the population health assessment report, a population health management strategy is developed, which includes identifying the health levels that require priority intervention and designing targeted health intervention programs. The strategy formulation process takes into account group characteristics and resource constraints, and outputs strategy documents and action plans; Set up a tracking mechanism and update the evaluation regularly to monitor changes.
[0045] In this embodiment, the method further includes: developing a group health management strategy based on the group health assessment report. This strategy development includes identifying priority health levels for intervention, typically targeting poor and average levels; designing targeted health intervention programs, such as dietary recommendations, probiotic supplementation, or lifestyle modifications; considering group characteristics and resource constraints during strategy development, including age distribution, gender ratio, and available resources; outputting a strategy document and action plan, the strategy document including intervention goals, a list of measures, and a timeline; and establishing a tracking mechanism to regularly update the assessment to monitor changes. This tracking mechanism includes setting review cycles and indicator tracking tables to ensure the effective implementation of the strategy.
[0046] According to a second embodiment of the present invention, the present invention claims protection for a gut microbiota-based population health assessment system, comprising: One or more processors; A memory storing one or more programs, which, when executed by one or more processors, enable the processors to implement the gut microbiota-based population health assessment method.
[0047] The following is a specific example: The evaluation subjects were selected from employees of a large enterprise, totaling 800 participants. Detailed standardized operating procedures were developed before sampling, including sampling time arrangements, sample labeling specifications, and storage and transportation requirements. Sterile fecal collection containers of uniform specifications were used, with sample stabilizing solution pre-filled inside the containers to maintain microbial activity. Sampling is conducted at fixed times in the mornings from Monday to Wednesday each week. Participants are required to fill out a simple diet record form before sampling. After sampling, the container is immediately sealed, labeled with a number including the sample number, collection time, and participant number, and placed in a dedicated biological sample transport box. The temperature inside the box is always maintained below -80°C. Temperature recorders are used to monitor the entire transportation process to ensure the integrity of the cold chain.
[0048] Upon arrival at the laboratory, samples undergo a quality check to exclude those with leaks, unclear labels, or abnormal temperatures. The DNA extraction is then performed using a certified DNA extraction kit. Accurately weigh 0.2g of fecal sample, add lysis buffer containing glass beads, homogenize thoroughly using a vortex mixer, then incubate at a specific temperature to lyse the cells, centrifuge, take the supernatant and pass it through a silica membrane adsorption column, wash several times to remove impurities, and finally collect and purify DNA with elution buffer. The extracted DNA underwent quality testing, including concentration determination and purity analysis. PCR amplification was performed on the V3-V4 hypervariable region of the 16S rRNA gene using universal primers with sequencing adapters. The reaction system included DNA template, primers, polymerase, and buffer. The amplification program was optimized, including initial denaturation, a denaturation-annealing-extension cycle with a specific number of cycles, and a final extension. After purification, the amplified products were sequenced using a high-throughput sequencing platform to obtain a sufficient number of sequence reads for each sample.
[0049] The raw sequencing data was quality filtered to remove low-quality sequences and adapter sequences. The high-quality sequences were then compared with a carefully constructed gut microbiota reference database, which contains complete 16S rRNA sequence information of various gut microbiota species that have been manually verified. Strict similarity thresholds and sequence length requirements are set during the comparison process, and an optimized comparison algorithm is used to ensure the accuracy of bacterial species identification; For each sample, the sequence counts of each bacterial species were counted. When calculating the relative abundance, a standardized method was used, which divided the sequence count of a single bacterial species by the total number of valid sequences in the sample to obtain the abundance value in percentage form. At the same time, a sample quality assessment system was established to exclude samples with insufficient sequencing depth or poor quality.
[0050] Based on the strain identification results, the system extracts two types of characteristic indicators: Microbial diversity characteristics include the statistical analysis of the number of unique species and the assessment of the uniformity of the relative abundance distribution of species; By observing the morphological characteristics of the abundance distribution curve, we can analyze the distribution patterns of dominant and rare species and record the proportion of high-abundance and low-abundance species. The core microbiota characteristics are determined by comparing them against a predefined list of core microbial species, which contains microbial species that have been validated by extensive research and are closely related to human health. The overall relative abundance of these core species was calculated, and detailed information on missing core species was recorded, including their taxonomic location and functional characteristics.
[0051] Establish a health reference database, which is derived from rigorously screened healthy population sample data and includes the normal range and distribution characteristics of various microbial community feature parameters; For each participant's microbial community characteristic data, each item is compared with the corresponding parameter in the reference database, and the degree of characteristic deviation is calculated. Based on the pre-set importance weights of the features, the deviation values of each feature are weighted and integrated to obtain a comprehensive health status score; The weighting was based on a large number of previous studies, taking into account the strength of the association between different microbial community characteristics and health status and their biological significance.
[0052] A multi-level health grading system is established based on the population distribution characteristics of health status scores. By analyzing the distribution patterns of the scoring data, reasonable thresholds for grading were determined, and participants were divided into four health levels: Excellent level indicates an ideal gut microbiota state with all characteristic parameters within the optimal range; Good level indicates a relatively good gut microbiota state with key characteristic parameters meeting health requirements; Average level indicates a gut microbiota state with some deviations and some characteristic parameters needing improvement; Poor level indicates a significantly abnormal gut microbiota state requiring close attention and intervention. Detailed descriptive definitions and characteristic criteria were developed for each level.
[0053] Multidimensional statistical analysis was conducted on the health level distribution of the target group, calculating the absolute number and relative percentage of each level, analyzing the central tendency and dispersion of the distribution, examining the distribution characteristics of different subgroups (such as age-segmented or gender-grouped), comparing the difference patterns between groups, and conducting in-depth analysis of the distribution pattern, including distribution symmetry assessment, outlier detection, and distribution pattern identification.
[0054] Based on the statistical analysis results, a detailed population health assessment report is compiled, which includes several components: First, a visual chart is used to show the distribution of levels, including a percentage chart and a quantity chart; second, a detailed textual analysis is provided, which describes the overall health status of the group, its main characteristics, existing problems, and suggestions for improvement. It also includes comparative analyses of various subgroups and explanations of special cases. The report is available in multiple output formats, providing both printable document formats and interactive electronic versions to meet different user needs.
[0055] A multi-level validation system is established to ensure the reliability of the evaluation results; internal validation uses resampling technology to assess the stability of the results through random subsampling; external validation analyzes the consistency of the results by comparing them with other health assessment data; at the same time, a full-process quality control record is established, including quality parameters of each step of sample processing, quality control data of the experimental process, and quality assessment reports of data analysis.
[0056] Based on the evaluation results, targeted health management plans were developed for different groups. Differentiated intervention measures were designed according to the characteristics and needs of different health levels. For the excellent health level group, the focus was on maintaining the status quo and regular monitoring; for the good health level group, optimization suggestions and preventative guidance were provided; for the average health level group, clear improvement goals and specific measures were formulated; and for the poor health level group, targeted interventions and personalized guidance were implemented. Simultaneously, an effectiveness evaluation mechanism was established to regularly track the intervention effects.
[0057] Examples of applications for health assessment targeting specific populations: The elderly residents of a specific community were selected as the evaluation subjects, with participants aged 65-85. Considering the physiological characteristics and medication use of the elderly, detailed information on each participant's basic information, health status, and medication history were recorded before sample collection. The sampling process was specifically arranged to take place at the community health service center, under the guidance of professional medical staff to ensure standardized sampling procedures. Home sampling services were provided for some participants with mobility difficulties. Expedited cold chain logistics were used for sample transportation to minimize sample processing time.
[0058] The evaluation methods were appropriately adjusted based on the characteristics of the elderly population; the list of core microbial strains was expanded to include strains closely related to the health of the elderly, with particular attention paid to strains related to immune function and metabolic health. Regarding the health reference database, a specially established database of healthy gut microbiota for the elderly is used. This database contains a large amount of rigorously screened data on healthy elderly individuals. When setting feature weights, the unique physiological state and health needs of the elderly are taken into account, and the weight ratios of different features are appropriately adjusted.
[0059] To address the characteristics of samples from the elderly population, laboratory processing procedures were optimized. During DNA extraction, the lysis time was appropriately increased to ensure full release of microbial DNA. During PCR amplification, the number of cycles was optimized to accommodate potentially lower microbial loads. In the data analysis phase, special attention was paid to excluding influencing factors such as antibiotic use to ensure the accuracy of the evaluation results.
[0060] After obtaining the health level distribution of the elderly population, an in-depth analysis was conducted. The characteristics of the lower-risk groups were highlighted, and possible factors leading to poor gut microbiota status were analyzed. The correlation between gut microbiota status and health status was explored by combining participants' health record information. Based on the analysis results, a health intervention program suitable for the characteristics of the elderly was developed, emphasizing nutritional support, appropriate exercise, and guidance on rational medication use.
[0061] Examples of applications for long-term health monitoring A group of volunteers was selected as the long-term monitoring subject, and a twelve-month follow-up evaluation plan was established. A comprehensive gut microbiota test and health status assessment should be conducted every three months. A detailed schedule should be established to ensure consistent testing conditions for each session. Complete participant profiles should be created, recording the results of each test and related information such as environment, diet, and lifestyle.
[0062] Each test is conducted strictly according to standard operating procedures to ensure the comparability of results. Sample collection is carried out at fixed time intervals, using the same batch of sampling materials and reagents; Laboratory processes were kept under consistent conditions and operated by the same team of researchers; data analysis employed the same parameter settings and evaluation criteria to ensure the consistency of the result sequence.
[0063] By comparing data from multiple tests, the dynamic changes in gut microbiota at both the individual and group levels were analyzed. Patterns of health level transitions were observed to identify trends of improvement or deterioration. The changing patterns of gut microbiota characteristic parameters were analyzed to explore factors influencing gut microbiota stability. In conjunction with information such as participants' lifestyle changes, possible causes of gut microbiota changes were interpreted.
[0064] Based on long-term monitoring data, a gut microbiota health early warning mechanism will be established. For individuals showing a deteriorating trend, timely intervention recommendations will be proposed; effective intervention measures will be analyzed, and effective methods to promote gut microbiota health improvement will be summarized. A predictive model for gut microbiota health changes will be established to provide a scientific basis for health management.
[0065] This series of examples fully demonstrates the application of gut microbiota-based population health assessment methods in different scenarios.
[0066] The implementation examples were designed to consider various scenarios in practical applications, including standardized requirements for sample processing, key quality control points for laboratory operations, rigor of data analysis, and targeted application of results. Demonstration of multiple application scenarios fully demonstrates the practicality, adaptability, and innovation of this technical solution.
[0067] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0068] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. The above are merely embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
[0069] The specific embodiments of the invention have been described in detail above, but they are only examples, and this application is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications or substitutions to the invention are also within the scope of this application. Therefore, all equivalent changes, modifications, and improvements made without departing from the spirit and principles of this application should be covered within the scope of this application.
Claims
1. A method for population-based health assessment based on gut microbiota, characterized in that, Includes the following steps: Step S1: Obtain gut microbiota sample data of the target population, wherein the gut microbiota sample data is obtained through high-throughput sequencing technology and contains gut microbial sequence information of each individual in the target population; Step S2: Analyze the gut microbial sequence information to identify the gut microbial composition of each individual and calculate the relative abundance of each species based on sequence count; Step S3: Based on the gut microbiota composition and relative abundance, extract the microbiota characteristics of each individual. The microbiota characteristics include microbiota diversity characteristics and core microbiota characteristics. Microbiota diversity characteristics are obtained by counting the number of unique microbiota and evaluating the uniformity of the relative abundance distribution of microbiota. Core microbiota characteristics are obtained by checking the existence and overall relative abundance of predefined core microbiota. Step S4: Compare the microbial characteristics of each individual with the pre-stored health reference characteristics, which are derived from the healthy population and include the average values of microbial diversity characteristics and core microbial characteristics. Calculate the health status score of each individual based on the comparison results. Step S5: Based on the health status score, divide each individual into multiple health levels, count the number of individuals in each health level in the target group, and calculate the percentage of individuals in each health level to obtain the health level distribution. Step S6: Based on the health level distribution, generate a group health assessment report, including a health level distribution chart and a text summary. The report describes the overall health status and potential risk trends of the group.
2. The method for population-based health assessment based on gut microbiota according to claim 1, characterized in that, Step S1 further includes: Fecal samples were collected from each individual in the target population using sterile sampling tools. The fecal samples were immediately frozen and stored in an ultra-low temperature environment. Intestinal microbial DNA was extracted from the fecal samples using a DNA extraction kit. The extracted intestinal microbial DNA was amplified by PCR targeting the 16S rRNA gene region and then sequenced using the Illumina sequencing platform to obtain intestinal microbial sequence information. The sampling process employs standardized sampling protocols to ensure sample consistency and comparability, including the use of uniform sampling containers, fixed sampling times, and storage conditions.
3. The method for population-based health assessment based on gut microbiota according to claim 1, characterized in that, Step S2 also includes: The gut microbial sequence information is compared with a pre-constructed gut microbiota reference database, which contains complete sequence information of a variety of known gut microbiota species. Sequence alignment software was used for similarity matching. Based on the matching results, the species identity corresponding to each sequence was determined, and the number of sequences for each species was counted. When calculating the relative abundance of each species, the number of sequences for each species is divided by the total number of sequences to obtain a relative abundance value in percentage form. In particular, setting a minimum sequence length threshold and quality filtering conditions during sequence alignment ensures the accuracy of identification.
4. The method for population-based health assessment based on gut microbiota according to claim 1, characterized in that, Step S3 also includes: The total number of unique bacterial species in each individual is counted as a richness index. The uniformity of the relative abundance distribution of bacterial species is assessed by analyzing the distribution of relative abundance values, where the uniformity of distribution is indirectly represented by calculating the coefficient of variation of relative abundance values. Extracting core microbial community features includes identifying the core microbial species present in each individual from a pre-stored list of core microbial species, which is defined based on species that are common in healthy individuals and positively correlated with health, and calculating the sum of the relative abundance of all core microbial species as the overall relative abundance; The absence of the core bacterial species is recorded as an auxiliary feature.
5. The method for population-based health assessment based on gut microbiota according to claim 1, characterized in that, Step S4 also includes: A large number of gut microbiota samples were collected from representative healthy individuals. After sequence analysis and feature extraction, the mean, median, and range of microbiota diversity characteristics and core microbiota characteristics were calculated, and these statistics were stored in a database. When comparing microbial community characteristics, for each microbial community characteristic of each individual, the absolute deviation of its value from the corresponding average value in the health reference characteristics is calculated, and the weighted sum of the deviation values of all characteristics is used to obtain the health status score. The weights are based on the predefined importance of the characteristics.
6. The method for population-based health assessment based on gut microbiota according to claim 1, characterized in that, Step S5 also includes: Set thresholds for health status scores. Individuals with health status scores above the first threshold are classified as excellent, those with scores between the first and second thresholds are classified as good, those with scores between the second and third thresholds are classified as average, and those with scores below the third threshold are classified as poor. The score thresholds are predefined based on the score distribution of healthy individuals and validated using historical data. After classification, a descriptive label is attached to each health level, including the level of health risk and directions for improvement.
7. The method for population-based health assessment based on gut microbiota according to claim 1, characterized in that, Step S6 also includes: Calculate the percentage of individuals at each health level relative to the total number of individuals in the target population, and calculate the cumulative percentage; Analyze the distribution pattern of the health levels, including identifying the main health levels and abnormal levels, and calculating distribution statistics such as mode and standard deviation; Compare the distribution of health grades among different subgroups, where subgroups are divided based on age, sex, or geographic factors.
8. The method for population-based health assessment based on gut microbiota according to claim 1, characterized in that, Step S6 also includes: Use data visualization tools to create pie charts and bar charts showing the distribution of health levels, displaying the percentage and number of individuals at each health level; The written summary includes a description of the overall health status of the group, identification of major health problems, risk level assessment, and trend analysis. Report output formats include electronic documents and interactive dashboards, allowing users to drill down into detailed data; The report generation process also includes quality control steps to ensure data consistency and report accuracy.
9. The method for population-based health assessment based on gut microbiota according to claim 1, characterized in that, It also includes step S7: Verify the reliability of the population health assessment results, including internal consistency checks and comparison with external benchmarks. Internal consistency checks assess the stability of the health grade distribution through resampling methods, while external benchmark comparisons compare the results with independent health survey data. After verification, a verification report is generated, including consistency indicators and deviation analysis, to optimize the subsequent evaluation process; It also includes step S8: Based on the population health assessment report, a population health management strategy is developed, which includes identifying the health levels that require priority intervention and designing targeted health intervention programs. The strategy formulation process takes into account group characteristics and resource constraints, and outputs strategy documents and action plans; Set up a tracking mechanism and update the evaluation regularly to monitor changes.
10. A gut microbiota-based population health assessment system, characterized in that, include: One or more processors; A memory having stored one or more programs that, when executed by one or more processors, cause the one or more processors to implement a gut microbiota-based population health assessment method according to any one of claims 1 to 9.
Citation Information
Cited By
Intelligent analysis and natural language report generation system and method for traffic simulation result
CN121835646A
Intelligent analysis and natural language report generation system and method for traffic simulation results
CN121835646B