Wildlife non-invasive genomic monitoring and population health assessment system
By collecting and analyzing microbial and environmental data from wild animal fecal samples, and utilizing microbial succession models and a health baseline database, the initial gut microbiota map was reconstructed. This solved the problem of misdiagnosis in health assessment of wild fecal samples over time, enabling accurate diagnosis and early warning of the physiological state of wild animals.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHAANXI INST OF ZOOLOGY NORTHWEST INSTOF ENDANGERED ZOOLOGICAL SPECIES
- Filing Date
- 2026-03-25
- Publication Date
- 2026-06-23
AI Technical Summary
Existing technologies cannot effectively overcome the problem of misdiagnosis in health assessment caused by changes in the microbial composition of fecal samples collected in the field over time. In particular, the accuracy and reliability of old samples are insufficient, and false positive signals from environmental bacteria cannot be ruled out.
Raw sequencing data and microenvironment data are acquired through the biological data acquisition module. Using a pre-built microbial succession time model and core gut microbiota baseline library, time indicator features are generated, the duration of sample environmental exposure is calculated, environmental background microbiota noise is removed, the initial gut microbiota map is reconstructed, and physiological health assessment is performed in conjunction with a health baseline.
It enables reliable health assessment of old fecal samples, eliminates environmental noise interference, improves the accuracy and reliability of health diagnosis, and supports long-term, low-interference wildlife health monitoring and conservation management.
Smart Images

Figure CN122266770A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of animal health monitoring technology and relates to a non-invasive genome monitoring and population health assessment system for wild animals. Background Technology
[0002] The gut microbiota, as a complex "invisible organ" in animals, plays an irreplaceable role in maintaining individual health and physiological homeostasis due to its composition and functional balance directly participating in the host's nutritional metabolism, immune regulation, and energy supply. Especially for wild animals that are difficult to observe directly in the wild, the structure of the gut microbiota can indirectly reflect their dietary adaptability, energy acquisition efficiency, and potential immune status, serving as a key biological window for assessing their physiological health. Meanwhile, non-invasive fecal monitoring, which does not require capturing or disturbing individual animals, has become an important tool in wildlife health ecology research. By analyzing the microorganisms, metabolites, and genetic material in feces, crucial information on animal physiological status, pathogen carriage, and environmental adaptability can be obtained under natural conditions, which is of great significance for species conservation and population management.
[0003] It is noteworthy that once feces are removed from the animal's internal environment, their internal microecological system is exposed to external conditions. The original anaerobic environment is disrupted, and the microbial community structure undergoes continuous and dramatic succession. Obligate anaerobic bacteria rapidly die under oxygen exposure, while aerobic or facultative anaerobic bacteria from soil, air, or humic environments proliferate in large numbers. This process is driven by multiple factors such as environmental temperature, humidity, and ultraviolet radiation, causing the microbial composition of fecal samples to deviate significantly from their original intestinal state within hours.
[0004] Currently, the mainstream method for collecting fecal samples in the field still involves directly performing microbial DNA sequencing and bioinformatics analysis, comparing the results with a healthy baseline or pathogen database. However, this method has a fundamental flaw: traditional analysis relies on static reference databases and lacks modeling of the decay / growth dynamics of microorganisms over time in vitro, failing to reconstruct the original microbial community state at the moment of sample excretion. This makes health assessment results heavily dependent on sample "freshness" and cannot eliminate false positive signals from environmental microbial colonization, especially for samples exposed for extended periods, easily leading to misdiagnosis, such as misjudging normal individuals as having dysbiosis or malnutrition. This limits the reliability and applicability of non-invasive surveillance in long-term field research, disease early warning, and conservation decision-making. Summary of the Invention
[0005] In view of the problems existing in the prior art, the present invention provides a non-invasive genome monitoring and population health assessment system for wild animals to solve the above-mentioned technical problems.
[0006] To achieve the above and other objectives, the technical solution adopted by the present invention is as follows:
[0007] This invention provides a non-invasive genome monitoring and population health assessment system for wild animals, which includes the following modules:
[0008] Biological data acquisition module: Acquires raw sequencing data of wild animal fecal samples and microenvironment data of sampling points. The raw sequencing data includes the raw relative abundance of various microorganisms.
[0009] Sample exposure calculation module: Based on the original relative abundance, the abundance data of the predetermined indicator genera are screened and the ratio is calculated to generate time indicator marker features; the preset microbial succession time model is called, and the time indicator marker features and the microenvironmental data of the sampling point are used as input variables to calculate the sample environmental exposure duration;
[0010] Sample biological classification module: Based on the original relative abundance, the microorganisms are classified using a pre-set core gut microbiota baseline library and a typical environmental saprophytic microbiota library to determine the gut native microbiota to be corrected and the environmental background microbiota to be removed;
[0011] The gut microbiota generation module performs inverse abundance attenuation compensation calculations on the original gut microbiota to be corrected based on the sample environmental exposure duration and the microenvironment data of the sampling points, and performs proliferation noise removal calculations on the background microbiota to be removed, generating an initial gut microbiota map.
[0012] Individual health assessment module: Calculates physiological health indicators based on the initial gut microbiota profile, compares the physiological health indicators with the preset health baseline, and generates a physiological health diagnosis report for the individual wild animal.
[0013] Another aspect of the present invention provides a non-invasive genome monitoring and population health assessment device for wild animals, including a processor, a memory, and a communication bus;
[0014] The memory stores a computer-readable program that can be executed by the processor;
[0015] The communication bus enables communication between the processor and the memory;
[0016] When the processor executes the computer-readable program, it performs modules to implement a non-invasive genome monitoring and population health assessment system for wild animals as described in any one of the present invention.
[0017] As described above, the non-invasive genome monitoring and population health assessment system for wild animals provided by this invention has at least the following beneficial effects:
[0018] 1. The non-invasive genome monitoring and population health assessment system for wild animals provided by this invention generates time indicator markers based on the ratio of predetermined indicator bacterial genera, and combines this with a pre-trained microbial succession model using microenvironmental data input to intelligently infer the actual exposure duration of samples in the wild. Then, using this key parameter, the abundance of native gut bacteria that have decreased due to exposure is inversely compensated, and background bacteria proliferating in the environment are noise-removed, thereby mathematically reconstructing the initial gut microbiota map at the moment of sample excretion. This process fundamentally overcomes the core problem of microbiota distortion caused by environmental exposure, transforming the traditional passive constraint of relying on sample freshness into an active analytical capability that can be corrected by algorithms. This allows even old samples exposed for several days to be used to obtain highly reliable host physiological information, greatly expanding the practical window and application scope of non-invasive monitoring.
[0019] 2. This invention further deeply couples the reconstructed high-preservation fungal community map with health assessment, enabling scientific diagnosis and early warning of the physiological state of individual wild animals. Based on the restored initial microbial community map, the system calculates multi-dimensional physiological health indicators, including diversity indices, functional microbial community proportions, and metabolic pathway abundance. By comparing these indicators with a species-specific preset health baseline, a quantitative health diagnosis report is generated. This effectively eliminates environmental interference noise, making it possible for the first time to unbiasedly assess the true state of the gut microbiota using aged fecal samples. This significantly improves the accuracy of health assessment, nutritional status analysis, and pathogen detection, avoiding major misdiagnosis such as misidentifying environmental bacteria as pathogens or misdiagnosing natural microbial decay as ecological imbalance. It provides a stable and reliable technical tool for wildlife conservation and management, supporting long-term, large-scale, and low-interference health monitoring of inaccessible rare and endangered species. It assists in early warning of population disease risks, habitat adaptability assessment, and analysis of the effectiveness of conservation measures, and has profound significance for biodiversity conservation and ecosystem health management. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of the system logic connection of the present invention.
[0022] Figure 2 This is a schematic diagram of the operational logic connection of the microbial community map generation module in this invention.
[0023] Figure 3 This is a schematic diagram of the preset logical connection of the microbial environmental response coefficient table in this invention. Detailed Implementation
[0024] The following description, in conjunction with the implementation of this invention, is merely an example and illustration of the concept of this invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the inventive concept or exceed the scope defined in these claims, all of which should fall within the protection scope of this invention.
[0025] Example 1:
[0026] Please see Figures 1-3 As shown, a non-invasive genome monitoring and population health assessment system for wild animals includes a biological data acquisition module, a sample exposure calculation module, a sample biological classification module, a microbial community map generation module, and an individual health assessment module.
[0027] The various modules are connected via wired and / or wireless means to enable data transmission between them;
[0028] Biological data acquisition module: Acquires raw sequencing data of wild animal fecal samples and microenvironmental data of sampling points. The raw sequencing data includes the raw relative abundance of various microorganisms, and the microenvironmental data of sampling points includes at least local temperature data and relative humidity data of sampling points.
[0029] Specifically, in this embodiment, the system first uses a high-throughput sequencing platform to amplify and sequence the V3-V4 hypervariable region of the 16S rRNA gene in the collected wild animal fecal samples to obtain raw binary sequencing data. Then, the DADA2 noise reduction clustering algorithm is used to cluster the raw sequences into amplicon sequence variants and perform standardization to generate raw relative abundance data that does not depend on sequencing depth.
[0030] In this process, to ensure the comparability of microbial abundance among different samples, the system uses a "per million reads" standardization algorithm to calculate the original relative abundance of various microorganisms. The calculation formula is as follows: ,in This represents the relative abundance value of the i-th type of microbial genus. This data is usually processed using standardization methods such as RPM (representing the relative abundance of the genus in the sample community as a dimensionless ratio.
[0031] This represents the absolute number of reads belonging to the i-th genus identified by sequencing. This represents the total number of reads obtained for this sample. The sequencing quality control factor is set to a value range of [0.8, 1.0]. This parameter is dynamically determined by the Q30 base ratio output by the sequencer. The calculation logic is: if the Q30 ratio is greater than 90%, then... =1, if it is between 80% and 90%, then Linear decreasing is intended to impose a mathematical penalty on the abundance confidence of low-quality sequencing samples, preventing low-quality data from being over-amplified in subsequent models.
[0032] At the same time, the system synchronously captures micro-environmental data at sampling points through an integrated sensor array deployed on the sampling equipment, and constructs an environmental state vector. .in Local temperature data, acquired using a PT1000 platinum resistance thermometer, is expressed in degrees Celsius. The relative humidity data at the sampling points is obtained through a capacitive humidity sensor and is expressed as a percentage. The ultraviolet radiation intensity is expressed in microwatts per square centimeter. To eliminate sensor noise fluctuations, the system uses a sliding window mean filter on the above environmental data, with the window size preset to a 5-minute interval before and after the sampling time.
[0033] Specifically, when generating the final input dataset for analysis, the system also performs a "low-abundance noise filtering" preprocessing operation to remove spurious bacterial communities that may be introduced by sequencing errors or aerosol contamination. The filtering threshold determination formula is as follows: ,in The abundance truncation threshold; The average abundance of background noise bacteria detected in a pre-set blank control sample; The standard deviation of the background noise; This is the stringency coefficient, preset to 3, ensuring the removal of 99.7% of random sequencing noise. Only when the calculated abundance of a certain genus... Only then will the data be retained and stored in the original sequencing database.
[0034] Based on the above technical solution, the following example illustrates a specific application scenario:
[0035] Suppose that in a field sampling of golden snub-nosed monkey feces, the sequencer outputs the total number of reads for that sample. =50,000 reads, of which the absolute number of reads identified was Bacteroides. =4,500 sequences. The sequencing report shows that the Q30 ratio for this sequencing run was 95%, therefore the quality control coefficient was... The value is 1.
[0036] The original relative abundance of Bacteroides was calculated using the standardized formula described above: This means that out of every million standardized sequences, 90,000 belong to the Bacteroides genus, or 9%.
[0037] Meanwhile, it is assumed that the mean background noise of this genus in the blank control sample is... =50, standard deviation =10, then the filtering threshold .
[0038] Due to the calculation > The system determined that the Bacteroides data was valid and included it as part of the "raw sequencing data," along with the data recorded by the sensor. and They are packaged together and transmitted to the subsequent time succession model for processing.
[0039] Sample exposure calculation module: Based on the original relative abundance, the abundance data of the predetermined indicator genera are screened and the ratio is calculated to generate time indicator marker features; the preset microbial succession time model is called, and the time indicator marker features and the microenvironmental data of the sampling point are used as input variables to calculate the sample environmental exposure duration.
[0040] Preferably, the construction step of the pre-set microbial succession time model includes:
[0041] Fresh excrement from the target wild animals was collected as a standard reference sample. The standard reference sample was divided into several equal parts and placed in different preset constant temperature and humidity chambers for continuous observation.
[0042] The standard reference samples in the environmental chamber were continuously sampled at multiple points according to the preset time intervals, and the collected samples were subjected to high-throughput sequencing to obtain microbial community evolution data at different time sections.
[0043] By analyzing the microbial community evolution data, aerobic bacteria that showed a monotonically increasing trend over time and anaerobic bacteria that showed a monotonically decreasing trend were identified and established as positive and negative time-related indicators, respectively.
[0044] The abundance ratio of positively correlated time indicators to negatively correlated time indicators was calculated. A regression analysis algorithm was used to establish a numerical mapping relationship between the abundance ratio and the sample storage time. This numerical mapping relationship was then encapsulated into a pre-set microbial succession time model.
[0045] Specifically, the construction process of the pre-set microbial succession time model first focuses on establishing a high-precision benchmark for microbial temporal changes. This process begins with the precise preparation of standard reference samples, namely, collecting fresh fecal samples during the golden window period after excretion by the target wild animal (preferably within 0.5 hours after excretion) to ensure that the initial microbial community structure has not been significantly disturbed by environmental factors. Subsequently, the samples are divided into several standard aliquots of 50g each and placed in multiple pre-set constant temperature and humidity environmental chambers, with the temperature gradient of the environmental chambers set from 5°C to 35°C. (Step size 5℃), relative humidity gradient is set from 30% to 90% (step size 10%) to simulate complex and variable microclimate conditions in the field; then high-density time-series sampling and sequencing are performed, and samples in each environmental chamber are continuously sampled at multiple points according to preset time intervals (e.g., every 2 hours) until the 72-hour observation endpoint is reached. Total DNA is extracted from all collected samples and high-throughput sequencing is performed on the V3-V4 variable region of the 16S rRNA gene. After quality control and OTU clustering, a microbial abundance matrix containing the time dimension is generated.
[0046] Based on this, the system screens indicator bacteria and constructs mathematical features through correlation analysis. Specifically, Spearman rank correlation analysis is performed on the relative abundance of microorganisms at the genera level in the matrix over time to calculate the correlation coefficient ρ. Microorganisms with ρ > 0.85 and P < 0.05 are defined as a set of aerobic bacteria that proliferate over time, while microorganisms with ρ < -0.85 and P < 0.05 are defined as a set of anaerobic bacteria that decay over time. The logic behind setting the threshold of 0.85 is to ensure that the selected indicator bacteria have extremely high sensitivity and monotonicity to changes over time, excluding interference from random fluctuations. At the same time, the P value is a core concept in statistical hypothesis testing, used to determine whether the "correlation" is statistically significant.
[0047] Then the characteristics of the time indicator markers were calculated. The calculation formula is: ,in , where X and Y are dimensionless time indicator features, and X and Y are the total number of anaerobic and aerobic bacteria, respectively; Let represent the relative abundance of the y-th aerobic bacterium. denoted as the relative abundance of the xth anaerobic bacterium. The preset smoothing factor is used to prevent the denominator from being zero and to stabilize numerical fluctuations under low abundance. The exponential bacterial growth decay rate is transformed into a linear trend through natural logarithmic transformation, which facilitates regression fitting.
[0048] Finally, a multivariate nonlinear regression algorithm was used to construct the model. Considering the regulatory effects of temperature and humidity on microbial metabolic rates, a kinetic model modified by the Arrhenius equation was established. The mapping relationship between the sample environmental exposure duration Δt and microenvironment data is shown in the model fitting formula: ,in The initial characteristic intercept at time zero represents the baseline state of fresh feces. is the basic succession rate constant under standard conditions, expressed as the reciprocal of the hour; and These are the reference temperature and reference humidity, respectively. For temperature sensitivity coefficient, 1 represents the humidity correction factor, and all parameters are dimensionless. These parameters were obtained by iteratively solving the multidimensional environmental chamber data using the least squares method. , , and The value is used to complete the encapsulation of the model.
[0049] Preferably, the step of selecting abundance data of predetermined indicator genera based on the original relative abundance and performing ratio calculations to generate time indicator marker features specifically includes:
[0050] Identify the abundance values of Bacteroidetes and Proteobacteria from the raw relative abundance;
[0051] Calculate the numerical ratio of Bacteroidetes to Proteobacteria and its rate of change relative to the baseline;
[0052] Numerical ratios and rates of change are labeled as time-indicating markers, which are used to quantify the degree of microbial community succession over time.
[0053] Specifically, the system first traverses the original relative abundance data matrix, and based on annotation information from the NCBI or Silva taxonomic databases, identifies and extracts the abundance values of "Bacteroidetes" and "Proteobacteria" and all their genera. Since Bacteroidetes are mostly dominant obligate anaerobes in the animal gut, they undergo oxidative stress death after being exposed to atmospheric oxygen in feces, causing their abundance to decrease exponentially over time. Proteobacteria, on the other hand, contain a large number of facultative anaerobes and environmentally derived aerobic bacteria, which can survive under environmental conditions and even proliferate using undigested organic matter in feces. The system uses a weighted aggregation algorithm to calculate the total abundance at the phylum level for both, and constructs a feature extraction model based on logarithmic decay kinetics. The calculation process is as follows:
[0054] First, the system calculates the current instantaneous phylum ratio, and the calculation formula is constructed as follows: ,in The instantaneous phylum ratio of the sample to be tested is a dimensionless value. The normalized relative abundance of the i-th identified Bacteroidetes genus; The normalized relative abundance of the j-th Proteobacteria genus identified; n and m are the number of genera of the corresponding phylum detected, respectively; To calculate the stability smoothing factor, the default value is... Its function is to prevent computational overflow anomalies caused by the denominator approaching zero due to extremely low abundance of Proteobacteria (e.g., environmental bacteria have not yet proliferated in extremely fresh samples).
[0055] Next, in order to quantify the degree of deviation of this ratio from "time zero", the system calls the "species-specific core baseline library" stored in the database to obtain the baseline ratio of freshly collected feces for the target wild animal species. The rate of change of community succession was calculated using a log-difference model, and the calculation formula is as follows: .in, The succession index is a time indicator characteristic that represents the degree of aging of microbial community structure over time; the larger the value, the longer the exposure time. , where is the species attenuation sensitivity coefficient, is a dimensionless correction parameter used to compensate for the influence of different species' fecal physical properties (such as moisture content and fiber density) on oxygen permeability. For herbivores, due to the high fiber content and porosity of their feces, oxygen permeability is rapid, resulting in a high rate of microbial attenuation. The range is set to 0.8 to 1.0; however, for carnivorous animals, the feces are dense, and the internal anaerobic environment is maintained for a longer period of time, usually... The preferred setting is within the range of 1.2 to 1.5, with a default value of 1.0. Finally, the system will calculate the instantaneous phylum ratio. With succession rate of change Combined and labeled as a time indicator feature vector .
[0056] It should be added that the construction logic of the species-specific core baseline library is as follows:
[0057] First, high-reliability "time-zero" samples are obtained through two pathways. Pathway one involves controlled environment collection, specifically manual collection within 5 minutes of the target species' excrement in zoos, wildlife rescue centers, or semi-free-range habitats, ensuring the samples are free from environmental oxidation and contamination. Pathway two involves extremely fresh samples from the wild, obtained by tracing fresh footprints to obtain still-warm fecal samples. Deep metagenomic sequencing is performed on the collected samples, and species composition analysis algorithms are used to identify stable bacterial genera comprising over 80% of the healthy population of that species, i.e., the "core flora." The system focuses on extracting the relative abundance distribution range of these core flora at time-zero. For each target species, the system performs a normality test on the abundance ratio of its core flora, calculating the mean, variance, and 95% confidence interval. Finally, these statistical parameters are stored in the database as the species-specific core baseline.
[0058] Preferably, the steps for calculating the exposure duration of the sample environment by calling a pre-set microbial succession time model and using time indicator characteristics and microenvironmental data at sampling points as input variables include:
[0059] Local temperature data and relative humidity data of the sampling points are extracted from the microenvironment data of the sampling points.
[0060] The characteristics of time indicators are used as independent variables, and local temperature data and relative humidity data at sampling points are used as environmental correction parameters, which are then mapped to the multidimensional regression equation of the microbial succession time model.
[0061] The output of the multidimensional regression equation is analyzed to obtain the environmental exposure duration of the sample, measured in time units.
[0062] Specifically, the system first parses local temperature data from the synchronously collected microenvironment data packets of the sampling points. relative humidity data at sampling points The humidity data was then normalized to obtain... Subsequently, the system constructs a multidimensional nonlinear regression model based on the modified Arrhenius equation, which incorporates the succession index. Using environmental parameters as the independent variables of the correction factor, the environmental exposure duration Δt of the sample is calculated, and the model calculation formula is constructed as follows:
[0063] ;
[0064] Where △t is the environmental exposure duration of the sample to be determined, and the output unit is hours; The succession index is the input, which is a dimensionless positive number; the denominator represents the "comprehensive succession rate under the current environment".
[0065] This is the baseline succession rate constant for the species' feces under standard conditions, expressed as the reciprocal of the rate per hour. This parameter was obtained during the system's pre-training phase by substituting at least 200 sets of sample data with known exposure times into the least squares method. For primate feces, The preferred range is 0.2 to 0.3;
[0066] The temperature sensitivity coefficient is a dimensionless parameter that follows van der Hoff's rule. It characterizes the factor by which the metabolic rate of microorganisms increases for every 10-degree increase in temperature. It is preset for the microbial community in the field environment. =2.0;
[0067] The standard reference temperature is usually set to 20°C.
[0068] β is the humidity deviation penalty coefficient, a dimensionless parameter used to correct the nonlinear effect of humidity on oxygen permeability and bacterial proliferation. A value of 0.5 is recommended.
[0069] The optimal relative humidity for microbial succession is preset to 0.8. When the ambient humidity deviates from this value (whether it is too dry leading to dehydration or too wet leading to stomatal blockage), it will affect the succession rate through the absolute value term.
[0070] Based on the above technical solution, the following example illustrates a specific application scenario:
[0071] To verify the computational accuracy of the model, the succession index was input. =1.386.
[0072] Assume the sensor records the current environmental data as follows: local air temperature =30℃, relative humidity =60%.
[0073] System preset parameters: reference rate Reference temperature , =2.0, optimal humidity =0.8, humidity penalty coefficient β=0.5.
[0074] Step 1: Calculate the temperature correction term:
[0075] ;
[0076] This means that at 30 degrees Celsius, the rate of bacterial succession is twice that at 20 degrees Celsius.
[0077] The second step is to calculate the humidity correction term:
[0078] ;
[0079] This means that lower humidity has a slight effect on the succession rate.
[0080] The third step is to calculate the overall succession rate:
[0081] ;
[0082] Step 4: Determine the exposure duration.
[0083] ;
[0084] The calculation results show that although the microbial community structure changed significantly (a succession index of 1.386 usually corresponds to a relatively long time at room temperature), the high ambient temperature (30 degrees Celsius) greatly accelerated the metabolic and death processes of the microorganisms. Therefore, the system determined that the sample was actually exposed for only about 2.52 hours. This information is used to guide the calculation of index decay compensation, ensuring that the microbial abundance is not overcompensated due to incorrect time estimation.
[0085] Sample biological classification module: Based on the original relative abundance, the microorganisms are classified using a pre-set core gut microbiota baseline library and a typical environmental saprophytic microbiota library to determine the gut native microbiota to be corrected and the environmental background microbiota to be removed.
[0086] Preferably, the construction steps of the core gut microbiota baseline library and the typical environmental saprophytic microbiota library include:
[0087] Multiple fresh fecal baseline samples of the target wild animal species and ecological environment reference samples around the sampling points were obtained, and reference sequence identifiers of microorganisms and their corresponding reference relative abundances were extracted from each sample.
[0088] The detection frequency of each microorganism in all fresh fecal reference samples was statistically analyzed, and microorganisms with a detection frequency higher than the preset detection threshold were identified as candidate core bacteria.
[0089] The average relative abundance of each candidate core bacterial genus was extracted in all fresh fecal reference samples. The abundance stability value of each candidate core bacterial genus was calculated. Candidate core bacterial genus whose abundance stability value met the preset stability requirements were selected as intestinal native marker bacteria. They were then mapped and associated with the corresponding reference sequence markers to construct the core intestinal flora baseline library.
[0090] Extract the unique microbial set from the ecological environment reference sample and calculate the invasion frequency of each unique microorganism in the fresh fecal baseline sample;
[0091] Microorganisms with an invasion frequency lower than a preset invasion threshold and a reference relative abundance in ecological environment reference samples higher than a preset distribution threshold are screened out and identified as environmental source indicator bacteria;
[0092] By aggregating environmental source indicator bacteria and their corresponding reference sequence identifiers into taxonomic clusters, an association index between environmental source indicator bacteria and typical saprophytic degradation functions is established, thereby constructing a library of typical environmental saprophytic bacteria.
[0093] It should be explained that the system collects no fewer than 50 fresh fecal baseline samples of the target wild animal species (sampling time is limited to within 10 minutes after excretion to ensure no environmental contamination) and no fewer than 20 soil and air ecological environment reference samples from the surrounding area of the sampling point. After completing sequencing and amplicon sequence variant (ASV) clustering, the system extracts the reference sequence identifiers of microorganisms in each sample and their corresponding normalized reference relative abundances.
[0094] To construct a baseline database of the core gut microbiota, the system first executes a high-frequency detection screening strategy. The processor iterates through the fresh fecal baseline sample set, calculating the detection frequency of each taxonomic unit (OTU) across all samples using the following formula: ,in This represents the number of samples where the abundance of this OTU is greater than zero. This represents the total number of samples. The system has a preset detection threshold of 80%. The microorganisms identified as “candidate core genus” are based on the definition of core microbial communities in microbial ecology, which are symbiotic bacteria that are widely present in the host population.
[0095] Subsequently, the system introduced an "abundance stability screening" mechanism to eliminate transient transient bacteria that, while common, exhibit significant fluctuations. The system calculated the coefficient of variation for each candidate core genus across all samples. The calculation formula is: ,in Let be the standard deviation of the relative abundance of the i-th candidate genus in all samples. Its average relative abundance. The system is set with a preset stability requirement of... Candidate bacteria that meet this condition will be designated as "intestinal native marker bacteria".
[0096] To construct a typical environmental saprophytic microbial community library, the system employs a "reverse exclusion screening" strategy. The processor first extracts the top 100 dominant species by abundance from ecological environment reference samples, then calculates the invasion frequency (i.e., the probability of detection in fresh feces) of these species in fresh fecal baseline samples. The system then filters out microorganisms with invasion frequencies below a preset invasion threshold (set to 5%, meaning they are rarely found in fresh feces) and reference relative abundance in environmental samples above a preset distribution threshold, identifying them as "environmental source indicator bacteria."
[0097] Preferably, the operational logic of the sample biological classification module is as follows:
[0098] The original relative abundance of various microorganisms contained in the raw sequencing data is traversed, and the gene sequence feature identifiers corresponding to each type of microorganism are extracted.
[0099] Gene sequence feature identifiers were compared with the pre-set typical environmental saprophytic bacteria database and the core gut microbiota baseline database for homology, and the sequence similarity values of various microorganisms with the above two databases were calculated.
[0100] Based on the preset sequence similarity threshold, microorganisms with sequence similarity values higher than the sequence similarity threshold of the typical environmental saprophytic microbial community library are selected and included in the initial screening set of environmental background microbial communities.
[0101] Microorganisms with sequence similarity values higher than the sequence similarity threshold of the core gut microbiota baseline library were selected and included in the initial screening set of gut native microbiota.
[0102] For overlapping microorganisms that exist simultaneously in the initial screening set of environmental background microbiota and the initial screening set of gut native microbiota, the sequence similarity values between them and the two libraries are compared. The overlapping microorganisms are then classified into the set corresponding to the one with the higher sequence similarity value, thereby eliminating the intersection between the sets.
[0103] The initial set of processed environmental background microbial communities is defined as the environmental background microbial communities to be removed, and their corresponding original relative abundance is marked as the background noise abundance to be deducted.
[0104] The initial set of processed gut microbiota was defined as the gut microbiota to be corrected.
[0105] Specifically, the system first traverses the raw sequencing data in the memory and extracts the representative nucleotide sequence corresponding to each non-zero abundance amplicon sequence variant (ASV) as the query sequence. The sequence length is typically between 400bp and 460bp. Then, the system calls the Smith-Waterman local sequence alignment algorithm to align each sequence... The sequences were compared in pairs with the reference sequences in the pre-set "typical environmental saprophytic bacteria library" and "core gut microbiota baseline library", and the sequence similarity values were calculated.
[0106] The similarity calculation formula is constructed as follows ,in The similarity score between the query sequence and the reference sequence in the database, with a value range of [0, 100]. Levenshtein edit distance represents the minimum number of single-character edit operations (including insertion, deletion, and replacement) required to transform a query sequence into a reference sequence. The effective arrangement length of the alignment region is represented by base pairs. This ratio represents the proportion of edit operations to the effective length, directly reflecting the degree of difference in the sequence. The smaller the ratio, the smaller the difference.
[0107] By traversing and comparing data, the system obtains the maximum similarity between the microorganism and a typical environmental saprophytic bacteria database. and the maximum similarity to the core gut microbiota baseline database. .
[0108] Next, the system introduces a preset sequence similarity determination threshold. Perform initial screening. The system executes logical judgment: If... If so, then the ASV is marked in the "initial screening set of environmental background microbiota"; if If so, it will be marked into the "Initial Screening Set of Gut Native Microbiota".
[0109] For the overlapping microorganism problem, the system executes conflict resolution logic based on maximum likelihood. The similarity difference is calculated. .like If the value is >0, the microorganism is determined to be more closely related to gut colonizing bacteria in terms of genomic characteristics, and is forcibly classified into the gut native flora set and removed from the environmental background flora set; if If <0, it is determined to be of environmental origin and classified into the environmental background microbial community set; if If the abundance is 0, then abundance weighting is introduced, and the element is preferentially classified into the category with the higher original relative abundance.
[0110] The gut microbiota generation module performs inverse abundance attenuation compensation calculations on the original gut microbiota to be corrected based on the sample's environmental exposure duration and the microenvironmental data at the sampling points, and performs proliferation noise removal calculations on the background microbiota to be removed, generating an initial gut microbiota map. The initial gut microbiota map represents the zero-time microbial community structure at the time of sample excretion.
[0111] Preferably, the microbial community mapping module includes:
[0112] Based on local temperature data and relative humidity data at sampling points, the data is matched against a pre-set microbial environment response coefficient table to obtain the attenuation rate coefficient for the gut microbiota to be corrected and the proliferation rate coefficient for the background microbiota to be eliminated.
[0113] The attenuation correction factor is obtained by multiplying the sample environmental exposure time with the attenuation rate coefficient. Based on the attenuation correction factor, the biomass loss value generated by the gut microbiota to be corrected during the exposure period is calculated, and the biomass loss value is added to the original relative abundance of the gut microbiota to be corrected, so as to obtain the restored abundance of the gut microbiota.
[0114] The proliferation noise factor is obtained by multiplying the sample environmental exposure duration with the proliferation rate coefficient. Based on the proliferation noise factor, the biomass redundancy value generated by the background microbial community to be removed during the exposure period is calculated. The biomass redundancy value is then subtracted from the original relative abundance of the background microbial community to be removed to obtain the net abundance of the environmental microbial community.
[0115] A corrected microbial data matrix was constructed by integrating the restored abundance of native flora and the net abundance of environmental flora. The abundance values of each flora in the data matrix were normalized to calculate the true proportion distribution of each flora at time zero. This true proportion distribution was determined as the initial gut microbiota map characterizing the microbial community structure at time zero of sample excretion.
[0116] Preferably, the preset logic of the microbial environmental response coefficient table is as follows:
[0117] Representative indicator strains from the core gut microbiota baseline library and representative interference strains from the typical environmental saprophytic microbiota library were selected, and a multidimensional environmental simulation matrix was constructed based on preset temperature gradient values and preset humidity gradient values.
[0118] Representative indicator strains and representative interference strains were placed under the conditions of each node in the multidimensional environmental simulation matrix for in vitro culture for a preset duration, and the abundance changes of each strain were monitored at a preset sampling frequency.
[0119] For representative indicator strains, the abundance loss per unit time is calculated based on the abundance change value, and the abundance loss is determined as the reference decay rate coefficient under the corresponding environmental conditions.
[0120] For representative interfering strains, the abundance increment per unit time is calculated based on the abundance change value, and the abundance increment is determined as the reference proliferation rate coefficient under the corresponding environmental conditions.
[0121] The reference decay rate coefficient and reference proliferation rate coefficient are respectively indexed and mapped to the preset temperature gradient value and preset humidity gradient value in the multidimensional environmental simulation matrix to form a microbial environmental response coefficient table.
[0122] Preferably, the steps for obtaining the restored abundance of the native microbial community specifically include:
[0123] The original relative abundance of the gut microbiota to be corrected in the original sequencing data was extracted and set as the basic stock data. The input attenuation correction factor was set as the linear loss coefficient.
[0124] The product operation is performed on the basic stock data and the linear loss coefficient to quantify the biomass loss of the microbial community due to DNA degradation caused by environmental stress during the exposure time in the sample environment.
[0125] The biomass loss values are summed with the original relative abundance to generate a theoretically corrected abundance value that covers both the original retention and the estimated loss.
[0126] A preset biological abundance saturation threshold is obtained. The theoretically corrected abundance value is compared with the preset biological abundance saturation threshold, and the smaller value is selected as the final restored abundance of the native microbial community to prevent the restored abundance value from exceeding the reasonable biological range due to calculation errors.
[0127] Preferably, the steps of calculating the biomass redundancy value generated by the background microbial community to be eliminated during exposure based on the proliferation noise factor, and subtracting the biomass redundancy value from the original relative abundance of the background microbial community to be eliminated to obtain the net abundance of the environmental microbial community include:
[0128] The original relative abundance of the background microbial community to be removed was set as the mixed total amount data, and the proliferation noise factor was used as the environmental disturbance weight.
[0129] By multiplying the mixed total data with the environmental disturbance weight, the biomass redundancy of the microbial community due to in vitro proliferation during the sample environmental exposure time is estimated.
[0130] The primary net value of the bacterial community is obtained by subtracting the biomass redundancy value from the original relative abundance.
[0131] Non-negative logic verification is performed on the primary net value data. The primary net value data is compared with the zero baseline. When the primary net value data is lower than the zero baseline, the environmental microbial community net abundance is corrected to the preset minimum detection limit value. When the primary net value data is higher than or equal to the zero baseline, the primary net value data is directly confirmed as the environmental microbial community net abundance.
[0132] Specifically, the core of the microbial community mapping module lies in eliminating the "signal distortion" caused by environmental exposure after the sample is removed from the body. This process first establishes a precise numerical correction benchmark, that is, based on the measured local temperature data and relative humidity data of the sampling points, bilinear interpolation matching is performed in a pre-set microbial environmental response coefficient table. This coefficient table is a lookup table constructed based on a multidimensional environmental simulation matrix, covering microbial kinetic parameters from 5℃ to 40℃ and from 30% to 95% humidity. Through matching, the decay rate coefficients for the intestinal flora to be corrected are obtained respectively. And the proliferation rate coefficient of the background microbial community to be eliminated. Both units are ;
[0133] Next, reverse compensation calculations for the native microbial community are performed using the formula. Calculate the abundance of the restored native microbial community, among which The relative abundance is obtained from the original sequencing, Δt is the duration of environmental exposure of the sample, and the product term is... This is the attenuation correction factor (dimensionless), which quantifies the proportion of cumulative loss during exposure. The preset biological abundance saturation threshold is set to 1.0 to prevent the calculation results from exceeding the physical limits due to model extrapolation.
[0134] Simultaneously, noise removal calculations for the environmental background microbial community are performed. This step aims to remove noise "mixed in" due to in vitro proliferation, using a formula... Calculate the net abundance of environmental microbial communities, where The product term represents the original relative abundance of the background microbial community in the environment. This is the proliferation noise factor, which characterizes the proportion of biomass redundancy caused by suitable environment. The preset minimum detection limit value serves as the baseline for non-negative logic verification, preventing negative values from appearing in the calculation results.
[0135] Finally, a global normalization reconstruction was performed, integrating all the corrected original microbial community restoration abundances. With environmental microbial community net abundance A revised microbial data matrix was constructed, and the abundance of all microbial communities in each row (i.e., each sample) was normalized using the following formula: ,in This represents the true proportion distribution of the i-th bacterial community at time zero. Represents the corrected abundance value (i.e. or N represents the total number of detected bacterial communities. This step eliminates the proportional imbalance caused by errors in total biomass estimation, and the resulting initial gut microbiota map accurately reflects the microbial community structure of the sample at the moment of excretion (time zero), providing a pure and unbiased biological data foundation for subsequent health assessments. Here, j has no numerical or semantic specialness; it is merely a general symbol used to iterate through all bacterial communities and is the loop variable in the summation operation. The physical meaning of the entire formula is: dividing the corrected abundance of a single bacterial community (i) by the sum of the corrected abundances of all bacterial communities (j=1 to N) yields the true relative proportion of that community in the community at time zero.
[0136] Individual health assessment module: Calculates physiological health indicators based on the initial gut microbiota profile, compares the physiological health indicators with the preset health baseline, and generates a physiological health diagnosis report for the individual wild animal.
[0137] Preferably, the operational logic of the individual health assessment module includes:
[0138] The distribution information of each microbial community in the initial gut microbiota map is analyzed. Based on the preset diversity calculation logic, the community diversity index, which represents the richness of the community, is obtained. At the same time, the abundance of potential pathogenic bacteria in the map is identified to obtain the proportion of pathogenic bacteria load. The community diversity index and the proportion of pathogenic bacteria load are combined into physiological health indicator data.
[0139] Call the preset species health reference library to obtain the diversity standard range and pathogen safety threshold corresponding to the wild animal species, and set the diversity standard range and pathogen safety threshold as the preset health baseline;
[0140] The community diversity index in the physiological health indicator data is compared with the diversity standard interval in the preset health baseline to determine the diversity deviation level. At the same time, the difference between the pathogenic bacteria load ratio and the pathogenic bacteria safety threshold is calculated to determine the infection risk level.
[0141] Based on the diversity deviation level and infection risk level, the corresponding health status description text is matched in the pre-set diagnostic terminology mapping table, and the health status description text is packaged into a physiological health diagnosis report for the wild animal individual.
[0142] Specifically, when executing the individual health assessment module, the system normalizes the abundance values of each bacterial community in the data matrix to calculate the true proportion distribution of each bacterial community at time zero. Since this true proportion distribution mathematically represents the proportion of each bacterial community in the total microbial community, and its numerical sum is always equal to 1, this proportion distribution is biologically logically equivalent to the relative abundance of the bacterial community at the time of excretion of the wild animal individual.
[0143] Based on this, in the calculation of physiological health indicator data, directly using The relative abundance parameter is substituted into the Shannon-Wiener index formula for calculation. This transformation logic ensures that the community stability index (diversity index) relied upon for subsequent health assessment is based on the true biomass proportion after environmental exposure correction, thus eliminating the interference of sampling delay on health assessment results. Therefore, the community diversity index, characterizing community richness, is calculated using the Shannon-Wiener index formula. The calculation formula is: Where S is the total number of bacterial genera detected. The relative abundance of the i-th bacterial genus is represented by this index, which is a dimensionless value. A higher value indicates a more complex and stable bacterial community structure. Simultaneously, based on a pre-set "list of potential pathogens," the system identifies bacterial communities in the map that belong to this list and performs an accumulation calculation to determine the pathogen load percentage. The calculation formula is: ,in The relative abundance of the kth pathogenic bacterium is expressed as a percentage.
[0144] Next, a dynamic baseline is established by calling a pre-built species health reference library. The construction logic of this reference library is based on big data statistics of more than 1,000 healthy individual samples of the target species accumulated over the past 5 years, and the diversity standard interval is determined by fitting a normal distribution. The 95th percentile of the total number of pathogenic bacteria was set as the safe threshold for pathogenic bacteria. This pre-set logic ensures that the baseline is both statistically representative and excludes the interference of extreme outliers.
[0145] Subsequently, quantitative deviation analysis and risk rating are performed, rather than fuzzy comparisons. For diversity, the system calculates the diversity deviation. The formula is ,in and These are the mean and standard deviation from the reference database, respectively. The numerical range determines the "diversity deviation level": if A value less than -2 (i.e., below 2 standard deviations from the mean) is considered a "significant decrease". If -2 ≤ <-1, judged as "mildly reduced", if A value ≥-1 is considered "normal"; for pathogenic bacteria, an infection risk index is calculated. Determine the "infection risk level": If ≤0 indicates "safe (level 0)"; if 0 < ≤2% is considered "low risk (Level 1)". >2%, which is considered "high risk (level 2)".
[0146] Finally, a diagnostic report is generated based on a multidimensional logical matrix. The system uses a pre-set diagnostic terminology mapping table to combine and parse the above levels. For example, if an animal sample is calculated to... =1.8 ( =-2.3), and =4.0% ( =2.5%, Level 2 high risk), the system will trigger the combined logic "Level III diversity + Level 2 risk", automatically retrieving the corresponding pathological description text: "Alert: Severe intestinal flora imbalance detected, accompanied by invasion of highly abundant pathogenic bacteria. Immediate isolation and investigation of gastrointestinal bacterial infection are recommended." This outputs a physiological health diagnosis report that can intuitively guide veterinary intervention. Through the above end-to-end quantitative calculations, complex gene sequencing data is transformed into actionable medical decision-making data, achieving standardization and intelligentization of non-invasive health monitoring of wild animals.
[0147] Example 2:
[0148] A non-invasive genome monitoring and population health assessment device for wild animals, including a processor, memory, and communication bus;
[0149] The memory stores a computer-readable program that can be executed by the processor;
[0150] The communication bus enables communication between the processor and the memory;
[0151] When the processor executes the computer-readable program, it performs modules to implement a non-invasive genome monitoring and population health assessment system for wild animals as described in any one of the present invention.
[0152] It should be noted that the interval and threshold sizes are set for ease of comparison. The size of the threshold depends on the amount of sample data and the base number set by those skilled in the art for each set of sample data, as long as it does not affect the proportional relationship between the parameter and the quantized value. Furthermore, the above formulas are all dimensionless calculations, and the formulas are derived from software simulations using a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0153] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0154] It should be understood that determining B based on A does not mean determining B solely based on A; it also means determining B based on A and / or other information.
[0155] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0156] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A non-invasive genome monitoring and population health assessment system for wild animals, characterized in that, The system includes: Biological data acquisition module: Acquires raw sequencing data of wild animal fecal samples and microenvironment data of sampling points. The raw sequencing data includes the raw relative abundance of various microorganisms. Sample exposure calculation module: Based on the original relative abundance, the abundance data of the predetermined indicator genera are screened and the ratio is calculated to generate time indicator marker features; the preset microbial succession time model is called, and the time indicator marker features and the microenvironmental data of the sampling point are used as input variables to calculate the sample environmental exposure duration; Sample biological classification module: Based on the original relative abundance, the microorganisms are classified using a pre-set core gut microbiota baseline library and a typical environmental saprophytic microbiota library to determine the gut native microbiota to be corrected and the environmental background microbiota to be removed; The gut microbiota generation module performs inverse abundance attenuation compensation calculations on the original gut microbiota to be corrected based on the sample environmental exposure duration and the microenvironment data of the sampling points, and performs proliferation noise removal calculations on the background microbiota to be removed, generating an initial gut microbiota map. Individual health assessment module: Calculates physiological health indicators based on the initial gut microbiota profile, compares the physiological health indicators with the preset health baseline, and generates a physiological health diagnosis report for the individual wild animal.
2. The non-invasive genome monitoring and population health assessment system for wild animals according to claim 1, characterized in that, The steps for calculating the exposure duration of a sample environment by calling a pre-set microbial succession time model and using time indicator characteristics and microenvironmental data from sampling points as input variables include: Local temperature data and relative humidity data of the sampling points are extracted from the microenvironment data of the sampling points. The characteristics of time indicators are used as independent variables, and local temperature data and relative humidity data at sampling points are used as environmental correction parameters, which are then mapped to the multidimensional regression equation of the microbial succession time model. The output of the multidimensional regression equation is analyzed to obtain the sample's environmental exposure duration in time units.
3. The non-invasive genome monitoring and population health assessment system for wild animals according to claim 1, characterized in that, The operational logic of the sample biological classification module is as follows: The original relative abundance of various microorganisms contained in the raw sequencing data is traversed, and the gene sequence feature identifiers corresponding to each type of microorganism are extracted. Gene sequence feature identifiers were compared with the pre-set typical environmental saprophytic bacteria database and core gut microbiota baseline database for homology, and the sequence similarity values between various microorganisms and the two databases were calculated. Based on the preset sequence similarity threshold, microorganisms with sequence similarity values higher than the sequence similarity threshold of the typical environmental saprophytic microbial community library are selected and included in the initial screening set of environmental background microbial communities. Microorganisms with sequence similarity values higher than the sequence similarity threshold of the core gut microbiota baseline library were selected and included in the initial screening set of gut native microbiota. For overlapping microorganisms that exist simultaneously in the initial screening set of environmental background microbiota and the initial screening set of gut native microbiota, the sequence similarity values between them and the two libraries are compared. The overlapping microorganisms are then classified into the set corresponding to the one with the higher sequence similarity value, thereby eliminating the intersection between the sets. The initial set of processed environmental background microbial communities is defined as the environmental background microbial communities to be removed, and their corresponding original relative abundance is marked as the background noise abundance to be deducted. The initial set of processed gut microbiota was defined as the gut microbiota to be corrected.
4. The non-invasive genome monitoring and population health assessment system for wild animals according to claim 2, characterized in that, The microbial community map generation module includes: Based on local temperature data and relative humidity data at sampling points, the data is matched against a pre-set microbial environment response coefficient table to obtain the attenuation rate coefficient for the gut microbiota to be corrected and the proliferation rate coefficient for the background microbiota to be eliminated. The attenuation correction factor is obtained by multiplying the sample environmental exposure time with the attenuation rate coefficient. Based on the attenuation correction factor, the biomass loss value generated by the gut microbiota to be corrected during the exposure period is calculated, and the biomass loss value is added to the original relative abundance of the gut microbiota to be corrected, so as to obtain the restored abundance of the gut microbiota. The proliferation noise factor is obtained by multiplying the sample environmental exposure duration with the proliferation rate coefficient. Based on the proliferation noise factor, the biomass redundancy value generated by the background microbial community to be removed during the exposure period is calculated. The biomass redundancy value is then subtracted from the original relative abundance of the background microbial community to be removed to obtain the net abundance of the environmental microbial community. A corrected microbial data matrix was constructed by integrating the restored abundance of native flora and the net abundance of environmental flora. The abundance values of each flora in the data matrix were normalized to calculate the true proportion distribution of each flora at time zero. This true proportion distribution was determined as the initial gut microbiota map characterizing the microbial community structure at time zero of sample excretion.
5. The non-invasive genome monitoring and population health assessment system for wild animals according to claim 4, characterized in that, The default logic for the microbial environmental response coefficient table is as follows: Representative indicator strains from the core gut microbiota baseline library and representative interference strains from the typical environmental saprophytic microbiota library were selected, and a multidimensional environmental simulation matrix was constructed based on preset temperature gradient values and preset humidity gradient values. Representative indicator strains and representative interference strains were placed under the conditions of each node in the multidimensional environmental simulation matrix for in vitro culture for a preset duration, and the abundance changes of each strain were monitored at a preset sampling frequency. For representative indicator strains, the abundance loss per unit time is calculated based on the abundance change value, and the abundance loss is determined as the reference decay rate coefficient under the corresponding environmental conditions. For representative interfering strains, the abundance increment per unit time is calculated based on the abundance change value, and the abundance increment is determined as the reference proliferation rate coefficient under the corresponding environmental conditions. The reference decay rate coefficient and reference proliferation rate coefficient are respectively indexed and mapped to the preset temperature gradient value and preset humidity gradient value in the multidimensional environmental simulation matrix to form a microbial environmental response coefficient table.
6. The non-invasive genome monitoring and population health assessment system for wild animals according to claim 4, characterized in that, The specific steps to obtain the restored abundance of the native microbial community include: The original relative abundance of the gut microbiota to be corrected in the original sequencing data was extracted and set as the basic stock data. The input attenuation correction factor was set as the linear loss coefficient. The product operation is performed on the basic stock data and the linear loss coefficient to quantify the biomass loss of the microbial community due to DNA degradation caused by environmental stress during the exposure time in the sample environment. The biomass loss value is summed with the original relative abundance to generate the theoretically corrected abundance value. A preset biological abundance saturation threshold is obtained. The theoretical corrected abundance value is compared with the preset biological abundance saturation threshold, and the smaller value between the two is selected as the final restored abundance of the native microbial community.
7. The non-invasive genome monitoring and population health assessment system for wild animals according to claim 4, characterized in that, The specific steps to obtain the net abundance of environmental microbiota include: The original relative abundance of the background microbial community to be removed was set as the mixed total amount data, and the proliferation noise factor was used as the environmental disturbance weight. By multiplying the mixed total data with the environmental disturbance weight, the biomass redundancy of the microbial community due to in vitro proliferation during the sample environmental exposure time is estimated. The primary net value of the bacterial community is obtained by subtracting the biomass redundancy value from the original relative abundance. Non-negative logic verification is performed on the primary net value data. The primary net value data is compared with the zero baseline. When the primary net value data is lower than the zero baseline, the environmental microbial community net abundance is corrected to the preset minimum detection limit value. When the primary net value data is higher than or equal to the zero baseline, the primary net value data is directly confirmed as the environmental microbial community net abundance.
8. The non-invasive genome monitoring and population health assessment system for wild animals according to claim 1, characterized in that, The steps for constructing the pre-set microbial succession time model include: Fresh excrement from the target wild animals was collected as a standard reference sample. The standard reference sample was divided into several equal parts and placed in different preset constant temperature and humidity chambers for continuous observation. The standard reference samples in the environmental chamber were continuously sampled at multiple points according to the preset time intervals, and the collected samples were subjected to high-throughput sequencing to obtain microbial community evolution data at different time sections. By analyzing the microbial community evolution data, aerobic bacteria that showed a monotonically increasing trend over time and anaerobic bacteria that showed a monotonically decreasing trend were identified and established as positive and negative time-related indicators, respectively. The abundance ratio of positively correlated time indicators to negatively correlated time indicators was calculated. A regression analysis algorithm was used to establish a numerical mapping relationship between the abundance ratio and the sample storage time. This numerical mapping relationship was then encapsulated into a pre-set microbial succession time model.
9. The non-invasive genome monitoring and population health assessment system for wild animals according to claim 1, characterized in that, The construction steps of the core gut microbiota baseline library and the typical environmental saprophytic microbiota library include: Multiple fresh fecal baseline samples of the target wild animal species and ecological environment reference samples around the sampling points were obtained, and reference sequence identifiers of microorganisms and their corresponding reference relative abundances were extracted from each sample. The detection frequency of each microorganism in all fresh fecal reference samples was statistically analyzed, and microorganisms with a detection frequency higher than the preset detection threshold were identified as candidate core bacteria. The average relative abundance of each candidate core bacterial genus was extracted in all fresh fecal reference samples. The abundance stability value of each candidate core bacterial genus was calculated. Candidate core bacterial genus whose abundance stability value met the preset stability requirements were selected as intestinal native marker bacteria. They were then mapped and associated with the corresponding reference sequence markers to construct the core intestinal flora baseline library. Extract the unique microbial set from the ecological environment reference sample and calculate the invasion frequency of each unique microorganism in the fresh fecal baseline sample; Microorganisms with an invasion frequency lower than a preset invasion threshold and a reference relative abundance in ecological environment reference samples higher than a preset distribution threshold are screened out and identified as environmental source indicator bacteria; By aggregating environmental source indicator bacteria and their corresponding reference sequence identifiers into taxonomic clusters, an association index between environmental source indicator bacteria and typical saprophytic degradation functions is established, thereby constructing a library of typical environmental saprophytic bacteria.
10. A non-invasive genome monitoring and population health assessment device for wild animals, characterized in that, Includes processor, memory, and communication bus; The memory stores a computer-readable program that can be executed by the processor; The communication bus enables communication between the processor and the memory; When the processor executes the computer-readable program, it performs a module to implement the non-invasive genome monitoring and population health assessment system for wild animals as described in any one of claims 1 to 9.