Dose-dependent omics effect-oriented toxic substance identification method
By employing a dose-dependent omics approach combined with multiple analytical techniques, this study addresses the challenges of identifying unknown pollutants and linking them to biotoxicity in traditional toxic substance identification methods. It achieves highly sensitive and specific identification of toxic substances, providing accurate risk assessment support.
Patent Information
- Application Number
- CN202511738821.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-02-27
AI Technical Summary
Traditional methods for identifying toxic substances are insufficient to cover unknown or low-abundance micro-pollutants in the environment, and cannot effectively link the quantitative relationship between pollutants and biotoxic effects, resulting in a lack of accuracy in risk assessment, as well as insufficient sample pretreatment and integration of omics data.
Using a dose-dependent omics-guided approach, organic micropollutants were enriched using HLB columns, combined with RNA purification, reverse transcription, PCR amplification, RNA-seq sequencing, GO functional enrichment analysis, and GC-QTOF/MS non-targeted screening. Targeted quantitative analysis was then performed using a CTD database and liquid chromatography-mass spectrometry (LC-MS) to establish a complete technology chain for identifying toxic substances.
It significantly improves the sensitivity and specificity of toxic substance identification, accurately identifies key toxic substances, breaks through the limitations of traditional targeted analysis in identifying unknown pollutants, and provides accurate risk assessment support.
Smart Images

Figure CN121577778A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of toxic substance identification technology, and particularly relates to a toxic substance identification method guided by dose-dependent omics effects. Background Technology
[0002] Currently, surface water, as an important carrier of water resources, is susceptible to the residual organic micropollutants from industrial emissions and agricultural activities. These pollutants are low in concentration, complex in type, but significantly toxic, posing a potential threat to ecosystems and human health. Traditional identification of toxic substances largely relies on targeted chemical analysis, which requires prior identification of target pollutants and reliance on standards. This approach struggles to cover unknown or low-abundance micropollutants in the environment and cannot effectively correlate the quantitative relationship between pollutants and their biotoxic effects, resulting in a lack of precision in risk assessment.
[0003] With the development of omics technologies, transcriptomics and other methods have been gradually applied to toxicology research. However, existing methods often fail to fully integrate dose-dependent effects for systematic analysis, making it difficult to pinpoint key toxic pathways and substances through dose-response relationships in gene expression changes. Furthermore, traditional methods have shortcomings in sample pretreatment enrichment efficiency, standardization of omics data screening, and integration and correlation of chemical screening and biological effect data. This prevents the formation of a complete technological chain from sample processing to accurate identification of toxic substances, necessitating a toxic substance identification scheme that balances sensitivity, systematicity, and correlation. Summary of the Invention
[0004] The purpose of this invention is to address the aforementioned technical problems by providing a method for identifying toxic substances based on dose-dependent omics effects.
[0005] In view of this, the present invention provides a method for identifying toxic substances based on dose-dependent omics effects, comprising the following steps: Step 1: Store the surface water in a container, and clean the container in sequence with tap water, methanol, and deionized water, and then rinse it with the surface water to be collected before filling it with the sample. Step 2: After the water sample is coarsely filtered through a 0.45μm filter membrane, organic micropollutants are enriched using a hydrophilic-lipophilic equilibrium column activated with ultrapure water and HPLC-grade methanol, with the column flow rate controlled at 3-5 mL / min. Step 3: After extraction, freeze-dry the HLB column for 24 hours to remove moisture, elute the adsorbed organic micro-pollutants with LC-MS grade methanol, and concentrate the eluent using a nitrogen analyzer. Step 4: Select total RNA with RIN number > 7.0, purify mRNA and divide it into short fragments, reverse transcribe to synthesize cDNA and prepare U-labeled second-strand DNA, and then amplify by PCR after treatment with thermosensitive UDG enzyme; Step 5: Sequencing was performed using the Illumina paired-end RNA-seq method, and the reads were filtered using Cutadapt software; Step 6: Calculate the FPKM value of mRNA, standardize the transcriptional expression level and calculate log2FC, fit gene expression with 9 dose-response models, screen for significantly fitted differentially expressed genes and calculate PODDEG; Step 7: Perform GO functional enrichment analysis on differentially expressed genes, calculate and rank the PODGO of GO terms, and fit the bioefficacy ranking distribution curve; Step 8: Using differentially expressed genes as the test set, match the background set of the CTD database to obtain a list of suspected chemical substances, deconstruct the structure of the chemical substances, and screen for highly enriched structures using a t-test; Step 9: The sample was analyzed using a two-dimensional gas chromatograph, quadrupole mass spectrometer, and time-of-flight mass spectrometer. After matching with the NIST17 mass spectral library, substances were screened according to the screening criteria. Step 10: Extract the peak area of non-targeted screening substances, the POD value of the KEGG pathway, and the LC50 of acute toxicity test. After data conversion, perform correlation analysis and use liquid chromatography-mass spectrometry (LC-MS) to perform targeted quantitative analysis of standard chemicals.
[0006] Preferably, in step four, RNA purification is performed by Dynabeads Oligo to purify 5 μg of total RNA in two rounds of mRNA purification, the mRNA is split at high temperature using divalent cations, and the RNA fragments are reverse transcribed using SuperScript™ II reverse transcriptase.
[0007] Preferably, in step four, the PCR amplification conditions are: initial denaturation at 95℃ for 3 min; denaturation at 98℃ for 15 s, annealing at 60℃ for 15 s, extension at 72℃ for 30 s, for a total of 8 cycles; and a final extension at 72℃ for 5 min.
[0008] Preferably, in step five, the filtering conditions for sequencing reads are: removing reads containing aptamers, removing reads containing polyA and polyG, removing reads containing more than 5% unknown nucleotides, and removing reads containing more than 20% low-quality bases.
[0009] Preferably, in step six, the nine dose-response models include three types: S-shaped curve, linear curve, and U-shaped curve. The model with the smallest AIC value is selected as the best-fit model. The calculation method of PODDEG is as follows: the S-shaped curve corresponds to the concentration of the highest 50% response, the linear curve corresponds to the concentration of a 1.5-fold change in expression level, and the U-shaped curve takes the monotonically rising or falling part corresponding to the concentration of a 1.5-fold change in expression level.
[0010] Preferably, in step seven, the GO functional enrichment analysis retains only GO terms that match ≥3 differentially expressed genes, and the significantly enriched GO terms are determined by hypergeometric test; PODGO is the geometric mean of the PODDEG of the differentially expressed genes matched by the GO term.
[0011] Preferably, in step nine, the chromatographic columns for gas chromatography-mass spectrometry non-targeted chemical screening are two Agilent 15m×250μm×0.25μm columns, with the following column temperature program: initial 60℃ for 1 min, 40℃ / min ramp to 120℃ for 2 min, 15℃ / min ramp to 240℃ for 10 min, and 5℃ / min ramp to 300℃ for 10 min; the mass spectrometer uses an LE-EI ion source with an ion source temperature of 200℃ and an electron energy of 70eV.
[0012] Preferably, in step nine, the screening criteria are: retaining substances with a matching degree > 60, deleting substances containing silicon, and retaining those with higher matching degree and peak area when deleting duplicate substances.
[0013] Preferably, in step ten, the liquid chromatography-mass spectrometry (LC-MS) instrument is a SCIEX Triple Quad™ 6500+ LC-MS / MS, using a BEHC18 1.7μm 2.1*50mm column, with a positive ion mode running time of 7 min and a negative ion mode running time of 5 min. The ion source is an electrospray ionization source, and the detection mode is MRM (Multiple Ion Reaction Monitoring).
[0014] Preferably, in step four, the RNA reverse transcription conditions are: 65℃ for 5 min, 25℃ for 5 min, 42℃ for 60 min, and 72℃ for 5 min, with one cycle for each.
[0015] The beneficial effects of this invention are: Organic micropollutants were enriched using HLB columns at a flow rate of 3-5 mL / min, followed by 24-hour freeze-drying to remove moisture and nitrogen concentration. This effectively removed suspended solids and impurities from the water sample, significantly improving the recovery rate and purity of the target pollutants. The samples were then used for both chemical and biological analysis, laying a high-quality foundation for subsequent multi-dimensional detection. High-quality RNA with a RIN number > 7.0 was selected for RNA purification. Cutadapt software was used to rigorously filter reads containing aptamers and low-quality bases, ensuring the accuracy of the transcriptome sequencing data. Nine dose-response models were fitted, and the optimal model was selected using AIC values. This allowed for the accurate identification of significantly differentially expressed genes and the calculation of PODDEG, clearly establishing the dependence of gene expression on pollutant dose and solving the problem of difficulty in quantifying dose-response associations in traditional omics analysis.
[0016] By calculating PODGO using GO functional enrichment analysis and plotting biopotency ranking distribution curves, key biological pathways affected by pollutants can be quickly located, and the mechanisms of toxic effects can be clarified. Combined with the CTD database, key toxic structures can be predicted using hypergeometric distribution and t-tests. With GC-QTOF / MS non-targeted screening, different types of pollutants in winter and summer can be covered, overcoming the limitations of traditional targeted analysis in identifying unknown pollutants. Furthermore, targeted quantification using LC-MS / MS and correlation of transcriptome KEGG pathway POD values with zebrafish embryo 96hLC50 data achieves deep coupling between chemical detection and biological effects, significantly improving the specificity and reliability of key toxic substance identification and providing precise technical support for the risk assessment of toxic substances in surface water. Attached Figure Description
[0017] Figure 1 Schematic diagram of the predicted key toxic structures; Figure 2 Technology roadmap; Figure 3 Statistical chart of the number of differentially expressed genes; Figure 4 Overall biopotency ranking distribution curves of summer surface water samples based on dose-dependent transcriptomics; Figure 5 Key toxic pathway results diagram; Figure 6 The ToxPi score ranking results for summer vibes. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0019] It should be noted that all directional and positional terms used in this invention, such as "up," "down," "left," "right," "front," "back," "vertical," "horizontal," "inner," "outer," "top," "lower," "lateral," "longitudinal," and "center," are only used to explain the relative positional relationships and connections between components in a specific state (as shown in the accompanying drawings). They are merely for the convenience of describing the invention and do not require the invention to be constructed and operated in a specific orientation; therefore, they should not be construed as limitations on the invention. Furthermore, descriptions involving "first," "second," etc., are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated.
[0020] In the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0021] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0022] A method for identifying toxic substances based on dose-dependent omics effects includes the following steps: Sample collection: The collected surface water is stored in a container. Before use, the container is cleaned with tap water, methanol, and deionized water in sequence, and then rinsed with the surface water sample to be collected before filling it with surface water.
[0023] Sample extraction: After coarse filtration of the collected water samples to remove suspended solids using a 0.45 μm pore size filter membrane, organic micropollutants in the water samples were enriched using a hydrophilic-lipophilic equilibrium column activated with ultrapure water and HPLC-grade methanol. The flow rate of the HLB column was controlled at 3–5 mL / min.
[0024] Drying and elution: After extraction, the water in the HLB column was removed by a freeze dryer. After freeze-drying each batch of samples for 24 hours, the organic micro-pollutants adsorbed in the HLB column were eluted with LC-MS grade methanol. The eluted samples were then concentrated by a nitrogen analyzer.
[0025] RNA purification: First, the quantity and purity of total RNA were analyzed, and sequencing libraries were constructed using high-quality RNA samples with RIN numbers > 7.0. After total RNA extraction, mRNA in the total RNA (5 μg) was purified twice using Dynabeads Oligo. After purification, the mRNA was fragmented into short fragments using divalent cations at high temperature. The fragmented RNA fragments were then reverse transcribed using SuperScript™ II reverse transcriptase to create cDNA, and U-labeled second-stranded DNA was synthesized using E. coli DNA polymerase I, RNase H, and dUTP solution. After treating the U-labeled second-stranded DNA with the thermosensitive UDG enzyme, the ligation product was amplified by PCR.
[0026] Amplification conditions: initial denaturation at 95℃ for 3 min; denaturation at 98℃ for 15 s, annealing at 60℃ for 15 s, extension at 72℃ for 30 s, 8 cycles; final extension at 72℃ for 5 min.
[0027] (5) Sequencing: The transcriptome was sequenced using the Illumina paired RNA-seq method, generating millions of 2×150bp paired reads. The reads were further filtered using Cutadapt (https: / / cutadapt.readthedocs.io / en / stable / , version:cutadapt-1.9).
[0028] Filtering conditions: 1. Remove readings containing aptamers; 2. Remove readings containing polyA and polyG; 3. Remove readings containing more than 5% unknown nucleotides (N); 4. Remove low-quality readings containing more than 20% low-quality (Q value ≤ 20) bases.
[0029] (6) Differentially expressed gene analysis: First, the expression levels of all transcripts were estimated, and the mRNA expression abundance was accurately assessed by calculating the FPKM (number of fragments per thousand bases per million mapped transcripts). Then, genes were screened. For these retained genes, the "TMM" algorithm of the edgeR package in R software was used to normalize the transcriptional expression levels between samples. Next, the log2 foldchange (log2FC) of each gene under different exposure concentrations was calculated.
[0030] Next, using the drc and DoseFinding packages in R software, nine dose-response models were applied to each gene individually. These models primarily included three curve types: "S-shaped curve," "linear curve," and "U-shaped curve." The dose-response model with the lowest Akaike's Information Criterion (AIC) value was selected as the best-fit model for that gene. Genes with significantly fitted best-fit models (p-value < 0.05) were defined as differentially expressed genes (DEGs). Then, the initial response concentration value (PODDEG) for each DEG was calculated using the concentration-response relationship of the best-fit models. If the best-fit model was an "S-shaped curve," PODDEG represented the concentration corresponding to the 50% highest response; if the best-fit model was a "linear curve," PODDEG represented the concentration corresponding to a 1.5-fold change in expression; and if the best-fit model was a "U-shaped curve," PODDEG represented the concentration corresponding to a 1.5-fold change in expression on the monotonically rising or falling portion of the non-monotonic curve.
[0031] (7) Quantitative analysis of biological processes: First, GO functional enrichment analysis was performed on DEGs. GO enrichment analysis provided all GO terms that were significantly enriched in DEGs compared to the background genome. First, all DEGs were mapped to GO terms in the GO database (http: / / www.geneontology.org / ), the number of genes for each GO term was calculated, and only GO terms with ≥3 matching DEGs were retained. The GO terms that were significantly enriched in DEGs compared to the background genome were determined by hypergeometric test. Subsequently, quantitative analysis of biological processes was performed. The geometric mean of the PODDEG of the DEG to which the GO term was matched was calculated as the initial response concentration (PODGO) of the GO term. The PODGO of the GO terms enriched in each environmental sample was sorted by percentage, and a four-parameter dose-response curve was fitted in GraphPad software to obtain the biological potency distribution curve (BDC), and the top ten GO terms affected by interference in each environmental sample were screened out.
[0032] The formula for calculating percentage sorting is: GO terms, percentage sort = GO terms, sort / maximum value of all GO terms sorted × 100% Predicting Key Toxic Structures: First, a test gene set was compiled. DEGs from various environmental samples obtained through transcriptome sequencing were used as the test set. Second, a background gene set was compiled. 1,613,109 chemical interactions from the Comparative Toxicogenomics Database (CTD) were collected and compiled as the background set, with chemical relationships derived from published articles. Third, the test set was matched to the background set, and a list of chemical substances significantly enriched in the test gene set was obtained using hypergeometric distribution calculations (phyper function). Chemical substances with a P-value < 0.05 were compiled into a list of suspected chemical substances based on the CTD database. Then, ChemoTyper software was used to deconstruct the substances in the CTD background chemical substance list and the CTD suspected chemical substance list to obtain the names and numbers of structures contained in each chemical substance. Twenty chemical substances were randomly selected from the CTD background chemical substance list and the CTD suspected chemical substance list, and the frequency of each structure in both lists was obtained. This process was repeated 20 times. A t-test was used to compare the frequency of each structure in the CTD background chemical substance list and the CTD suspected chemical substance list. Finally, structures with significantly higher frequency of occurrence in the CTD suspected chemical substances list than in the CTD background chemical substances list, and with a P-value < 0.05, were included in the list of key toxic structures for transcriptomics prediction. The flowchart is shown below. Figure 1 .
[0033] Non-targeted chemical screening by gas chromatography-mass spectrometry (GC-MS): The surface water extract from the Qiantang River, enriched at 2000 REF, was filtered through a 0.22 μm organic filter membrane and transferred to a 1.5 mL brown vial containing a 250 μL inner tube. The vial was stored at -20°C and then subjected to non-targeted chemical screening using a two-dimensional GC-MS / QTOF system (Agilent 8890-7250 GC / Q-TOF). The specific experimental conditions are as follows: Gas chromatography column 1 (15m × 250μm × 0.25μm, Agilent), column 2 (15m × 250μm × 0.25μm, Agilent). Column temperature program: initial temperature 60℃, hold for 1 min; ramp to 120℃ at a rate of 40℃ / min, hold for 2 min; ramp to 240℃ at a rate of 15℃ / min, hold for 10 min; ramp to 300℃ at a rate of 5℃ / min, hold for 10 min. Column 1 pressure: 11.712 psi; carrier gas: helium (purity ≥99.999%), flow rate: 1 mL / min. Column 2 pressure: 6.1495 psi; carrier gas: helium (purity ≥99.999%), flow rate: 1.1 mL / min. Injection volume: 1 μL; split ratio: 300:1; split flow rate: 300 mL / min. Inlet temperature: 250°C; MSD transfer line temperature: 290°C. The mass spectrometer used an LE-EI ion source with an ion source temperature of 200℃, an electron energy of 70eV, a sampling rate of 1Hz, and a sampling time of 1000ms / Spectrum.
[0034] Data processing workflow: After data acquisition, peak extraction, and data processing, the data is matched with the NIST17 mass spectrometry library to export a table containing information such as CAS number, substance name, retention time, similarity, probability, peak area, signal-to-noise ratio, and ion mass number. The table is then filtered to remove interference items such as impurities and column bleed.
[0035] Screening criteria: 1. Retain substances with a matching degree > 60. 2. Delete substances containing silicon. 3. When deleting duplicate substances, retain those with higher matching degrees and higher peak areas.
[0036] (10) Combining Targeted Analysis with Transcriptomics: First, peak areas were extracted from the non-targeted screening results, and the original data were log-transformed. Second, pathway POD values were extracted from the KEGG pathway results from transcriptomics, and the original data were transformed using the formula y=-x+Max (x= GeoMean). Third, the LC50 values of the 96-hour developmental acute toxicity test results for different loci were extracted. Finally, the above results were integrated into an Excel spreadsheet, and correlation analysis was performed using the cor function in R software to obtain the correlation coefficient values of the peak areas of the non-targeted screening substances, the POD values of the key KEGG pathway, and the LC50 values of the 96-hour developmental acute toxicity test.
[0037] Targeted quantitative analysis of chemical substances selected from the list of critical toxic substances that are readily available standard chemicals was performed using a liquid chromatography-mass spectrometry (SCIEX Triple Quad™ 6500+ LC-MS / MS) instrument.
[0038] Chromatographic conditions: Chromatographic conditions (positive ion mode)
[0039] Chromatographic conditions (negative ion mode)
[0040] Mass spectrometry conditions:
[0041] Example 1; This example describes a method for identifying toxic substances based on dose-dependent omics effects, with the technical route as follows: Figure 1 The steps are as follows: (1) Sample collection: The collected surface water is stored in a container. Before use, the container is cleaned with tap water, methanol and deionized water in sequence, and it is rinsed with the surface water sample to be collected before filling it with surface water. Then, it is stored in a refrigerator at 4°C until it is sent to the laboratory for pretreatment.
[0042] (2) Sample extraction: After the collected Qiantang River water samples were coarsely filtered through a 0.45 μm pore size filter membrane to remove suspended solids, organic micropollutants in the Qiantang River surface water samples were enriched using a hydrophilic-lipophilic balance (HLB, 200 mg, 6 mL) column activated with ultrapure water and HPLC-grade methanol. The flow rate of the HLB column was controlled at 3-5 mL / min.
[0043] Drying and elution: After extraction, the water in the HLB column was removed using a freeze dryer. Each batch of samples was freeze-dried for 24 hours, and then the adsorbed organic micro-contaminants in the HLB column were eluted with LC-MS-grade methanol. The eluted sample was then concentrated using a nitrogen analyzer. The sample was divided into two parts according to requirements, for chemical analysis and biological analysis, respectively.
[0044] Total RNA extraction: First, remove dead embryos from each sample well. Then, transfer all zebrafish embryos from each sample well to the corresponding RNAase-free 1.5 mL centrifuge tube, aspirate the exposure liquid, and add 1 mL of RNAiso Plus and two 2 mm diameter sterile ZrO2 grinding beads to each centrifuge tube. Temporarily place the tubes in an ice box for cryogenic storage. After adding RNAiso Plus and grinding beads to all samples, thoroughly grind them using a high-speed cryogenic tissue homogenizer. Subsequently, add an equal volume of 400 μL of isopropanol to each centrifuge tube, shake vigorously to mix the solution thoroughly, and let stand at room temperature for at least 20 min to precipitate the RNA. After precipitation, balance the mixture and centrifuge for 10 min. Wash the RNA by aspirating all the supernatant with a pipette tip, leaving only the white RNA precipitate. Add 700 μL of 80% anhydrous ethanol along the wall of the centrifuge tube, invert the tube to allow the RNA to float, and repeat the step once, washing the RNA twice in total. Centrifuge for 5 minutes, aspirate the supernatant, open the centrifuge tube cap, and wait for the RNA to dry completely and become transparent. Then, add 20 μL of DEPC water to the dried precipitate, gently blow and agitate to dissolve the RNA evenly, and temporarily store it in an ice box.
[0045] After briefly shaking to mix the sample PCR tubes, place them in a PCR instrument and incubate for 5 minutes at 65°C to complete the RNA melting reaction. Remove the PCR tubes and place them in an ice pack to cool. Once cooled, add 8 μL of the reverse transcription kit reagents to the original 12 μL system to prepare a 20 μL reverse transcription system. The 8 μL reagents include: 5 × Reaction Buffer (4 μL), 10 mM dNTP Mix (2 μL), RNA polymerase (1 μL), and RNA Inhibitor (1 μL). After briefly shaking to mix the sample PCR tubes, place them in the PCR instrument and start the reverse transcription program.
[0046] RNA reverse transcription conditions:
[0047] Amplification conditions: initial denaturation at 95℃ for 3 min; denaturation at 98℃ for 15 s, annealing at 60℃ for 15 s, extension at 72℃ for 30 s, 8 cycles; final extension at 72℃ for 5 min.
[0048] Sequencing: The transcriptome was sequenced using Illumina paired-end RNA-seq, generating millions of 2×150bp paired-end reads. The reads were further filtered using Cutadapt (https: / / cutadapt.readthedocs.io / en / stable / , version:cutadapt-1.9). A total of 659.16 Gb of clean data was obtained.
[0049] Filtering conditions: 1. Remove readings containing aptamers; 2. Remove readings containing polyA and polyG; 3. Remove readings containing more than 5% unknown nucleotides (N); 4. Remove low-quality readings containing more than 20% low-quality (Q value ≤ 20) bases.
[0050] Differential gene expression analysis: StringTie and ballgown (http: / / www.bioconductor.org / packages / release / bioc / html / ballgown.html) were used to estimate the expression levels of all transcripts, and mRNA expression abundance was accurately assessed by calculating FPKM (number of fragments per kilobase per million mapped transcripts). Genes were then screened. For these retained genes, the TMM algorithm in the edgeR package of the R software was used to normalize the transcriptional expression levels between samples. Next, the log2 fold change (log2FC) of each gene under different exposure concentrations was calculated.
[0051] Next, using the drc and DoseFinding packages in R software, nine dose-response models were applied to each gene individually. These models primarily included three curve types: "S-shaped curve," "linear curve," and "U-shaped curve." The dose-response model with the lowest Akaike's Information Criterion (AIC) value was selected as the best-fit model for that gene. Genes with significantly fitted models (p-value < 0.05) were defined as differentially expressed genes (DEGs). Then, the initial response concentration value (PODDEG) for each DEG was calculated using the concentration-response relationship of the best-fit models. Each Qiantang River surface water extract sample had an average of 1503 differentially expressed genes, with most genes showing linear fit as the best-fit model. The statistical results of the number of differentially expressed genes are shown below. Figure 2 As shown.
[0052] (7) Quantitative analysis of biological processes: Differentially expressed genes were further annotated using the GO database to explore key biological processes. The analysis yielded 84 to 1252 GO terms per sample, with an average of 330 GO terms per Qiantang River surface water extract sample. Among the tested summer Qiantang River surface water extracts, Q3 had the most GO terms, while Q1 had the fewest.
[0053] Based on the initial bioavailability concentration (POD) of zebrafish embryo enrichment GO terminology, overall biological pathway dose-response curves were plotted for 14 samples and a Blank sample from the Qiantang River during summer and winter. The PODGO20 values for all samples were also obtained. This value reflects the overall bioavailability level of the samples; a lower value indicates higher sensitivity and greater susceptibility to environmental pollutants. The bioavailability levels and ranking of all samples in summer are as follows: Q3 (POD...) GO20 =0.723 REF)>Q6 (POD) GO20 =0.702 REF)>Q2(POD) GO20 =0.558 REF)>Q4(POD) GO20 =0.527 REF)>Q7 (POD) GO20 =0.480 REF) > Q5 (POD) GO20 =0.465 REF)>Q1(POD) GO20 =0.399 REF). The dose-response curves of the overall biological pathways in summer are as follows: Figure 3 As shown.
[0054] Key toxic substances were screened using gas chromatography-quadrupole time-of-flight mass spectrometry (GC-QTOF / MS) based on non-targeted gas chromatography-mass spectrometry. The screening was based on matching with the NIST17 mass spectrometry library. After preliminary screening, the number of consensus-recognized chemical substances in summer samples was 131, and the number in winter samples was 275.
[0055] Export the ToxPi scoring results of key toxic substances in the surface water extract of the Qiantang River using non-targeted gas chromatography-mass spectrometry, draw a ranking score chart, and include the highest-scoring substances in the list of key toxic substances. The list of key toxic substances consists of the Top 20 substances, totaling 40 substances.
[0056] This invention elucidates the structures of key toxic substances identified through non-targeted gas chromatography screening. It reveals that the structures of key toxic substances in summer samples include aromatic rings, heterocycles, fused rings, ester groups, amino groups, nitro groups, carbonyl groups, alkenyl groups, alcohol hydroxyl groups, and ethers; while those in winter samples also include aromatic rings, heterocycles, fused rings, ester groups, carbonyl groups, and hydroxyl groups. These structures also appear in the dose-dependent transcriptomics prediction list of key toxic structures. These commonly occurring structures, such as aromatic rings, heterocycles, fused rings, and ester groups, are likely key factors contributing to the toxic effects of surface water extracts from the Qiantang River. The key toxic pathway results are shown in the diagram below. Figure 5 As shown in the figure, the scoring ranking chart of the key toxic substance ToxPi is as follows: Figure 6 As shown.
[0057] Targeted analysis combined with transcriptomics: Seven chemical substances that were readily available as standard chemicals from the list of key toxic substances were selected for targeted validation.
[0058] In this invention, the concentration of methyl pentadecanoate detected by targeted chemical analysis was significantly higher than that of other substances. The detection range of this substance in summer was 23,150 to 27,950 ng / L, with an average of 25,492.86 ng / L; the detection range of this substance in winter was 21,050 to 30,550 ng / L, with an average of 26,021.43 ng / L.
[0059] The embodiments of this application have been described above with reference to the accompanying drawings. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. This application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for identifying toxic substances based on dose-dependent omics effects, characterized in that: Includes the following steps: Step 1: Store the surface water in a container, and clean the container in sequence with tap water, methanol, and deionized water, and then rinse it with the surface water to be collected before filling it with the sample. Step 2: After the water sample is coarsely filtered through a 0.45μm filter membrane, organic micropollutants are enriched using a hydrophilic-lipophilic equilibrium column activated with ultrapure water and HPLC-grade methanol, with the column flow rate controlled at 3-5 mL / min. Step 3: After extraction, freeze-dry the HLB column for 24 hours to remove moisture, elute the adsorbed organic micro-pollutants with LC-MS grade methanol, and concentrate the eluent using a nitrogen analyzer. Step 4: Select total RNA with RIN number > 7.0, purify mRNA and divide it into short fragments, reverse transcribe to synthesize cDNA and prepare U-labeled second-strand DNA, and then amplify by PCR after treatment with thermosensitive UDG enzyme; Step 5: Sequencing was performed using the Illumina paired-end RNA-seq method, and the reads were filtered using Cutadapt software; Step 6: Calculate the FPKM value of mRNA, standardize the transcriptional expression level and calculate log2FC, fit gene expression with 9 dose-response models, screen for significantly fitted differentially expressed genes and calculate PODDEG; Step 7: Perform GO functional enrichment analysis on differentially expressed genes, calculate and rank the PODGO of GO terms, and fit the bioefficacy ranking distribution curve; Step 8: Using differentially expressed genes as the test set, match the background set of the CTD database to obtain a list of suspected chemical substances, deconstruct the structure of the chemical substances, and screen for highly enriched structures using a t-test; Step 9: The sample was analyzed using a two-dimensional gas chromatograph, quadrupole mass spectrometer, and time-of-flight mass spectrometer. After matching with the NIST17 mass spectral library, substances were screened according to the screening criteria. Step 10: Extract the peak area of non-targeted screening substances, the POD value of the KEGG pathway, and the LC50 of acute toxicity test. After data conversion, perform correlation analysis and use liquid chromatography-mass spectrometry (LC-MS) to perform targeted quantitative analysis of standard chemicals.
2. The method for identifying toxic substances based on dose-dependent omics effects according to claim 1, characterized in that: In step four, RNA purification was performed using Dynabeads Oligo to purify 5 μg of total RNA in two rounds of mRNA purification. The mRNA was split at high temperature using divalent cations, and the RNA fragments were reverse transcribed using SuperScript™ II reverse transcriptase.
3. The method for identifying toxic substances based on dose-dependent omics effects according to claim 1, characterized in that: In step four, the PCR amplification conditions are as follows: initial denaturation at 95℃ for 3 min; denaturation at 98℃ for 15 s, annealing at 60℃ for 15 s, extension at 72℃ for 30 s, for a total of 8 cycles; and final extension at 72℃ for 5 min.
4. The method for identifying toxic substances based on dose-dependent omics effects according to claim 1, characterized in that: In step five, the filtering conditions for sequencing reads are as follows: remove reads containing aptamers, remove reads containing polyA and polyG, remove reads containing more than 5% unknown nucleotides, and remove reads containing more than 20% low-quality bases.
5. The method for identifying toxic substances based on dose-dependent omics effects according to claim 1, characterized in that: In step six, the nine dose-response models include three types: S-shaped curve, linear curve, and U-shaped curve. The model with the smallest AIC value is selected as the best-fit model. The PODDEG is calculated as follows: the S-shaped curve corresponds to the concentration at which the highest response is 50%, the linear curve corresponds to the concentration at which the expression level changes by 1.5 times, and the U-shaped curve takes the monotonically rising or falling part corresponding to the concentration at which the expression level changes by 1.5 times.
6. The method for identifying toxic substances based on dose-dependent omics effects according to claim 1, characterized in that: In step seven, GO functional enrichment analysis retains only GO terms that match ≥3 differentially expressed genes, and significantly enriched GO terms are determined by hypergeometric test; PODGO is the geometric mean of PODDEG of differentially expressed genes matched by the GO term.
7. The method for identifying toxic substances based on dose-dependent omics effects according to claim 1, characterized in that: In step nine, the chromatographic columns for gas chromatography-mass spectrometry (GC-MS) non-targeted chemical screening were two Agilent 15m×250μm×0.25μm columns. The column temperature program was as follows: initial 60℃ for 1 min, ramping up to 120℃ at 40℃ / min and holding for 2 min, ramping up to 240℃ at 15℃ / min and holding for 10 min, and ramping up to 300℃ at 5℃ / min and holding for 10 min. The mass spectrometer used an LE-EI ion source with an ion source temperature of 200℃ and an electron energy of 70 eV.
8. The method for identifying toxic substances based on dose-dependent omics effects according to claim 1, characterized in that: In step nine, the screening criteria are: retain substances with a matching degree > 60, delete substances containing silicon, and when deleting duplicate substances, retain those with higher matching degree and peak area.
9. The method for identifying toxic substances based on dose-dependent omics effects according to claim 1, characterized in that: In step ten, the liquid chromatography-mass spectrometry (LC-MS) instrument used was a SCIEX Triple Quad™ 6500+ LC-MS / MS, employing a BEH C18 1.7μm 2.1 nm filter. A 50 mm column was used, with a run time of 7 min in positive ion mode and 5 min in negative ion mode. The ion source was an electrospray ionization source, and the detection mode was MRM (Multiple Ion Reaction Monitoring).
10. The method for identifying toxic substances based on dose-dependent omics effects according to claim 1, characterized in that: In step four, the conditions for RNA reverse transcription are: 65℃ for 5 min, 25℃ for 5 min, 42℃ for 60 min, and 72℃ for 5 min, with one cycle for each.